Electronic nose using LLM shows potential in few-shot lung cancer classification

5 hours ago
Electronic nose using LLM shows potential in few-shot lung cancer classification

A natural language processing-pretrained large language model (LLM) has shown its capabilities for few-shot lung cancer classification in a real-world mixed clinical cohort using electronic nose (eNose) breathprints, resulting in lesser dependence on large training datasets, according to a study.

The investigators obtained eNose breathprints of lung cancer and nonlung cancer patients from two medical centres in Taiwan. They compared a GPT-2‒backbone LLM with parameter-efficient adaptation with convolutional neural networks (CNN) trained from scratch or pretrained on CIFAR-100 and assessed few-shot protocols (2‒6 shots per class) and full-data training.

A total of 432 eNose breathprints were collected from two sites (S1 and S2). On S1, LLM yielded an area under the curve (AUC) of 0.79 (95 percent confidence interval [CI], 0.71‒0.87), sensitivity of 0.74 (95 percent CI, 0.63‒0.83), and specificity of 0.77 (95 percent CI, 0.67‒0.87), with six labelled samples per class. On S2, LLM had an AUC of 0.76 (95 percent CI, 0.69‒0.82), sensitivity of 0.77 (95 percent CI, 0.69‒0.84), and specificity of 0.61 (95 percent CI, 0.51‒0.70).

LLM performed better than scratch CNN models (S1: AUC, 0.44; p=0.0002; S2: AUC, 0.63; p=0.0198) and CNN pretrained on CIFAR-100 images (S1: AUC, 0.57; p=0.0100; S2: AUC, 0.61; p=0.0248).

“LLM or a CNN model trained on the source site fails to improve performance after transferring to the target site for fine-tuning,” the investigators said “[F]or the LLM, performance even deteriorates.”

Respirology 2026;31:912-921