Tuberculosis (TB) remains one of India’s most significant public-health challenges, yet many people are diagnosed late because conventional testing can be costly, facility-dependent, or difficult to access. Acoustic biomarkers for non-invasive tuberculosis screening offer an emerging approach: analysing cough, breathing, speech, or other respiratory sounds with signal-processing and machine-learning systems to identify patterns associated with pulmonary TB.
This technology is not a replacement for microbiological confirmation. Its strongest near-term role is likely to be triage—identifying people who should receive confirmatory testing such as sputum molecular testing or chest imaging. If developed responsibly, acoustic screening could extend case finding to primary-health centres, mobile clinics, workplaces, schools, and community settings.
What are acoustic biomarkers?
An acoustic biomarker is a measurable feature in an audio signal that may correlate with a disease state, physiological process, or treatment response. For TB screening, researchers may study sounds generated by:
- Coughing, including cough intensity, duration, repetition, and spectral characteristics
- Breathing, including wheeze-like components, crackles, rhythm, and airflow-related changes
- Speech and sustained vowels, which may capture breath control, vocal-tract effects, or systemic illness
- Combined audio events, such as cough followed by breathing or short clinical prompts
These signals are converted into numerical representations. Common features include Mel-frequency cepstral coefficients (MFCCs), spectral centroid, spectral bandwidth, zero-crossing rate, formants, energy envelopes, temporal statistics, and learned embeddings from neural networks.
A biomarker is useful only when it is reproducible, clinically meaningful, and sufficiently robust across populations and recording conditions. A model that performs well on one curated dataset may fail when deployed on low-cost phones in noisy Indian communities.
Why acoustic screening is relevant to tuberculosis
Pulmonary TB can affect the airways and lung tissue, producing symptoms such as persistent cough, sputum, chest discomfort, fever, fatigue, and weight loss. These symptoms are not specific to TB; asthma, pneumonia, chronic obstructive pulmonary disease, COVID-19, smoking, and other infections can generate overlapping acoustic patterns.
That overlap makes acoustic analysis a difficult classification problem, but it also creates a practical opportunity. Audio is inexpensive to capture, can be collected remotely, and may support screening where radiography or laboratory infrastructure is limited. A smartphone-based system could potentially guide a person toward confirmatory testing without requiring specialised hardware at the first point of contact.
The realistic clinical question is therefore not “Can audio diagnose TB?” but rather:
> Can an acoustic model identify people at sufficiently high risk of pulmonary TB to justify prompt confirmatory testing?
How an AI-powered acoustic screening pipeline works
A credible system usually includes six stages.
1. Guided audio acquisition
The application should provide simple instructions, for example asking the participant to record several voluntary coughs, normal breathing, a deep breath, or a short spoken phrase. The protocol must account for fatigue, language diversity, age, privacy, and the possibility that some people cannot produce a strong cough.
Recording metadata should include device type, microphone characteristics where available, sampling rate, environment, distance from the phone, and whether a mask was worn. These variables are essential for analysing dataset shift.
2. Quality control and event detection
Before classification, the system should detect speech, cough, breathing, background noise, clipping, silence, and overlapping voices. Low-quality recordings can either be rejected with a clear retake instruction or passed through a model trained specifically for noisy conditions.
Quality control should never become an invisible exclusion mechanism. People in crowded homes, rural areas, or low-connectivity settings may systematically produce noisier recordings. Excluding them can create biased performance and reduce public-health value.
3. Signal preprocessing
Typical preprocessing may involve resampling, amplitude normalisation, denoising, voice-activity detection, segmentation, and conversion to time-frequency representations such as log-Mel spectrograms. Care is required: aggressive denoising may remove clinically relevant information or create artificial patterns.
4. Feature extraction
Traditional machine-learning systems may use engineered features such as MFCCs, spectral roll-off, harmonicity, pitch variation, cough duration, and energy distribution. Deep-learning systems can learn representations directly from spectrograms or raw waveforms using convolutional neural networks, recurrent networks, transformers, or self-supervised audio encoders.
A hybrid approach can be valuable. Learned embeddings may capture complex patterns, while interpretable temporal and spectral features help researchers understand model behaviour and monitor data quality.
5. Risk prediction
The model may produce a probability or risk score rather than a binary diagnosis. A clinical decision threshold should be selected according to the intended use. Screening generally prioritises sensitivity, because missed TB cases can continue transmission, although excessive false positives can overload diagnostic services.
The output should be phrased carefully: “Recommend confirmatory TB testing” is more appropriate than “You have tuberculosis.”
6. Referral and confirmatory testing
An acoustic result has value only if it connects to care. The workflow should support referral for a molecular test, sputum examination where clinically appropriate, chest radiography, and professional evaluation. In India, integration with existing public-health pathways and the National Tuberculosis Elimination Programme (NTEP) would be more useful than creating a standalone consumer app.
Model architectures and technical design choices
Researchers can evaluate several modelling strategies:
- Classical models: logistic regression, random forests, support-vector machines, and gradient-boosted trees trained on engineered acoustic features
- Spectrogram-based deep learning: CNNs or vision transformers applied to log-Mel or other time-frequency images
- Sequence models: CNN-LSTM, temporal convolutional networks, or audio transformers for cough and breathing sequences
- Self-supervised learning: pretrained audio encoders fine-tuned on TB-labelled data, potentially reducing the need for large labelled datasets
- Multimodal models: audio combined with age, symptoms, smoking status, oxygen saturation, or chest-imaging features
The best architecture is not necessarily the largest one. In field settings, latency, battery use, offline operation, model size, and calibration may matter more than a small improvement in laboratory accuracy. Quantisation, pruning, knowledge distillation, and on-device inference can help deploy models on affordable Android phones.
Dataset requirements: the central challenge
Acoustic TB research is highly vulnerable to dataset bias. A useful dataset should include confirmed TB-positive participants and appropriate controls, preferably with microbiological reference standards. Controls should represent realistic alternatives, including healthy individuals and people with pneumonia, asthma, COPD, upper-respiratory infections, post-TB lung disease, and smoking-related cough.
Important diversity dimensions include:
- Age, sex, pregnancy status, and comorbidities
- Pulmonary versus extrapulmonary disease context
- Smear status, bacterial burden, and drug-resistant TB where relevant
- HIV status and immunosuppression, subject to ethical safeguards
- Regional languages, accents, and cultural recording practices
- Rural, peri-urban, and urban environments
- Smartphone brands, microphones, operating systems, and network conditions
- Indoor, outdoor, clinic, and community noise profiles
Data splitting must happen at the participant level, not at the individual cough level. Otherwise, the model may memorise a person’s voice or recording environment and produce inflated results. External validation on a geographically and technically separate cohort is essential.
Evaluation metrics that matter clinically
Accuracy alone is inadequate for TB screening, especially when prevalence is low. Researchers should report:
- Sensitivity and specificity
- Positive and negative predictive values at realistic prevalence levels
- Receiver operating characteristic area under the curve (ROC-AUC)
- Precision-recall curves and area under the precision-recall curve
- Calibration, including reliability curves and Brier score
- Likelihood ratios and decision-curve analysis
- Performance by subgroup and recording condition
- Referral yield: confirmed TB cases identified per number screened
- Number needed to screen and downstream diagnostic workload
A model with high AUC can still be poorly calibrated or operationally unusable. Thresholds should be chosen with clinicians and public-health teams, considering confirmatory-test capacity and the consequences of false negatives and false positives.
India-specific deployment considerations
India’s TB burden, linguistic diversity, uneven connectivity, and wide variation in healthcare access make it a demanding but important test environment. A deployable solution should be designed for low-bandwidth operation, local language instructions, affordable Android devices, and assisted screening by community health workers.
Privacy is especially important because voice recordings are biometric-like data. Systems should minimise collection, encrypt data in transit and at rest, define retention periods, and obtain informed consent in understandable language. Where feasible, feature extraction and risk inference should occur on-device, with only necessary results transmitted to a secure clinical system.
Indian developers should also consider the Digital Personal Data Protection framework, applicable health-data governance requirements, institutional ethics review, and medical-device or software-as-a-medical-device obligations where the product makes regulated clinical claims. Regulatory classification depends on intended use, claims, risk, and implementation; legal and regulatory review should occur before deployment.
Most importantly, the tool must not widen disparities. Performance should be tested across genders, dialects, socioeconomic groups, disabilities, and device types. A system that works only for clear recordings from premium phones is unlikely to deliver equitable screening.
Clinical validation and regulatory pathway
A promising retrospective result is not enough for adoption. A robust validation programme can progress through:
1. Technical feasibility: assess recording quality, reproducibility, and robustness to noise.
2. Retrospective clinical validation: test locked models against confirmed cases and realistic controls.
3. Prospective observational validation: evaluate performance in the intended screening population before clinicians act on the score.
4. Implementation study: measure referral completion, turnaround time, diagnostic yield, user acceptance, and health-worker workload.
5. Impact evaluation: determine whether the tool increases earlier detection and reduces transmission or diagnostic delay.
Pre-specifying the analysis plan, locking the model before prospective testing, reporting missing data, and publishing subgroup results improve credibility. Independent replication is particularly important because audio datasets can contain hidden site-specific cues.
Limitations and safety risks
Acoustic biomarkers may be affected by smoking, age-related voice changes, seasonal infections, language, microphone placement, background conversations, and coexisting lung disease. TB is biologically heterogeneous, and some infectious individuals may have little or no cough. A negative acoustic result must not delay care when symptoms or clinical risk remain high.
Other risks include automated stigma, unauthorised reuse of recordings, overdiagnosis, referral-system overload, and model drift as devices and populations change. Human oversight, transparent patient communication, audit logs, threshold monitoring, and periodic revalidation should be built into the product from the beginning.
Practical roadmap for founders and researchers
Teams developing acoustic TB screening should prioritise the following:
- Define the intended use as screening or triage, not unverified diagnosis.
- Build a clinically grounded data-collection protocol with microbiological reference standards.
- Capture realistic negative controls and environmental variation.
- Use participant-level splits and external validation sites.
- Measure calibration, subgroup fairness, and referral workload—not just AUC.
- Design for offline or intermittent-connectivity operation.
- Include clinicians, microbiologists, public-health experts, community workers, and affected communities.
- Treat recordings as sensitive personal data and document governance decisions.
- Plan prospective validation and regulatory consultation early.
- Connect every positive risk score to an actionable confirmatory-testing pathway.
What the future may hold
Acoustic biomarkers are unlikely to replace sputum molecular testing, radiology, or clinical judgement in the near term. Their potential lies in making the first step of the TB pathway more accessible: screening more people, identifying risk earlier, and directing limited diagnostic capacity toward those most likely to benefit.
Progress will depend less on impressive demonstrations and more on high-quality datasets, transparent validation, responsible data governance, and integration with India’s existing TB services. Multimodal systems that combine audio with symptoms and other low-cost measurements may eventually improve sensitivity, but each additional signal introduces new privacy, bias, and implementation questions.
FAQ: Acoustic Biomarkers for Non-Invasive Tuberculosis Screening
Can acoustic biomarkers diagnose tuberculosis?
Not on their own. Acoustic models should currently be treated as screening or triage tools that recommend confirmatory testing. A diagnosis requires appropriate clinical and laboratory evaluation.
What sounds are analysed for TB screening?
Research may analyse cough, breathing, sustained vowels, speech, or combinations of these signals. The useful patterns may involve timing, frequency content, intensity, and learned audio representations.
Can a smartphone record the necessary data?
Potentially, yes. Smartphones can capture audio at low cost, but performance depends on microphone quality, distance, noise, software design, and model robustness. Field validation is essential.
Will the technology work for all types of tuberculosis?
It is primarily being explored for pulmonary TB, where respiratory symptoms may produce measurable audio signals. It is not expected to identify extrapulmonary TB reliably through sound alone.
What should happen after a high-risk result?
The person should be referred promptly for confirmatory testing and clinical assessment. The app should provide clear, non-stigmatising instructions and should not present the risk score as a definitive diagnosis.
Apply for AI Grants India
Are you an Indian AI founder building responsible technology for TB screening, digital health, or public-health access? Apply through AI Grants India to explore support for validating and scaling your AI innovation.