Healthcare AI teams often describe ATC ICD-10 LLM training as if it were a single classification problem. It is not. ATC and ICD-10 represent different clinical concepts, use separate hierarchies, and require different evidence from medical records. A dependable system must distinguish medicines from diagnoses, preserve uncertainty, and support a trained coder rather than silently replacing one.
For Indian hospitals, health-tech companies, insurers, and research teams, the opportunity is practical: reduce repetitive chart review, improve coding consistency, and create structured data for claims, analytics, pharmacovigilance, and clinical research. The hard part is not merely fine-tuning a model. It is building a governed workflow around authoritative code sets, representative Indian documentation, and measurable human oversight.
ATC and ICD-10 are different targets
The Anatomical Therapeutic Chemical (ATC) system classifies medicines according to the organ or system they act on, therapeutic purpose, pharmacological properties, and chemical identity. Its hierarchy has five levels, moving from broad anatomical groups to specific chemical substances. ATC is useful for drug-utilisation analysis, formulary work, pharmacovigilance, and population-level studies.
ICD-10 classifies diseases, disorders, symptoms, injuries, and other health conditions. The correct code depends on documented clinical facts such as diagnosis, site, acuity, laterality, encounter context, and complications. A medicine mentioned in a note does not automatically prove the diagnosis it is commonly used to treat.
That distinction should shape the model architecture. A system may use one shared language encoder but should generally maintain separate ATC and ICD-10 prediction heads, retrieval sources, confidence thresholds, and review queues. Teams working specifically on diagnosis coding should also consult this practical guide to ICD-10 codes for LLM training.
Define the task before selecting a model
“Coding from clinical text” can mean several different tasks:
- Entity extraction: identify medicines, active ingredients, diagnoses, symptoms, procedures, and negations.
- Normalisation: map a phrase such as a brand name or misspelling to a standard clinical concept.
- Code ranking: return likely ATC or ICD-10 candidates with evidence.
- Multi-label coding: assign several valid codes to one encounter.
- Code validation: detect unsupported, contradictory, outdated, or incomplete codes.
- Abstention: route ambiguous cases to a human instead of forcing a prediction.
Start with one narrow, measurable use case—for example, ranking ICD-10 candidates for discharge summaries or mapping documented active ingredients to ATC codes. Avoid training a general-purpose “medical coder” until the team has established reliable labels, audit procedures, and an escalation policy.
Build a trustworthy training dataset
The dataset should combine authoritative terminology resources with de-identified, locally representative clinical text. Public code descriptions alone are insufficient: they teach the model what a code means, not how Indian clinicians document the underlying condition.
A robust dataset plan includes:
- Source inventory: discharge summaries, outpatient notes, prescriptions, laboratory narratives, claims fields, and coding audit records.
- Provenance: source system, document type, timestamp, code-set version, annotator, and approval status.
- Language coverage: English, Indian English, abbreviations, transliterated terms, and regional-language fragments where they occur in production.
- Clinical context: negation, uncertainty, historical conditions, family history, ruled-out diagnoses, and copied-forward text.
- Medication detail: brand and generic names, strength, route, frequency, combination products, and discontinued medicines.
- Split discipline: patient-level and facility-level separation to prevent leakage between training and evaluation.
India-specific notes often contain shorthand, mixed units, inconsistent spelling, and brand-heavy prescriptions. These are not noise to remove automatically; they are part of the deployment environment. Keep the original text, create normalised fields separately, and record every transformation. For broader dataset governance, use the principles in How to audit AI training data integrity.
Annotation and model training workflow
Create an annotation guide before hiring or assigning reviewers. It should specify when a code is supported, how to handle multiple diagnoses, what counts as a principal diagnosis, how to treat suspected conditions, and when the correct action is “insufficient documentation.” Have qualified medical coders review disagreements and periodically measure inter-annotator agreement.
A practical pipeline is:
1. De-identify and classify documents while preserving clinically relevant dates and relationships.
2. Detect sections and entities such as assessment, plan, medication list, and past history.
3. Resolve context using negation, temporality, dosage, and relation extraction.
4. Retrieve candidate codes from versioned ATC and ICD-10 terminology stores.
5. Rank candidates with a fine-tuned encoder, instruction-tuned model, or hybrid retrieval-generation system.
6. Return evidence—the text span, terminology definition, code version, and rationale category.
7. Apply thresholds that determine auto-accept, coder review, or mandatory escalation.
Retrieval-augmented systems are often preferable to relying on model memory. Code sets change, local billing rules vary, and a model may produce plausible but nonexistent codes. Store terminology outside the model and require the output to reference an approved code entry.
Evaluate accuracy and safety—not just F1
A useful evaluation suite should report performance by document type, facility, language pattern, specialty, and code frequency. Include rare and difficult cases rather than reporting only aggregate accuracy.
Track:
- Exact-match and top-k code accuracy.
- Precision, recall, and F1 at the code and encounter levels.
- Unsupported-code rate and hallucinated-code rate.
- Performance on negation, uncertainty, abbreviations, and co-morbidities.
- Calibration: whether confidence scores correspond to actual correctness.
- Abstention quality and percentage of cases safely routed to review.
- Time saved per chart and coder correction rate.
- Claim-denial, audit-finding, and downstream documentation impact where available.
Run a silent pilot first: generate suggestions without changing production decisions. Compare the model with existing coder outcomes, investigate disagreements, and test drift after terminology or workflow changes. Never treat a high benchmark score as permission for unsupervised coding.
Privacy, security, and governance in India
Clinical notes are sensitive personal data. Establish a lawful processing basis, minimum-necessary access, retention limits, role-based permissions, encryption, and an incident-response process. Prefer de-identification before experimentation and keep training, validation, and production environments separate.
For cloud or external model providers, document where data is processed, whether prompts are retained, how subcontractors are controlled, and whether customer data is used for further training. Maintain a model card, dataset register, versioned code references, change log, and human-override record. Add adversarial tests for prompt injection in imported notes and for leakage of patient identifiers.
Teams building their own infrastructure should balance latency and cost against privacy and reliability. Efficient fine-tuning and quantisation may make private deployment more feasible; this overview of post-training quantization explains the main trade-offs.
Deployment pattern for healthcare teams
A safe first release is a coder-assist interface, not an autonomous billing engine. Display the proposed code, supporting text span, terminology version, confidence band, and “why not” warnings. Let the coder accept, modify, reject, or request more documentation. Capture these actions as feedback, but do not automatically treat every correction as a gold label.
Use a staged rollout:
- Phase 1: retrospective evaluation on locked, de-identified records.
- Phase 2: silent production monitoring.
- Phase 3: limited coder-assist deployment in one specialty or facility.
- Phase 4: controlled expansion with periodic revalidation.
Set explicit stop conditions for unexplained accuracy drops, privacy incidents, code-set changes, or rising override rates. Assign ownership across clinical coding, compliance, data science, IT security, and operations.
Bottom line
ATC ICD-10 LLM training is valuable when treated as a clinical decision-support and information-quality project—not a generic chatbot exercise. Use separate task definitions for drug and diagnosis coding, version every terminology source, train on realistic Indian documentation, expose evidence, measure abstention, and keep qualified human review in the loop. A narrow system that is auditable and reliable will create more value than a broad model that produces confident but unsupported codes.