0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · atc icd10 codes for llm

ATC ICD-10 Codes for LLMs: Mapping, Validation and Use

  1. aigi

    Large language models can extract medication and diagnosis information from clinical notes, but they should not be allowed to invent or casually equate coding systems. ATC classifies medicines; ICD-10 classifies diseases, symptoms, and health conditions. The phrase “ATC ICD-10 codes for LLM” usually refers to using an LLM to identify, normalize, validate, or relate codes from both systems—not to a single combined code set.

    For Indian healthcare teams, this distinction matters across electronic medical records, claims workflows, pharmacovigilance, hospital analytics, and clinical research. It also matters for AI governance: a model-generated code is a suggestion requiring evidence and review, not a diagnosis or billing decision.

    ATC and ICD-10: Different systems, different questions

    The WHO Anatomical Therapeutic Chemical (ATC) classification groups medicines by:

    • The organ or anatomical system they act on
    • Their therapeutic, pharmacological, and chemical properties
    • A hierarchical code structure, typically ending in a specific chemical substance

    For example, metformin is classified as A10BA02. That code describes the medicine and its pharmacological classification; it does not establish why a particular patient received it.

    ICD-10 records diseases, disorders, symptoms, and related health conditions. A code such as E11 indicates type 2 diabetes mellitus at a broad category level, while a country-specific implementation may require a more detailed subcode. ICD-10 coding conventions vary by jurisdiction and use case, so teams must use the applicable national or payer rule set rather than assume that a WHO category is sufficient for reimbursement.

    The two systems can be linked in a clinical workflow, but the relationship is contextual. One medicine may be prescribed for several conditions, and one condition may be treated with many medicines. A model should therefore preserve the distinction between:

    • Medication identity: generic name, strength, formulation, and route
    • ATC classification: the medicine’s standard therapeutic classification
    • Clinical indication: the reason documented by the clinician
    • ICD-10 diagnosis: the supported disease or condition code
    • Evidence and certainty: where the relationship appears in the record and how reliable it is

    What LLMs can do with ATC and ICD-10 data

    An LLM is useful as a language-processing layer around authoritative terminology services. Common tasks include:

    • Extracting medicines, diagnoses, symptoms, routes, and dosages from notes
    • Normalizing brand names and misspellings to generic substances
    • Suggesting ATC candidates from medication descriptions
    • Suggesting ICD-10 candidates from documented diagnoses
    • Detecting missing specificity, conflicting dates, or unsupported diagnoses
    • Producing a structured review queue for coders and clinicians
    • Explaining why a candidate code was selected, with a citation to the source text

    For implementation details, teams working on clinical extraction can compare this workflow with ATC ICD-10 code extraction and validation, while teams preparing model-training datasets should separate it from ICD-10 codes for LLM training.

    The model should not independently decide that a medication proves a diagnosis. Metformin may be used in type 2 diabetes, prediabetes, polycystic ovary syndrome, or other contexts. Insulin use does not, by itself, establish type 1 diabetes. Likewise, a medicine’s ATC code cannot replace the clinician’s documented indication.

    A practical mapping workflow

    A dependable pipeline should treat mapping as a sequence of constrained decisions rather than a single prompt.

    1. Define the coding target. Decide whether the output is an ATC code, an ICD-10 code, a medication-to-indication relationship, or all three.
    2. Capture the source span. Store the exact sentence, section, date, and author context supporting each extracted fact.
    3. Normalize the medicine. Resolve brand names, abbreviations, combination products, strength, formulation, and route before looking up ATC.
    4. Normalize the diagnosis. Distinguish confirmed diagnoses from suspected conditions, symptoms, history, rule-outs, and negated statements.
    5. Query authoritative terminology sources. Use current ATC and locally applicable ICD-10 references. Do not rely on model memory for code validity.
    6. Generate constrained candidates. Ask the model to choose only from retrieved terminology entries, or to return “no suitable candidate.”
    7. Validate the relationship. Check whether the note explicitly supports the medication, diagnosis, indication, encounter date, and level of specificity.
    8. Route uncertain cases to review. Human coders should handle ambiguity, conflicting documentation, and consequential billing or reporting decisions.

    A useful output schema might include medicine_text, normalized_medicine, atc_code, diagnosis_text, icd10_code, evidence_span, negation_status, confidence, terminology_version, and review_status. Versioning is essential because terminology releases and local coding rules change.

    Validation controls that reduce hallucinations

    LLM coding systems need controls at three levels.

    Terminology validation: Confirm that every returned code exists in the selected release, has the expected format, and maps to the retrieved label. Reject invented codes and obsolete entries.

    Clinical validation: Check for negation, uncertainty, temporal context, and contradictions. “No history of asthma” must not become an asthma diagnosis. “Rule out pneumonia” is not equivalent to confirmed pneumonia.

    Workflow validation: Require review for low-confidence outputs, high-impact diagnoses, incomplete notes, and claims submission. Log who accepted or changed a suggestion. In India, deployments should also align with applicable health-data protection, hospital, payer, and clinical-governance requirements.

    A retrieval-augmented design is usually safer than fine-tuning alone: retrieve the approved terminology record, expose the relevant definition to the model, constrain the response, and run deterministic checks afterward. Evaluate not only exact-match accuracy but also false-positive rate, unsupported-code rate, abstention quality, and performance across Indian English, regional language translations, abbreviations, and brand names.

    Common mistakes to avoid

    • Treating ATC and ICD-10 as interchangeable code sets
    • Inferring a diagnosis solely from a prescribed medicine
    • Returning broad ICD-10 categories when documentation supports—or requires—a more specific code
    • Mixing WHO ICD-10 with a national modification without labeling the version
    • Ignoring combination products, paediatric formulations, or route differences
    • Removing negation and temporality during text preprocessing
    • Allowing the model to submit claims or update a patient record without review
    • Training on copied code lists without checking licensing, provenance, and release date

    Teams building an LLM coding assistant should also measure frontier model evaluation practices such as reproducible test sets, adversarial examples, and regression testing. If the system calls terminology APIs, database searches, or review tools, explicit Claude model tool orchestration patterns are relevant even when a different model is used.

    Recommended implementation checklist

    Before production, confirm that your system:

    • Uses a documented ATC and ICD-10 release
    • Separates extraction, normalization, code lookup, and validation
    • Preserves evidence spans and provenance
    • Supports abstention instead of forced predictions
    • Tests negation, uncertainty, duplicates, and conflicting notes
    • Measures results by specialty, hospital, language, and note type
    • Applies role-based access, encryption, audit logs, and retention controls
    • Keeps a human reviewer in the loop for consequential decisions
    • Revalidates outputs when terminology releases change

    The strongest use case is not an autonomous coder. It is a transparent assistant that reduces search and clerical effort while making uncertainty visible. With authoritative terminology lookup, evidence-linked outputs, deterministic validation, and trained human review, LLMs can make ATC and ICD-10 workflows faster without weakening clinical accountability.

    Last updated 24 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.