0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · extract atc icd10 codes

How to Extract ATC and ICD-10 Codes from Healthcare Data

  1. aigi

    Why ATC and ICD-10 extraction needs a careful workflow

    Extracting ATC and ICD-10 codes is not simply a matter of finding alphanumeric strings in a medical record. The process must connect a medication or diagnosis to the correct controlled vocabulary, preserve its clinical context, and record how the code was selected. This matters for hospitals, health-tech companies, insurers, researchers, and Indian AI builders working with fragmented clinical data.

    ATC codes classify medicines by anatomical, therapeutic, pharmacological, and chemical properties. ICD-10 codes classify diseases, symptoms, injuries, and other health conditions. They answer different questions: ATC helps explain what medicine was prescribed or administered, while ICD-10 helps describe why care was provided. A reliable pipeline should not treat the two systems as interchangeable.

    Teams building broader clinical NLP systems may also benefit from this guide to ICD-10 codes for LLM training, particularly when designing datasets, labels, and evaluation rules.

    Understand the source data before extracting codes

    Start by inventorying where the information lives. Common sources include:

    • Structured EHR fields for diagnoses, prescriptions, allergies, and procedures
    • Pharmacy and hospital information systems
    • Insurance claims and pre-authorisation records
    • Discharge summaries, progress notes, referral letters, and scanned documents
    • Laboratory or clinical research databases
    • FHIR, HL7, CSV, JSON, and proprietary exports

    The source determines the extraction method. A structured medication field may already contain a product identifier, ingredient, strength, and route. A clinical note may mention a brand name, a misspelling, or a drug only indirectly—for example, “continue metformin.” A scanned prescription may require OCR before any terminology matching is possible.

    For Indian deployments, assess language and script variation early. Clinical notes may combine English with Hindi, Malayalam, Tamil, or regional abbreviations. OCR quality, transliteration, and local brand names can materially affect recall. If your pipeline handles Indian-language documents, review practices described in AI for Malayalam document extraction and extracting data from Malayalam PDF documents.

    A practical extraction pipeline

    1. Ingest and normalize the document

    Convert files into a consistent representation while retaining the original record. Capture document ID, patient or encounter key, source system, timestamp, author, and document type. For PDFs and images, run OCR and store page coordinates so a reviewer can locate the evidence later.

    Normalize encoding, whitespace, punctuation, dosage formats, and common abbreviations. Do not remove negation terms such as “no,” “without,” or “denies”; they are essential for diagnosis interpretation.

    2. Detect medication and diagnosis mentions

    Use a hybrid approach rather than relying on one model:

    • Dictionary matching for known medicines, ingredients, diagnoses, and abbreviations
    • Regular expressions for candidate ICD-10 patterns and structured ATC values
    • Named-entity recognition for free-text mentions
    • Context rules for negation, uncertainty, family history, and historical conditions
    • Section detection to distinguish active diagnoses from past medical history

    A candidate mention is not yet a confirmed code. “Rule out pneumonia” should not automatically become an active diagnosis. Similarly, “allergic to penicillin” identifies a medication-related concept but not a current prescription.

    3. Map mentions to the right terminology

    For medicines, prefer mapping through the active ingredient and, where available, strength and formulation—not just a brand name. Then map the ingredient to the appropriate ATC level required by your use case. For diagnoses, map the documented clinical concept to the applicable ICD-10 release and national modification, if one is being used.

    Keep terminology versions in the output. A code without its code system, release, and description is difficult to audit. A useful record might include:

    • source_text
    • concept_type — medication or diagnosis
    • normalized_concept
    • code_system — ATC or ICD-10
    • code
    • coding_version
    • certainty
    • negation_status
    • evidence_span
    • confidence_score
    • review_status

    Do not silently convert ICD-10 into ICD-10-CM, ICD-10-AM, or another national adaptation. These variants can differ in granularity and meaning. Confirm the terminology expected by the hospital, payer, regulator, or research protocol.

    4. Validate against authoritative sources

    Validation should include both syntactic checks and clinical checks. Confirm that a candidate code exists in the selected release, has the expected format, and corresponds to the extracted concept. For ATC, check that the substance and classification level are consistent. For ICD-10, check inclusion, exclusion, laterality, encounter, and complication guidance where applicable.

    Use licensed or authoritative terminology services where required. Public examples and open-source libraries are useful for prototyping, but production systems need a documented source, update process, and licence review. Never assume that a code found through a general web search is current or appropriate for billing.

    5. Add human review for uncertain cases

    Route low-confidence or clinically consequential cases to a trained coder, pharmacist, clinician, or health-information specialist. Display the original evidence, proposed code, alternatives, and reason for uncertainty. This is more useful than presenting a single unexplained model prediction.

    Active-learning workflows can prioritize examples that will improve the model most: brand-name ambiguity, short abbreviations, multilingual text, negated diagnoses, and codes with frequent manual corrections. Teams building larger extraction systems can apply the same principles used in AI research agents for complex data extraction, while maintaining stricter clinical governance.

    SQL and AI approaches: when to use each

    SQL is appropriate when codes already exist in structured tables. Use explicit joins, terminology-version filters, encounter constraints, and duplicate handling. Preserve raw values alongside normalized values so corrections remain traceable.

    AI or NLP is useful when codes must be inferred from narrative text, images, or inconsistent exports. A robust architecture often combines OCR, retrieval from a terminology database, deterministic rules, and a language model constrained to return only valid candidates. Avoid asking a general-purpose model to invent codes from memory.

    If you automate extraction with agents, define narrow tools and approval gates. The broader principles in how to automate data extraction using AI agents apply, but healthcare pipelines require additional controls for privacy, auditability, and clinical escalation.

    Quality checks that belong in production

    Track performance separately for ATC and ICD-10, and report more than overall accuracy:

    • Precision and recall by document type and language
    • Exact-code accuracy versus correct higher-level classification
    • Negation and historical-condition errors
    • Brand-to-ingredient mapping accuracy
    • Version and terminology compliance
    • Human-review rate and correction rate
    • Duplicate, missing, and conflicting code frequency

    Build a gold-standard sample reviewed by qualified professionals. Re-test it after every terminology update, OCR change, prompt change, or model release. Monitor drift: new brands, changed documentation habits, and new clinical terminology can reduce performance without any software failure.

    Privacy, security, and Indian deployment considerations

    Clinical records are sensitive personal data. Apply data minimisation, role-based access, encryption, retention controls, audit logs, and secure key management. De-identify data for model development where possible, and document whether processing occurs on-premises, in a private cloud, or through an external API.

    Before deployment in India, align the workflow with the organisation’s legal, procurement, information-security, and clinical-governance requirements. Do not send identifiable records to an AI provider merely because an API is convenient. Establish incident response, vendor review, and a process for correcting patient-linked records.

    A rollout plan for healthcare teams

    Begin with one document type and a narrow set of high-value use cases, such as medication reconciliation or chronic-disease cohort identification. Measure baseline manual effort and error rates. Then:

    • Create a terminology registry with versions and owners
    • Build a labelled evaluation set
    • Launch in shadow mode without affecting billing or care decisions
    • Review errors with coders and clinicians
    • Add confidence thresholds and escalation paths
    • Expand only after accuracy, security, and operational metrics are stable

    The goal is not maximum automation. It is reliable, explainable extraction that reduces repetitive work without obscuring clinical judgement.

    Last updated 24 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.