ICD-10 codes are often treated as simple labels for diseases. In healthcare AI, they are more useful—and more dangerous—than that. They encode diagnoses, symptoms, encounters, complications, and administrative context in a structured vocabulary that can connect clinical notes to analytics, claims workflows, registries, and machine-learning datasets.
For a large language model (LLM), an ICD-10 code is not a diagnosis by itself. It is a context-dependent representation created for a particular documentation, billing, reporting, or classification purpose. A reliable system must therefore understand the note, identify evidence, distinguish confirmed conditions from suspected ones, and avoid inventing specificity that the clinician did not document.
This guide explains how Indian healthcare builders, researchers, and product teams can use ICD-10 codes with LLMs responsibly in 2026.
What ICD-10 codes represent
The International Classification of Diseases, 10th Revision (ICD-10), is maintained internationally by the World Health Organization. Countries and healthcare systems adapt it for local reporting, reimbursement, and operational requirements. India’s coding environment may involve ICD-10 alongside payer, hospital, government-program, and specialty-specific requirements, so a global code list should not be assumed to be the only source of truth.
Codes are hierarchical. A broad category can be refined with additional characters and decimal extensions. For example, E11.9 generally represents Type 2 diabetes mellitus without complications in commonly used ICD-10-CM terminology. However, the exact meaning, permitted use, and billing implications depend on the adopted national or clinical modification.
For an LLM application, store more than the code string:
- Code system and version: ICD-10, ICD-10-CM, or another local modification.
- Description and hierarchy: The human-readable label and parent category.
- Encounter date: The code may describe a condition at one point in a patient’s care.
- Evidence and provenance: The note, result, order, or clinician who supported the code.
- Status: Confirmed, suspected, historical, ruled out, or resolved.
- Role: Primary diagnosis, secondary diagnosis, symptom, complication, or reason for encounter.
Why ICD-10 matters for healthcare LLMs
ICD-10 provides a bridge between unstructured clinical language and structured healthcare workflows. An LLM can read phrases such as “poorly controlled diabetes with neuropathic symptoms,” identify relevant evidence, and suggest candidate codes for review. Structured outputs can then support population health dashboards, chart review, cohort discovery, prior-authorisation preparation, and coding quality audits.
The value is not simply automation. Consistent coding can make longitudinal records easier to search and compare, especially when clinicians use different terminology over time. It can also support evaluation: a team can measure whether a model identifies the right diagnosis family, captures complications, or abstains when documentation is insufficient.
Teams building datasets should also review this practical guide to ICD-10 codes for LLM training, particularly before using codes as labels for supervised fine-tuning or retrieval.
Common LLM use cases
Candidate code suggestion
The model proposes one or more codes from a clinical note, while a qualified coder or clinician makes the final decision. This is safer than allowing the model to post codes directly to a claim.
Clinical data normalisation
An LLM can map variations such as “high blood sugar,” “T2DM,” and “type two diabetes” into a controlled concept, then preserve the original wording and confidence. Mapping should not erase distinctions such as gestational diabetes, drug-induced diabetes, or diabetes with a documented complication.
Cohort discovery
Researchers can use codes to identify potential patient cohorts for studies or service planning. Code-only cohorts often miss patients whose conditions are documented differently, while note-only cohorts may include false positives. Combining structured codes with text evidence produces a more defensible screening workflow.
Record summarisation
A longitudinal summary can group coded conditions by active, historical, and resolved status. The model should show dates and supporting evidence rather than presenting every code as a current diagnosis.
Coding and documentation quality review
Models can flag missing specificity, conflicting documentation, duplicate diagnoses, or codes unsupported by the note. These are review prompts—not automatic findings of fraud or clinician error.
A safer implementation architecture
A production workflow should separate interpretation, terminology lookup, validation, and action.
1. Ingest governed data. Define which notes, encounters, claims, and laboratory records the model may access. Apply role-based access and remove unnecessary identifiers.
2. Extract evidence. Ask the model to return quoted spans, dates, negation, temporality, and certainty before requesting a code.
3. Retrieve from an approved terminology service. Do not rely on memorised code descriptions. Keep the code system, local edition, effective dates, and synonyms explicit.
4. Generate constrained candidates. Use structured JSON or function calling with fields such as code, description, evidence, certainty, and rationale.
5. Validate deterministically. Check that the code exists, is billable where relevant, matches the selected terminology version, and is compatible with encounter rules.
6. Route for review. Require human approval for claims, high-risk diagnoses, new chronic conditions, paediatric cases, and codes with financial or clinical consequences.
7. Log every decision. Preserve model version, prompt or policy version, retrieved terminology, output, reviewer action, and timestamp.
This separation is especially important when a team is building an AI-native product; the principles in AI-native product development are useful for designing review queues, observability, and failure handling from the start.
Prompt and output design
Avoid prompts that ask, “What is the ICD-10 code?” without context. A stronger instruction specifies the code system, local version, task, and abstention policy:
- Return no code when the condition is only suspected or contradicted.
- Do not infer a complication from medication use alone.
- Distinguish current, historical, and family history.
- Return several candidates when documentation is ambiguous.
- Quote the supporting text and identify missing information.
A useful output schema might include code, system, description, diagnosis_role, certainty, evidence_span, documentation_gaps, and requires_review. Numeric confidence should not be presented as clinical certainty unless it has been calibrated against an appropriate validation set.
Evaluation: what to measure
Accuracy alone is inadequate. Evaluate at several levels:
- Code validity: Does the output exist in the approved release?
- Concept accuracy: Does it represent the documented condition?
- Specificity: Did the model avoid both undercoding and unsupported detail?
- Negation and temporality: Did it distinguish “no evidence of pneumonia” from pneumonia?
- Abstention quality: Does it defer when evidence is incomplete?
- Fairness and language coverage: Does performance hold across Indian English, regional-language notes, transliteration, and different facility types?
- Operational impact: Review time, denial rates, correction rates, and clinician burden.
Build a locked test set with expert adjudication. Include confusing abbreviations, comorbidities, copied-forward text, discharge summaries, emergency notes, and multilingual or code-switched documentation. Monitor performance after every model, prompt, terminology, or workflow change; model evaluation practices such as frontier model evaluation can inform the testing discipline, but clinical acceptance criteria must remain domain-specific.
Privacy, governance, and Indian deployment concerns
Clinical notes are sensitive personal data. Apply data minimisation, encryption, access controls, retention limits, audit logs, and contractual safeguards for vendors. Do not send identifiable records to an external model without an approved legal, security, and clinical-governance pathway.
For Indian deployments, map the workflow to the organisation’s obligations under applicable data-protection, health-record, medical-device, and payer requirements. Keep a human accountable for consequential decisions. Explain to users whether an output is a suggestion, a coded record, or a claim-ready value.
Practical checklist for builders
Before launch, confirm that you have:
- Selected the correct ICD-10 variant and release for the use case.
- Obtained an authoritative terminology source and documented licensing.
- Created labelled examples from the target hospitals and specialties.
- Implemented negation, temporality, uncertainty, and provenance handling.
- Added deterministic validation and a clear abstention path.
- Tested Indian language variation and low-resource clinical settings.
- Established reviewer training, escalation, and audit procedures.
- Measured clinical and operational harm, not just model accuracy.
FAQ
Are ICD-10 codes the same in every country?
No. National and clinical modifications differ in descriptions, specificity, reporting rules, and billing use. Always identify the code system and version.
Can an LLM assign ICD-10 codes automatically?
It can suggest candidates, but automatic posting should be limited to low-risk, validated workflows with strong monitoring. Claims and consequential clinical records generally require qualified review.
Should ICD-10 codes be used as the only training labels?
No. Codes may reflect billing or administrative choices and may omit clinical nuance. Pair them with expert-reviewed text evidence, dates, and certainty labels.
What is the best first project?
Start with retrospective coding assistance or chart-review prioritisation. Measure correction rates and reviewer time before expanding to live clinical or claims workflows.