0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to get consistent json from llms

How to Get Consistent JSON from LLMs

  1. aigi

    LLMs are useful for extracting fields, classifying requests, routing workflows, and transforming documents. But a model response is not automatically a dependable API payload. It may add markdown fences, omit a field, return the wrong type, invent a value, or produce valid JSON that still violates your application’s business rules.

    The reliable approach is to treat an LLM as a probabilistic component inside a typed software pipeline. Constrain the response where your provider supports it, validate every result, and make failure handling explicit.

    Define the contract before writing the prompt

    Start with the consumer, not the model. Decide exactly what downstream code needs:

    • Required and optional fields
    • Primitive types, allowed values, and maximum lengths
    • Whether missing information should be null, an empty array, or an error
    • Date, currency, phone-number, and identifier formats
    • Rules for nested objects and arrays
    • Behaviour when the source text does not contain an answer

    A schema should express structure, while application code should enforce domain rules. For example, JSON Schema can require an integer between 0 and 120, but your business logic may still need to check whether a date is plausible or whether a GSTIN matches your expected format.

    Keep the contract small. A narrow extraction schema is easier to validate and more stable than a broad object containing speculative fields. If your service exposes the result to other systems, document it like an API; this pairs well with a practical workflow for generating API specifications with AI LLMs.

    Use structured output features first

    In 2026, many commercial and open-source model stacks offer one or more of these controls:

    • JSON mode: asks the model to return syntactically valid JSON.
    • Structured outputs: constrains generation against a supplied schema.
    • Tool or function calling: asks the model to populate typed arguments.
    • Grammar-constrained decoding: limits tokens to a formal grammar, often useful with self-hosted models.

    Prefer the strongest control supported by your chosen model. JSON mode is not the same as schema adherence: it can prevent a prose answer while still allowing missing fields or incorrect types. Tool calling can be effective for routing and extraction, but validate the arguments before executing any tool.

    Do not assume that a feature works identically across models. Test nested arrays, optional properties, unions, enums, long strings, and refusal or empty-input cases. If you deploy an open model, compare the available constrained-decoding libraries and model-specific limitations before selecting an inference server. Teams evaluating several models can use open-source frameworks for evaluating LLMs to make this comparison repeatable.

    Write prompts that remove ambiguity

    A good prompt complements the schema rather than replacing it. State the task, input boundaries, missing-data policy, and output restrictions in plain language:

    Extract the invoice fields from the supplied text.
    Return only the structured response required by the schema.
    Use null when a field is not present; do not infer values.
    Dates must use YYYY-MM-DD. Amounts must be numbers, not strings.

    Useful prompt practices include:

    • Put the untrusted document between clear delimiters.
    • Separate instructions from source text to reduce prompt injection risk.
    • Give one or two representative examples only when ambiguity remains.
    • Explain distinctions such as unknown versus not applicable.
    • Avoid asking for chain-of-thought; request concise evidence or a confidence field only when it serves a defined review process.
    • Use a stable system prompt and version it alongside the schema.

    Few-shot examples should cover difficult cases, not just ideal inputs. For Indian applications, include realistic variation in names, addresses, pin codes, rupee amounts, date formats, and multilingual text. If the task depends on local training data, review guidance on how to train LLMs on Indian datasets rather than trying to solve every limitation with prompting.

    Validate in layers

    Never pass raw model output directly into a database, payment flow, customer message, or privileged tool. Use a layered pipeline:

    1. Transport check: confirm the response arrived, was not truncated, and contains the expected provider object.
    2. Parsing check: parse JSON with a strict parser. Strip markdown only if your policy explicitly permits it; silent cleanup can hide regressions.
    3. Schema check: validate required fields, types, formats, enums, and additional-property rules.
    4. Business check: verify relationships and domain constraints, such as totals, date ranges, or permitted account actions.
    5. Security check: reject unsafe URLs, unexpected commands, oversized strings, and data that could trigger injection in later systems.

    Use typed models in your application language—such as Pydantic, Zod, Java records, or equivalent—and keep them aligned with the canonical schema. Return a structured error to the caller instead of guessing. For sensitive workflows, route failed or low-confidence records to human review.

    Design retries and fallbacks deliberately

    A retry should change something meaningful. First classify the failure:

    • Transient provider failure: retry with bounded exponential backoff and a request identifier.
    • Malformed JSON: retry with the original input and a concise correction message.
    • Schema failure: include the validation errors, but avoid exposing unnecessary sensitive data.
    • Business-rule failure: ask for clarification or send the item to review; repeated retries may only repeat the error.
    • Unsupported or unsafe request: refuse or use a deterministic fallback.

    Set a strict retry budget, such as one or two attempts, and record the final failure. A fallback can be a smaller model, deterministic parser, queue for later processing, or human operator. Do not let an automatic retry create duplicate side effects: separate extraction from action, and require an idempotency key before executing payments, messages, or infrastructure changes.

    Tune generation, but do not rely on temperature

    Lower temperature can reduce variation, but it cannot enforce a schema or correct factual extraction. Use deterministic settings where appropriate, while recognising that provider changes, model updates, tool routing, and sampling seeds can still affect results. Pin model versions when possible and record the provider, model, prompt version, schema version, temperature, token limits, and validation outcome.

    For local deployments, test quantised and full-precision variants separately. Latency and memory constraints matter when JSON generation runs on mobile or edge devices; deploying open-source LLMs for mobile apps offers relevant deployment considerations.

    Test consistency like an API

    Build a fixture set containing normal, ambiguous, empty, multilingual, adversarial, and very long inputs. For every model or prompt change, measure:

    • Parse success rate
    • Schema-validation success rate
    • Field-level accuracy and null correctness
    • Business-rule failure rate
    • Retry and human-review rate
    • Latency, token usage, and cost
    • Unsafe or unexpected-content rate

    Run repeated evaluations on the same inputs, but do not confuse identical output with correctness. A stable wrong answer is still a production defect. Use golden cases, property-based tests for schema invariants, and sampled human review. Keep personally identifiable information out of logs; redact documents and store only the fields needed for debugging.

    A production pattern

    A robust request path looks like this:

    input → prompt assembly → constrained generation → parse
          → schema validation → business validation → approved action
                                  ↘ retry / review / fallback

    Keep raw responses in a restricted store only when retention is justified. Expose metrics and alerts for sudden drops in validation success, increases in retries, and changes in field-level accuracy. For regulated or sensitive Indian use cases, document data residency, vendor access, retention, consent, and audit requirements before sending source data to an external model.

    FAQ

    Does JSON mode guarantee correct JSON? No. It usually addresses syntax, not required fields, types, factual accuracy, or business rules. Validate the result.

    Should every field be required? No. Make a field required only when the downstream process cannot operate without it. Define a clear missing-value policy for everything else.

    Is post-processing enough? No. Cleanup can repair formatting, but it cannot reliably recover omitted information or detect invented values. Combine constrained generation, validation, and review.

    What is the safest default for high-impact workflows? Separate model-generated data from irreversible actions, validate at every boundary, enforce least privilege, and require human approval for exceptions.

    Consistent JSON is achieved through engineering discipline, not a single prompt. Define a narrow contract, use native structured-output controls, validate in layers, test against realistic Indian data, and monitor the complete pipeline after deployment.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.