LLM agents can plan, call tools, retrieve information, and act on a user’s behalf. That autonomy makes a single pass through a language model risky: an agent may misunderstand a request, use the wrong source, invent a result, or take an irreversible action. A self-critique loop adds a structured review step before the agent answers or acts.
The goal is not to make a model “think harder” indefinitely. A useful loop checks specific failure modes, applies evidence and policy constraints, and escalates uncertain cases. For Indian teams, this must also account for multilingual input, code-switching, uneven data quality, privacy obligations, and high-volume, cost-sensitive production environments.
What a self-critique loop should do
A self-critique loop separates generation from verification. The agent first creates a draft plan, answer, or tool action. A critic then evaluates it against explicit criteria and returns structured findings. The system either revises the draft, asks the user for clarification, routes the case to a human, or proceeds.
A practical loop contains five stages:
- Specify: Convert the user request into an objective, constraints, and a defined completion condition.
- Generate: Produce a proposed answer, plan, or tool call.
- Check: Test factual support, policy compliance, tool arguments, uncertainty, and task completion.
- Revise or escalate: Correct the output, request missing information, or require approval.
- Record: Store inputs, evidence, critiques, revisions, and final outcomes for evaluation.
This pattern is useful when building generative AI agents, but it is not a substitute for deterministic validation. A model should not be trusted to “self-approve” a payment, medical instruction, identity decision, or destructive database operation.
Design the critic as a contract
Avoid a vague prompt such as “review your answer carefully.” Define what the critic must inspect and how it must respond. A JSON schema makes the result easier to test and route:
{
"pass": false,
"issues": [
{
"type": "unsupported_claim",
"severity": "high",
"evidence": "No retrieved source supports the eligibility rule",
"fix": "Retrieve the current policy or ask for the relevant scheme"
}
],
"needs_human": false,
"confidence": 0.62
}Useful checks include:
- Grounding: Is each material claim supported by retrieved documents, tool output, or user-provided facts?
- Instruction following: Did the agent satisfy the request without adding unauthorised objectives?
- Safety and policy: Does the response expose sensitive data, provide unsafe advice, or bypass access controls?
- Action validity: Are tool names, parameters, permissions, and expected side effects correct?
- Completeness: Are required fields, edge cases, and next steps present?
- Uncertainty: Does the answer distinguish known information from assumptions?
Keep the critic focused. A single general-purpose critic often produces polished but shallow approval. Separate critics for factual grounding, policy, and action safety are easier to evaluate and can be invoked according to risk.
Choose the right loop architecture
The simplest architecture is draft–critique–revise. It works for customer support responses, document extraction, and research summaries. Limit it to one or two revisions unless evaluation shows a clear benefit.
For tool-using agents, use plan–validate–execute. The validator checks the proposed action before execution; after the tool returns, a second check verifies whether the result actually completed the task. This is safer than allowing an agent to execute first and explain later.
For high-risk workflows, use critic plus deterministic gates. Code should enforce spending limits, allow-lists, schema validation, authentication, rate limits, and approval requirements. The model can recommend an action, but policy code decides whether it is permitted.
In distributed or multi-agent systems, avoid letting every agent critique every other agent. Define ownership: one agent proposes, a specialised validator checks, and an orchestrator resolves conflicts. This is especially important when building distributed systems with AI agents, where retries and duplicated messages can otherwise multiply costs and side effects.
Make it work for India’s language and deployment context
Evaluate the loop on English, Hindi, and the languages your product actually serves. Include code-mixed prompts, transliterated Hindi, regional names, numerals in different formats, ambiguous dates, and speech-to-text errors. A critic that performs well on formal English may miss a safety issue in Hinglish or a local-language transcript.
Use local, consented test sets rather than translating an English benchmark and assuming equivalence. For a restaurant workflow, test menu names, addresses, delivery instructions, and noisy phone audio; relevant lessons can be found in guides to multilingual voice agents for restaurants in India. For healthcare, require clinician-reviewed test cases, conservative uncertainty handling, and human escalation. Patient follow-up systems should be assessed alongside practical voice-agent patient follow-up workflows in India.
Data governance matters as much as prompt design. Minimise personal data in critique logs, redact identifiers where possible, restrict access, define retention periods, and prevent sensitive conversations from entering training pipelines without a valid basis. Document where evaluation data came from and whether users consented to its reuse.
Measure whether critique actually helps
Track the system before and after adding the loop. At minimum, measure:
- Critical error rate and false approval rate
- Factuality or citation-support rate
- Successful task completion, not just response quality
- Human-escalation precision and recall
- Revision rate and improvement after revision
- Latency, token use, and cost per completed task
- Failure rates by language, channel, customer segment, and model version
Create a test set with labelled failure categories. Include adversarial prompts, incomplete information, conflicting documents, prompt injection in retrieved content, tool timeouts, duplicate events, and attempts to trigger unauthorised actions. Run it in continuous evaluation whenever prompts, models, retrieval indexes, or tools change.
Do not use the critic’s confidence score as a probability unless calibrated against real outcomes. A model can be highly confident and wrong. Sample passed cases for human review, and monitor production incidents rather than only benchmark scores.
Common failure modes
Infinite self-reflection increases latency without improving accuracy. Set a maximum number of rounds, a token budget, and a stop condition.
Critique by the same model alone can reproduce the original error. Add retrieval, deterministic checks, a different model, or human review for important decisions.
Rewarding verbosity encourages long explanations instead of corrections. Score the final outcome and evidence, not the length of the critique.
Hidden state and untraceable revisions make incidents impossible to investigate. Log versioned prompts, tool calls, retrieved documents, critique findings, and final decisions with appropriate redaction.
Over-blocking frustrates users and pushes them to unsafe workarounds. Distinguish between refusal, clarification, safe partial completion, and human escalation.
A practical implementation sequence
Start with one workflow and define its most expensive or harmful errors. Establish a baseline, add a structured critic, and introduce deterministic gates for irreversible actions. Test across the languages and channels used by real customers. Then run a limited pilot with shadow evaluation before enabling automatic revisions.
For production, add tracing, alert thresholds, replayable test cases, access controls, and a rollback path. Reassess the loop whenever the underlying model, retrieval corpus, business policy, or tool permissions change. Teams deploying open models can also review practices for deploying Llama 3 agents in production, particularly around serving, observability, and failure handling.
A strong self-critique loop is modest, measurable, and constrained. It catches known classes of mistakes, makes uncertainty visible, and routes consequential decisions to the right control—not merely back to another model-generated paragraph.