Custom AI agents are software systems that happen to use language models. They interpret requests, retrieve context, maintain state, call APIs, make decisions, and sometimes change records or move money. Custom AI agent testing must therefore evaluate more than whether a response sounds plausible: it must verify the agent’s decisions, permissions, system state, failure handling, and business outcomes.
For Indian teams, production realism matters. Users may switch between English, Hindi, Hinglish, and regional languages; send fragmented WhatsApp-style messages; provide Indian dates, addresses, and rupee amounts; or speak through noisy phone connections. Agents also interact with CRMs, payment gateways, ticketing systems, commerce platforms, and internal databases. A useful test programme defines what the agent may do, what evidence it needs, when it must ask a question, and when it must stop or escalate.
Define the agent’s contract before testing
Start with a written contract for every workflow. Document:
- supported intents and channels;
- required inputs and accepted formats;
- approved knowledge sources;
- tools the agent may call and the permissions attached to each;
- actions requiring confirmation or human approval;
- expected escalation conditions;
- latency, cost, and availability limits;
- data retention, redaction, and audit requirements.
Specify expected outcomes rather than ideal wording. A test case should say whether the agent must answer, ask for clarification, refuse, escalate, or complete an action. For example, a refund request may require identity verification, order lookup, policy retrieval, confirmation, and an idempotent refund call. A fluent answer without the correct state change is a failed test.
This contract also helps decide whether a voice interface is appropriate. Teams evaluating phone workflows should account for interruptions, silence, call transfers, pronunciation, background noise, and language switching; what a voice agent is and how voice AI works in 2026 provides useful context for testing the full speech-to-action pipeline.
Build a production-shaped evaluation set
A few hand-written prompts cannot represent real usage. Maintain a versioned dataset containing:
- De-identified production examples: support tickets, call transcripts, failed searches, corrections, and escalations.
- Business scenarios: routine, ambiguous, exceptional, prohibited, and high-value requests.
- Controlled variations: misspellings, slang, transliteration, code-switching, short replies, long messages, and different tones.
- Adversarial cases: prompt injection, malicious retrieved documents, requests for secrets, privilege escalation, and attempts to bypass approval.
- Boundary conditions: missing fields, duplicate requests, expired links, conflicting records, unavailable inventory, timeouts, and maximum input lengths.
Label each case with the expected action, required evidence, permitted tools, risk level, and acceptable alternatives. Preserve the supporting documents and expected system state so failures can be reproduced. Synthetic examples are valuable for coverage, but review them before using them as ground truth; generated cases can contain unrealistic assumptions.
For India-facing products, include DD/MM/YYYY and ambiguous date phrases, Indian numbering such as lakh and crore, GST terminology where relevant, PIN codes, state and city names, rupee formatting, and common Hindi-English transliterations. If the workflow handles restaurant reservations, test changes, cancellations, late arrivals, duplicate bookings, and unavailable slots—not only the initial booking. The restaurant table booking voice agent guide for India shows why operational edge cases deserve as much attention as conversation quality.
Test in layers
Deterministic unit and contract tests
Use conventional software tests for components that should behave predictably. Cover prompt construction, routing, schema validation, authentication, redaction, policy checks, retry logic, and idempotency. Mock model and external-service responses to test malformed JSON, rate limits, timeouts, partial failures, and provider outages.
Every tool should have a strict schema. Validate arguments outside the model, enforce permissions in application code, and reject unexpected fields. Test that a repeated model response cannot create duplicate tickets, bookings, payments, or messages.
Retrieval and grounding tests
Evaluate retrieval independently from generation. Measure whether the correct chunks are returned, whether irrelevant or outdated documents are excluded, and whether access controls are respected. Include no-answer and conflicting-document cases. The agent should identify insufficient evidence instead of filling gaps with a confident guess.
For each grounded response, inspect the source passages and verify that material claims are supported. Test index refreshes, document deletion, metadata filters, regional content, and tenant isolation. A retrieval system that exposes one customer’s documents to another is a critical security failure even if its answer reads well.
Tool and integration tests
Run tools in sandbox environments with realistic records. Check selection, parameter accuracy, authentication, permission boundaries, confirmation flows, retries, timeout behaviour, and rollback or compensation logic. Test contradictory information across systems—for example, a CRM lead marked active while the booking system shows no availability.
For irreversible operations, require explicit confirmation and record who or what authorised the action. For sensitive workflows, model the full path from identity verification through audit logging. Voice deployments serving hospitals or other regulated environments should use stronger controls; HIPAA-compliant voice agents for hospitals is a useful reference for access, escalation, and traceability principles.
End-to-end conversation tests
Run complete multi-turn scenarios, including corrections, interruptions, abandonment, resume journeys, handoffs, and repeated requests. Evaluate both the transcript and the resulting state. Test session isolation, memory expiry, consent, and account switching. A customer correcting an address should update the intended order only—not a previous order or another user’s record.
For voice agents, include noisy audio, silence, barge-in, accents, number and name pronunciation, DTMF fallback, call transfer, dropped calls, and language changes. If a team lacks internal speech-testing expertise, it can compare top-rated voice agent services for Indian businesses against the requirements before choosing an implementation partner.
Use a rubric, not one quality score
Score separate dimensions and define severity levels:
- Task success: Did the user’s intended outcome occur?
- Factuality and grounding: Are claims accurate and supported?
- Instruction following: Did the agent follow workflow, format, and policy rules?
- Tool correctness: Were the right tools called with valid arguments?
- Safety and privacy: Did it avoid disclosure, unsafe advice, and unauthorised actions?
- Conversation quality: Was it clear, concise, respectful, and appropriately localised?
- Recovery: Did it handle uncertainty, failures, and escalation correctly?
LLM-as-judge evaluations can provide scale, but calibrate them against human-labelled examples and periodically measure judge drift. Use human review for high-risk cases, sampled production traces, disagreements, and all critical failures. Keep a failure taxonomy—such as hallucination, wrong intent, retrieval miss, tool error, policy breach, language failure, or state corruption—so engineering work targets causes rather than symptoms.
Metrics and release gates
Report results by intent, language, channel, model, prompt version, retrieval index, and customer segment. Track:
- task completion, escalation, abandonment, and repeat-contact rates;
- factual error, unsupported-answer, refusal, and clarification rates;
- tool-selection accuracy, invalid-argument rate, duplicate-action rate, and rollback success;
- retrieval precision, recall, citation validity, and no-answer accuracy;
- p50, p95, and p99 latency, including speech and integration time;
- cost per conversation and cost per successful task;
- privacy, security, policy, and data-isolation violations.
Set thresholds before comparing releases. Use blocking gates for data leakage, unauthorised actions, broken tenant isolation, and unsafe regulated advice. Treat cost, latency, and quality as a bundle: a model upgrade that improves wording but doubles the cost or increases failed tool calls is not automatically better. Compare new versions with the current production version on identical cases and review statistically meaningful regressions, not only average scores.
Put testing into CI and production operations
Version-control evaluation cases, expected outcomes, prompts, model settings, retrieval indexes, tool schemas, and policy rules. A practical pipeline runs:
1. fast unit, schema, and contract tests;
2. deterministic safety and permission checks;
3. representative regression and multilingual suites;
4. retrieval, tool-use, and prompt-injection evaluations;
5. latency and cost checks;
6. sandbox end-to-end tests before deployment.
Use shadow traffic or a canary release for model and prompt changes. Monitor traces, tool events, latency, cost, user corrections, escalation outcomes, and final state while minimising personal data in logs. Define alerts for critical failures and maintain a rollback path that does not depend on the new agent. Feed confirmed production failures back into the regression set, with a clear owner and target fix.
Launch checklist for Indian teams
Before release, confirm that the agent:
- meets success thresholds for priority workflows and languages;
- handles Hindi-English and relevant regional-language variations promised to users;
- asks for clarification when critical information is missing;
- refuses or escalates unsupported and high-risk requests;
- requires confirmation for irreversible actions;
- preserves tenant, session, and role boundaries;
- meets p95 latency and cost targets;
- recovers safely from tool, network, and provider failures;
- produces auditable traces without unnecessary personal data;
- supports rollback, incident response, and human takeover.
Reliable custom AI agent testing is a continuous engineering discipline, not a pre-launch demo review. Combine software tests, adversarial evaluation, multilingual scenarios, human oversight, and production monitoring. The result is an agent whose limits are measurable and whose failures are contained—stronger foundations for Indian businesses moving from prototypes to dependable automation.