A production LLM application needs more than a model API and a prompt. It needs a control plane that decides which model to call, what context to provide, what data may leave your environment, how failures are handled, and whether the answer is good enough to show a user. That control plane is the AI harness.
For Indian startups, enterprises, and public-sector builders, a custom harness is especially useful when applications handle regulated information, support multiple Indian languages, operate over unreliable networks, or need predictable costs. The goal is not to build every component yourself. The goal is to own the decisions that affect reliability, security, and product quality.
What a custom AI harness should do
A harness sits between your product and one or more models. It should provide a consistent application interface while managing the complexity behind it:
- Request handling: authentication, rate limits, tenant isolation, input validation, and streaming.
- Model routing: selecting a provider or local model based on task, latency, quality, geography, and cost.
- Context assembly: conversation history, retrieved documents, user permissions, tools, and structured business data.
- Policy enforcement: PII detection, prompt-injection controls, content policies, and output validation.
- Reliability: retries, timeouts, fallbacks, circuit breakers, caching, and idempotency.
- Measurement: traces, token usage, latency, user feedback, and automated evaluations.
Do not confuse the harness with an orchestration library. LangChain, LlamaIndex, or a home-grown workflow engine may be components inside it. The harness also includes your API layer, data controls, model gateway, evaluation pipeline, and operational policies.
Start with a narrow architecture
A good first version should solve one high-value workflow rather than become a general-purpose agent platform. Define the request contract before choosing frameworks. For example:
{
"tenant_id": "acme",
"user_id": "u_123",
"task": "answer_policy_question",
"input": "What is the reimbursement limit?",
"locale": "en-IN",
"conversation_id": "c_456"
}The response contract should include the answer, citations where relevant, a request ID, model metadata, and a structured error state. Keep internal reasoning, raw prompts, and provider credentials out of the client response.
A practical request path is:
1. Authenticate the caller and identify the tenant.
2. Validate and classify the request.
3. Apply privacy transformations and policy checks.
4. Retrieve permitted context or select tools.
5. Assemble a versioned prompt or structured model request.
6. Route to a model with explicit timeout and token limits.
7. Validate the output, attach citations, and stream or return it.
8. Record telemetry and user feedback asynchronously.
Use a simple state machine for this flow. It is easier to debug than an unconstrained agent loop and gives you clear points for retries and human approval.
Model routing and cost controls
A multi-model harness should treat providers as interchangeable adapters. Define an internal interface for chat, embeddings, tool calling, structured output, and token accounting. Then map that interface to hosted APIs or local inference servers such as vLLM.
Route by task instead of sending every request to the most capable model. A small model can classify intent, rewrite a search query, extract fields, or produce a draft. A stronger model can handle ambiguous policy questions or high-risk decisions. For each route, set:
- maximum input and output tokens;
- timeout and retry policy;
- acceptable regions and data-processing terms;
- target quality and latency;
- fallback provider or a safe refusal path.
Track cost per successful task, not only cost per token. Cache deterministic operations such as embeddings, document classification, and repeated policy lookups. Never retry blindly: a timed-out request may still have been billed, and non-idempotent tool calls can create duplicate actions.
Build RAG as a permissioned data system
Retrieval-Augmented Generation is not simply “put documents in a vector database.” Treat it as a searchable data product with ownership, freshness, and access controls.
Your ingestion pipeline should extract text, preserve document structure, remove duplicates, attach metadata, and record source versions. Chunk by meaning—headings, clauses, tables, or FAQ entries—rather than using one universal character limit. Store document ID, page or section, language, effective date, department, and access scope with every chunk.
At query time:
- detect language and normalize the query without destroying names or domain terms;
- apply tenant and permission filters before or during retrieval;
- combine keyword and vector search for exact identifiers and semantic questions;
- rerank the shortlist when answer quality justifies the added latency;
- reject or clarify questions when evidence is weak;
- require citations for factual answers.
For Indic applications, test tokenization, transliteration, code-mixed queries, and spelling variation separately. The low-resource Indic NLP guide is useful when your corpus includes Hindi, Tamil, Bengali, Marathi, or other regional languages. A retrieval benchmark should contain real user phrasing, not only translated English questions.
Prompts, tools, and state
Store prompts as versioned configuration with an owner, change note, model compatibility, and evaluation results. Separate stable system policy from task instructions and user content. Treat retrieved text and tool results as untrusted data; delimit them clearly and instruct the model not to follow instructions found inside documents.
Conversation memory should be selective. Keep a short recent window, maintain a structured summary for older turns, and retrieve only facts needed for the current task. Do not place sensitive history into every prompt by default.
Tools need stricter controls than text generation. Define JSON schemas, validate arguments server-side, apply per-tool authorization, and require confirmation for irreversible actions such as payments, account changes, or outbound messages. Agentic workflows should have step limits, budgets, and a stop condition. For more complex multi-agent designs, study the trade-offs in building distributed systems with AI agents before introducing additional workers.
Guardrails for Indian production environments
Security belongs in the request path, not in a post-launch checklist. Detect and redact or tokenize sensitive fields such as Aadhaar numbers, PAN details, bank information, phone numbers, and health records according to the application’s legal and operational requirements. Maintain a reversible mapping only where the workflow genuinely needs it, and protect that mapping separately.
Add controls for prompt injection, data exfiltration, unsafe tool use, and unsupported claims. Output validation should check schema, citations, prohibited disclosures, and business rules. A second model can help classify risk, but deterministic checks are preferable for identifiers, permissions, limits, and transaction states.
Keep provider contracts, retention settings, encryption, access logs, and incident procedures documented. For legal or sensitive workflows, a private deployment may be justified; the private AI chatbot architecture for lawyers illustrates the additional isolation and audit requirements such systems need.
Evaluation and observability
Build an evaluation set before optimising prompts. Include normal questions, ambiguous requests, adversarial inputs, multilingual and code-mixed examples, long-context cases, and known failure modes. Score separate dimensions:
- retrieval relevance and citation coverage;
- factual faithfulness to approved sources;
- task completion and structured-output validity;
- refusal quality and policy compliance;
- latency, cost, and failure rate.
Use automated checks for every prompt, model, chunking, or retriever change. LLM-as-a-judge is useful for scale, but calibrate it against human ratings and retain representative review samples.
Instrument each request with a correlation ID and spans for classification, retrieval, prompt assembly, provider calls, tools, and validation. Capture token counts, cache hits, model version, status, latency, and error category. Redact secrets and sensitive content from logs by default. Dashboards should show p50/p95 latency, time to first token, cost per task, fallback rate, grounded-answer rate, and user correction rate.
Streaming and deployment checklist
Use Server-Sent Events or WebSockets for long responses, but do not stream unvalidated high-risk actions. Send a visible “checking sources” or status event only when it reflects a real pipeline stage. For mobile users, support reconnection, cancellation, and bounded output sizes.
Before production, verify that you have:
- provider adapters and a tested fallback path;
- tenant-aware retrieval and authorization tests;
- prompt and model versions recorded in every trace;
- timeouts, circuit breakers, quotas, and budget alerts;
- offline evaluation gates in CI;
- encrypted secrets and documented retention policies;
- human escalation for uncertain or high-impact cases;
- a rollback plan for prompts, models, indexes, and policies.
Start with a modular monolith and a small number of dependable components. Split services only when traffic, team boundaries, or isolation requirements justify the operational cost. A harness earns its complexity by making an AI product safer, measurable, and easier to change—not by containing the largest possible number of frameworks.