Custom LLM applications are no longer defined by the model API alone. The strongest products combine a capable model with reliable retrieval, controlled tool use, evaluation, observability, security, and a deployment plan that matches Indian users and budgets.
The best tools for building custom LLM apps depend on the job: a support assistant needs grounded answers and escalation, a finance workflow needs auditability, and a voice agent needs low latency and dependable speech infrastructure. Start with the product workflow, then choose the smallest stack that can support it in production.
Start with the application architecture
Most custom LLM products use some combination of five layers:
- Model layer: hosted APIs or self-hosted open-weight models for generation, reasoning, embeddings, speech, or vision.
- Application layer: prompts, structured outputs, business rules, conversation state, and tool calls.
- Knowledge layer: document ingestion, chunking, metadata, retrieval, reranking, and citations.
- Reliability layer: evaluations, tracing, guardrails, red-team tests, and human review.
- Product layer: authentication, billing, analytics, queues, APIs, and a user interface.
Do not add an agent framework, vector database, or GPU cluster before the workflow requires it. A narrow, deterministic pipeline is usually easier to test and cheaper to operate than an autonomous agent.
Model providers: test quality, latency, and unit economics
Hosted providers are the fastest route to a working prototype. OpenAI, Anthropic, Google, and other API vendors offer strong general-purpose models, structured generation, tool calling, and multimodal capabilities. Compare them on your own task set rather than relying on public benchmarks.
Measure:
- Answer accuracy and refusal behaviour on representative Indian-language and domain-specific queries.
- Time to first token and complete response latency.
- Input and output token prices, including retrieval context and retries.
- Context-window limits and support for JSON schemas, tool calls, vision, and batch processing.
- Data retention, regional processing options, enterprise controls, and contractual terms.
Open-weight models are useful when data residency, customisation, predictable cost, or offline operation matters. Hugging Face provides model and dataset discovery, while Ollama is convenient for local development. For production inference, evaluate managed GPU platforms and serving systems such as vLLM against your concurrency and latency requirements. A smaller model with good retrieval and constrained outputs can outperform a larger model used without controls.
For Indian products, test English alongside Hindi and other Indic languages, code-mixed speech, transliteration, names, addresses, and local formats such as GSTINs, IFSC codes, dates, and rupee amounts. Do not assume that English evaluation results transfer to these cases.
Orchestration: keep workflows explicit
LangChain is useful for composing model calls, tools, retrievers, and integrations. Its ecosystem is broad, but teams should keep business logic in ordinary application code rather than hiding critical decisions inside opaque chains.
LlamaIndex is particularly useful when the product is centred on private knowledge: PDFs, websites, databases, and enterprise repositories. It offers ingestion and retrieval abstractions that can accelerate a RAG prototype. Haystack is another modular option for teams that want explicit pipelines and open-source components.
For many production applications, you do not need a large framework. A typed Python or TypeScript service with a model SDK, a queue, and a small set of tested functions may be more maintainable. Use durable workflow systems when tasks run for minutes or hours, require retries, or involve human approval.
Agent frameworks are appropriate when the model must choose among tools or plan across steps. They are not a substitute for permissions. Every tool should have a narrow schema, explicit access control, timeouts, rate limits, and logs. For a deeper architecture treatment, see this guide to building distributed systems with AI agents.
Retrieval and knowledge: choose the simplest store that works
RAG is often the right starting point for private or frequently changing information. The pipeline matters more than the brand of vector database:
1. Collect and clean source documents.
2. Preserve titles, sections, dates, permissions, and source URLs as metadata.
3. Split content according to its structure, not an arbitrary character count.
4. Generate embeddings and retrieve candidates with metadata filters.
5. Rerank when precision matters.
6. Instruct the model to answer from retrieved evidence and cite the source.
Pinecone is a managed option for teams that want minimal infrastructure. Weaviate supports vector and hybrid search, which helps when exact terms, product codes, or legal references matter. Chroma is convenient for local prototypes, while pgvector lets teams keep embeddings inside PostgreSQL and simplify operations. Milvus is suited to larger-scale workloads where a dedicated vector system is justified.
A vector database is not mandatory for every app. A small document set may fit into a model context window, and a relational database or full-text search engine may be better for structured records. Test retrieval recall and answer faithfulness before selecting infrastructure.
If model behaviour must reflect proprietary examples, policies, or terminology, compare RAG with fine-tuning. This guide to fine-tuning LLMs on custom data explains when tuning is useful and when better retrieval or prompting is the safer fix.
Evaluation and observability are production requirements
A demo can look convincing while failing on the cases that matter. Create an evaluation set before launch, including normal requests, ambiguous questions, adversarial prompts, multilingual inputs, stale documents, missing permissions, and sensitive data.
Track:
- Retrieval recall, ranking quality, citation correctness, and groundedness.
- Task completion, structured-output validity, refusal quality, and escalation rate.
- Latency, token consumption, error rate, cache hits, and cost per completed task.
- Safety failures, prompt injection attempts, data leakage, and unauthorised tool calls.
LangSmith provides tracing and evaluation workflows for LangChain-based systems. Arize Phoenix offers open-source tracing and LLM evaluation capabilities, while Weights & Biases can help teams manage experiments and prompt versions. OpenTelemetry-compatible traces and structured application logs are valuable regardless of vendor.
Keep production traces privacy-aware. Redact phone numbers, financial information, health data, and identity documents; define retention periods; and restrict access by role. In India, align data handling with the Digital Personal Data Protection Act and your sector-specific obligations. For high-impact decisions, retain a human review path and an auditable reason for each automated action.
Deployment, latency, and cost control
A practical stack may use Next.js or another web framework for the interface, a Python or TypeScript API, PostgreSQL for application data, object storage for documents, a queue for long-running jobs, and a managed model endpoint. Vercel's AI tooling can simplify streaming interfaces, while Replicate and managed inference providers can speed up open-model experiments.
Control cost with:
- Smaller models for classification, routing, extraction, and summarisation.
- Prompt and embedding caches for repeated work.
- Retrieval limits and context compression.
- Batch processing for offline jobs.
- Timeouts, retries with backoff, and fallback models.
- Per-tenant budgets and usage dashboards.
For voice products, latency and turn-taking matter as much as language quality. Speech-to-text, interruption handling, text-to-speech, telephony integration, and consent recording need separate tests. Compare a traditional IVR with an agent before adding complexity; this voice agent versus IVR comparison covers the main trade-offs.
A sensible stack for Indian builders
For a first production release, consider:
- A hosted frontier model or reliable open model selected through task-based evaluations.
- LlamaIndex, LangChain, or plain SDK calls only where they reduce development time.
- PostgreSQL plus pgvector, or a managed vector database if scale demands it.
- Object storage with document versioning and access metadata.
- Phoenix, LangSmith, or an equivalent tracing and evaluation setup.
- A queue, rate limiter, secrets manager, and structured audit logs.
Use Bhashini and other Indic-language resources where they improve accessibility, but validate quality on your target dialects and domain. For founders working on multilingual or open-source systems, the guide for Indian student developers building open-source AI offers useful direction on datasets, community, and deployment constraints.
FAQ
Should I start with an agent? Usually not. Begin with a tested workflow and add tool selection only when fixed routing cannot meet the requirement.
Which vector database is best? Use pgvector when PostgreSQL is sufficient, Chroma for local experiments, and a managed or dedicated database when scale, hybrid search, or operational needs justify it.
Do I need fine-tuning? Not for most early RAG applications. Improve data quality, retrieval, prompts, and evaluations first.
How do I choose between hosted and open models? Compare total cost, quality, latency, privacy, customisation, and operational capability on your own evaluation set.
What should I build first? Pick one measurable workflow, such as document-grounded support or structured extraction, and define success, failure, escalation, and cost limits before writing the full platform.
Apply for AI Grants India
If you are building a serious LLM product from India, funding can support dataset creation, evaluation, cloud credits, security reviews, and early customer pilots. Apply to AI Grants India to explore support for your next AI application.