0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · large language models

Large Language Models: How They Work and How to Build With Them

  1. aigi

    Large language models (LLMs) are general-purpose AI systems that predict and generate language from context. They can summarise documents, answer questions, write code, extract structured data, translate text, and call tools. But an LLM is not a database or a reasoning engine with guaranteed accuracy. It is a probabilistic component that must be designed, evaluated, and supervised inside a product.

    For Indian builders, the central challenge is not simply choosing the model with the largest parameter count. It is delivering useful performance across English, Hindi, and other Indian languages; handling code-mixed queries; protecting sensitive data; and keeping latency and inference costs within budget.

    What large language models do

    Most modern LLMs use transformer architectures. During pretraining, the model sees large collections of text and learns to predict missing or subsequent tokens. That process builds statistical representations of grammar, facts, styles, and relationships. Later stages—such as supervised fine-tuning and preference optimisation—make the model more useful for instructions and conversations.

    At runtime, the model receives a sequence of tokens called the context window and calculates a probability distribution for the next token. Repeating this process produces an answer. The model may appear to understand a question, but its output remains a generated prediction. It can therefore produce fluent errors, unsupported citations, or confident answers to ambiguous prompts.

    Common capabilities include:

    • Generation: drafting text, code, emails, reports, and product copy.
    • Transformation: summarising, translating, rewriting, classifying, and extracting fields.
    • Conversation: answering questions over a controlled knowledge base.
    • Tool use: invoking search, databases, calculators, APIs, and business workflows.
    • Multimodal interaction: processing images, audio, video, and text in models that support those inputs.

    For Indian-language products, general benchmarks can be misleading. A model may perform well in English while struggling with Devanagari, transliteration, regional terminology, or code-mixed speech. Work with relevant low-resource Indic NLP methods and domain-specific datasets rather than relying only on English evaluations.

    The main ways to build with an LLM

    Prompting and structured outputs

    Start with a clear system instruction, representative examples, explicit constraints, and a defined output schema. JSON schemas or function-calling interfaces are safer than asking a model to return free-form text when the output feeds software.

    Prompting works well for summarisation, drafting, classification, and lightweight extraction. It is fast to iterate, but prompts alone do not add dependable new knowledge or guarantee consistent behaviour.

    Retrieval-augmented generation

    Retrieval-augmented generation (RAG) gives the model relevant documents at query time. A typical pipeline ingests documents, splits them into chunks, creates embeddings, retrieves candidates, reranks them, and places the best evidence in the prompt. The application should instruct the model to answer only from supplied sources and expose citations where appropriate.

    RAG is useful for Indian enterprises with changing policies, product catalogues, government schemes, or internal manuals. Measure retrieval quality separately from answer quality: a model cannot produce a grounded answer if the right document never reaches its context.

    Fine-tuning and adapters

    Fine-tuning changes model behaviour using task-specific examples. It can improve formatting, tone, classification, or specialised workflows, but it is not usually the best way to load frequently changing facts. Parameter-efficient methods such as LoRA can reduce training costs and make experimentation practical for smaller teams. See this guide to fine-tuning Llama for Indian regional languages before assembling a multilingual training pipeline.

    Local and open-weight deployment

    Running an open-weight model on your own infrastructure can improve privacy, control, and predictable access. It also transfers responsibility for hardware, quantisation, serving, updates, monitoring, and safety controls to your team. Compare hosted APIs with local LLM deployment options, especially when handling health, financial, identity, or public-sector data.

    A practical selection framework

    Choose a model against the workload, not its marketing label. Test at least three candidates on a private evaluation set containing real user inputs, including spelling variation, code-mixing, and difficult edge cases.

    Assess:

    • Task quality: correctness, instruction following, extraction accuracy, and citation faithfulness.
    • Language coverage: performance in English, Hindi, and the exact regional languages your users speak.
    • Reliability: refusal behaviour, format compliance, sensitivity to prompt changes, and long-context performance.
    • Operations: latency, rate limits, uptime, context length, streaming support, and observability.
    • Economics: input and output token prices, caching, batch support, GPU requirements, and engineering overhead.
    • Governance: data retention, training-on-input policies, regional hosting, auditability, and access controls.

    Small language models can be better for fixed classification, routing, extraction, and on-device use. The growing ecosystem of open-source small language models for Hindi is particularly relevant when latency, cost, or offline access matters more than broad general knowledge.

    Evaluation that reflects production

    A demo is not an evaluation. Create a test set of several hundred representative examples where possible, and label expected answers, acceptable variations, refusal cases, and evidence requirements. Keep a separate holdout set for final comparison.

    Track task-specific measures such as exact match, F1, recall, groundedness, citation accuracy, translation quality, and schema-valid output. Add human review for helpfulness, cultural fit, harmful content, and language naturalness. Test adversarial inputs, prompt injection, sensitive-data requests, and attempts to bypass business rules.

    After launch, log prompts and outputs safely, sample conversations for review, monitor drift, and provide a user feedback path. Redact personal information and define retention limits before collecting production traces.

    Reliability, safety, and privacy

    Treat model output as untrusted input. Validate generated JSON, enforce permissions in application code, and never allow a model to decide access rights by itself. Tool calls should use allowlists, parameter validation, timeouts, rate limits, and human approval for irreversible actions.

    Reduce hallucinations by retrieving authoritative evidence, asking for concise answers when appropriate, requiring citations, and returning an explicit “not enough information” response. Do not hide uncertainty behind polished language.

    For Indian deployments, map data flows carefully. Aadhaar details, health records, financial information, student data, and customer communications may require strict handling under organisational policy and applicable law. Obtain consent where required, minimise collected data, encrypt it in transit and at rest, and ensure vendors provide clear contractual commitments about retention and model training.

    Bias requires ongoing measurement, not a one-time checklist. Evaluate dialects, gendered language, caste-related terms, regional names, disability contexts, and minority communities with local reviewers. Establish escalation routes for harmful or discriminatory outputs.

    Cost and architecture decisions

    Total cost includes more than API tokens. Account for retrieval storage, embedding generation, reranking, observability, moderation, retries, engineering, and human review. Use smaller models for routing and simple transformations, cache repeated requests, stream long responses, limit unnecessary context, and batch offline workloads.

    A robust production architecture often separates the user interface, policy layer, model gateway, retrieval service, tool executor, and evaluation pipeline. This makes it easier to switch providers, compare models, enforce security, and diagnose failures. For image-heavy workflows, combine an LLM with specialist systems rather than expecting one model to do everything; teams working on visual products can also review open-source vision-language models for Indian languages.

    Where LLMs are heading

    As of 2026, progress is shifting from raw scale toward useful systems: smaller capable models, longer and better-managed context, multimodal inputs, tool orchestration, efficient inference, and stronger evaluation. Indian builders should prioritise language coverage, affordable deployment, and measurable outcomes over chasing the largest model.

    The winning product is rarely “a chatbot.” It is a well-scoped workflow in which the model performs a valuable step, reliable data supplies context, software enforces rules, and people can review consequential decisions.

    FAQ

    Are large language models databases?

    No. They store learned statistical patterns rather than providing guaranteed, current records. Use retrieval or a database for authoritative and changing information.

    Should a startup fine-tune an LLM?

    Only after testing prompting and RAG. Fine-tuning is most useful for repeatable behaviour, style, classification, or structured formats—not as a substitute for a live knowledge source.

    Are open-source models always cheaper?

    No. They may reduce API dependence but add GPU, serving, monitoring, upgrade, and engineering costs. Compare the full cost per successful task.

    How can teams reduce hallucinations?

    Use trusted retrieval, narrow the task, require evidence, validate outputs, measure groundedness, and give the model a safe way to decline when information is missing.

    What should an Indian-language AI team measure first?

    Measure task accuracy by language and script, code-mixed performance, latency, cost per interaction, safety failures, and user success—not just an aggregate benchmark score.

    Last updated 24 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.