0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build applications with claude ai models

How to Build Applications with Claude AI Models

  1. aigi

    Claude is most useful when treated as a component in a reliable software system, not as the entire product. A production application needs a clear user problem, controlled model access, structured outputs, observability, security, and a fallback path when the model is uncertain or unavailable.

    This guide explains how to build applications with Claude AI models using Anthropic’s API and current application patterns. The examples apply to support assistants, document workflows, coding tools, internal knowledge systems, and multilingual products for Indian users.

    Start with a narrow, measurable use case

    Begin with one workflow where language understanding creates measurable value. Good first projects include:

    • Summarising long documents into a fixed schema
    • Classifying support tickets and drafting replies
    • Extracting fields from invoices, contracts, or forms
    • Answering questions over a controlled knowledge base
    • Assisting employees with search, analysis, or code generation

    Define success before choosing a model. Useful metrics include answer accuracy, citation coverage, task completion rate, latency, cost per successful task, escalation rate, and user-edit rate. For India-facing products, also test performance across English, Hindi, Hinglish, and the regional languages relevant to your users. If language coverage is central to the product, pair Claude with the methods described in this guide to low-resource Indic NLP.

    Avoid starting with a general-purpose chatbot. A bounded workflow is easier to evaluate, secure, price, and improve.

    Choose the API integration pattern

    Create an Anthropic account, generate an API key, and keep it on the server. Do not expose the key in browser or mobile code. Anthropic’s API documentation should be your source of truth for model names, limits, headers, Messages API behaviour, streaming, and tool use.

    A minimal Python request using the Messages API looks like this:

    import os
    from anthropic import Anthropic
    
    client = Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
    
    message = client.messages.create(
        model="claude-sonnet-4-20250514",
        max_tokens=800,
        system="You are a concise support assistant. If uncertain, say so.",
        messages=[
            {"role": "user", "content": "Summarise this ticket in three bullets."}
        ],
    )
    
    print(message.content[0].text)

    Use environment variables or a secrets manager for credentials. In production, add request IDs, timeouts, retries with exponential backoff, rate-limit handling, and structured logs. Pin a tested model version where reproducibility matters, and keep model selection in configuration so you can evaluate upgrades without rewriting application code.

    For interactive interfaces, use streaming so users see partial output while Claude works. Streaming improves perceived latency, but your backend should still validate the completed response before displaying or storing it.

    Design prompts as application contracts

    A strong prompt defines the task, available context, constraints, and output format. Separate stable instructions from changing user content:

    System: You are an accounts-payable assistant. Extract only facts present in the document.
    
    Rules:
    - Never invent invoice numbers, dates, or tax amounts.
    - Return valid JSON matching the supplied schema.
    - Use INR for monetary values when the document states rupees.
    - Set a field to null when it is absent.
    
    User: <document text>

    Use examples for ambiguous tasks, specify what the model must do when information is missing, and require concise answers where latency and cost matter. Prefer structured outputs or JSON schemas for data that your software will consume. Parse and validate the result with a typed model; never assume that text which looks like JSON is valid JSON.

    Prompt quality is not a substitute for product logic. Enforce permissions, eligibility rules, calculations, and irreversible actions in code.

    Add retrieval for private or changing knowledge

    Claude does not automatically know your company’s latest policies, inventory, or case records. For those tasks, retrieve relevant content from your database or search index and pass only the necessary excerpts in the request. A practical retrieval-augmented generation flow is:

    1. Ingest and clean source documents.
    2. Split them into meaningful sections with metadata.
    3. Create embeddings and store them in a vector or hybrid search system.
    4. Retrieve a small set of relevant passages for each question.
    5. Ask Claude to answer only from those passages.
    6. Return citations or document references with the answer.

    Filter retrieval by tenant, user permission, language, and document freshness before sending context to the model. Keep source text separate from instructions so documents cannot easily override your system rules. For larger systems, scaling backend infrastructure for AI applications covers queues, caching, observability, and service boundaries that become important at volume.

    Use tools for actions, not just conversation

    Tool use lets Claude request a function such as check_order_status, search_policy, or create_ticket. Your application executes the function, validates its arguments, and sends the result back to Claude. The model should propose an action; your code decides whether it is permitted.

    Apply these controls:

    • Define narrow tools with typed arguments and clear descriptions.
    • Check identity, authorisation, and business rules on every call.
    • Require confirmation before payments, deletions, messages, or other irreversible actions.
    • Set timeouts and limit tool-call depth to prevent loops.
    • Record tool inputs, outputs, failures, and approvals for audit.
    • Redact personal, financial, and health information from logs.

    This architecture is useful for agents, but do not add autonomous loops until a single request-response workflow is reliable. For more complex orchestration, compare the design trade-offs in building generative AI agents and distributed systems with AI agents.

    Build evaluation before launch

    Create a test set from real, anonymised tasks and include difficult cases: incomplete inputs, conflicting documents, prompt injection, code-mixed language, ambiguous requests, and attempts to access another user’s data. Score both the model response and the complete workflow.

    Track:

    • Correctness: Is the answer supported by the supplied evidence?
    • Grounding: Are citations accurate and complete?
    • Safety: Does the system refuse or escalate risky requests?
    • Reliability: Does structured output pass validation?
    • Operations: What are latency, token usage, error, and retry rates?

    Use a small human-reviewed benchmark for every prompt, retrieval, or model change. In production, sample conversations for review with appropriate consent and privacy controls. Measure outcomes, not just thumbs-up ratings.

    Plan cost, privacy, and Indian deployment constraints

    Estimate cost from tokens per task, requests per user, retries, tool calls, and retrieval context. Reduce unnecessary context, cache stable instructions where supported, route simple tasks to smaller or cheaper models, and set per-user budgets. Include GST, cloud egress, observability, storage, and human-review costs in your unit economics.

    Send the minimum data required. Mask Aadhaar numbers, PAN details, phone numbers, and other sensitive fields unless the workflow genuinely needs them. Define retention periods, access controls, deletion procedures, and vendor-processing terms. For regulated sectors, review applicable Indian requirements and your organisation’s data-residency and contractual obligations before production use.

    Design for unreliable networks and low-end devices: keep model calls server-side, return compact responses, support retries without duplicate actions, and provide a human or deterministic fallback. Products intended for the next billion users should also account for language, literacy, voice, and bandwidth constraints; this India-focused AI app guide offers relevant product considerations.

    Deployment checklist

    Before launch, verify that you have:

    • API keys stored in a secrets manager
    • Authentication, authorisation, rate limits, and tenant isolation
    • Input and output validation, including schema checks
    • Prompt-injection and sensitive-data tests
    • Timeouts, retries, circuit breakers, and fallback responses
    • Token, latency, error, and cost monitoring
    • A versioned evaluation set and rollback plan
    • Human escalation for high-impact or uncertain decisions
    • Clear user disclosure when they are interacting with AI

    A Claude application is production-ready when it remains useful under incomplete data, adversarial inputs, provider errors, and changing demand—not merely when it produces impressive demos. Start with a narrow workflow, make every important boundary explicit in code, and expand only after measured reliability supports the next use case.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.