0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · building custom mcp servers for llm orchestration

Building Custom MCP Servers for LLM Orchestration

  1. aigi

    What custom MCP servers solve

    LLM applications are moving from isolated chat interfaces to systems that retrieve records, call APIs, update software, and coordinate multi-step work. The difficult part is rarely the model call itself. It is giving an agent dependable access to business systems without creating a separate integration for every client, model, and workflow.

    The Model Context Protocol (MCP) provides a common interface between an AI host and external capabilities. A custom MCP server packages your organisation’s data and actions behind typed, discoverable interfaces. The same integration can then serve an approved desktop client, an internal agent platform, or a production orchestration service—subject to the security and transport controls you implement.

    This is especially useful for Indian startups and enterprises working with proprietary fintech, healthtech, logistics, government, or multilingual datasets. MCP does not replace application architecture or access control. It gives those systems a consistent boundary that an LLM can use.

    Understand the MCP building blocks

    An MCP server typically exposes three kinds of capability:

    • Tools: Callable operations such as searching a knowledge base, checking an order, creating a support ticket, or initiating a payment review.
    • Resources: Addressable data that a client can read, such as a policy document, a database schema, or a generated incident report.
    • Prompts: Reusable instructions that help a client start a well-defined task with the right context.

    The server is responsible for validating requests, applying business rules, calling downstream systems, and returning useful results. The host application coordinates the conversation and decides when the model should use a capability. This separation is valuable when building distributed systems with AI agents, where reliability, retries, observability, and ownership must be explicit rather than hidden inside prompts.

    MCP messages use structured protocol exchanges, commonly over stdio for local processes and network transports for remotely hosted services. Transport choice should follow deployment needs, not convenience: stdio is simple and well suited to a developer workstation, while a remote server needs authentication, encryption, tenancy controls, and operational monitoring.

    Decide what belongs in a custom server

    Build a custom MCP server when an integration needs business-specific behaviour that a generic connector cannot safely provide. Strong candidates include:

    • A legacy database with organisation-specific schemas and permissions.
    • Internal search across policies, tickets, contracts, or technical documentation.
    • Narrow wrappers around banking, logistics, CRM, ERP, or government APIs.
    • Local engineering tools, simulators, containers, or hardware interfaces.
    • Composite operations that combine several APIs and return a business-level result.

    Avoid exposing an entire database or a general-purpose shell when a narrow operation will do. A tool such as find_customer_transactions is easier to secure and evaluate than execute_sql. Similarly, create_refund_request gives the server an opportunity to check limits and approval status before any side effect occurs.

    If your application depends on model adaptation rather than external actions, compare this integration approach with best practices for fine-tuning LLMs on custom data. Fine-tuning can change model behaviour; MCP changes what the model can access and do. They solve different problems and are often used together.

    Plan the server contract before writing code

    Start with a capability inventory. For each proposed tool, document:

    • Its purpose and expected user outcome.
    • Required and optional inputs, with types, limits, and examples.
    • Authentication and authorisation requirements.
    • Whether it reads data or causes a side effect.
    • Expected latency, retry behaviour, and failure modes.
    • The minimum information returned to the model.

    Design tool names and descriptions for both developers and models. A description should state what the tool does, when to use it, important constraints, and what it does not do. Keep schemas strict: constrain enum values, dates, identifiers, page sizes, and free-text lengths. Reject malformed or ambiguous input on the server; never assume that a model-generated argument is safe.

    For Indian deployments, explicitly map data classes and residency requirements. Personal data, financial information, health records, Aadhaar-related identifiers, and internal credentials should not be returned merely because a model requested them. Apply masking, field-level filtering, and purpose-based access where required by your organisation’s policies and applicable regulation.

    A minimal Python implementation

    The Python SDK and TypeScript ecosystem are practical starting points. The following FastMCP-style example illustrates the shape of a read-only tool; production code should replace the placeholder search with a controlled repository or retrieval service.

    from mcp.server.fastmcp import FastMCP
    
    mcp = FastMCP("internal-knowledge")
    
    @mcp.tool()
    def search_docs(query: str, limit: int = 5) -> str:
        """Find approved internal documents relevant to a user question."""
        if not query.strip():
            raise ValueError("query must not be empty")
        if limit < 1 or limit > 20:
            raise ValueError("limit must be between 1 and 20")
    
        # Call a permission-aware search service here.
        results = search_approved_documents(query, limit)
        return format_results(results)
    
    if __name__ == "__main__":
        mcp.run(transport="stdio")

    The important design decision is not the decorator. It is the boundary around search_approved_documents: it should receive the authenticated user or tenant context, enforce document permissions, apply filtering, and log the request without leaking sensitive content into logs.

    A local client configuration generally points to the executable and its arguments. Keep secrets out of configuration files; use the operating system’s secret store or an injected environment managed by your deployment platform.

    Build safe orchestration patterns

    A useful MCP server does more than expose isolated API wrappers. It can provide workflow-level tools while keeping each step observable and reversible.

    For high-impact actions, separate preview from commit. A preview_refund tool can calculate eligibility and show the proposed change; a separate submit_refund tool can require an approval token or human confirmation. Add idempotency keys to operations that create records, send messages, or move money. Use short timeouts, bounded retries, and circuit breakers for unreliable dependencies.

    Tool chaining should pass stable identifiers rather than large unverified text blobs. For example, a support workflow might search a CRM, retrieve a permitted customer record, draft a response, and then require approval before sending it. The server should verify that the same user is authorised at every stage, not just when the first tool is called.

    Resources are useful for large or changing context, such as a current policy, service status, or incident log. Return concise, well-labelled content and include timestamps or version identifiers. Do not treat resource access as a substitute for access control.

    Security and operations checklist

    Treat an MCP server like a production API, not a prompt extension.

    • Authentication: Use strong identity for remote clients, TLS, short-lived credentials, and tenant binding.
    • Authorisation: Enforce permissions server-side for every tool and resource. Do not rely on model instructions.
    • Input safety: Validate schemas, parameterise database queries, restrict filesystem paths, and prohibit arbitrary command execution.
    • Output control: Minimise returned fields, redact secrets, and prevent cross-tenant data exposure.
    • Abuse limits: Set request, token, concurrency, and downstream API quotas to contain loops and unexpected spend.
    • Auditability: Record caller identity, tool name, decision outcome, latency, correlation ID, and safe input summaries.
    • Reliability: Define timeouts, idempotency, retries, fallbacks, and clear error messages.
    • Human oversight: Require approval for irreversible actions, regulated decisions, payments, account changes, and external communications.

    Use an inspector or protocol-aware test client during development to inspect capability discovery, schemas, requests, responses, and errors. Automated tests should cover invalid arguments, permission boundaries, duplicate calls, dependency failures, and prompt-injection content returned by connected systems.

    Deploy and evaluate in stages

    Begin with a read-only server and a small set of high-value tools. Run it against synthetic or redacted data, then test with realistic workflows and adversarial cases. Measure task completion, tool-selection accuracy, latency, failure recovery, unauthorised-access attempts, and cost per successful task—not just model quality.

    For production, place remote servers behind an API gateway or service mesh where appropriate. Keep development, staging, and production credentials separate. Version tool schemas deliberately: renaming a field or changing semantics can break clients and agent plans. Publish deprecation windows and maintain compatibility when multiple clients are active.

    MCP is particularly promising for India’s vertical AI builders because it can connect specialised agents to regulated, fragmented systems without forcing every product team to rebuild integrations. Teams developing voice workflows can apply the same separation of model, tools, and permissions seen in voice agents for customer service; the channel changes, but the orchestration discipline remains.

    Frequently asked questions

    Do MCP servers need GPUs? No. Most servers are integration services. Inference may run elsewhere, although a server that performs local embedding or reranking could need CPU or GPU capacity.

    Can MCP servers be written in Go, Java, or Rust? Yes. Use a supported SDK where available or implement the protocol carefully with a compatible JSON-RPC and transport stack.

    Is MCP limited to one model provider? No. The protocol is designed for interoperability, but actual support, tool-selection behaviour, and security features vary by host.

    Should every API become an MCP tool? No. Expose narrow, stable, permission-aware capabilities that correspond to real user tasks. Fewer well-designed tools usually outperform a sprawling catalogue.

    What is the best first project? Choose a read-only workflow with measurable value, such as internal document search or ticket lookup. Add side effects only after identity, audit, approval, and failure handling are proven.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.