0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build custom python wrappers for llms

How to Build Custom Python Wrappers for LLMs

  1. aigi

    A custom Python wrapper is the boundary between your application and an LLM provider. Done well, it keeps provider-specific SDKs, prompts, retries, token settings, privacy controls, and response parsing out of business logic. Done poorly, it becomes another thin abstraction that hides failures and makes model changes harder.

    This guide shows how to design a wrapper that can support hosted APIs, self-hosted models, structured responses, tool calls, and production operations. The approach is useful whether you are building an internal copilot, an Indian-language product, or a high-volume customer workflow.

    What a production wrapper should do

    Start with a small, explicit contract. Application code should ask for an outcome; the wrapper should manage the mechanics of obtaining it.

    A practical wrapper can own:

    • Provider access: authentication, model selection, endpoint configuration, and SDK calls.
    • Request policy: system instructions, temperature, token limits, timeouts, and safety settings.
    • Output handling: text cleanup, JSON or schema validation, refusal handling, and tool-call parsing.
    • Reliability: retries for transient failures, exponential backoff, rate-limit handling, and fallbacks.
    • Observability: latency, token usage, model name, request identifiers, errors, and cost estimates.
    • Governance: redaction, retention rules, tenant isolation, and configurable logging.

    Keep business rules outside the wrapper. For example, an order-service should decide what counts as an eligible refund; the wrapper should only provide a reliable model interaction.

    Choose an interface before choosing an SDK

    Define the calls your product needs rather than copying a provider’s API into your codebase. A minimal interface might include:

    from dataclasses import dataclass
    from typing import Any, Mapping, Sequence
    
    @dataclass
    class LLMResult:
        text: str | None
        raw: Any
        model: str
        usage: Mapping[str, int]
    
    class LLMClient:
        def complete(
            self,
            messages: Sequence[Mapping[str, str]],
            *,
            model: str | None = None,
            temperature: float = 0.2,
            max_tokens: int = 800,
            response_schema: type | None = None,
        ) -> LLMResult:
            raise NotImplementedError

    This contract gives your application a stable seam. A second implementation can target another hosted provider, an on-premise inference server, or a smaller model for low-cost workloads without requiring a rewrite of every feature.

    Avoid excessive abstraction. If your wrapper exposes every provider-specific option, callers will still be coupled to that provider. Offer common controls directly and place uncommon settings in a clearly named provider_options object.

    Set up configuration and secrets safely

    Use a virtual environment and pin compatible dependencies. Store credentials in environment variables or a secret manager, never in source code, notebooks, or client-side JavaScript.

    python -m venv .venv
    source .venv/bin/activate
    pip install pydantic tenacity python-dotenv

    A configuration model makes defaults visible and testable:

    from pydantic import BaseModel, Field
    
    class LLMSettings(BaseModel):
        api_key: str
        model: str = "your-production-model"
        timeout_seconds: float = Field(default=30, gt=0)
        max_retries: int = Field(default=3, ge=0, le=6)

    Load settings at process startup, validate them, and fail fast if a required value is missing. For Indian deployments, also make region, data-retention, and logging settings explicit when provider or customer contracts require them.

    Implement a provider adapter

    Keep SDK imports and response translation in one module. The rest of the application should receive your own LLMResult, not a provider-specific dictionary.

    from tenacity import retry, retry_if_exception_type, stop_after_attempt, wait_exponential
    
    class ProviderError(Exception):
        pass
    
    class APIClient(LLMClient):
        def __init__(self, sdk_client, settings: LLMSettings):
            self.client = sdk_client
            self.settings = settings
    
        @retry(
            retry=retry_if_exception_type(ProviderError),
            wait=wait_exponential(multiplier=0.5, min=1, max=8),
            stop=stop_after_attempt(3),
            reraise=True,
        )
        def complete(self, messages, *, model=None, temperature=0.2,
                     max_tokens=800, response_schema=None):
            try:
                response = self.client.responses.create(
                    model=model or self.settings.model,
                    input=list(messages),
                    temperature=temperature,
                    max_output_tokens=max_tokens,
                )
            except Exception as exc:
                # Translate only retryable provider failures in real code.
                raise ProviderError(str(exc)) from exc
    
            return LLMResult(
                text=getattr(response, "output_text", None),
                raw=response,
                model=model or self.settings.model,
                usage=getattr(response, "usage", {}) or {},
            )

    Do not retry every exception. Retry timeouts, temporary unavailability, and explicitly identified rate limits. Do not automatically retry authentication errors, invalid requests, content-policy refusals, or schema failures. Add jitter to backoff and enforce a total request deadline so retries do not create hidden latency.

    Make structured output a first-class path

    String parsing is fragile when downstream code expects fields such as language, intent, amount, or eligibility. Define a schema and validate the response before returning it to application code.

    from pydantic import BaseModel, Field
    
    class Classification(BaseModel):
        intent: str
        confidence: float = Field(ge=0, le=1)
        language: str
    
    class Classifier:
        def __init__(self, llm: LLMClient):
            self.llm = llm
    
        def classify(self, text: str) -> Classification:
            result = self.llm.complete([
                {"role": "system", "content": "Return only valid classification data."},
                {"role": "user", "content": text},
            ], response_schema=Classification)
            return Classification.model_validate_json(result.text)

    In production, use the provider’s native structured-output feature where available. Otherwise validate JSON, capture the invalid payload for debugging under your privacy policy, and decide whether to repair, retry with a constrained prompt, or return a typed failure. Never silently convert malformed output into a plausible default.

    For Indic applications, test code-switching, transliteration, spelling variation, and mixed scripts. A wrapper that assumes English-only tokenisation or Unicode-normalised input can fail on Hindi-English, Tamil-English, or voice-transcribed requests. Teams building for broader access should also review low-resource Indic NLP techniques and the constraints discussed in AI apps for India’s next billion users.

    Handle prompts, tools, and context deliberately

    Treat prompts as versioned assets rather than scattered strings. Give each prompt an identifier, record its version in telemetry, and test it against a fixed evaluation set. Keep retrieved documents, user content, and instructions in separate message fields where the provider supports them.

    If the model can call tools, validate tool names and arguments against an allowlist before execution. Apply authentication and authorisation in application code; a model must never be the final authority for whether a user may perform an action. Set limits on tool-call count, payload size, and execution time.

    For long conversations, define a context policy: truncate old turns, summarise them, or retrieve only relevant records. Track prompt and completion tokens separately because input-heavy workloads can dominate cost.

    Add observability and cost controls

    Every request should produce structured telemetry, ideally with a correlation ID. Record:

    • model and wrapper version;
    • request latency and retry count;
    • input and output token counts;
    • success, refusal, timeout, or validation status;
    • estimated cost and tenant or feature identifier.

    Do not log raw prompts or responses by default. Redact phone numbers, Aadhaar-related data, financial information, health data, and customer messages unless there is a documented, access-controlled reason to retain them. Sample traces for debugging and separate operational metrics from sensitive payloads.

    Set per-feature budgets, token caps, concurrency limits, and fallback policies. A smaller model may be appropriate for routing or extraction, while a stronger model handles complex reasoning. For customer-facing voice systems, latency and interruption behaviour matter as much as text quality; compare the wrapper’s design with voice-agent architecture and deployment patterns.

    Test the wrapper like infrastructure

    Unit-test request construction, configuration validation, response mapping, retry classification, timeout behaviour, redaction, and schema failures. Use mocked SDK clients so tests do not depend on network access or live model output.

    Add contract tests against each provider adapter. Maintain an evaluation set containing normal, ambiguous, adversarial, multilingual, and empty inputs. Measure exact-match accuracy for structured tasks, groundedness for retrieval tasks, refusal correctness, p95 latency, and cost per successful request.

    Test failure scenarios explicitly:

    • provider returns a 429 or 5xx;
    • response is truncated or malformed;
    • model returns an unexpected tool call;
    • network succeeds but exceeds the deadline;
    • credentials are absent or expired;
    • a user submits prompt-injection content;
    • a fallback model produces a different schema.

    A practical rollout checklist

    Before exposing the wrapper to users, confirm that:

    • the public interface is provider-neutral and documented;
    • secrets are managed outside the repository;
    • timeouts, retries, rate limits, and circuit breakers are configured;
    • structured outputs are validated before use;
    • prompts and model versions are traceable;
    • sensitive logs are redacted and access-controlled;
    • cost and latency budgets are visible by feature;
    • evaluation tests run in CI;
    • fallback behaviour is safe and explicit;
    • model and SDK upgrades have a rollback path.

    The goal is not to hide the LLM. It is to make its boundaries dependable. A focused wrapper lets Indian product teams change providers, support local languages, control costs, and improve reliability without rewriting core application logic. For complex multi-step systems, keep orchestration separate and study patterns for distributed systems with AI agents rather than turning one wrapper into an entire platform.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.