0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · building serverless ai apps with modal

Building Serverless AI Apps with Modal

  1. aigi

    Modal is a strong fit for teams that need GPU-backed AI without operating Kubernetes, managing VM images, or keeping expensive servers running between requests. You define Python environments, remote functions, GPU requirements, storage, and HTTP endpoints in code; Modal provisions the underlying compute when work arrives and scales it down when demand falls.

    For Indian founders, student builders, and small research teams, this model shortens the path from notebook experiment to usable product. It is particularly effective for bursty workloads such as image generation, speech transcription, document processing, evaluation jobs, and batch inference. It is less suitable when you need a permanently warm service, strict data-residency guarantees in India, or highly predictable low-latency traffic that justifies reserved infrastructure.

    What Modal provides

    Modal is a Python-first serverless compute platform designed for workloads that exceed the limits of conventional function services. A Modal application typically combines:

    • Images that declare Python packages, system libraries, and runtime dependencies.
    • Functions that run remotely on CPU or GPU containers.
    • Web endpoints that expose functions over HTTP.
    • Volumes and cloud storage for model weights, checkpoints, and generated artefacts.
    • Secrets for API keys and credentials injected only at runtime.
    • Scaling controls for concurrency, timeouts, retries, and container reuse.

    The programming model is close to ordinary Python, but the execution boundary matters. Local code submits work; remote functions execute in Modal’s infrastructure. Design your application around that boundary instead of treating a remote function as a local method with unlimited access to local files, processes, or network connections.

    Modal can complement a broader architecture. For example, a frontend may run on a conventional web host, a relational database may store users and job status, and Modal may handle the GPU-heavy portion. Teams building high-performance AI applications with open-source tools can use this separation to keep product logic portable while outsourcing burst compute.

    A practical architecture

    A production-grade serverless AI application usually has four layers:

    1. Request layer: An API receives a prompt, file, or job request and validates authentication, size, and format.
    2. Compute layer: A Modal function loads the model and performs inference on the required CPU or GPU.
    3. State layer: A database records job status and metadata; object storage holds inputs and outputs that are too large for an API response.
    4. Client layer: A web or mobile application polls a job, receives a webhook, or streams a result.

    Keep synchronous endpoints for short operations. For large documents, video, image batches, and fine-tuning, return a job identifier and process the task asynchronously. Modal’s parallel execution primitives can replace much of the plumbing normally built with worker queues. This is useful for workflows such as building a voice agent with Whisper and ElevenLabs, where transcription, language processing, and speech generation may have different latency and hardware requirements.

    Create a GPU inference service

    Install the CLI and authenticate your account:

    pip install modal
    python -m modal setup

    A current Modal application can be structured like this:

    import io
    import modal
    
    app = modal.App("image-generation-api")
    image = (
        modal.Image.debian_slim(python_version="3.11")
        .pip_install("diffusers", "transformers", "accelerate", "torch")
    )
    model_volume = modal.Volume.from_name(
        "diffusion-models", create_if_missing=True
    )
    
    @app.function(
        image=image,
        gpu="A10G",
        volumes={"/models": model_volume},
        timeout=300,
        scaledown_window=60,
    )
    def generate_image(prompt: str) -> bytes:
        import torch
        from diffusers import StableDiffusionPipeline
    
        pipe = StableDiffusionPipeline.from_pretrained(
            "/models/sd-model", torch_dtype=torch.float16
        ).to("cuda")
        result = pipe(prompt, num_inference_steps=25).images[0]
        output = io.BytesIO()
        result.save(output, format="PNG")
        return output.getvalue()
    
    @app.function(image=image)
    @modal.fastapi_endpoint(method="POST")
    def generate(prompt: str):
        return {"image": generate_image.remote(prompt).decode("latin1")}

    The exact decorators and endpoint helpers can change as Modal evolves, so verify the syntax against the current Modal documentation before deployment. The important design choices are stable: define dependencies explicitly, attach the appropriate GPU, mount persistent storage, and keep model loading separate from request validation.

    For lower latency, load the pipeline once per container rather than reconstructing it for every request. Modal’s container reuse can preserve an in-memory model between calls, while a Volume prevents repeated downloads when new containers start. Do not assume a Volume is a substitute for a transactional database: use it for model artefacts and cacheable files, not user records or concurrent writes that require strong consistency.

    Model loading, caching, and cold starts

    Large model weights dominate startup time and often dominate operational cost indirectly. Use these practices:

    • Download and verify weights in a controlled setup job rather than during every request.
    • Pin model revisions and package versions so a redeploy cannot silently change outputs.
    • Store weights in a mounted Volume or object store close to the execution environment.
    • Keep the container image focused; unnecessary packages increase build and startup time.
    • Warm only the functions where latency justifies the additional spend.
    • Quantise or distil models when quality requirements allow it.

    For multilingual Indian products, benchmark the complete pipeline rather than only the model. Tokenisation, language detection, audio decoding, image preprocessing, and post-processing can become the real bottleneck. If you are serving Indic-language chat or support workflows, compare an open model with hosted APIs using the evaluation methods described in building multilingual chatbots for Indian startups.

    Batch jobs and parallel workloads

    Modal is especially valuable when work can be split into independent units:

    @app.function(image=image, gpu="T4", timeout=900)
    def transcribe(url: str) -> dict:
        # Download audio, run Whisper, and return structured segments.
        return {"url": url, "text": "..."}
    
    # In a remote Modal function, fan out work in parallel.
    results = list(transcribe.map(audio_urls))

    Use bounded concurrency rather than launching unlimited GPU jobs. Respect provider quotas, downstream API limits, and your own budget. For retries, make operations idempotent: a repeated job should not charge a customer twice, duplicate a database row, or overwrite a final artefact unexpectedly.

    Agentic applications add another consideration. If several model calls, tools, and workers collaborate, define explicit budgets for tokens, runtime, and retries. Patterns from building distributed systems with AI agents are relevant here: track correlation IDs, preserve intermediate state, and make failures visible instead of hiding them inside a single long-running function.

    Security and production readiness

    Create Modal Secrets for provider keys and read them from environment variables at runtime. Never commit credentials to the repository or bake them into an image. Add authentication at the public endpoint, validate uploaded files, restrict prompts and payload sizes, and avoid returning raw internal exceptions to users.

    Before launch, establish:

    • Observability: request IDs, latency, cold-start rate, GPU utilisation, failures, and cost per successful job.
    • Reliability: timeouts, retries with backoff, cancellation, and a dead-letter path for failed jobs.
    • Privacy: retention rules, redaction of personal data, encryption, and a review of where data is processed.
    • Quality: regression prompts, language-specific tests, hallucination checks, and model-version tracking.
    • Capacity: concurrency limits and a fallback response when GPUs or external APIs are unavailable.

    If your application handles health, finance, education, or government data, document the data flow before selecting a deployment region. “Serverless” does not remove compliance obligations; it changes who operates the compute layer.

    Controlling Modal costs

    GPU billing is usually driven by execution time, selected hardware, and idle container duration. Measure cost per request or per completed document, not just hourly GPU rates. A practical optimisation sequence is:

    • Start with the smallest GPU that meets latency and memory requirements.
    • Batch compatible requests where it improves utilisation without harming user experience.
    • Set function timeouts and short scale-down windows for bursty traffic.
    • Move long jobs to asynchronous execution so HTTP clients do not hold connections open.
    • Cache deterministic results and avoid re-embedding or re-transcribing identical inputs.
    • Compare quantised, distilled, and hosted models against your quality target.

    Convert dollar-denominated infrastructure costs into rupees per user action for planning in India, and include bandwidth, storage, observability, payment processing, and support. A low idle bill can still become expensive if a poorly bounded endpoint accepts unlimited traffic.

    When to choose Modal

    Choose Modal when your team is Python-heavy, GPU demand is intermittent, and speed of iteration matters more than controlling every layer of infrastructure. Consider a conventional cloud service or dedicated GPU nodes when traffic is steady, you need a specific private network topology, or long-lived in-memory state is central to the product.

    The best first project is a narrow vertical slice: one authenticated endpoint, one model, one representative Indian-language or domain-specific dataset, and measurable latency and cost targets. Once that works, add queues, larger models, streaming, and multi-region routing deliberately. Builders who want to keep the implementation open and inspectable can also explore open-source AI tools for Indian developers before committing to a proprietary serving stack.

    For founders building a serious AI product, AI Grants India can help connect technical progress with funding and ecosystem support. Treat Modal as an execution layer—not the product itself—and use the flexibility it provides to validate the product, economics, and user experience quickly.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.