0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · ai pipeline resource descriptors

AI Pipeline Resource Descriptors: Schema, Examples and Best Practices

  1. aigi

    AI pipelines fail for reasons that are often invisible in model code: a dataset changes without notice, a GPU job exhausts memory, an evaluation run cannot be reproduced, or a production endpoint uses a different model artifact from the one approved. AI pipeline resource descriptors provide a shared, structured record of what each pipeline stage needs, produces, and depends on.

    A descriptor can be a YAML or JSON file, a database record, or metadata managed by an orchestration platform. The format matters less than the discipline: every important input, output, resource constraint, and lineage decision should be explicit enough for another engineer—or an automated system—to validate and run the pipeline.

    What an AI pipeline resource descriptor should capture

    A useful descriptor connects five layers of information:

    • Inputs: datasets, schemas, prompts, feature tables, credentials references, and external services.
    • Processing: code version, container image, pipeline stage, framework, parameters, and dependencies.
    • Resources: CPU, GPU, memory, disk, network, runtime limits, and scaling behaviour.
    • Outputs: model files, embeddings, reports, metrics, logs, and deployment packages.
    • Controls: ownership, data classification, retention, approval status, lineage, and validation rules.

    This makes the descriptor more than documentation. It becomes an operational contract between data engineering, ML engineering, platform teams, and reviewers.

    For example, a training stage might require a versioned image, a Parquet dataset with a fixed schema, one NVIDIA GPU with 24 GB of VRAM, 64 GB of RAM, and a maximum runtime of six hours. It might produce a model artifact, a metrics file, and a data-quality report. Writing those requirements down lets an orchestrator schedule the job correctly and lets reviewers assess whether the run is reproducible.

    A practical descriptor structure

    There is no single universal standard for every AI workload. Start with a small schema and extend it only when a real operational need appears. The following fields work well for most batch, training, evaluation, and inference pipelines:

    name: support-call-summariser-training
    version: 1.3.0
    owner: ml-platform@organisation.in
    stage: training
    code:
      repository: git.example.in/ai/support-summariser
      commit: 8f31c2a
    runtime:
      image: registry.example.in/ai/train:2026.04
      python: "3.12"
    inputs:
      - name: transcripts
        uri: s3://data/support/transcripts/2026-03/
        format: parquet
        schema_version: 4
        sensitivity: restricted
    resources:
      cpu: 8
      memory_gb: 64
      gpu: "1x24GB"
      max_runtime_minutes: 360
    outputs:
      - name: model
        uri: s3://models/support-summariser/
        format: safetensors
    quality_gates:
      min_rouge_l: 0.42
      max_hallucination_rate: 0.03
    lineage:
      parent_run: run-2026-04-18-017

    The example separates stable identity from run-specific values. Keep the descriptor versioned in Git, while storing generated values—such as run ID, timestamps, and actual cost—in the orchestration system or experiment tracker.

    Resource fields that prevent expensive failures

    Compute and storage

    Specify minimum and preferred resources separately where possible. A training job may run on a CPU fallback for development but require a GPU in production. Record GPU memory, not just GPU count; a single 16 GB device cannot support the same batch size as a 48 GB device. Include ephemeral disk for dataset extraction, checkpointing, and temporary indexes.

    For Indian teams operating under tight budgets or intermittent access to accelerators, this distinction is critical. Descriptors can support queue priorities, spot or pre-emptible policies, and regional placement. If a model can run on modest hardware, link the descriptor to a design approach such as building lightweight ML models for low-resource hardware.

    Data contracts

    Record the dataset URI, owner, licence or access basis, schema version, expected row count range, partitioning, and freshness requirement. Add checks for null rates, duplicate IDs, label balance, language coverage, and personally identifiable information. A descriptor should state whether a failed check blocks the run or merely creates a warning.

    This is especially important for Indic datasets, where language, script, transliteration, and dialect coverage can change model behaviour. Teams working with scarce or unevenly distributed data should also consult low-resource language datasets for AI training in India and document sampling decisions rather than treating the dataset as homogeneous.

    Runtime and dependencies

    Pin the container image by digest where practical, record framework versions, and list system libraries that affect tokenisation, audio processing, or GPU kernels. Capture environment variables by reference—never embed secrets in the descriptor. Define network access explicitly, including whether a stage may call an external API and how retries are handled.

    Descriptors across the pipeline lifecycle

    A descriptor should evolve with the workload, but its changes must be reviewable.

    1. Plan: specify intended inputs, resource limits, acceptance criteria, and data controls.
    2. Build: bind the descriptor to a code commit and immutable runtime image.
    3. Run: record actual resources, duration, failure reason, cost, and output locations.
    4. Evaluate: attach metrics, test-set identity, fairness checks, and human review results.
    5. Deploy: define serving resources, latency targets, model signature, rollback artifact, and monitoring thresholds.
    6. Retire: record deprecation date, retained artifacts, and deletion obligations.

    For teams standardising orchestration, a descriptor should complement—not replace—the pipeline definition. See implementing scalable ML pipelines for predictive analytics for a broader production architecture, or build end-to-end ML pipelines in Python for an implementation-oriented workflow.

    Governance, security and Indian deployment realities

    Resource descriptors are valuable evidence for audits, but only if they identify ownership and policy decisions. Include data classification, lawful-use basis where relevant, retention period, approved regions, and access roles. For sensitive workloads, record whether data is encrypted at rest and in transit, whether logs may contain user content, and how deletion requests propagate to derived artifacts.

    Do not put API keys, personal data, or raw prompts into descriptors. Use references to a secrets manager and maintain access-controlled lineage. For public-sector, healthcare, education, and financial applications in India, add human-approval gates and an incident contact. A descriptor should make it possible to answer: which data trained this model, who approved it, what changed, and how can we reproduce or roll back the run?

    Cost and performance optimisation

    Once descriptors are machine-readable, teams can compare expected and actual resource use. Track cost per training run, cost per thousand inferences, GPU utilisation, queue time, storage growth, and data egress. Set budgets or alerts at the stage level instead of discovering overspend at the end of a billing cycle.

    Use descriptors to test alternatives: smaller batch sizes, quantisation, caching, CPU inference, or a different model family. For LLM systems, include context-window limits, retrieval index size, embedding model, and maximum tokens. If the workflow has several model calls, document each stage and its fallback path. A descriptor makes trade-offs visible before they become production incidents; for complex language workflows, compare it with a multi-stage LLM pipeline for developers.

    Common mistakes to avoid

    • Documentation without enforcement: validate descriptors in CI and fail builds when required fields or schemas are missing.
    • Vague resource claims: replace “high memory” with measurable CPU, RAM, VRAM, storage, and runtime values.
    • Mutable references: pin dataset snapshots, image digests, model versions, and code commits.
    • One giant schema: use a small core schema plus stage-specific extensions.
    • Ignoring actual usage: write runtime telemetry back to run metadata and compare it with requested resources.
    • Treating governance as an afterthought: classify data and define retention before the first production run.

    A 30-day adoption plan

    Week 1: inventory pipeline stages and agree on required fields, owners, naming, and versioning. Week 2: create templates for batch, training, evaluation, and serving; add schema validation to pull requests. Week 3: integrate descriptors with the orchestrator, registry, experiment tracker, and cost monitoring. Week 4: pilot on one production-bound workflow, review failures, and publish a reusable standard.

    Start with the pipeline that causes the most operational friction, not the one that is easiest to document. A good descriptor should reduce ambiguity, improve scheduling, accelerate incident response, and make responsible deployment routine rather than ceremonial.

    Last updated 24 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.