0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · optimizing python scripts for large scale ai data

Optimizing Python Scripts for Large-Scale AI Data

  1. aigi

    Large AI systems often fail before the model runs. A pipeline may spend most of its time parsing files, copying arrays, waiting on object storage, or repeatedly transforming the same data. Optimizing Python scripts for large-scale AI data means treating the script as a data system: measure each stage, control memory, minimise movement, and scale only where the workload justifies it.

    For Indian AI teams, this matters across multilingual datasets, document processing, medical records, satellite imagery, speech corpora, and high-volume inference. Cloud egress, GPU idle time, and duplicate storage can become larger costs than compute. The goal is not to remove Python; it is to use Python as an orchestration layer around efficient native libraries and well-designed data contracts.

    Start with a measurable pipeline

    Before changing code, define the pipeline stages and record their cost:

    • Input listing and metadata discovery
    • Download or object-storage reads
    • Decoding and parsing
    • Cleaning and validation
    • Feature or embedding generation
    • Serialisation and writes
    • Queue wait, worker utilisation, and failures

    Use representative data, not a tiny local sample. Profile both a warm run and a cold run, because filesystem cache and object-store latency can produce very different results. cProfile is useful for Python-level call counts; py-spy can inspect a running process with low overhead; Scalene helps separate CPU, memory, and copy costs. For distributed jobs, add stage-level metrics such as rows per second, bytes read, retry counts, and peak resident memory.

    A practical baseline records throughput, p95 latency, peak RAM, storage read volume, CPU utilisation, and cost per million records. Without these numbers, a faster benchmark may simply be doing less useful work.

    Use compact, typed data structures

    Python lists and dictionaries are convenient but carry substantial per-object overhead. A list of millions of strings or nested dictionaries can consume many times more memory than the underlying values. Use typed, contiguous representations wherever possible:

    • NumPy arrays for dense numerical data
    • Arrow arrays and tables for typed, interoperable columns
    • Categorical or dictionary-encoded values for repeated strings
    • int32, float32, and other suitable widths instead of defaulting to 64-bit types
    • Structured records only when their access pattern justifies them

    Avoid converting between formats at every stage. A pipeline that reads Parquet into Arrow, converts it to pandas, creates Python objects, and then converts it to tensors may spend more time copying than computing. Keep data in Arrow, NumPy, or tensor form until a Python-native object is genuinely required.

    For tabular workloads, Polars is often a strong single-machine option because its Rust execution engine supports parallel operations, predicate pushdown, and lazy query planning. Pandas remains useful for ecosystem compatibility, but avoid using it as the default for every stage. Teams handling regulated or high-stakes data should also pair performance work with a clear data veracity infrastructure approach, including provenance, validation rules, and reproducible transformations.

    Choose storage formats for the access pattern

    CSV is portable but expensive: it has no reliable schema, repeats delimiters and text representations, and requires full parsing. Convert raw inputs once into a typed format, then process the converted data repeatedly.

    Parquet works well for analytical and feature datasets because it is columnar, compressed, and supports predicate and column pushdown. Partition by fields commonly used for filtering—such as date, geography, or dataset version—but avoid creating millions of tiny files. As a practical starting point, target files large enough for efficient sequential reads and use a metadata layer so workers do not list object storage repeatedly.

    For images, audio, and video, millions of small objects can overwhelm metadata operations. Sharded archives such as WebDataset-style tar files, or a managed dataset format with sequential reads, can improve throughput. Store a manifest containing sample ID, checksum, label, source, licence, and split. This makes retries, deduplication, and audit checks much easier.

    Stream data instead of loading it all

    Let available memory determine batch size, not dataset size. A robust batch loop should:

    1. Read a bounded batch.
    2. Validate schema and required fields.
    3. Transform and write results.
    4. Release references and reuse buffers.
    5. Record progress and checkpoint state.

    Generators and iterators are useful for sequential processing, while numpy.memmap can expose large numeric files without loading them entirely into RAM. For Parquet, select only required columns and apply filters during the scan. For model training, align preprocessing batches with the accelerator’s expected batch size so the CPU does not generate data faster than the GPU can consume it.

    Do not confuse virtual memory with sufficient capacity. Memory mapping still incurs page faults, and random access over a network filesystem can be slower than a deliberate sequential read. Benchmark local NVMe, attached volumes, and object storage separately.

    Match parallelism to the bottleneck

    The Global Interpreter Lock limits concurrent execution of Python bytecode, but many numerical libraries release it. Choose the execution model based on the work:

    • Vectorisation: Use NumPy, PyTorch, or Arrow kernels for regular numerical operations.
    • Threads: Suitable for I/O-bound work and libraries that release the GIL.
    • Processes: Useful for Python-heavy CPU work, but account for serialisation and memory duplication.
    • Async I/O: Effective for many network requests, provided concurrency, timeouts, and backpressure are enforced.
    • Native compilation: Use Numba for stable numerical loops that cannot be vectorised.

    Never create unbounded tasks. Set worker counts from measured CPU, memory, file-descriptor, and storage limits. Use queues with bounded capacity so producers cannot fill RAM while consumers are waiting on a slow destination. When calling remote model or embedding APIs, implement rate limits, exponential backoff, idempotency keys, and a durable retry queue. For Python web services that feed AI workflows, the guidance on integrating LLM APIs in Python web apps is a useful complement.

    A Numba function can accelerate tight numeric loops:

    from numba import njit
    import numpy as np
    
    @njit
     def row_sum(values):
        total = 0.0
        for i in range(values.shape[0]):
            total += values[i]
        return total

    Benchmark compilation warm-up separately from steady-state execution, and confirm that the data types passed to the function are stable.

    Scale out only after the single-node path is sound

    Dask can extend familiar array and dataframe patterns across workers. Ray is useful when the workload combines data processing, actors, model inference, and task scheduling. Spark may be the better fit when your organisation already operates a mature JVM-based data platform. The tool is less important than partition design and operational discipline.

    Use partitions that are large enough to amortise scheduling overhead but small enough to retry cheaply. Avoid shipping large Python objects between workers; place immutable data in shared or object storage and pass references where the framework supports it. Watch for skew: one oversized customer, language, document, or time window can leave a worker processing long after others finish. Hash partitioning, smaller shards, and dynamic work assignment can help.

    In multilingual Indian datasets, do not assume language or script distributions are uniform. A Hindi, Tamil, or low-resource-language shard may have different tokenisation and processing costs. Measure per-partition work rather than relying only on row counts; teams building low-resource language datasets for AI training should track bytes, tokens, and decode time as well.

    Make correctness part of optimisation

    A fast pipeline that silently drops records is not efficient. Add checks for row counts, null rates, schema drift, duplicate IDs, checksums, label distributions, and train-validation leakage. Write outputs atomically and retain the input version, code revision, dependency lockfile, and configuration used for each run.

    For sensitive domains, minimise data copies and redact or tokenise personal information before broad distribution. Medical teams should align processing and verification with the relevant governance requirements; the ICMR-compliant medical AI data verification guide provides domain-specific context.

    A practical optimisation sequence

    Use this order for most projects:

    • Profile a representative end-to-end run.
    • Remove unnecessary columns, conversions, and duplicate reads.
    • Convert repeated inputs to Parquet, Arrow, or an appropriate shard format.
    • Add batching, streaming, and bounded queues.
    • Vectorise hot paths or compile only the measured bottleneck.
    • Tune worker counts and partition sizes.
    • Move to distributed execution when one machine remains the limiting factor.
    • Re-run quality, reproducibility, and cost checks after every major change.

    The best Python data pipeline is not the one with the most workers. It is the one that spends its resources on useful records, keeps memory predictable, produces auditable outputs, and delivers stable throughput as the dataset grows.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.