0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to optimize python code for ai performance

How to Optimize Python Code for AI Performance

  1. aigi

    Python remains the control layer for much of India’s AI stack, but readable code is not automatically fast code. In a training pipeline, recommendation service, document-processing workflow, or LLM application, performance depends on the full path from data ingestion to model output. A slow Python loop may matter less than repeated data copies, inefficient tokenisation, CPU–GPU transfers, or a model that is too large for the latency target.

    This guide explains how to optimize Python code for AI performance in a repeatable way. The focus is practical: measure first, improve the highest-impact bottleneck, and verify that the change improves latency, throughput, memory use, or cloud cost without reducing accuracy.

    Start with a performance target

    Define the metric before changing the code. Different AI workloads require different optimisations:

    • Batch training: samples per second, step time, GPU utilisation, and peak memory.
    • Real-time inference: p50, p95, and p99 latency, plus requests per second.
    • Batch inference: total processing time, cost per million records, and failure rate.
    • Data preparation: records per second, storage read speed, and worker utilisation.
    • Edge deployment: model size, cold-start time, power consumption, and device memory.

    Set a baseline using a representative dataset and fixed hardware. Include warm-up runs, because the first request may include model loading, compilation, or memory allocation. For production services, track performance by model version and input size rather than relying on a single average.

    Profile before you optimise

    Optimisation without measurement often produces harder-to-maintain code with little benefit. Use a layered approach:

    • `cProfile` or `py-spy`: identify functions consuming the most wall-clock time.
    • `line_profiler`: inspect expensive lines inside known hot functions.
    • `memory_profiler` and `tracemalloc`: locate allocation spikes and retained objects.
    • PyTorch Profiler or TensorBoard: separate Python overhead, kernel execution, data loading, and GPU idle time.
    • Application monitoring: measure queue time, database calls, API latency, and model inference separately.

    Profile realistic batches. A function that is slow on 100 rows may be irrelevant when the real job processes millions; conversely, a small serialisation overhead can dominate a low-latency API. For LLM systems, monitor token generation speed, time to first token, prompt length, and external provider latency. The principles in LLM application performance monitoring in India are useful when Python is only one part of the serving path.

    Replace Python loops with array and tensor operations

    Native Python loops have substantial per-element overhead. Prefer operations implemented in NumPy, pandas, PyTorch, or another compiled library. These libraries process contiguous blocks of data in C, C++, or GPU kernels.

    import numpy as np
    
    scores = np.asarray(scores, dtype=np.float32)
    normalised = (scores - scores.mean()) / (scores.std() + 1e-8)

    This is generally faster and more memory-efficient than repeatedly updating individual list elements. Use boolean masks, broadcasting, reductions, and matrix operations where possible. Avoid converting between lists, NumPy arrays, pandas objects, and tensors inside a hot loop.

    Vectorisation is not always the answer. A large temporary array can increase peak memory and trigger swapping. Measure both runtime and memory, and use chunked vectorisation when the complete dataset does not fit comfortably in RAM.

    Control data movement and memory use

    In AI workloads, data movement frequently costs more than computation. Keep data in an appropriate format and avoid unnecessary copies:

    • Store tabular datasets in Parquet or another columnar format instead of repeatedly parsing CSV.
    • Select only required columns and apply filters while reading.
    • Use float32 or lower precision when model quality permits; do not default to float64.
    • Reuse preallocated buffers in repeated numerical operations.
    • Release references to large intermediate objects and inspect memory growth in long-running workers.
    • Pin and batch data when transferring from CPU to GPU, but verify that pinned memory is helping the actual workload.

    For data-heavy projects, reusable preprocessing utilities can make the biggest difference. See Python scripts for automating data preprocessing for patterns covering validation, cleaning, feature creation, and repeatable dataset preparation.

    Make training pipelines feed the accelerator

    A powerful GPU cannot improve throughput if it spends most of its time waiting for Python or storage. For PyTorch, tune DataLoader settings such as num_workers, pin_memory, persistent_workers, and prefetch_factor. Benchmark combinations on the deployment machine rather than copying settings from another environment.

    Use sufficiently large batches to improve device utilisation, subject to memory limits. Gradient accumulation can simulate a larger batch, but it may increase wall-clock time. Mixed precision, such as FP16 or BF16 where supported, can improve throughput and reduce memory consumption. Always validate loss curves and final model quality after changing precision.

    Move transfers out of the critical path where possible, overlap data preparation with computation, and avoid calling synchronisation methods such as unnecessary GPU-to-CPU conversions during every step. Logging every batch can also force synchronisation and distort measurements.

    Compile only the right functions

    JIT compilation is effective for stable numerical kernels that remain difficult to vectorise. Numba can accelerate CPU loops over NumPy arrays, while framework-level compilation can optimise portions of a neural network. Use explicit signatures or appropriate modes when available, and exclude unsupported Python objects from compiled functions.

    Compilation has startup overhead and may create specialised versions for different input shapes. It is therefore most useful for repeated workloads, not one-off scripts. Benchmark warm and cold execution separately. For model serving, consider exporting to an optimised runtime such as ONNX Runtime or a vendor-supported inference engine when compatibility and accuracy are confirmed.

    Improve inference latency and throughput

    Inference optimisation should match the service’s traffic pattern:

    • Batch requests when throughput matters and users can tolerate queueing.
    • Dynamic batching when requests arrive continuously and the serving framework supports it.
    • Cache repeated embeddings, retrieval results, or deterministic preprocessing outputs.
    • Use smaller models, quantisation, or distillation when accuracy remains within the product target.
    • Keep models loaded rather than reinitialising them per request.
    • Limit serialisation overhead by using compact payloads and avoiding repeated conversions.

    For an LLM API integration, stream responses only when it improves perceived latency, cap unnecessary context, and avoid making serial external calls in a request path. The engineering considerations in Integrating LLM APIs in Python web apps complement low-level Python tuning.

    Parallelise carefully

    Use multiprocessing or job queues for CPU-bound work that releases little benefit from Python threads. Threads are often appropriate for I/O-bound tasks, such as downloading files or calling services. Joblib and Dask can help with independent preprocessing or batch workloads, but parallelism introduces serialisation, memory, scheduling, and coordination costs.

    Do not create a worker per small task. Partition work into sufficiently large batches, limit the number of processes to available CPU and memory, and prevent nested parallelism—for example, a process pool whose workers each start their own BLAS threads. Test on the same container or cloud instance used in production.

    Build a benchmark into the workflow

    Keep a small, repeatable benchmark in the repository. It should include fixed inputs, warm-up iterations, representative batch sizes, and checks for output correctness. Record:

    • Runtime and throughput
    • p50 and p95 latency for services
    • Peak resident memory and GPU memory
    • CPU and GPU utilisation
    • Cost per job or request
    • Accuracy, output drift, and error rates

    Run it in CI for significant changes, and use a performance budget to catch regressions. Automated review tools can help flag inefficient patterns, but they should not replace measurement; automated production-grade code reviews with AI works best when paired with concrete benchmarks and test data.

    A practical optimisation sequence

    For most Python AI projects, use this order:

    1. Establish a representative baseline and correctness tests.
    2. Profile end to end, then inspect the hottest functions.
    3. Reduce data reads, conversions, copies, and serial API calls.
    4. Vectorise numerical work and batch model calls.
    5. Tune data loading, device transfers, precision, and memory use.
    6. Apply JIT, compilation, quantisation, or a faster runtime to proven hotspots.
    7. Parallelise only after measuring task size and coordination overhead.
    8. Re-run quality, latency, throughput, and cost checks before shipping.

    This approach keeps optimisation tied to product outcomes. For teams building complete systems rather than isolated scripts, build end-to-end ML pipelines in Python provides the broader structure for reproducible data, training, evaluation, and deployment.

    FAQ

    Does vectorisation always make Python AI code faster?
    No. It usually improves numerical workloads, but large temporary arrays or frequent conversions can increase memory use. Benchmark the complete operation.

    Should every Python AI workload use a GPU?
    No. Small models, irregular preprocessing, and low-volume jobs may run faster and cheaper on CPUs. Measure utilisation and total cost before selecting hardware.

    When should I use Numba?
    Use it for repeated, numerical CPU functions that are difficult to express with existing vectorised operations. It is less suitable for code dominated by Python objects, strings, or external calls.

    How do I avoid optimising the wrong part of an application?
    Set a target, profile end to end, and change one major factor at a time. Confirm improvements with representative inputs and regression tests.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.