Performance work is most effective when treated as an engineering loop: measure, identify the bottleneck, change one variable, and verify the result. The goal is not to make every component faster in isolation. It is to deliver a responsive experience at the required reliability and cost, including on mobile networks and lower-end devices common across India.
This guide explains how to optimize system performance for web apps in 2026, with a practical focus on latency, throughput, resilience, and operating cost. It applies to conventional SaaS products as well as AI-enabled applications that combine web requests with model inference, retrieval, file processing, or asynchronous jobs.
1. Define performance targets before changing code
Start with service-level objectives (SLOs), not a list of fashionable tools. Set separate targets for the user experience and backend behaviour:
- Latency: Track p50, p95, and p99 response times rather than averages. A good average can hide a poor experience for a significant minority of users.
- Availability: Define acceptable downtime and error rates for critical journeys such as login, checkout, search, or an AI response.
- Throughput: Measure requests per second, jobs per minute, database transactions, and concurrent users.
- Resource efficiency: Monitor CPU, memory, database connections, bandwidth, and cost per active user or completed task.
- Regional performance: Test from Indian networks and regions, including mobile connections. A result from a single cloud region is not a complete picture.
Instrument key journeys with browser telemetry, API metrics, structured logs, and distributed traces. A performance budget—such as a JavaScript limit, maximum API payload size, or p95 latency target—turns optimisation into an enforceable product constraint. Teams building complex architectures can use a system design learning resource to reason about trade-offs before introducing more infrastructure.
2. Find the bottleneck with profiling and traces
Do not begin by adding Redis, rewriting a service, or moving to a new programming language. First establish where time is spent:
- Break request duration into DNS, connection setup, TLS, server processing, database calls, external APIs, and response transfer.
- Add correlation IDs so one user request can be followed across services, queues, and database calls.
- Use continuous profiling or targeted profilers to identify CPU-heavy functions, lock contention, excessive allocations, and garbage-collection pauses.
- Capture slow-query logs and execution plans, not just query duration.
- Load-test realistic traffic patterns, including bursts, large payloads, cache misses, and concurrent writes.
Synthetic tests reveal predictable regressions; real-user monitoring reveals device, browser, ISP, and geography-specific problems. For AI products, trace token generation time, time to first token, retrieval latency, queue wait, model retries, and tool-call duration separately. A streaming interface can improve perceived responsiveness, but it does not fix an overloaded inference service.
3. Optimise the database and connection path
Databases are frequent bottlenecks because every request competes for finite CPU, memory, locks, and connections. Begin with the highest-volume and slowest queries:
- Add indexes for common filters, joins, and sort orders, but validate them with
EXPLAINor the database’s query-plan tools. - Prefer composite indexes that match the actual access pattern; column order matters.
- Avoid selecting unused columns and paginate large result sets with keyset pagination where offset pagination becomes expensive.
- Eliminate N+1 queries in ORM code and batch related reads.
- Use connection pooling, with limits sized for the database rather than copied from application replicas. PostgreSQL deployments may benefit from PgBouncer.
- Keep transactions short and choose isolation levels deliberately to reduce lock contention.
- Archive or partition high-volume tables when retention requirements make unbounded growth costly.
Read replicas can increase read capacity, but replication lag creates consistency surprises. Route only operations that tolerate stale data to replicas. Denormalisation is useful when it removes repeated joins on a read-heavy path, but document the update strategy and test write amplification.
4. Build a cache hierarchy with explicit freshness rules
Caching works when the data’s lifetime and invalidation behaviour are understood. Use several layers selectively:
- Browser and CDN cache: Cache immutable versioned assets for a long time. Use Brotli where supported, correct
Cache-Controlheaders, and responsive image variants. - Edge caching: Cache public pages and safe API responses near users. Choose a CDN with useful Indian points of presence and verify routing from major networks rather than assuming proximity.
- Application cache: Redis is suitable for short-lived objects, rate limits, locks, and carefully designed session data. Set TTLs, size limits, and eviction policies.
- Database-level tactics: Use efficient indexes and query plans first. Treat query-result caching as an optimisation with invalidation costs, not a substitute for a healthy schema.
Define whether each cached item is immutable, stale-while-revalidate, or must be strongly consistent. Prevent cache stampedes with request coalescing, jittered expirations, or a single-flight mechanism. Never cache personalised or sensitive responses without a clear key design and privacy review.
5. Reduce frontend work, not just download size
A smaller bundle helps, but users also pay for parsing, compiling, executing, layout, and rendering. Improve the critical path by:
- Server-rendering or statically generating content that does not require client-side interactivity.
- Splitting JavaScript by route and loading non-critical components on demand.
- Removing unused dependencies and measuring the cost of third-party scripts.
- Reserving image and video dimensions to prevent layout shifts; serve AVIF or WebP with responsive
srcsetvariants. - Preloading only truly critical resources and using
deferorasyncappropriately. - Virtualising long lists and reducing unnecessary state updates and re-renders.
- Keeping accessibility and low-power device performance in the acceptance criteria.
Track Core Web Vitals, including LCP, INP, and CLS, alongside backend metrics. A fast API cannot compensate for a page that blocks the main thread for several seconds.
6. Make APIs predictable and asynchronous
API design directly affects both latency and infrastructure cost. Return only the fields a client needs, enforce payload limits, paginate collections, and use compression for suitable text responses. Idempotency keys are essential for retryable payments, provisioning, and other mutation endpoints.
Move email, notifications, document generation, media processing, and model inference into queues when they do not need to complete inside the request. Workers should have bounded concurrency, visibility timeouts, retry policies, and a dead-letter path. Backpressure matters: accepting unlimited work merely shifts failure from the API to the queue or database.
For AI features, consider streaming output, batching compatible inference requests, caching embeddings where valid, and using smaller or quantised models for routine tasks. If your product combines Python APIs with language models, this guide to integrating LLM APIs in Python web apps covers implementation concerns that affect latency and reliability.
7. Scale infrastructure without losing control
Keep application servers stateless where practical so instances can scale horizontally behind a load balancer. Store durable state in managed databases or object storage, and use health checks that test meaningful dependencies rather than merely confirming that a process is alive.
Autoscaling should respond to demand signals such as request concurrency, queue depth, or latency—not CPU alone. Set minimum and maximum capacity, scale-up speed, cooldown periods, and budget alerts. Vertical scaling can be the simplest short-term fix; horizontal scaling improves fault tolerance but adds coordination and observability overhead.
Serverless is useful for bursty, event-driven work, but account for cold starts, execution limits, outbound networking, and vendor costs. For heavier AI workloads, compare managed APIs, GPU instances, and serverless GPU platforms against workload duration and utilisation. A practical serverless AI apps guide can help evaluate this choice.
8. Turn optimisation into an operating practice
Performance regresses through dependency upgrades, new features, data growth, and traffic changes. Protect gains with automated load tests, performance budgets in CI, query-plan reviews, and dashboards that show latency by endpoint, region, status code, and release.
Run failure drills for database exhaustion, cache loss, queue backlogs, CDN misconfiguration, and model-provider outages. Use feature flags and gradual rollouts so a performance regression can be isolated and reversed quickly. Review cost per request alongside latency: an improvement that doubles the bill may be unsuitable for an early-stage Indian startup.
A practical optimisation sequence
For most teams, this order produces the fastest learning:
1. Instrument the critical user journeys and establish p95 targets.
2. Fix obvious errors, N+1 queries, unbounded work, and oversized payloads.
3. Optimise the slowest database queries and connection usage.
4. Add browser, CDN, and application caching with explicit invalidation rules.
5. Reduce frontend execution and rendering work.
6. Move non-critical work to bounded asynchronous workers.
7. Load-test peak traffic, then tune autoscaling and capacity.
8. Re-measure after every material change and document the trade-off.
High-performance engineering is a repeatable discipline, not a one-time infrastructure purchase. Teams building demanding AI products can also review open-source tools for high-performance AI applications when selecting components, but should benchmark each choice against their own traffic, data, and reliability requirements.
FAQ
What should I optimise first? Start with measurement and the slowest high-volume path. In many applications, that means an inefficient database query, excessive backend work, or a large frontend bundle—not a wholesale architecture change.
Is TTFB the most important metric? TTFB is valuable for diagnosing server and network delay, but it is not the whole user experience. Pair it with LCP, INP, CLS, API p95 latency, error rate, and resource cost.
Should every web app use microservices? No. A well-structured modular monolith is often easier to operate and can scale far enough for an early product. Split services when independent scaling, ownership, isolation, or deployment needs justify the added complexity.
How do I handle AI latency? Separate queue wait, retrieval, model time, tool calls, and output streaming in your traces. Use the smallest reliable model, cache reusable work, stream where appropriate, and keep heavy inference off the synchronous request path when possible.