High-density AI infrastructure turns power and heat into first-order architecture constraints. A conventional enterprise rack may draw a few kilowatts, while GPU training and inference systems can exceed 30–100 kW per rack depending on accelerator generation, server configuration and utilisation. In India, this challenge is amplified by grid variability, high ambient temperatures, water stress in some regions, monsoon conditions and uneven availability of data-centre capacity.
Power and thermal design for high-density Indian AI clusters must therefore be treated as one engineering problem. Every additional kilowatt delivered to a rack becomes almost exactly one kilowatt of heat that must be removed. The right design balances electrical capacity, cooling technology, uptime targets, operating cost, land, water, deployment schedule and future GPU generations.
Why AI clusters require a different design approach
AI clusters are not simply larger versions of CPU-oriented data centres. Their electrical and thermal profiles are more concentrated and dynamic:
- High rack density: AI servers combine multiple GPUs, high-speed networking, memory and local storage in a small footprint.
- Large step loads: Training jobs can start or stop many servers simultaneously, creating rapid changes in demand.
- High heat flux: Modern accelerators generate concentrated heat that may exceed the practical limits of air cooling.
- Network-driven layouts: GPU fabrics require short, low-latency cable paths, often constraining rack placement and containment.
- Continuous operation: Training runs can last days or weeks, making thermal excursions and power interruptions particularly costly.
A design based only on average IT load can fail during peak training, GPU boost operation or simultaneous equipment restart. Engineers should model real workload envelopes, including sustained load, transient load, idle-to-peak behaviour and future expansion.
Start with the electrical load model
The first step is a bottom-up load model. Avoid sizing the facility from a single nameplate number or an assumed average utilisation percentage.
A practical rack-level estimate is:
Rack IT load = GPU server load + CPU/storage load + network load + rack-level auxiliariesAt cluster level:
Total IT load = sum of rack loads + storage and management systems + test and staging capacityThen account for electrical losses:
Facility power = IT load / electrical-chain efficiencyThe electrical-chain efficiency includes UPS conversion, transformers, switchgear, distribution losses and power-distribution units. A design with 10 MW of IT load does not require only 10 MW at the utility connection; the facility must also cover cooling, pumps, controls, lighting and electrical losses.
Nameplate, measured and design load
Use three separate values:
- Nameplate load: The maximum rating printed on equipment. It is useful for protection and worst-case checks but may overstate normal operation.
- Expected operating load: The measured or modelled consumption under the intended AI workload.
- Design load: The value used for infrastructure sizing, including uncertainty, growth and failure scenarios.
For new GPU platforms, obtain vendor power data for the complete server, not just the accelerator. Include GPU power limits, CPU TDP, memory, fans, NICs, drives and motherboard consumption. If the platform supports dynamic power capping, model both capped and uncapped operation.
Rack density and floor planning
Rack density determines whether a hall can use conventional air cooling, hybrid cooling or liquid-dominant systems. It also affects structural loading, busway placement, maintenance clearances and fire protection.
Indicative planning bands are useful, but they are not substitutes for equipment data:
- Below 15 kW per rack: Generally compatible with conventional air cooling if airflow management is strong.
- 15–30 kW per rack: May require enhanced containment, higher airflow, rear-door heat exchangers or partial liquid cooling.
- 30–60 kW per rack: Typically needs a carefully engineered hybrid or liquid-assisted design.
- Above 60 kW per rack: Direct-to-chip liquid cooling, immersion or a comparable high-capacity solution is usually required.
The exact threshold depends on supply-air temperature, allowable GPU inlet temperature, server fan capability, rack geometry and local climate. Do not mix low-density general-purpose racks and high-density AI racks without modelling airflow and cooling zones separately.
Reserve space for:
- Electrical busways and tap-off units
- Coolant distribution units (CDUs)
- Manifolds and dripless quick disconnects
- Pump and control equipment
- Service aisles and extraction paths
- Future rack expansion
- Fire detection and suppression equipment
A compact layout may reduce construction cost but increase maintenance risk and restrict future GPU upgrades. For an Indian AI startup, modular halls or dedicated high-density pods can reduce the risk of overbuilding the entire facility.
Power architecture for Indian deployments
A resilient AI cluster usually includes utility supply, medium-voltage transformation, UPS systems, static transfer or automatic transfer equipment, low-voltage distribution and rack-level power delivery. The architecture should reflect the required uptime rather than automatically copying a hyperscale design.
Utility and grid considerations
Indian sites may experience voltage variation, interruptions, feeder constraints and scheduled maintenance. Before selecting a location, assess:
- Contracted demand and sanctioned load
- Dual-feed availability and substation capacity
- Power-quality history, including sags and harmonics
- Renewable power procurement options
- Diesel-generator permissions and fuel logistics
- Local data-centre, electrical and environmental requirements
A dual utility feed is valuable only if the feeds are genuinely independent. Trace them back through substations, transformers and cable routes before assigning redundancy credit.
UPS topology and autonomy
AI workloads are sensitive to interruptions, but the required battery autonomy depends on generator start time, utility reliability and operational policy. Online double-conversion UPS systems provide isolation and stable output, while lithium-ion batteries can reduce footprint and maintenance compared with traditional valve-regulated lead-acid systems.
Key decisions include:
- N, N+1 or 2N capacity
- Distributed versus centralised UPS architecture
- Battery chemistry and fire protection
- Maintenance bypass arrangements
- Generator ride-through time
- Recovery after a cluster-wide power event
Redundancy should be evaluated at the failure-domain level. A nominally N+1 UPS system may still have a single switchboard, control panel or cooling loop that can interrupt the entire AI hall.
Generator and fuel strategy
Backup generation is often necessary where utility reliability cannot support the required service level. Size generators for the actual critical load, including cooling systems, pumps, controls and starting transients. Large GPU clusters can create demanding step-load behaviour; generator selection should be validated with dynamic load-step analysis rather than steady-state kW alone.
Consider emissions controls, noise, fuel storage, refuelling during extended outages and local approvals. Where possible, use staged generator loading and power-management controls to avoid unnecessary operation at very low loads.
Cooling options for high-density AI racks
Enhanced air cooling
Air remains attractive because it is familiar, serviceable and compatible with existing data-centre operations. It can support moderate densities when combined with:
- Hot-aisle or cold-aisle containment
- Proper blanking panels and sealed cable openings
- Adequate raised-floor or overhead supply paths
- High-efficiency EC fans
- Variable-speed computer-room air handlers
- Accurate airflow balancing
Air cooling becomes inefficient when fan power rises sharply or when supply air must be excessively cold to protect a small hot spot. Lowering room temperature is not a substitute for correct airflow design.
Rear-door heat exchangers
A rear-door heat exchanger captures heat at the rack exhaust and transfers it to a water loop. This can extend the useful range of air cooling without modifying every server. It is suitable for mixed-density halls, but door weight, hose routing, leak detection and maintenance access must be addressed.
Direct-to-chip liquid cooling
Direct-to-chip systems attach cold plates to GPUs and often CPUs, removing most of the heat at its source. A CDU separates the facility water loop from the technology loop and controls flow, pressure, filtration and heat exchange.
Design considerations include:
- Coolant chemistry and material compatibility
- Supply and return temperature
- Flow rate and pressure drop
- CDU redundancy
- Leak detection and automatic isolation
- Quick-disconnect reliability
- Service procedures for replacing servers
- Residual air load from memory, drives, fans and power supplies
Warm-water cooling can improve chiller efficiency because it permits higher return temperatures and, in suitable climates, more hours of economisation. However, the allowable temperature must match the GPU and server vendor specifications.
Immersion cooling
Single-phase or two-phase immersion can achieve very high density and reduce fan energy. It also changes maintenance, fluid handling, server compatibility and supply-chain requirements. Immersion is most appropriate when density or acoustic constraints justify the operational complexity. Validate fluid availability, warranty conditions, fire protection and technician training before committing.
India-specific thermal and environmental design
Indian climate data should drive the cooling design. A site in a hot, humid coastal city has different economiser potential and corrosion risk from a dry inland location. Peak design conditions must account for dry-bulb temperature, wet-bulb temperature, humidity, dust and monsoon operation.
Water strategy
Cooling towers and evaporative systems can reduce energy consumption but consume water and may face restrictions during scarcity. A credible water plan should include:
- Annual and peak-day water demand
- Cooling-tower cycles of concentration
- Blowdown treatment and disposal
- Make-up water quality
- Rainwater and recycled-water opportunities
- Drought and municipal supply contingencies
- Water usage effectiveness (WUE) targets
Air-cooled chillers reduce water dependence but may increase peak electrical demand. Liquid cooling does not automatically mean low water use; the facility heat-rejection system still determines much of the water profile.
Dust, humidity and corrosion
Filtration, positive pressure and appropriate fresh-air control are essential near construction zones, industrial areas or dusty environments. Excessive humidity can cause condensation risk, while aggressive dehumidification increases energy use. Coastal sites may need enhanced corrosion controls for outdoor coils, electrical equipment and metalwork.
Measuring efficiency: PUE, WUE and carbon
Power Usage Effectiveness is calculated as:
PUE = total facility energy / IT equipment energyA lower PUE is desirable, but it should not be pursued by compromising availability or operating temperature limits. Measure PUE over time and by operating condition rather than relying on a design estimate.
Also track:
- WUE: Water consumption per unit of IT energy
- CUE: Carbon emissions associated with energy use
- Cooling kW per IT kW: Helps identify cooling-system performance
- UPS and distribution losses: Reveals electrical inefficiency
- GPU energy per training run or inference request: Connects facility efficiency to business output
For Indian AI companies, renewable procurement, open-access power, green tariffs and behind-the-meter solar-plus-storage may reduce emissions or energy cost, subject to state-specific rules and availability. Solar generation alone does not replace firm capacity for night-time or monsoon-period AI workloads.
Controls, monitoring and operational readiness
High-density clusters need granular telemetry. Monitor at minimum:
- Utility, generator, UPS and busway power
- Rack-level real power and apparent power
- GPU and CPU temperatures
- Coolant supply and return temperature
- Flow, pressure, conductivity and leak sensors
- Room temperature, humidity and differential pressure
- Chiller, pump and fan efficiency
Integrate building management, data-centre infrastructure management and cluster orchestration where practical. Power-aware scheduling can defer non-urgent training, cap GPU power during grid constraints and avoid exceeding a rack or electrical branch limit.
Set alert thresholds based on action, not just alarm volume. A rising coolant temperature should trigger a defined workload-throttling or migration procedure. Test controls during commissioning, including sensor failures, network loss, pump failure, utility loss and emergency shutdown.
Commissioning and capacity planning
Commissioning should validate the complete system under realistic AI loads. Recommended tests include:
1. Verify electrical protection, phase balance and harmonic performance.
2. Apply progressive rack loads to confirm distribution and cooling capacity.
3. Test generator step response and UPS ride-through.
4. Simulate a failed CDU, pump, chiller, CRAH unit or electrical path.
5. Confirm leak detection, isolation and recovery procedures.
6. Run sustained GPU workloads at expected peak power.
7. Measure rack inlet temperatures and temperature uniformity.
8. Validate restart sequencing after a total power interruption.
Use a digital capacity register that tracks installed, reserved and available power and cooling by hall, row, rack and branch circuit. This prevents sales or engineering teams from allocating capacity that exists only on paper.
Cost and procurement considerations
The lowest capital-cost design may create the highest lifetime cost if it limits rack density or requires early retrofit. Compare options using total cost of ownership, including:
- Electrical infrastructure and utility upgrades
- Cooling plant and distribution
- Water treatment and disposal
- UPS batteries and replacement cycles
- Generator fuel and maintenance
- Technician training and spare parts
- Energy under realistic utilisation
- Downtime and workload interruption risk
- Future GPU compatibility
Specify measurable performance requirements in procurement documents. Ask vendors for full-load and part-load efficiency curves, maximum coolant temperature, minimum and maximum flow, acoustic data, service intervals, warranty exclusions and failure-mode behaviour.
Practical design checklist
Before approving a high-density Indian AI cluster, confirm that the project has:
- A measured or vendor-validated rack power model
- Separate normal, peak and future design loads
- A defined failure-domain and redundancy strategy
- Verified utility and backup-generation capacity
- Cooling selected against peak heat, not average heat
- Liquid-cooling compatibility confirmed with server warranties
- Water and heat-rejection plans suited to the site
- Protection against dust, humidity and monsoon conditions
- Rack-level power and thermal telemetry
- Tested emergency, leak and restart procedures
- A phased expansion plan with reserved electrical and cooling capacity
- Documented operating procedures and trained technicians
FAQ
What rack density requires liquid cooling?
There is no universal threshold. Many facilities begin evaluating liquid cooling around 20–30 kW per rack, while racks above 50–60 kW commonly need direct-to-chip or immersion solutions. Server design and site conditions are decisive.
Is air cooling unsuitable for Indian AI data centres?
No. Air cooling can work well at moderate density with containment, filtration and correctly sized airflow systems. It becomes less practical as GPU heat flux and rack power increase.
How much power does cooling add?
The answer depends on climate, cooling technology and operating point. Model cooling as part of facility power and use measured PUE rather than applying a fixed percentage to IT load.
Should startups build their own AI data centre?
Often, no. Colocation or managed GPU infrastructure can reduce capital exposure and accelerate deployment. Owning infrastructure may make sense when utilisation is high, workloads are predictable and the company needs specialised density, security or control.
What is the most important early design decision?
Establish a credible rack-level power and heat model before choosing the site, electrical architecture or cooling plant. Incorrect assumptions at this stage are expensive to correct later.
Apply for AI Grants India
Building efficient, resilient AI infrastructure can strengthen an Indian startup’s technical and commercial case. Apply through AI Grants India to explore support and opportunities for your AI venture.