Open-weight models make serious AI development accessible without requiring a hyperscale budget. But the hardware decision is easy to get wrong: a GPU may have excellent compute performance yet fail to load a model because it lacks VRAM, has weak software support, or cannot fit comfortably into your power and cooling setup.
For Indian developers, the practical choice is usually between an existing workstation, a new consumer GPU, a used high-memory card, or rented cloud capacity. The right answer depends on whether you are running inference, fine-tuning, evaluating models, or building a production service.
Start with the model, not the GPU
“Open-weight” means that model weights are available for download, but it does not imply that every model can be freely used commercially. Check the model licence, hardware requirements, context length, and supported inference stack before buying equipment.
Estimate memory needs using these rough figures:
- FP16 or BF16 weights: approximately 2 bytes per parameter.
- INT8 weights: approximately 1 byte per parameter, plus runtime overhead.
- 4-bit weights: approximately 0.5 bytes per parameter, plus metadata and the KV cache.
- Runtime overhead: reserve additional VRAM for activations, CUDA kernels, batching, and the context window.
A 7B model in 4-bit precision may run comfortably on a 8–12 GB card for single-user inference. A 13B model is more comfortable with 16–24 GB, while 30B–34B models generally need 24 GB or multiple GPUs after quantisation. Larger models require multi-GPU systems, CPU offload, or cloud instances.
The KV cache becomes important when serving long prompts, multiple users, or large batches. A model that fits during a short test may run out of memory in production. Leave headroom rather than sizing a card to the exact file size of the weights.
GPU recommendations by workload
12–16 GB: local experimentation and smaller models
A current mid-range NVIDIA card with 12–16 GB VRAM is a sensible entry point for students, independent developers, and teams testing 7B–14B models. It can handle quantised chat models, embeddings, reranking, image generation at moderate resolutions, and parameter-efficient fine-tuning on smaller datasets.
This tier is also suitable for developers following practical projects such as open-source AI projects for student developers. Prioritise a modern architecture, efficient cooling, and at least 32 GB of system RAM.
24 GB: the strongest single-card workstation tier
A 24 GB GPU remains one of the most useful configurations for local AI work. It provides room for larger quantised language models, LoRA or QLoRA fine-tuning, computer vision training, vision-language experiments, and higher context lengths.
Older 24 GB cards such as the RTX 3090 can still offer strong value on the used market, but inspect condition carefully. Check memory temperatures, fan noise, warranty status, power connectors, and whether the card has spent years in sustained workloads. Newer 24 GB cards may offer better performance per watt and improved software features, but their price in India can vary significantly between retailers.
For computer vision work, GPU memory is only one part of the equation. Dataset pipelines, augmentation, storage speed, and CPU performance also matter. If you are building models from repositories, this guide to building computer vision models on GitHub provides a useful workflow alongside the hardware decision.
48–80 GB: fine-tuning, evaluation and production serving
Professional cards such as NVIDIA RTX 6000-class GPUs, A100s, H100s, and newer data-centre accelerators are designed for sustained workloads, ECC memory, virtualisation, and multi-GPU deployment. Their value is not simply faster token generation. They offer substantially more memory, better reliability, and support for workloads that cannot fit on a consumer card.
These cards make sense when you need to fine-tune larger models, run long-context evaluation, serve several concurrent users, or avoid splitting a workload across multiple consumer GPUs. They are expensive to purchase, so Indian startups often begin with cloud rental, measure utilisation, and only then consider a dedicated server.
NVIDIA, AMD, Apple and cloud alternatives
NVIDIA remains the safest default for open-weight model development because CUDA, cuDNN, PyTorch integrations, quantisation libraries, and deployment frameworks are widely tested on it. This reduces setup time, especially when using libraries such as vLLM, llama.cpp with CUDA, TensorRT-LLM, bitsandbytes, or popular fine-tuning frameworks.
AMD hardware can be attractive where price or availability is favourable. ROCm support has improved, but compatibility remains workload-specific. Confirm that your exact GPU, operating system, PyTorch version, and chosen inference library are supported before purchasing. A cheaper card that requires substantial debugging may cost more in engineering time.
Apple Silicon systems are useful for quiet local prototyping and unified-memory experiments. They are not a direct replacement for a CUDA workstation when you need mature multi-GPU training or maximum inference throughput.
Cloud GPUs are often the best option for irregular workloads. Compare the full hourly price, storage, data-transfer fees, setup time, and minimum billing period. Indian teams should also account for data residency, latency, GST treatment, and whether sensitive datasets can leave the organisation. Use local hardware for frequent development and cloud capacity for occasional large runs when the economics support it.
Build the system around the GPU
A GPU upgrade is rarely just a GPU purchase. Plan for:
- Power: high-end cards can draw hundreds of watts. Use a reliable PSU with suitable transient headroom and correct power cables.
- Cooling: sustained training exposes weak airflow quickly. Choose a case with clear intake and exhaust paths, and monitor hotspot and memory temperatures.
- System RAM: 32 GB is a practical floor for local model work; 64 GB or more is preferable for larger models, preprocessing, and CPU offload.
- Storage: use a fast NVMe SSD for model files, caches, datasets, and checkpoints. Large model collections can consume hundreds of gigabytes.
- Motherboard layout: verify slot spacing, PCIe lanes, physical clearance, and support for multiple cards before buying.
- Operating system: Linux usually offers the smoothest experience for CUDA, containers, and production inference. Windows can work well for local experimentation but may require more troubleshooting.
Keep model environments reproducible with Docker or pinned Python dependencies. Record driver, CUDA, PyTorch, quantisation, and serving-library versions so an experiment can be repeated or moved to a cloud instance.
A practical buying framework
Choose hardware in this order:
1. Define the model sizes, quantisation formats, context lengths, and concurrency you need.
2. Calculate VRAM with overhead rather than relying on the model download size.
3. Decide whether your workload is occasional or continuous.
4. Compare purchase cost with cloud rental for your expected monthly GPU hours.
5. Check software compatibility for the exact stack you plan to use.
6. Benchmark a representative prompt, batch size, context length, and generation target.
For a first workstation, a reliable 24 GB NVIDIA GPU is usually more versatile than a faster card with 12 or 16 GB. For an API serving many users, prioritise memory capacity, throughput, observability, and predictable uptime. For a research project with uncertain requirements, rent before committing to an expensive multi-GPU machine.
The same discipline applies when moving from experiments to production. Teams deploying open-source agents should assess queueing, batching, fallback models, rate limits, and monitoring; this production guide for open-source AI agents covers the operational side.
Bottom line
The best GPU for open-weight models is the one that fits the model in memory, supports your software stack, and remains affordable at your expected utilisation. In 2026, NVIDIA is still the lowest-friction choice for most builders, while AMD, Apple Silicon, and cloud accelerators can be sensible alternatives for specific workloads.
Start with a quantised model and a clear benchmark. Measure tokens per second, latency, VRAM use, power draw, and cost per useful run—not just headline specifications. For teams building Indian-language applications, also test the actual scripts, context lengths, and evaluation data you will use. Hardware that performs well on a generic benchmark may behave differently on a real multilingual workload, including projects involving open-source vision-language models for Indian languages.
FAQ
How much VRAM do I need? 12–16 GB is enough for many smaller quantised models. Choose 24 GB for a flexible single-card workstation, and 48 GB or more for larger fine-tuning and multi-user serving.
Is a used RTX 3090 still worth considering? It can be, particularly for 24 GB of VRAM at a lower price. Test memory stability, temperatures, fans, power delivery, and warranty before buying.
Can I run a model across multiple GPUs? Yes, but multi-GPU inference introduces communication overhead, configuration complexity, and additional power costs. A single card with enough VRAM is simpler when available.
Should I buy or rent? Buy when usage is frequent and predictable. Rent when you need occasional large experiments, specialised accelerators, or a quick capacity test.
Can CPUs replace GPUs? CPUs are useful for small models, preprocessing, and fallback inference, but GPUs usually deliver much better latency and throughput for interactive workloads.
Apply for AI Grants India
If you are building an AI product, research prototype, or Indian-language application, apply to AI Grants India for potential support, visibility, and connections to the wider builder ecosystem.