Start with the system, not the model
AI model architecture guidance is most useful when it begins with the product requirement rather than a favourite framework or model family. Before designing layers, answer five questions:
- What decision or user interaction must the system support?
- What data is available, and who is permitted to use it?
- What accuracy, latency, availability, and cost targets matter?
- Where will inference run: cloud, private infrastructure, a device, or a hybrid?
- What happens when the model is uncertain or wrong?
For an Indian product, also account for multilingual inputs, code-mixed speech and text, intermittent connectivity, regional accents, privacy expectations, and infrastructure budgets. A model that performs well on an English benchmark may fail on Hindi-English queries, low-quality mobile images, or domain terminology used outside major metros.
Write these constraints into a short architecture brief. It becomes the reference point for model selection, experiments, procurement, and later production reviews.
Choose the simplest architecture that meets the target
The right architecture is not necessarily the largest one. Start with a baseline that is easy to test and operate, then add complexity only when measurements justify it.
Common choices
- Classical machine learning: Strong for tabular data, ranking, forecasting, and credit or operations workflows when features are reliable and explainability matters.
- Convolutional or vision transformers: Useful for image classification, detection, segmentation, document processing, and quality inspection. Teams building vision prototypes can review computer vision models on GitHub for practical starting points.
- Sequence models: RNNs and temporal convolutional networks remain relevant for compact time-series and streaming workloads, although transformers often dominate larger language and multimodal tasks.
- Transformers and small language models: Suitable for text classification, extraction, summarisation, retrieval, and generation. For Hindi-first or multilingual products, compare models on your own data; open-source small language models for Hindi offer a useful evaluation path.
- Retrieval-augmented generation (RAG): Appropriate when answers must reflect changing company, legal, academic, or public information. Retrieval does not fix a weak source corpus, so document quality and citation checks are part of the architecture.
- Hybrid systems: Combine rules, search, classifiers, and generative models. This is often the best option for regulated or operational workflows because deterministic checks can constrain model output.
Do not use a generative model for a task that a classifier, SQL query, or rules engine can solve more cheaply and predictably.
Design the data path before tuning the network
Model quality is usually limited by data quality, coverage, and labels—not by a missing layer. Document the complete path from input to output:
1. Collection: Record source, consent or licence, geography, language, and timestamp.
2. Validation: Check schema, duplicates, corrupt files, missing values, unsafe content, and leakage between training and test sets.
3. Representation: Define tokenisation, image resolution, audio sampling, categorical encoding, or embedding strategy.
4. Labelling: Create clear guidelines, measure annotator agreement, and retain difficult examples for review.
5. Splitting: Use time-based or group-based splits where random splitting would leak information.
6. Monitoring: Track drift in language, devices, users, locations, and class balance after launch.
For Indian deployments, evaluate performance by language, script, state or region where relevant, device type, network quality, and user segment. Aggregate accuracy can conceal serious failures for smaller language communities.
Match architecture to inference constraints
Architecture decisions should include a cost and latency budget from the first experiment. Measure end-to-end performance, not just model inference time.
- Cloud inference supports larger models and centralised updates, but introduces network latency, recurring usage charges, and data-governance questions.
- On-device or edge inference improves responsiveness and can keep sensitive data local. Quantisation, pruning, distillation, and smaller backbones may be necessary. See this guide to AI model optimisation for mobile devices when Android, low-memory hardware, or offline operation is part of the brief.
- Hybrid inference can route easy cases locally and escalate uncertain cases to a server model.
- Batch inference is economical for reports, recommendations, and back-office processing; real-time inference is justified only when the user experience requires it.
Record parameter count, memory use, throughput, p50 and p95 latency, energy consumption where relevant, and cost per 1,000 requests. These numbers make architecture trade-offs explicit.
Build evaluation around failure modes
Accuracy alone is not an acceptance criterion. Select metrics that correspond to the decision being made:
- Precision and recall for detection and moderation
- F1 or macro-F1 for imbalanced classification
- Calibration and abstention rate when confidence affects action
- Word error rate for speech systems
- Retrieval recall, groundedness, citation accuracy, and refusal quality for RAG
- Human preference or task-completion rate for assistants
- Latency, uptime, cost, and memory for production operations
Keep a fixed evaluation set, a recent holdout set, and a deliberately difficult challenge set. Test prompt injection, personally identifiable information, unsupported claims, adversarial inputs, and out-of-distribution examples. Establish a fallback: ask for clarification, return “not found,” use a deterministic workflow, or send the case to a human.
For model selection, reproduce results with fixed seeds where possible and log dataset versions, code commits, configuration, and hardware. Framework choice matters, but it should follow the team’s deployment needs. A practical comparison of AI frameworks for Indian student entrepreneurs can help teams choose between PyTorch, TensorFlow, JAX, and higher-level tooling.
Plan the production architecture
A production model is one component in a larger system. Separate training, evaluation, serving, and monitoring so each can be changed safely.
A useful minimum design includes:
- Versioned datasets, model artefacts, prompts, and configuration
- A reproducible training and evaluation pipeline
- An API or serving layer with authentication, rate limits, timeouts, and retries
- Input and output validation, logging, and redaction of sensitive fields
- A model registry and rollback mechanism
- Monitoring for quality proxies, drift, latency, cost, and abuse
- Human review for high-impact or low-confidence decisions
For agentic applications, constrain tool access, validate tool arguments, cap loops, and maintain audit logs. Teams building conversational systems should treat voice, retrieval, orchestration, and observability as separate layers; the voice agent architecture and deployment guide provides a useful reference model.
A practical decision sequence for 2026
Use this sequence to avoid premature scaling:
1. Define the user, task, risk, and measurable success threshold.
2. Build a simple non-generative or small-model baseline.
3. Establish representative Indian-language and real-world evaluation data.
4. Compare hosted, open-weight, and fine-tuned options on quality, latency, privacy, and total cost.
5. Add retrieval, tools, or multimodal inputs only when they solve a measured gap.
6. Pilot with shadow traffic or a limited cohort before full release.
7. Monitor failures and retrain or revise the system using versioned evidence.
The strongest architecture is the one your team can evaluate, secure, operate, and improve. Treat model choice as an engineering decision tied to data and product constraints—not as a permanent commitment to a particular architecture.