Python and GitHub make a strong foundation for AI products, but a working demo is only the beginning. A production application also needs reproducible dependencies, reliable data flows, evaluation, secure secrets, observable APIs, and a deployment process that a small Indian team can afford to operate.
This guide explains how to build AI applications with Python and GitHub in 2026. It covers the path from an initial repository to a tested service, with practical choices for traditional machine learning, retrieval-augmented generation (RAG), computer vision, and agentic applications.
Start with the product boundary
Before installing a model, define one user-facing task. “Build an AI chatbot” is too broad; “answer questions from a company’s GST policy documents and cite the source” is testable. Write down:
- The input and expected output.
- Acceptable latency and monthly request volume.
- Accuracy, refusal, and safety requirements.
- Whether data may leave India or be sent to a third-party API.
- The cost limit per request.
Choose the smallest architecture that meets the requirement. A scikit-learn classifier may be better than an LLM for structured prediction. A hosted model API may be more practical than self-hosting when traffic is uncertain. For complex tool-using workflows, study the design trade-offs in this guide to building distributed systems with AI agents before adding multiple agents.
Create a reproducible Python repository
Use a predictable repository structure so another developer can run the project without reverse-engineering your laptop:
ai-app/
├── src/app/
│ ├── api.py
│ ├── model.py
│ ├── retrieval.py
│ └── settings.py
├── tests/
├── scripts/
├── pyproject.toml
├── Dockerfile
├── README.md
└── .github/workflows/ci.ymlUse Python 3.11 unless your selected framework requires another version. Manage the environment with venv, uv, Poetry, or Conda; the tool matters less than committing a lockfile and documenting the exact setup. A minimal start might be:
python -m venv .venv
source .venv/bin/activate
pip install fastapi uvicorn pytest ruff pydantic-settings
pip freeze > requirements.txtFor a maintained project, prefer pyproject.toml with pinned or bounded dependencies and separate development dependencies. Add .env to .gitignore, commit .env.example, and make the README explain local, staging, and production commands.
Select the model and application pattern
Your Python architecture should follow the task rather than the popularity of a framework.
- Predictive ML: Use pandas or Polars for data preparation, scikit-learn for baselines, and PyTorch for custom deep learning.
- RAG: Parse and chunk documents, create embeddings, retrieve relevant passages, and generate an answer constrained by those passages.
- Computer vision: Define image quality checks, preprocessing, inference, and confidence thresholds separately. See this workflow for building computer vision models on GitHub.
- Agents: Keep tools explicit, validate tool inputs, set timeouts, and limit the number of steps. Do not allow an agent unrestricted shell, database, or network access.
- Voice: Separate speech recognition, orchestration, text-to-speech, interruption handling, and call-state management. A voice agent architecture and deployment guide covers these boundaries in greater detail.
Start with a baseline and an evaluation set before optimising. For an Indian-language application, include representative Hindi, Tamil, Bengali, or code-mixed examples—not only translated English prompts. Record expected answers, citations, refusal cases, and sensitive inputs in version-controlled evaluation files, with personal data removed or synthetic.
Version code, data, and models correctly
Git should remain the source of truth for application code, configuration templates, schemas, tests, and evaluation definitions. It is not a suitable home for large datasets or every model checkpoint.
- Use Git LFS only when the files are small enough for your repository policy and genuinely need Git-style versioning.
- Use DVC or an object store for large datasets and training artefacts; commit immutable references and checksums to GitHub.
- Store model versions in a registry or release artefact with metadata: training data version, commit SHA, metrics, licence, and known limitations.
- Never commit API keys, customer records, raw call recordings, or private documents.
A pull request should make it possible to answer: what changed, which data or prompt changed, how quality changed, and whether the cost or latency moved. This discipline is especially important for regulated use cases such as legal or financial applications; a private AI chatbot for lawyers illustrates why access control and data handling must be designed early.
Serve inference through a tested API
FastAPI is a practical choice for Python inference services. Keep model loading outside the request handler so the model is initialised once per worker. Validate inputs with Pydantic, return stable response schemas, and add request IDs for tracing.
from fastapi import FastAPI
from pydantic import BaseModel
app = FastAPI()
class Query(BaseModel):
text: str
@app.post("/predict")
def predict(query: Query):
# Call a separately tested model or service here.
return {"label": "example", "request_text": query.text}Add health endpoints, timeouts, rate limits, structured logs, and graceful handling for provider errors. For long-running jobs, use a queue rather than holding an HTTP connection open. Keep synchronous and asynchronous workloads separate, and measure p50, p95, and p99 latency. Teams planning larger workloads should review guidance on scaling backend infrastructure for AI applications.
Add GitHub Actions before deployment
A useful CI pipeline should run on every pull request:
1. Format and lint with Ruff.
2. Run unit tests and API contract tests with pytest.
3. Scan dependencies and secrets.
4. Execute a small model smoke test.
5. Run the evaluation suite and fail when quality falls below a defined threshold.
6. Build the Docker image and scan it for vulnerabilities.
Deploy only from a protected branch or tagged release. Use GitHub Actions environments for staging and production approvals, and store credentials in environment-specific secrets or, preferably, short-lived cloud identity federation. Do not print prompts, documents, tokens, or full user inputs in CI logs.
Control compute and cost in India
Prototype on CPU when possible, then benchmark the actual bottleneck. For GPU workloads, compare hosted APIs, rented cloud GPUs, and Indian providers on total cost, data residency, support, and minimum commitment—not hourly price alone. Quantisation, batching, caching, smaller embeddings, and prompt limits often deliver larger savings than premature infrastructure changes.
For production, track cost per successful request, not only infrastructure spend. Add fallback models for provider outages, circuit breakers for repeated failures, and a queue for burst traffic. If your application is an agent, cap tool calls and token budgets explicitly.
Security, evaluation, and launch checklist
Before inviting real users, verify:
- Secrets are stored outside the repository and rotated.
- Authentication, authorisation, tenant isolation, and deletion workflows work.
- Logs do not expose personal or confidential information.
- Prompt injection, unsafe uploads, data exfiltration, and abusive usage are tested.
- Outputs are evaluated on accuracy, groundedness, toxicity, latency, and cost.
- Model and dependency licences permit your intended use.
- Rollback to the previous model and application version has been rehearsed.
Document limitations in the README and product UI. An AI feature that says when it is uncertain is more useful than one that confidently invents an answer.
FAQ
Do I need a GPU? No. Use a CPU for API integration, tests, small models, and orchestration. Rent or access a GPU only when benchmarks show it is necessary.
Can GitHub host the live AI application? GitHub hosts code, workflows, and selected artefacts. The API, model, database, and storage must run on appropriate infrastructure.
Should I use an LLM framework? Only when it removes real complexity. Direct Python calls are often easier to test; add an orchestration framework when you need retrieval, tools, memory, or tracing that you can govern.
What should an Indian student or early founder build first? Choose a narrow workflow with measurable value, publish a reproducible repository, use synthetic or permissioned data, and collect evaluation results before seeking scale. Open-source contributors can also start with this guide to contributing to AI GitHub repositories in India.
Apply for AI Grants India
If you have a working prototype, evaluation results, and a clear plan for responsible deployment, apply for an AI grant. Strong applications explain the user problem, technical approach, early evidence, compute requirements, and how funding will accelerate the next product milestone.