What you are building
A custom AI agent is an application that interprets a request, decides which steps are needed, uses approved tools or data, and returns an outcome. Unlike a simple chatbot, an agent can retrieve documents, call an API, update a record, or hand a task to a human. GitHub provides the collaboration and engineering layer: repositories, issues, pull requests, automated tests, secrets management, and deployment workflows.
The most reliable approach is to build a narrow agent around one measurable job. Examples include triaging support tickets, extracting fields from invoices, checking application documents, or answering questions from an internal knowledge base. Avoid starting with a general-purpose “do everything” assistant.
For Indian products, design for multilingual inputs, uneven connectivity, local compliance requirements, and cost-sensitive inference from the beginning. If your use case involves Indic languages, the guidance on low-resource Indic natural language processing is a useful companion.
Choose the agent architecture before the framework
Start with the simplest architecture that can meet the requirement:
- Deterministic workflow: fixed steps, validation, and API calls. Best for predictable business processes.
- Retrieval-augmented generation (RAG): searches approved documents before generating an answer. Best for policies, product information, and internal support.
- Tool-using agent: selects from a defined set of functions such as search, calculation, or ticket creation.
- Human-in-the-loop agent: pauses for approval before irreversible or sensitive actions.
- Multi-agent system: separates roles across specialist agents. Use only when a single orchestrator becomes difficult to test; distributed designs add coordination and failure modes, as discussed in building distributed systems with AI agents.
Write a one-page specification containing the user, trigger, allowed actions, prohibited actions, input sources, expected output, latency target, and success metric. A strong metric might be “correctly routes 90% of support requests with fewer than 5% unsafe escalations,” rather than “sounds intelligent.”
Set up a maintainable GitHub repository
Create a private repository for sensitive or commercial work. A practical structure is:
agent-project/
├── app/ # API, orchestration, and business logic
├── tools/ # Explicit tool definitions and permission checks
├── prompts/ # Versioned system and task prompts
├── data/ # Schemas and local test fixtures, not secrets
├── evals/ # Golden examples and scoring scripts
├── tests/ # Unit and integration tests
├── .github/workflows/
├── pyproject.toml
├── README.md
└── .env.examplePin dependencies, use a virtual environment, and keep configuration separate from code. Never commit API keys, personal data, production exports, or model credentials. Store secrets in GitHub Actions Secrets or a managed secret vault, and use separate credentials for development, staging, and production.
Your README should explain local setup, supported models, environment variables, tool permissions, test commands, deployment steps, and known limitations. Add a SECURITY.md file with a responsible disclosure process. Use branch protection and require review for changes to prompts, tools, authentication, and data handling.
Select the model and agent framework
Choose based on task quality, latency, context length, hosting constraints, and total cost—not popularity. Hosted models can shorten time to market; open-weight models may offer greater control, data residency, or predictable economics. For every candidate, test the same representative dataset and record:
- task accuracy and citation quality;
- tool-call accuracy and malformed arguments;
- latency and token usage;
- refusal and escalation behaviour;
- cost per successful task;
- performance across English, Hindi, and relevant regional languages.
Frameworks can help with state, tool calling, tracing, and orchestration, but they do not replace application design. Keep the core business rules in ordinary, testable code. Use a framework behind a small adapter so you can change models or libraries without rewriting the product.
Build the first agent loop
A safe agent loop has five explicit stages:
1. Validate input: authenticate the user, apply rate limits, and reject unsupported requests.
2. Plan or classify: decide whether the request can be answered, requires retrieval, needs a tool, or must go to a person.
3. Execute approved tools: validate structured arguments and enforce permissions server-side.
4. Verify the result: check schemas, calculations, citations, and business constraints.
5. Respond or escalate: return a concise answer with sources, or explain why human review is required.
A simplified Python shape might look like this:
def handle_request(user, request):
validate_request(request)
decision = router.classify(request)
if decision == "human_review":
return create_review_task(user, request)
context = retriever.search(request) if decision == "retrieve" else None
plan = planner.create(request, context=context)
result = tool_runner.execute(plan, user_permissions=user.permissions)
verified = verifier.check(result, request)
return formatter.respond(verified)Do not let a model generate arbitrary shell commands, database queries, or network destinations. Expose narrow functions such as lookup_order(order_id) or create_callback_task(phone_number), then validate every argument and log every invocation.
Add evaluation before adding features
Create an evaluation set before launch. Include normal cases, ambiguous requests, prompt injection attempts, missing information, multilingual phrasing, long inputs, and tool failures. Store inputs, expected properties, and labels in version-controlled fixtures without personal information.
Run evaluations in pull requests whenever prompts, retrieval logic, model versions, or tools change. Combine automated checks with human review. Useful checks include exact field validation, citation presence, prohibited-action detection, groundedness, and escalation accuracy. Track regressions by model and repository commit so a cheaper model cannot quietly reduce quality.
Observability matters in production. Record request IDs, model and prompt versions, latency, token counts, tool names, error classes, and escalation outcomes. Redact phone numbers, Aadhaar details, health information, payment data, and other personal identifiers before logs are stored.
Secure and deploy the agent
Threat-model the system before exposing it to users. Prompt injection can arrive through a webpage, email, uploaded file, or retrieved document. Treat all external content as untrusted. Separate instructions from data, restrict tools by role, limit network access, and require confirmation for payments, deletion, messages, or changes to official records.
Use GitHub Actions for linting, tests, dependency scanning, container builds, and staged deployment. A sensible release path is:
- pull request checks on every change;
- deployment to a staging environment after review;
- evaluation and smoke tests in staging;
- gradual production rollout with rollback support;
- weekly review of cost, errors, unsafe outputs, and user feedback.
For voice-first products, the same principles apply to speech recognition, turn-taking, tool permissions, and escalation. Compare the architecture with a conventional IVR using this voice agent versus IVR guide, or review the broader voice agent architecture and deployment guide.
A practical 30-day build plan
Days 1–5: define the job, users, risks, baseline process, and evaluation set.
Days 6–12: create the repository, implement one workflow, add one or two tools, and use synthetic data.
Days 13–20: add retrieval if necessary, implement authentication, logging, validation, and human escalation.
Days 21–26: run multilingual and adversarial tests, measure cost and latency, and fix the largest failure modes.
Days 27–30: deploy to a limited cohort, document operations, configure rollback, and review grant or pilot-readiness evidence.
Common mistakes to avoid
- Building a multi-agent system before proving a single workflow.
- Treating a prompt as a security boundary.
- Giving tools broad permissions or unrestricted internet access.
- Measuring fluent answers instead of completed, correct tasks.
- Committing secrets or real customer data to GitHub.
- Skipping human escalation for high-impact decisions.
- Failing to pin model and dependency versions.
- Ignoring Indian language, accent, connectivity, and cost requirements.
GitHub is most valuable when it supports disciplined iteration: clear issues, small pull requests, reproducible evaluations, and visible ownership. Build the smallest agent that solves a real operational problem, then expand only when evidence justifies the added autonomy.