What a custom AI agent is
A custom AI agent is software that uses a language model or other AI model to pursue a defined goal by planning work, calling tools, reading context, and deciding what to do next. It is more than a chatbot prompt: a useful agent has a clear scope, controlled access to systems, observable actions, and a way to recover when a model is uncertain.
GitHub is useful because it brings the entire engineering workflow together: source control, issues, pull requests, automated tests, secrets management, documentation, and deployment integrations. The platform does not automatically make an agent reliable. Your repository structure, evaluation method, permissions, and operating discipline do that.
Start with a narrow, measurable use case
Avoid beginning with “build a general-purpose agent”. Choose one workflow with a clear input, output, and success condition. Good first projects include:
- Classifying and routing support tickets.
- Extracting fields from invoices or government forms.
- Searching an internal knowledge base and drafting an answer.
- Summarising GitHub issues and proposing next actions.
- Collecting restaurant bookings or patient follow-up details through a voice interface.
For India-focused products, define language, channel, and compliance requirements at the start. A customer-support agent may need English, Hindi, and one or more regional languages; an agent handling health information needs stronger access controls and audit trails. For low-resource language work, see this practical guide to Indic natural language processing.
Write an agent contract before writing code:
- Goal: what outcome must it achieve?
- Allowed actions: which tools may it call?
- Forbidden actions: what must always require human approval?
- Inputs and outputs: what formats are accepted and returned?
- Success metrics: accuracy, task completion, latency, cost, or escalation rate.
Design the repository before the agent
Create a small repository that makes behaviour easy to inspect. A practical structure is:
agent-project/
├── app/ # orchestration and application code
├── tools/ # typed integrations and validation
├── prompts/ # versioned system and task instructions
├── evals/ # test cases and scoring scripts
├── tests/ # unit and integration tests
├── docs/ # architecture, runbooks, and threat model
├── .env.example
├── pyproject.toml
└── README.mdKeep prompts, tool definitions, model settings, and business rules separate. Store model responses and evaluation results as versioned test artefacts, but remove personal data and secrets. Use branches and pull requests for changes to prompts and agent policies; these changes can alter production behaviour just as significantly as code changes.
A strong README should explain local setup, supported models, required environment variables, tool permissions, known limitations, test commands, and the process for adding an evaluation case. Pin dependency versions and add a licence before inviting outside contributors.
Choose the simplest suitable architecture
Most first agents need only four components:
1. Model layer: a hosted or self-hosted language model.
2. Orchestrator: a loop that supplies context, requests a response, executes approved tools, and repeats within a limit.
3. Tools: small functions for search, databases, APIs, or calculations.
4. State and observability: conversation state, traces, errors, latency, and token usage.
Use a plain tool-calling loop for a focused workflow. Add retrieval when the agent must answer from changing documents. Use a planner-executor pattern only when tasks genuinely require decomposition. Multi-agent or swarm designs introduce coordination failures, higher cost, and harder debugging; adopt them only after a single-agent baseline is measured. If your system needs independent workers, queues, retries, and durable state, compare the design with distributed systems built with AI agents.
Framework choice should follow requirements rather than popularity. Evaluate SDK quality, structured-output support, streaming, tracing, model portability, community health, and how easily you can inspect the underlying request. A lightweight Python service is often enough for a first production pilot.
Build tools as strict APIs
Tools are the agent’s authority. Define each one with a narrow purpose and a typed schema. Validate every argument on the server, even if the model produced it. Return concise, machine-readable results and predictable error types.
For example, a get_order_status(order_id) tool should verify the caller’s access to that order, limit the query, and return only required fields. A refund_order tool should never be available in the same unrestricted loop as a read-only search tool. Separate tools by risk:
- Read: search, retrieve, calculate, or inspect.
- Write: create, update, or send.
- High impact: refund, delete, approve, publish, or change permissions.
Require explicit human approval for high-impact actions. Apply timeouts, rate limits, idempotency keys, and retries with backoff. Treat external webpages, uploaded files, and retrieved documents as untrusted input: prompt injection can try to make an agent reveal secrets or misuse tools.
Add retrieval and memory carefully
Retrieval-augmented generation can ground answers in company documents, policies, or product data. Start with clean document ownership, metadata, access filters, chunking, and citations. Test whether the right passages are retrieved before changing prompts. Do not use “memory” as a vague database of everything a user has said. Define what is stored, why it is needed, how long it remains, and how a user can correct or delete it.
For voice workflows, the architecture also includes speech recognition, turn-taking, text-to-speech, interruption handling, and fallback to a human. Review the voice-agent architecture and deployment guide before adding a phone channel.
Test behaviour, not just code
Unit tests should cover tool validation, permission checks, parsing, retries, and failure handling. Agent evaluations should use a fixed dataset of realistic tasks, including ambiguous requests, missing information, adversarial instructions, and unsupported languages.
Track at least:
- Task completion and factual accuracy.
- Correct tool selection and argument validity.
- Unnecessary tool calls and escalation quality.
- Latency, token usage, and cost per task.
- Leakage of personal data or hidden instructions.
Run evaluations in pull requests when prompts, tools, models, or retrieval settings change. Use deterministic fixtures where possible, and manually review a sample of production traces. A passing test is not proof of safety; it is evidence that a defined case still works.
Secure the GitHub workflow
Never commit API keys, customer data, private certificates, or .env files. Use GitHub Actions secrets or an external secrets manager, apply least-privilege tokens, and restrict who can approve deployments. Protect the default branch, require reviews, scan dependencies, and enable secret and code scanning where available.
Create separate credentials for development, staging, and production. Log tool calls and decisions without logging raw sensitive content by default. For healthcare use cases, map data flows and retention rules before launch; a voice workflow involving clinical or patient information needs more than a generic chatbot disclaimer. You can use the patient follow-up with voice agents guide to think through operational boundaries.
Deploy incrementally
Begin with a local command-line prototype, then expose a versioned API behind authentication. Add staging, synthetic test data, dashboards, alerting, and a kill switch before production. Set maximum loop iterations, budgets, response timeouts, and fallback messages. If the agent cannot complete a task safely, it should state the limitation and route the case to a human or deterministic workflow.
Roll out to a small cohort, compare results with the existing process, and review failures weekly. Treat model, prompt, retrieval, and tool changes as releases with rollback plans. Keep a changelog that records what changed and which evaluation results justified the release.
A practical build sequence
1. Define one workflow and measurable success criteria.
2. Create the repository, README, licence, and development environment.
3. Implement a deterministic baseline before adding model autonomy.
4. Add one model and one or two read-only tools.
5. Validate structured outputs and enforce permissions server-side.
6. Build an evaluation set from real, anonymised examples.
7. Add retrieval, memory, or additional tools only when evidence requires them.
8. Add tracing, cost limits, security checks, and human escalation.
9. Deploy to staging, run a limited pilot, and document failures.
10. Review every production change through GitHub pull requests.
If you want to contribute rather than start from scratch, learn how to contribute to AI GitHub repositories in India. The fastest way to improve an agent is usually not adding more autonomy; it is making its scope, tools, tests, and failure paths clearer.