General-purpose coding assistants are useful, but they rarely understand the decisions that make a particular codebase work: internal APIs, deployment rules, service ownership, security controls, naming conventions, and years of accumulated engineering knowledge. Building custom AI developer productivity tools means turning that context into reliable software that fits the way your team already builds and ships.
For Indian startups and engineering organisations, the opportunity is especially practical. A focused internal tool can reduce onboarding time, improve review quality, make legacy systems easier to change, and help a small team operate at the scale of a much larger one. The strongest products are not chatbots bolted onto a repository. They are workflow systems with controlled access to code, documentation, developer tools, and CI/CD.
Start with a costly engineering bottleneck
Do not begin by choosing a model. Begin with a measurable problem. Interview developers, review pull-request data, and observe where work stalls. Common opportunities include:
- Finding the correct service, owner, or internal API
- Writing repetitive tests, migrations, and boilerplate
- Reviewing pull requests for security and policy violations
- Explaining legacy Java, .NET, or Python code
- Keeping documentation and runbooks aligned with production
- Diagnosing failed builds, deployments, and data pipelines
Choose one workflow where the input and expected output are clear. “Help developers code faster” is too broad for an MVP; “generate a first-pass integration test for a changed endpoint and run it in a sandbox” is specific enough to evaluate.
Teams exploring agent-based workflows can also study patterns from building distributed systems with AI agents, particularly around task boundaries, retries, observability, and failure handling.
A practical architecture
A production-grade tool usually has six layers:
1. Developer interface: VS Code or JetBrains extension, command-line interface, pull-request bot, Slack integration, or an internal web application.
2. Orchestration service: Authenticates users, selects tools, manages conversation state, applies policy, and records traces.
3. Context service: Retrieves relevant code, documentation, tickets, ownership information, schemas, and recent changes.
4. Model gateway: Routes requests to an approved hosted or self-hosted model, with controls for cost, latency, region, and fallback behaviour.
5. Execution sandbox: Runs tests, linters, static analysis, migrations, or previews without giving the model unrestricted production access.
6. Evaluation and telemetry: Measures correctness, acceptance, latency, cost, security events, and developer outcomes.
Keep these layers separate. It should be possible to change the model without rewriting repository indexing, permission checks, or the developer interface.
Build codebase context that can be trusted
Retrieval-augmented generation is usually the right starting point. Index source code, design documents, API specifications, runbooks, incident reports, and ownership metadata—but do not treat every file as equally authoritative.
Use syntax-aware parsing where possible. ASTs, symbols, imports, call relationships, and language-server information provide more useful units than arbitrary text chunks. Combine several retrieval methods:
- Symbol and exact search for class names, endpoints, error codes, and configuration keys
- Semantic search for conceptual questions and unfamiliar terminology
- Dependency and ownership graphs for impact analysis
- Recency and authority signals to favour maintained documentation and current code
Apply repository and user permissions before retrieval, not after the model has received the context. Exclude secrets, credentials, generated artefacts, and files that the user is not authorised to view. Redact sensitive values in logs and set retention periods for prompts, outputs, and tool traces.
A useful answer should cite file paths, symbols, commit references, or documentation pages. “The model said so” is not an acceptable provenance model for production engineering.
Choose prompting, RAG, or fine-tuning deliberately
These techniques solve different problems:
- Prompting controls format, instructions, and behaviour. Use it for response structure, coding standards, and tool-use rules.
- RAG supplies changing company knowledge. Use it for repositories, tickets, APIs, and runbooks.
- Fine-tuning teaches repeatable output patterns, terminology, or classification behaviour. Use it when you have a clean dataset and a stable task.
Fine-tuning will not reliably keep a model aware of a fast-changing codebase. Start with retrieval and structured prompts, then fine-tune only after you can show a persistent quality gap. The best practices for fine-tuning LLMs on custom data are particularly relevant when preparing accepted code changes, review comments, or internal support examples.
Give agents limited, testable tools
An agent should have the smallest permission set needed to complete its task. Useful tools include repository search, symbol lookup, documentation retrieval, test execution, linting, static analysis, issue creation, and pull-request drafting.
Avoid unrestricted shell access and direct production credentials. Run generated commands in an isolated, resource-limited environment. Require explicit approval for destructive actions, external communication, database changes, and merges. Every tool call should have a timeout, audit record, and clear error response.
A safe workflow for code changes is:
1. Inspect the relevant files and tests.
2. Propose a plan and identify assumptions.
3. Generate a patch in a temporary branch.
4. Run formatting, type checks, security scans, and targeted tests.
5. Present the diff, evidence, and unresolved risks.
6. Let a human approve the pull request.
This makes the assistant an accountable engineering system rather than an autocomplete feature with excessive privileges.
Design the MVP for adoption
Start with the surface where developers already work. A pull-request reviewer is easier to deploy than a new portal; a CLI can validate an idea before a full IDE extension. For an Indian startup, a sensible first release might answer repository questions with citations, generate tests for changed code, or explain failed CI jobs.
Set a narrow success criterion, such as reducing time spent triaging build failures by 30% or increasing meaningful test coverage for new endpoints. Invite a small group of engineers to use the tool on real work. Capture rejected suggestions and the reason for rejection—incorrect context, unsafe change, poor style, excessive verbosity, or simply the wrong workflow.
Open-source ecosystems can provide useful implementation patterns and contributors. For context, see Indian open-source AI developer projects and open-source AI projects for student developers, while keeping production access and proprietary data under your organisation’s controls.
Evaluate quality beyond lines of code
Generated code volume is a weak metric. Track outcomes across four dimensions:
- Engineering speed: lead time to merge, time to resolve CI failures, and onboarding time
- Quality: escaped defects, rollback rate, test effectiveness, and review rework
- Trust: acceptance rate, correction rate, citation accuracy, and unsafe-action blocks
- Economics: model cost per task, infrastructure cost, latency, and engineer time saved
Create a fixed evaluation set from real historical tasks. Include easy, ambiguous, and adversarial examples. Test retrieval separately from generation: a model cannot produce a correct answer if the relevant code or policy was never retrieved. Re-run evaluations after changing chunking, prompts, models, permissions, or tools.
Security, privacy, and Indian operating realities
Treat repository content as sensitive intellectual property. Use enterprise model agreements or controlled deployments with clear data-retention terms. Encrypt data in transit and at rest, integrate identity with your existing provider, and enforce repository-level access. For regulated sectors such as fintech and health-tech, document data flows, retention, auditability, and human approval points before rollout.
Plan for regional reliability and predictable costs. Cache safe, repeated retrieval results; route simple classification tasks to smaller models; reserve larger models for complex changes; and monitor token growth caused by oversized context windows. Latency matters: developers will abandon a tool that interrupts their flow, even if its answers are excellent.
A 90-day delivery plan
Days 1–30: select one workflow, define permissions, collect evaluation examples, and build read-only retrieval with citations.
Days 31–60: add one controlled tool, such as test execution or CI-log analysis; integrate with the IDE or pull-request workflow; measure baseline and pilot outcomes.
Days 61–90: add approval gates, audit logs, cost controls, regression evaluations, and a documented support process. Expand only if the first workflow shows measurable value.
The durable advantage is not access to a particular model. It is the combination of high-quality organisational context, safe tool execution, strong evaluation data, and a workflow developers choose to use. Indian teams that build those foundations can create internal leverage today—and potentially turn a proven engineering system into a defensible product tomorrow.