Code smells are not bugs by themselves. They are recurring signals—such as excessive complexity, duplicated logic, or misplaced responsibilities—that make software harder to change and easier to break. Traditional linters and static-analysis tools catch many of these signals reliably, but they can miss intent, relationships across files, and project-specific design problems.
Automated code smell detection using large language models adds a semantic layer to that workflow. An LLM can examine code, surrounding documentation, test names, commit context, and architectural conventions to explain why a pattern may be risky. The strongest implementations do not replace deterministic analysis; they combine it with model-assisted triage and recommendations.
For Indian startups and engineering teams, this distinction matters. A useful system must work across fast-moving repositories, mixed-language stacks, legacy services, and strict requirements around customer data and source-code confidentiality.
What counts as a code smell?
A code smell is an observable pattern that suggests future maintenance, reliability, or design trouble. It is a prompt for investigation, not an automatic verdict. Common examples include:
- Long methods and large classes: Responsibilities become difficult to understand and test.
- Duplicated code: Fixes must be repeated, creating inconsistent behaviour over time.
- Deep nesting and high cyclomatic complexity: Small changes require reasoning through too many paths.
- Feature envy: A function depends heavily on another module’s data or methods.
- Shotgun surgery: One conceptual change requires edits across many unrelated files.
- Dead code and stale abstractions: Unused paths increase cognitive load and deployment risk.
- Inconsistent error handling: Similar failures are logged, retried, or surfaced differently.
Severity depends on context. A long method in generated code may be harmless; the same pattern in a payment, health, or identity service may deserve immediate review.
Where LLMs add value
Conventional tools are excellent at measurable rules: line length, duplication thresholds, unreachable code, type errors, and known vulnerability patterns. LLMs are more useful when interpretation is required.
They can:
- Explain a finding in plain language for a mixed-experience team.
- Compare a suspicious function with local repository conventions.
- Identify likely responsibility leakage across classes or services.
- Suggest a small refactoring plan rather than a wholesale rewrite.
- Group repeated findings into one underlying architectural problem.
- Use issue descriptions, tests, and recent diffs to estimate practical impact.
This is similar to how teams use AI for automated user feedback categorization in Indian SaaS: the model creates useful structure, but a human-defined taxonomy and review process determine whether the output is trustworthy.
A robust detection pipeline
1. Collect the right context
Send the model only what it needs: the changed files, relevant imports, interfaces, tests, configuration, and selected repository guidance. A full repository dump is expensive, slow, and more likely to expose sensitive material.
Include metadata such as:
- Programming language and framework
- Service ownership and criticality
- Whether the code is generated or hand-written
- Recent diff and linked issue
- Existing lint and test results
- Performance, security, or compliance constraints
2. Run deterministic checks first
Use a conventional toolchain for syntax, types, security rules, dependency issues, duplication, and complexity. These findings provide anchors for the LLM and prevent it from spending tokens rediscovering basic facts.
3. Ask for structured judgements
Require machine-readable output with fields such as:
- Smell category
- File and line range
- Evidence from the code
- Confidence score
- Likely impact
- Suggested next action
- Whether the finding is blocking, advisory, or informational
Prompts should instruct the model to say “insufficient evidence” rather than inventing a problem. Ask for the smallest safe refactoring and require preservation of public interfaces unless explicitly approved.
4. Validate before creating work
A finding should pass checks before becoming a ticket or pull-request comment. Re-run tests, compare against repository conventions, and let developers mark findings as valid, invalid, or not actionable. That feedback can improve prompts and thresholds without blindly fine-tuning on noisy labels.
Teams building broader AI workflows can apply the same evaluation discipline used in low-resource Indic natural language processing: define representative data, measure performance by category, and test failure modes instead of relying on impressive examples.
Measuring quality
Accuracy alone is not enough. Track:
- Precision: How many reported smells are confirmed by reviewers?
- Recall: How many known, high-value smells does the system find?
- Actionability: How often does a finding lead to a useful change?
- Developer acceptance: Do engineers keep the integration enabled?
- Review time: Does the system reduce, rather than increase, review effort?
- Regression rate: Do suggested refactors introduce defects?
Maintain a labelled evaluation set from your own codebase. Separate generated code, legacy code, test code, and production-critical paths. Review results by language and repository; a model that performs well on Python services may behave differently on Java, Go, or SQL-heavy systems.
Security and privacy controls
Source code can contain credentials, personal data, proprietary algorithms, and regulated information. Before sending code to an external model, establish clear controls:
- Redact secrets and customer data before inference.
- Prefer private endpoints, self-hosted models, or approved enterprise providers for sensitive repositories.
- Define retention, training-use, and deletion terms contractually.
- Log prompts, model versions, findings, and reviewer decisions without storing unnecessary source.
- Restrict access by repository and team.
- Prevent model-generated patches from merging without normal tests and approval.
For Indian organisations, also map the workflow to internal security policies and applicable obligations under the Digital Personal Data Protection framework when personal data appears in code, logs, or test fixtures.
Deployment pattern for engineering teams
Start with pull-request analysis on changed lines. This limits cost and noise while producing feedback where developers already work. After two to four weeks of review data, expand to scheduled repository scans for architectural smells and technical-debt trends.
A practical rollout looks like this:
1. Establish baseline findings from static analysis.
2. Select two or three smell categories with clear remediation patterns.
3. Run the LLM in advisory mode for a pilot group.
4. Review precision and false-positive rates weekly.
5. Add blocking thresholds only for high-confidence, high-impact rules.
6. Publish ownership and remediation guidance for recurring findings.
Teams serving multilingual users should not assume that English-only tooling is sufficient for documentation or support workflows. Lessons from open-source small language models for Hindi can inform local deployment choices, especially when engineering instructions, comments, and issue reports use Indian languages.
Common failure modes
Avoid treating an LLM as an autonomous code-quality authority. Common mistakes include:
- Using vague prompts: “Find bad code” produces inconsistent results.
- Ignoring repository context: Generic best practices may conflict with local architecture.
- Posting every finding: Alert fatigue quickly destroys adoption.
- Accepting generated patches automatically: A plausible explanation is not proof of correctness.
- Evaluating only on examples chosen by the vendor: Internal, adversarial, and legacy cases matter more.
- Blocking on subjective smells: Reserve gates for evidence-backed issues.
The practical outlook
As of 2026, the best use of LLMs in code-quality workflows is not replacing linters, reviewers, or tests. It is reducing the cost of interpretation: explaining findings, connecting symptoms across files, prioritising technical debt, and proposing bounded next steps.
The winning architecture is hybrid: deterministic rules for repeatability, LLMs for semantic reasoning, repository-specific evaluation for trust, and engineers accountable for final decisions. Start narrowly, protect source code, measure accepted findings, and expand only where the system demonstrably improves maintainability.