Large language models become substantially more useful when they can plan multi-step work, retain the right context, and interact with external tools. These three capabilities—LLM planning, memory, and tool use—turn a text generator into an agent that can research, reason over data, call APIs, execute workflows, and recover from errors.
A robust agent is not created by simply adding more prompts. It needs an explicit control loop, carefully scoped memory, typed tool interfaces, safeguards, and evaluation against real tasks. This guide explains the core concepts, practical architectures, implementation patterns, and production considerations for building LLM agents.
What Are LLM Planning, Memory, and Tool Use?
The three capabilities solve different problems:
- Planning: Decides what steps are needed to achieve a goal.
- Memory: Stores and retrieves information useful across turns or tasks.
- Tool use: Lets the model perform actions or access information outside its context window.
For example, an AI assistant asked to analyse an Indian startup’s grant eligibility might:
1. Clarify the startup’s sector, incorporation status, and stage.
2. Retrieve relevant government schemes from a knowledge base.
3. Use a calculator or spreadsheet tool to check financial thresholds.
4. Remember the founder’s preferences and previously supplied documents.
5. Produce a cited recommendation and identify missing evidence.
Planning determines the sequence, memory supplies context, and tools connect the model to current or authoritative systems.
Why These Capabilities Must Be Designed Together
Planning without memory repeatedly asks for the same information. Memory without planning creates a large but poorly used context. Tool use without either can lead to random calls, incorrect parameters, or unsafe actions.
A useful agent architecture treats the model as a decision component inside a controlled runtime:
User request
↓
Task interpretation and policy checks
↓
Planner / model decision
↓
Memory retrieval or tool selection
↓
Tool execution and observation
↓
State update and verification
↓
Final answer or next planned stepThe model should not be given unrestricted authority. The runtime should validate tool arguments, enforce permissions, record events, limit budgets, and decide whether a result requires human approval.
LLM Planning Patterns
ReAct: Reasoning and Acting in a Loop
The ReAct pattern alternates between an internal decision and an external action. The agent identifies what it needs, calls a tool, observes the result, and continues.
Goal → decide → call tool → observe → update plan → repeatIt works well for research, troubleshooting, and tasks where the next step depends on the previous result. However, unrestricted loops can become expensive or repetitive. Add a maximum step count, timeout, duplicate-call detection, and a completion criterion.
Plan-and-Execute
In plan-and-execute systems, one component creates a sequence of steps and another executes them. This is useful for long workflows such as preparing a market report or processing a set of documents.
The plan should be represented as structured data rather than prose alone:
{
"goal": "Assess grant eligibility",
"steps": [
{"id": "s1", "action": "collect_company_facts", "status": "pending"},
{"id": "s2", "action": "retrieve_matching_schemes", "status": "pending"},
{"id": "s3", "action": "check_requirements", "status": "pending"}
],
"constraints": {"max_tool_calls": 12, "requires_citations": true}
}Structured plans are easier to inspect, resume, evaluate, and modify when a step fails.
Hierarchical Planning
Complex objectives can be decomposed into tasks and subtasks. A top-level planner defines outcomes, while specialist workers handle retrieval, extraction, calculation, or drafting.
For example:
- Research agent: finds authoritative sources.
- Extraction agent: converts documents into structured facts.
- Verification agent: checks contradictions and dates.
- Writing agent: produces the final response with citations.
Hierarchical planning improves modularity but introduces coordination overhead. Use it when tasks have clear boundaries; avoid adding multiple agents merely because the workflow sounds complex.
Reflection and Self-Critique
A verification step can inspect whether the output satisfies requirements, cites evidence, and avoids unsupported claims. Reflection is most useful when it is tied to concrete checks, such as schema validation or source comparison.
A vague instruction like “think again” is weaker than a checklist:
- Are all mandatory fields present?
- Does every factual claim have a source?
- Are dates and eligibility conditions current?
- Did the tool output actually support the conclusion?
Memory in LLM Agents
Memory is not one database. It is a set of storage and retrieval mechanisms designed for different time horizons.
Working Memory
Working memory is the current context: the user’s request, recent messages, active plan, tool results, and constraints. It should be compact and relevant. Passing the entire conversation into every call increases cost and can distract the model.
Use summarisation or state compaction to preserve:
- Current objective
- Decisions already made
- Unresolved questions
- Important entities and identifiers
- Evidence and source references
- Tool errors and retry status
Episodic Memory
Episodic memory records events from previous interactions, such as a user’s past request or a completed workflow. It is useful when an agent must resume a task or personalise future interactions.
Store event metadata, not just text:
{
"event_type": "document_review_completed",
"user_id": "redacted-id",
"timestamp": "2026-01-15T10:30:00Z",
"outcome": "missing incorporation certificate",
"confidence": 0.91,
"source_refs": ["doc_184"]
}Semantic Memory
Semantic memory stores durable facts, concepts, and documents. Vector search can help retrieve semantically similar content, but embeddings alone do not guarantee correctness. Combine vector retrieval with metadata filters, keyword search, access controls, and reranking.
For India-focused applications, metadata may include:
- State or union territory
- Government department
- Scheme category
- Publication and expiry dates
- Startup stage
- Industry or technology area
- Source authority
Procedural Memory
Procedural memory represents how a task should be performed: policies, workflows, templates, and tool instructions. It should generally be version-controlled and reviewed like software rather than learned casually from conversation.
Retrieval and Memory Design
A practical retrieval pipeline usually contains these stages:
1. Query construction: Rewrite the user’s request into searchable concepts.
2. Candidate retrieval: Use vector, keyword, or hybrid search.
3. Filtering: Apply permissions, geography, date, and document type.
4. Reranking: Prioritise results by relevance and authority.
5. Context assembly: Include concise passages with source identifiers.
6. Grounded generation: Require the model to cite or quote evidence.
Memory writes need equal care. Do not automatically save every user statement. Classify candidate memories by sensitivity, usefulness, confidence, and retention period. Personal data should have a clear purpose, access policy, deletion process, and audit trail.
Tool Use: Turning Model Decisions into Actions
A tool is an external capability exposed through a controlled interface. Common examples include:
- Search and retrieval
- SQL queries
- Calculators and code execution
- CRM or ticketing systems
- Email and calendar APIs
- Document parsers
- Payment or workflow systems
- Internal business APIs
Design Tools with Typed Schemas
Tool definitions should specify required fields, data types, allowed values, and side effects. A model should not have to infer an API contract from a paragraph.
{
"name": "search_schemes",
"description": "Find currently active public funding schemes",
"parameters": {
"type": "object",
"properties": {
"sector": {"type": "string"},
"state": {"type": "string"},
"as_of_date": {"type": "string", "format": "date"}
},
"required": ["sector", "as_of_date"]
}
}The execution layer must validate the model’s output independently. Never rely on the model to enforce authorization, SQL safety, transaction limits, or irreversible-action rules.
Separate Read and Write Tools
Read-only tools can often run automatically. Write tools—sending messages, modifying records, approving applications, or initiating payments—should use stricter permissions and frequently require confirmation.
A useful risk classification is:
- Low risk: Search public documentation or calculate a value.
- Moderate risk: Create a draft, update a non-critical record, or run a bounded query.
- High risk: Send external communication, change financial data, delete records, or submit a legal application.
Return Useful Observations
Tool outputs should be concise, structured, and explicit about errors. Include status, relevant data, source or transaction identifiers, and whether the result is complete. Avoid returning huge raw payloads to the model; summarise them while preserving a path to the original record.
A Reference Agent Control Loop
A production control loop can be implemented as follows:
state = load_task_state(task_id)
for step in range(MAX_STEPS):
context = assemble_context(
goal=state.goal,
plan=state.plan,
recent_events=state.events[-8:],
memories=retrieve_relevant_memory(state.goal),
)
decision = model.choose_action(context, tools=allowed_tools(state.user))
validate_decision(decision, state.policy)
if decision.type == "final":
verify_output(decision.content, state.requirements)
return decision.content
result = execute_tool_safely(decision.tool, decision.arguments)
state.events.append(record_result(decision, result))
state.plan = update_plan(state.plan, result)
raise AgentLimitError("Task exceeded the allowed number of steps")Important production controls include idempotency keys, retries with backoff, circuit breakers, timeouts, rate limits, and durable state. If a process crashes, the agent should resume from a recorded checkpoint rather than repeat an external action.
Common Failure Modes
Hallucinated Tool Results
The model may claim that a tool succeeded when it did not run. The runtime should treat only verified execution results as observations and attach unique call IDs to every tool event.
Context Overload
More context is not always better. Retrieve fewer, higher-quality passages, summarise old events, and separate instructions from untrusted retrieved content.
Infinite or Repetitive Loops
Detect repeated queries, identical arguments, unchanged state, and excessive retries. Establish budgets for tokens, time, cost, and tool calls.
Stale Memory
Facts such as scheme rules, prices, policies, and API responses expire. Store timestamps, validity periods, and source URLs. Prefer fresh retrieval for volatile information.
Prompt Injection
External documents and web pages may contain instructions aimed at the agent. Treat retrieved content as data, not authority. Use content isolation, allowlisted tools, permission checks, and explicit instruction hierarchy.
Over-Automation
An agent should not autonomously take high-impact actions merely because it can. Use human approval for financial, legal, employment, medical, or irreversible decisions.
Evaluating LLM Planning and Tool Use
Evaluation should measure the complete workflow, not just the final wording. Useful metrics include:
- Task success rate: Did the agent achieve the intended outcome?
- Plan validity: Were steps relevant, complete, and correctly ordered?
- Tool accuracy: Were the correct tools and arguments selected?
- Groundedness: Are conclusions supported by retrieved evidence?
- Recovery rate: Can the agent handle tool errors and missing data?
- Efficiency: How many tokens, seconds, and tool calls were required?
- Safety: Did it respect permissions and confirmation policies?
- Consistency: Does it behave reliably across equivalent inputs?
Build a test set containing normal cases, ambiguous prompts, stale documents, API failures, prompt injection attempts, and adversarial parameters. Log traces with privacy protections so engineers can inspect the exact plan, memory retrieval, tool calls, observations, and final answer.
India-Specific Considerations
Indian AI products often operate across multiple languages, fragmented data sources, variable connectivity, and regulated sectors. Design for these realities from the start.
- Support multilingual queries while preserving names, addresses, identifiers, and legal terms accurately.
- Treat Aadhaar, PAN, health information, financial records, and other personal data as sensitive.
- Apply purpose limitation, access control, retention rules, and deletion workflows consistent with applicable Indian data-protection obligations and sector regulations.
- Verify government scheme information against official department sources and publication dates.
- Plan for GST, INR, Indian date formats, state-specific eligibility, and local business terminology.
- Use regional hosting, encryption, audit logs, and vendor assessments where contractual or regulatory requirements demand them.
- Keep a human review path for lending, insurance, healthcare, recruitment, and public-sector decisions.
For startups building in India, a narrow agent that reliably solves one operational problem is usually more valuable than a general-purpose autonomous system with unclear accountability.
A Practical Build Roadmap
Phase 1: Define the Task
Choose a workflow with measurable success criteria. Document inputs, expected outputs, tools, risks, and escalation conditions.
Phase 2: Build a Deterministic Baseline
Implement retrieval, validation, and business rules without an agent where possible. This establishes a reliable foundation and clarifies where model flexibility is actually needed.
Phase 3: Add Structured Tool Use
Expose a small set of typed tools. Validate arguments, separate read and write operations, and record every execution.
Phase 4: Add Planning and State
Introduce a plan representation, checkpoints, step limits, retry logic, and resumable task state. Test failures before expanding scope.
Phase 5: Add Memory Carefully
Start with working memory and task state. Add long-term memory only when a clear user benefit exists, with consent, retention, correction, and deletion mechanisms.
Phase 6: Evaluate and Operate
Create offline tests, monitor production traces, measure cost and latency, and review safety incidents. Improve prompts, tools, retrieval, and policies separately so changes are attributable.
FAQ: LLM Planning, Memory and Tool Use
What is the difference between planning and chain-of-thought?
Planning is an observable task structure—goals, steps, dependencies, and status. Chain-of-thought refers to hidden or internal reasoning. Production systems should log concise rationales, structured plans, and tool traces rather than depend on exposing private reasoning.
Does every LLM agent need a vector database?
No. Simple agents may need only conversation state, SQL, keyword search, or a curated document store. Use vector retrieval when semantic matching solves a demonstrated requirement, and combine it with metadata and authority filters.
How can tool use be made safe?
Use typed schemas, allowlisted tools, independent argument validation, least-privilege credentials, timeouts, audit logs, rate limits, and human confirmation for high-impact actions.
How much memory should an AI agent retain?
Only what is necessary for a defined purpose. Retain task state and useful preferences when justified, but avoid storing sensitive data by default. Add timestamps, confidence, provenance, retention limits, and user controls.
What is the best first use case for an Indian startup?
Begin with a bounded workflow such as support-ticket classification, document extraction, internal knowledge search, compliance checks, or grant discovery. These use cases provide measurable value without requiring unrestricted autonomy.
Apply for AI Grants India
Building an AI product around planning, memory, or tool use? Indian AI founders can explore support, funding opportunities, and ecosystem guidance through AI Grants India. Apply today and take the next step toward turning your agent idea into a reliable, deployable product.