0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build custom knowledge graphs with ai assistant

How to Build Custom Knowledge Graphs with an AI Assistant

  1. aigi

    Large language models are useful at summarising and generating text, but they are not a dependable system of record. They can confuse similarly named companies, miss relationships spread across several documents, or present an outdated policy with unwarranted confidence. A custom knowledge graph gives an AI assistant a structured, queryable layer of entities, relationships, attributes, evidence, and permissions.

    The objective is not to replace search or a vector database. It is to make the assistant better at questions that require multi-step reasoning, consistent definitions, provenance, and controlled updates. For an Indian enterprise, that may mean connecting a customer to a GST registration, a lending product, a state-specific policy, and the documents that support each fact.

    When a knowledge graph is the right choice

    Use a graph when the value lies in connections rather than isolated passages. Strong candidates include:

    • Multi-hop questions: Which supplier serves plants that use a particular component?
    • Entity consistency: Are “Reliance Industries”, “RIL”, and a legal-entity identifier the same organisation?
    • Time-aware facts: Which director held a role during a specific financial year?
    • Explainability: Which source documents support an answer?
    • Policy and compliance: Which rule applies to a customer, product, location, or transaction?

    A graph is not automatically better than RAG. Vector retrieval remains effective for manuals, case notes, and long-form explanations. The practical architecture is usually hybrid: retrieve relevant text for nuance, then use graph traversal for entities and relationships. Teams building agentic applications should also review best practices for fine-tuning LLMs on custom data before deciding whether fine-tuning is needed at all.

    1. Define a narrow, useful ontology

    Start with one decision or workflow, not an attempt to model the entire business. Write down the questions the assistant must answer and work backwards to the minimum schema required.

    For a fintech onboarding assistant, the first version might contain:

    • Entities: Applicant, Company, PAN, GSTIN, Bank Account, Product, Document, State, Risk Rule.
    • Relationships: OWNS, REGISTERED_IN, SUBMITTED, ISSUED_BY, ELIGIBLE_FOR, TRIGGERS.
    • Attributes: legal name, identifier, effective date, status, confidence, source, and last verified time.

    Separate identity from descriptive text. A company name can change; a stable internal identifier should not. Define cardinality and constraints early: one GSTIN should map to the expected legal entity, while an applicant may submit many documents. Include temporal fields such as valid_from, valid_to, and observed_at; otherwise, the assistant may combine facts from different periods.

    Keep the ontology versioned. Store who changed a class or relationship, why it changed, and which records require reprocessing. This is especially important when rules vary by state, regulator, language, or financial year.

    2. Ingest documents with evidence attached

    A production graph should not contain unsupported assertions. Build an ingestion pipeline that preserves the original document, page or paragraph location, extraction model, timestamp, and confidence score.

    A practical flow is:

    1. Collect PDFs, spreadsheets, APIs, CRM records, emails, and approved web sources.
    2. Parse text while retaining page numbers, tables, headings, and document identifiers.
    3. Classify the document and route it to the relevant extraction schema.
    4. Extract candidate entities, relationships, dates, and citations.
    5. Validate candidates against ontology rules.
    6. Resolve entities and write only approved or reviewable facts to the graph.

    Use an LLM for candidate extraction, not unquestioned truth. Ask for strict JSON with a fixed schema, required evidence spans, and an explicit unknown value. Reject malformed outputs and flag low-confidence records for review. For Indian-language material, plan for code-switching, transliteration, and OCR errors. A graph that ignores local language variation will create duplicate entities and weak retrieval; the guide to low-resource Indic natural language processing is relevant when Hindi, Tamil, Marathi, or mixed-language documents are part of the corpus.

    3. Choose storage and query patterns

    Your database choice should follow graph size, query complexity, operational skills, and budget—not popularity alone. Neo4j, Amazon Neptune, ArangoDB, and FalkorDB can all support viable designs, while RDF stores may suit standards-heavy domains that need SPARQL.

    For an MVP, compare:

    • Managed service: less database operations work, but recurring cloud cost and potential vendor lock-in.
    • Self-hosted graph: more control and predictable infrastructure, but you own backups, upgrades, and failover.
    • Property graph: intuitive for application teams using node properties and Cypher-like queries.
    • RDF and ontologies: useful when interoperability, formal semantics, or linked-data standards are priorities.

    Create indexes for stable identifiers and frequently filtered properties. Do not return an unbounded neighbourhood around a node. Limit traversal depth, relationship types, result count, and execution time. For every answer, return a compact subgraph plus citations rather than dumping raw database output into the model.

    4. Build a guarded Graph-RAG loop

    A reliable assistant separates interpretation, retrieval, and generation:

    1. Interpret the question. Detect intent, entities, date range, geography, and required permissions.
    2. Resolve entities. Match names to canonical IDs, showing alternatives when ambiguity remains.
    3. Generate a constrained query. Permit only approved labels, relationships, properties, and read operations.
    4. Traverse the graph. Retrieve the smallest subgraph that answers the question.
    5. Fetch supporting text. Use vector or keyword search for definitions, exceptions, and source passages.
    6. Generate with citations. Instruct the model to answer only from retrieved evidence and state when evidence is insufficient.
    7. Log the trace. Save the question, query, retrieved facts, model version, answer, and user feedback.

    Never allow an LLM to execute arbitrary Cypher, Gremlin, or SPARQL with production write access. Use a read-only service account, query templates or an allow-list validator, row-level filters, and timeouts. Treat prompt-injected documents as untrusted input. Sensitive fields such as PAN, Aadhaar-related information, health records, and financial data require minimisation, encryption, access logging, and retention controls.

    5. Solve entity resolution before scaling

    Entity resolution is often harder than extraction. “ABC Pvt Ltd”, “ABC Private Limited”, and a supplier code may refer to one organisation—or three. Combine deterministic and probabilistic methods:

    • Match trusted identifiers first, such as GSTIN or an internal customer ID.
    • Normalise case, punctuation, whitespace, legal suffixes, and common transliterations.
    • Compare names, addresses, phone numbers, directors, and source systems.
    • Use embeddings only to generate candidates, never as the final decision.
    • Route uncertain matches to human review and retain the decision as labelled data.

    Maintain aliases and rejected matches. A merge should be reversible, with an audit trail showing which facts were affected.

    6. Evaluate factuality, retrieval, and business impact

    A graph project needs more than a demo that produces fluent answers. Create a test set from real user questions, including ambiguous names, missing facts, conflicting documents, and date-sensitive queries. Measure:

    • Entity resolution accuracy and false merges.
    • Extraction precision and recall by relationship type.
    • Graph query success rate and latency.
    • Citation completeness and evidence correctness.
    • Answer groundedness, abstention quality, and permission violations.
    • Workflow outcomes, such as reduced review time or fewer support escalations.

    Test Hindi and other relevant languages separately. Review failures by category: parser errors, ontology gaps, stale data, wrong entity resolution, unsafe query generation, or hallucinated synthesis. Only expand the schema after the first workflow is stable.

    7. A practical 30-day implementation plan

    Week 1: Select one workflow, collect representative documents, define ten to twenty core entities, and write a labelled evaluation set.

    Week 2: Build ingestion, evidence storage, extraction validation, and a small graph. Add identifiers and an approval queue.

    Week 3: Implement entity resolution, read-only query generation, hybrid retrieval, citations, and access controls.

    Week 4: Run accuracy and latency tests, conduct security review, instrument costs, and pilot with a small group of domain users.

    Keep humans in the loop for high-impact decisions. The assistant can assemble evidence and explain a recommendation, but loan approval, legal advice, medical action, or regulatory filing should follow the organisation’s authorised review process. If your architecture uses several specialised agents for ingestion, validation, and research, apply the same permission and observability principles described in building distributed systems with AI agents.

    Common mistakes to avoid

    • Building a broad ontology before validating one user workflow.
    • Storing facts without source, timestamp, confidence, or provenance.
    • Treating generated triples as ground truth.
    • Using vector similarity as final entity identity.
    • Returning whole neighbourhoods instead of bounded evidence.
    • Giving the model write access to the production graph.
    • Ignoring deletion, correction, retention, and schema migration requirements.
    • Measuring only answer fluency instead of factual and operational outcomes.

    FAQ

    Do I need a large dataset? No. A small, clean graph for one department is more valuable than a large graph full of duplicate and unsupported facts.

    Should I replace my vector database? Usually not. Use the graph for structure and multi-hop reasoning, and vector or keyword retrieval for relevant passages and semantic similarity.

    Which language should I use? Python is a practical choice because it has mature tooling for document processing, model APIs, graph clients, evaluation, and orchestration. Choose the database query language your team can operate safely.

    Can an AI assistant update the graph? It can propose updates, but production writes should pass schema validation, confidence thresholds, deduplication, and—where risk warrants it—human approval.

    Build with a measurable first use case

    Custom knowledge graphs are infrastructure for trustworthy AI, not a shortcut around data quality. Start with a bounded question, preserve evidence, constrain every query, and measure whether the assistant improves a real workflow. Indian founders building such systems can explore AI Grants India for funding and support as they move from prototype to production.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.