GraphRAG is retrieval-augmented generation that retrieves from a knowledge graph, alone or alongside a vector index. Instead of handing the LLM the few text chunks that look most similar to the question, GraphRAG gives it entities and the explicit relationships between them, so it can answer questions that need several hops, span many documents or ask about connections. It costs more to build and maintain than vector RAG, so the real skill is knowing which questions justify it.
If you are new to retrieval-augmented generation itself, read what RAG is and how the pipeline works first.
Knowledge graph basics: entities, relationships, properties
A knowledge graph stores facts as a network rather than as rows or paragraphs. It has three building blocks:
- Entities (nodes). The things your business talks about: a supplier, a contract, a plant, a part, an incident, a person, a regulation. Each node usually has a type, or label, such as
SupplierorContract. - Relationships (edges). Typed, directed connections between entities:
(Supplier)-[:PARTY_TO]->(Contract),(Incident)-[:INVOLVES]->(Supplier),(Supplier)-[:SUBSIDIARY_OF]->(Supplier). - Properties. Key-value attributes on nodes and edges: a contract's start and end dates, an incident's severity, a supplier's country, or the source document and page an edge was extracted from.
Most GraphRAG tooling uses property graphs, where nodes and edges both carry properties. The alternative, RDF, stores subject-predicate-object triples and suits formal ontologies and data integration.
Graph databases
Production systems store the graph in a graph database. Neo4j is a widely used property graph database and the origin of the Cypher query language. Amazon Neptune is AWS's managed graph service and supports property graph queries (Gremlin and openCypher) as well as RDF with SPARQL. Other options include Memgraph, TigerGraph, ArangoDB, Azure Cosmos DB's Gremlin API, and the Apache AGE extension, which adds openCypher queries to PostgreSQL. GQL, an ISO standard graph query language published in 2024, draws heavily on Cypher.
What vector RAG struggles with
Vector RAG retrieves the top-k chunks most similar to the question. That works when the answer sits in one or two passages, and breaks down in three predictable ways.
Multi-hop questions
"Which of our critical suppliers share a parent company with a vendor we put on a watch list last year?" The answer needs a chain: supplier, parent company, other subsidiaries, watch-list entries. Each link lives in a different document and none looks much like the question, so similarity search finds one hop and misses the rest.
Aggregation across documents
"What are the main themes in two years of customer complaints?" or "How many incidents involved suppliers in one region?" Top-k retrieval returns a handful of chunks, never the whole corpus, so the model summarises a sample and presents it as the whole picture.
Relationship questions
"How is this vendor connected to that outage?" Vector search finds text about the vendor and text about the outage; it does not know whether they are linked or how.
Before reaching for a graph, check that a simpler fix is not enough. Many "vector RAG failures" are really chunking or keyword-matching problems that hybrid search and reranking solve at much lower cost.
The main GraphRAG approaches
"GraphRAG" is used for several different designs. Most production systems combine two or more of the following.
1. Building a graph from documents with LLM extraction
The first job is turning unstructured text (contracts, incident reports, audit findings) into nodes and edges:
- Chunk documents into passages, keeping document ID, section and date (the same discipline as in RAG chunking strategies).
- Extract. Prompt an LLM, ideally with structured output, to return entities (name, type, description) and relationships (source, target, type, description) found in each chunk. Constrain it with a schema: the entity and relation types you care about.
- Resolve entities. Merge "Acme Components Pvt Ltd", "ACME Components" and "Acme (Pune)" into one node. This is the hardest step: combine normalisation rules, embedding similarity, master data (vendor IDs from the ERP) and human review.
- Load nodes and edges into the graph store, with provenance on every edge: which chunk, which document, which extraction run.
Where structured sources already exist (vendor master, CMDB, HR systems), load them directly as the graph's backbone and use LLM extraction only to attach facts from text.
2. Community summaries for global questions (Microsoft's GraphRAG)
Microsoft Research described an approach, published with an open-source implementation also called GraphRAG, aimed at "global" questions about a whole corpus, such as "what are the main themes in this dataset?". In broad terms its indexing pipeline works like this:
- An LLM extracts entities, relationships and short descriptions from every text chunk, building a graph of the corpus.
- A community detection algorithm (the Leiden algorithm) partitions the graph into clusters of closely connected entities, in a hierarchy from broad communities down to small ones.
- An LLM writes a summary report for each community, describing its key entities, relationships and themes.
At query time, global search works map-reduce style: the question is put to many community summaries in parallel, each produces a partial answer with a relevance score, and the most useful partial answers are combined into a final response. Local search handles entity-specific questions by starting at the entities that match the question and pulling in their neighbours, relationships, related source text and relevant community summaries.
The trade-off is cost: indexing calls an LLM over every chunk and again for every community, which is expensive for large or fast-changing corpora. Microsoft has since described lower-cost variants that defer much of the summarisation to query time.
3. Graph traversal combined with vector retrieval
This is the most common production pattern, sometimes called hybrid or graph-enhanced retrieval:
- Use vector (or hybrid) search to find entry points: the chunks and entities most relevant to the question.
- Map those chunks to their graph nodes and traverse one or two hops along relevant relationship types.
- Collect the connected entities, the relationship facts and the source chunks those edges came from.
- Give the LLM a context made of both text passages and a compact list of facts, each with a citation.
The vector side supplies recall on fuzzy language; the graph supplies the links similarity search cannot see. Limit the traversal by relationship type, depth and node count.
4. Text-to-Cypher querying
When the question is precise ("list active contracts with suppliers involved in severity-1 incidents this year"), the best retrieval is an exact query. An LLM is given the graph schema (node labels, relationship types, property names, a few example queries) and generates Cypher, which runs against the database; the results go back to the LLM to phrase the answer. It is the graph equivalent of text-to-SQL, and the same engineering discipline applies, as in our text-to-SQL agent project:
- Run generated queries with a read-only database role, and reject write clauses such as
CREATE,MERGE,SETandDELETEbefore execution. - Validate the query against the schema, enforce a
LIMITand a timeout, and cap variable-length path patterns. - Apply access control in the query layer, not in the prompt.
- Keep a library of tested query templates for frequent questions and let the LLM fill parameters; free-form generation is the fallback.
Text-to-Cypher only works if the graph's schema is clean and stable.
A GraphRAG reference architecture
INDEXING (batch + incremental)
docs -> chunk -> LLM extract -> entity resolve
| |
v v
vector index <--- chunk IDs ---> graph DB
(+ ERP/CMDB
master data)
QUERY
question -> router
|-- fact lookup -> text-to-Cypher -> graph DB
|-- open question -> vector search -> entry
| nodes -> 1-2 hop traverse
|-- corpus theme -> community summaries
v
context (facts + chunks + citations)
-> LLM -> answer -> user
The router can be a small classifier or an LLM call that picks a strategy. Log the chosen route with every request. The vector index and the graph should share chunk IDs, which is what makes citations work across both. The choice of vector database matters less than keeping those IDs consistent.
Want to build retrieval systems like this hands-on, from chunking and embeddings to agents that pick the right tool? Cloudsoft's AI, GenAI and Agentic AI course covers RAG pipelines, evaluation and agent design with guided labs.
Illustrative example: a supplier-risk question
Consider a manufacturer whose procurement team in a Hyderabad GCC supports plants in several countries. A risk analyst asks: "Which suppliers on contracts renewing in the next six months have been involved in quality or safety incidents, and do any of them share a parent company?"
With vector RAG alone, the assistant retrieves a few contract chunks that mention renewal and a few incident reports that mention suppliers. It may miss suppliers whose incidents used a trading name, and say nothing reliable about parent companies, because ownership sits in a due-diligence annex.
With GraphRAG, the indexing pipeline has already built:
Suppliernodes anchored to vendor IDs from the ERP, with trading names resolved onto them;Contractnodes with renewal dates, linked byPARTY_TOedges extracted from contract text;Incidentnodes with type and severity, linked byINVOLVESedges extracted from incident reports;SUBSIDIARY_OFedges from the due-diligence annexes.
The router sends the question down the text-to-Cypher path. The generated query, checked against the schema and run read-only, finds contracts renewing in the window, follows PARTY_TO to suppliers, follows INVOLVES back from quality and safety incidents, and groups the matching suppliers by parent. The result is a small table of facts, each edge carrying its source document. The LLM writes the answer and attaches the incident passages through shared chunk IDs.
Two caveats make it honest. First, the answer is only as complete as extraction and entity resolution: if an incident report names a supplier the resolver failed to match, that incident is invisible. Second, what counts as "involved in an incident" is a business definition the risk team must agree, not a model setting.
Costs, complexity and keeping the graph fresh
GraphRAG moves cost from query time to build time and adds components to own.
- Extraction cost. Every chunk passes through an LLM at least once, often with long structured prompts, and community summarisation adds more calls. Pilot on a representative slice and measure tokens per document before indexing everything.
- Schema design. Somebody must decide the entity and relation types with domain experts.
- Entity resolution and quality review. Plan for human review queues for low-confidence merges and spot checks of extracted edges.
- A second datastore to back up, secure and monitor.
Freshness is where many graph projects quietly decay. Design incremental updates from day one:
- Track a content hash per document; on change, re-extract only the affected chunks.
- Because every edge records its source chunk, you can delete edges from a removed or superseded document precisely, instead of rebuilding.
- Recompute community summaries on a schedule, or only for communities whose membership changed, since they are the most expensive artefact.
- Put effective dates on facts (a contract that ended, an ownership change) so the graph can answer "as of" questions rather than mixing past and present.
The patterns in data pipelines for RAG (change detection, idempotent loads, lineage) apply directly.
When GraphRAG is worth it, and when it is not
| Situation | GraphRAG worth it? | Why |
|---|---|---|
| FAQ, policy or runbook Q&A where answers sit in one passage | No | Hybrid search plus reranking is cheaper and usually enough |
| Questions that chain entities across documents (supplier, owner, incident, contract) | Yes | Explicit edges make multi-hop retrieval reliable |
| "Main themes" or summary questions over a large corpus | Often | Community summaries address what top-k retrieval cannot |
| Facts already in a relational database | Usually not | Text-to-SQL over existing tables avoids building a graph |
| Domains with a stable vocabulary of entities (suppliers, assets, drugs, accounts) | Yes | A clear schema keeps extraction and querying accurate |
| Small corpus, or a quick proof of concept | No | Build and maintenance cost outweighs the gain |
| Rapidly changing content with no incremental pipeline | Not yet | The graph goes stale; build freshness first |
| Regulated answers that must show how a conclusion was reached | Often | Edges with provenance give an auditable reasoning path |
How to evaluate GraphRAG
Evaluate the graph and the answers separately, so you can tell whether a wrong answer came from extraction, retrieval or generation. The general metrics are covered in our RAG evaluation metrics guide; GraphRAG adds a few layers.
- Extraction quality. Hand-label entities and relations in a sample of chunks and measure precision and recall of the extractor. Check entity resolution separately.
- Retrieval quality by question type. Build a test set with labelled categories (single-hop, multi-hop, aggregation, relationship, global theme) and expected supporting facts. Measure whether the retrieved context contains them, per category and per route.
- Query correctness. For text-to-Cypher, compare execution results with expected results rather than comparing query strings; two different queries can both be right.
- Answer quality. Faithfulness to retrieved context, completeness on list questions (did it find all the matching suppliers?) and citation accuracy. Tools such as Ragas help with faithfulness; completeness needs checks against a known answer set.
- A baseline. Run the same test set through your best vector or hybrid pipeline. If GraphRAG only wins on question types users rarely ask, the extra cost is not justified.
For global, summary-style questions there is often no single correct answer. Use side-by-side comparisons by domain experts, or an LLM judge with a clear rubric calibrated against human ratings.
Common GraphRAG mistakes
- Building a graph before proving vector RAG fails. Start with hybrid retrieval and an evaluation set.
- Extracting without a schema. The result is a hairball of vague relation types that nobody can query.
- Skipping entity resolution. Duplicate nodes split one supplier's facts, and multi-hop queries silently miss them.
- Losing provenance. An edge without a source chunk cannot be cited, audited or deleted when its document changes.
- Ignoring permissions. Summaries and edges built from restricted documents can leak facts to users who could not open those documents. Carry access labels onto nodes, edges and community summaries, and filter at query time.
- Letting generated queries run unchecked. Use a read-only role, limits and validation.
Turning a pattern like this into a system a customer actually relies on, with their data, their identity system and their definition of risk, is the kind of work Forward Deployed Engineers do: from AI demo to enterprise outcome.
Frequently asked questions
What is GraphRAG?
GraphRAG is retrieval-augmented generation that uses a knowledge graph as a retrieval source. Entities and the relationships between them are retrieved, often together with text chunks from a vector index, and given to the LLM so it can answer multi-hop, relationship and corpus-wide questions that similarity search alone handles poorly.
What is the difference between graph RAG and vector RAG?
Vector RAG retrieves the text chunks most similar to the question. Graph RAG retrieves entities and explicit relationships, and can follow connections across documents. Vector RAG is cheaper and works well when answers sit in one passage; graph RAG helps when answers depend on links between facts spread over many documents.
Does GraphRAG replace vector search?
Usually not. Most production systems use both: vector or hybrid search finds the relevant starting points and passages, and the graph adds connected facts.
What is Microsoft GraphRAG?
It is an approach from Microsoft Research, with an open-source implementation, that uses an LLM to extract entities and relationships from a corpus, groups them into a hierarchy of communities with the Leiden algorithm, and writes a summary for each community. Global search answers corpus-wide questions by combining partial answers from many community summaries; local search answers entity-specific questions.
Which graph database should I use for GraphRAG?
Common choices include Neo4j, Amazon Neptune, Memgraph, TigerGraph, ArangoDB and PostgreSQL with the Apache AGE extension. Choose on query language, managed hosting on your cloud, vector support, access control and team familiarity.
Is building a knowledge graph with an LLM expensive?
It can be, because every chunk is processed by an LLM during extraction, and community summarisation adds more calls. Control it with a tight schema, a smaller extraction model, direct loading of structured data and incremental updates.
How do I keep a knowledge graph up to date?
Detect changed documents with content hashes, re-extract only affected chunks, and use per-edge provenance to delete facts from superseded documents. Refresh community summaries when membership changes and store effective dates on facts.
GraphRAG rewards engineers who can design a schema with domain experts, run a reliable extraction pipeline and prove with evaluation that the graph earns its cost. To practise RAG, knowledge retrieval and agent design end to end, explore Cloudsoft's AI, GenAI and Agentic AI training in Hyderabad, in our Ameerpet classroom or live online. Call +91 96660 19191 for a free demo.



