GraphRAG
Standard RAG retrieves text chunks similar to a question. That works well for questions whose answer lives in one or a few passages, and poorly for questions that require connecting information across many documents ("how are these three suppliers related to last year's recall?") or summarizing a whole corpus ("what are the main themes in these 5,000 customer interviews?"). GraphRAG addresses both by building a knowledge graph of entities and relationships from the documents, then retrieving over that structure.
Popularized by Microsoft Research's GraphRAG project, and implemented in graph databases like Neo4j and many frameworks, the approach trades higher indexing cost and complexity for better multi-hop reasoning, corpus-level answers, and explainable connections between facts.
TL;DR
- GraphRAG extracts entities (people, products, organizations, concepts) and relationships from text, usually with an LLM, into a knowledge graph.
- Community detection (for example Leiden) groups related entities; LLM-written community summaries describe each cluster at several levels.
- Local search starts from entities in the question and expands to their neighbors, relationships, and source chunks.
- Global search answers corpus-wide questions by map-reducing over community summaries.
- It excels at multi-hop, relationship, and "what are the themes" questions; vector RAG remains better and cheaper for direct lookups.
- Costs: expensive indexing (many LLM calls), graph maintenance, and extraction quality issues. Hybrid graph + vector designs are common.
Quick Example
Extracting a graph with an LLM, and a local-search style query in Neo4j Cypher:
The retrieved entities, relationship descriptions, and evidence chunks form the context for the LLM's answer.
Core Concepts
Building the Graph
- Chunk documents (see chunking).
- Extract entities and relationships per chunk with an LLM using structured outputs, optionally with a predefined schema (ontology) of entity and relation types.
- Resolve entities: merge duplicates ("IBM", "International Business Machines", "I.B.M.") with normalization, embeddings, and LLM-assisted matching. This step is critical for graph quality.
- Store in a graph database or graph structure, keeping links back to source chunks for citations.
- Detect communities: cluster the graph (Leiden or Louvain algorithms, see graph algorithms) hierarchically.
- Summarize communities: an LLM writes a report per community at each level, describing key entities, relationships, and claims.
Local Search
For questions about specific entities ("What issues affected the X200 battery supplier?"):
- Identify entities in the question (extraction or embedding match against entity descriptions).
- Traverse to related entities and relationships (1–2 hops), ranked by relevance and relationship weight.
- Gather descriptions, community summaries, and linked source chunks.
- Generate an answer from that structured context, with citations.
Global Search
For corpus-wide, sensemaking questions ("What are the major risk themes across all audit reports?"):
- Select a level of the community hierarchy.
- Map: ask the LLM to answer the question from each community summary independently, with a relevance score.
- Reduce: combine the highest-rated partial answers into a final response.
Vector RAG handles these questions poorly, because no single chunk contains "the themes of the whole corpus".
Variants and Hybrids
- Hybrid graph + vector: vector search finds entry chunks or entities, and graph traversal expands context. It's the most common production pattern.
- Text2Cypher / Text2SPARQL: an LLM translates questions into graph queries over an existing, curated knowledge graph.
- Lighter-weight approaches (LightRAG, and the newer "fast" and "drift" GraphRAG modes): reduce indexing cost by extracting less or summarizing lazily.
- Agentic graph exploration: an agent calls graph query tools iteratively. See agentic RAG.
When to Use GraphRAG
Best Practices
Define a Schema for Your Domain
Unconstrained extraction yields inconsistent entity types and relationship names. A small ontology (the entity and relationship types you actually need) improves consistency and makes queries reliable.
Invest in Entity Resolution
Duplicate or fragmented entities break traversal: the same company as five disconnected nodes can't show connections. Combine normalization rules, embedding similarity, and LLM verification, and review high-impact merges.
Keep Provenance
Link every entity and relationship to the chunks it came from. Answers need citations, and extraction errors need to be traceable. See LLM hallucinations.
Evaluate Against Vector RAG
Measure both approaches on your questions, including multi-hop and global ones. GraphRAG's benefits are real for certain question types, but they don't justify its cost everywhere. See RAG evaluation.
Common Mistakes
Underestimating Indexing Cost
Extracting entities and relationships from every chunk, then summarizing every community, can require LLM calls numbering in the hundreds of thousands for large corpora. Estimate token costs up front, use smaller models for extraction, and index incrementally.
Treating Extracted Graphs as Ground Truth
LLM extraction misses relationships, invents weak ones, and misattributes facts. Treat the graph as a noisy index into the source text, and ground final answers in the original chunks.
Using GraphRAG for Simple Lookups
"What's our parental leave policy?" doesn't need graph traversal. Route simple questions to vector or hybrid retrieval, and reserve graph queries for relational and global questions.
FAQ
What's the difference between GraphRAG and regular RAG?
Regular RAG retrieves text chunks by semantic or keyword similarity to the question. GraphRAG first builds a knowledge graph of entities, relationships, and community summaries from the corpus, then retrieves by traversing that structure, or by summarizing communities. That enables multi-hop and corpus-wide questions that chunk similarity handles poorly.
Do I need a graph database for GraphRAG?
Not strictly. Microsoft's GraphRAG reference implementation stores its graph and summaries in files and tables. Graph databases like Neo4j help when you need flexible traversal queries, incremental updates, or to combine the extracted graph with existing structured data.
How expensive is GraphRAG?
Indexing is the main cost: LLM calls for extraction, entity resolution, and community summaries across the whole corpus, often an order of magnitude more than embedding-based indexing. Query costs vary: local search is modest, and global search, which reads many community summaries, can be expensive per question.
Can GraphRAG and vector search be combined?
Yes, and that's the most common production design. Vector or hybrid search finds relevant entities or chunks, graph traversal adds connected context, and a reranker selects the final context. Routing by question type (lookup vs relational vs global) keeps costs down.
Related Topics
- RAG — Retrieval-augmented generation overview
- Neo4j — A graph database for knowledge graphs
- Graph Algorithms — Community detection and traversal
- Agentic RAG — Agents exploring graphs and indexes
- Hybrid Search — Combining with vector and keyword retrieval
- Structured Outputs — Reliable entity extraction