Plain RAG retrieves passages; GraphRAG retrieves a map. The project at microsoft/graphrag, MIT-licensed and currently at version 3.3.0, describes itself as a data transformation suite that extracts meaningful, structured data from unstructured text using the power of LLMs, and the structured part is the whole point. Instead of stuffing top-k chunks into a prompt, GraphRAG first builds a knowledge graph: an LLM pass pulls out entities and relationships, a clustering pass organizes that graph into hierarchical communities, a summarization pass writes a report for every community at every level, and only then do questions get asked, either by walking from a specific entity outward or by marching up the community hierarchy for a global answer. The technique came out of Microsoft Research, first released in July 2024, with the arXiv paper behind it and a research blog post titled around unlocking LLM discovery on narrative private data. Python 3.11 through 3.14 is supported, and the README carries an unusual warning worth quoting: the project is largely in maintenance mode, and indexing can be an expensive operation, so start small.
The source layout rewards that caution with clarity. The repository is a monorepo of eight packages under packages, and the decomposition is a lesson in itself: graphrag-llm, graphrag-storage, graphrag-vectors, graphrag-cache, graphrag-input, and graphrag-chunking are each swappable libraries, while the main graphrag package composes them into a command-line application with an indexing side and a query side. Everything is wired through factories with a shared factory helper in graphrag-common, so the code reads like a kit of parts rather than a monolith. As always in this series, what follows is an educational tour of published source code.
GraphRAG at a glance: the CLI loads config and fans out to indexing or querying, indexing reads documents through input readers and chunkers, extracts a graph and its communities with LLM help, and finalizes parquet and vector artifacts, which the query engines then load to answer questions through the LLM middleware stack.
Reading the overview from left to right:
- The front door is packages/graphrag/graphrag/cli/main.py, which exposes init, index, query, and prompt-tune commands.
- Settings live in the config package at packages/graphrag/graphrag/config, defining models, input, storage, and output.
- Documents enter through the readers in packages/graphrag-input/graphrag_input/input_reader_factory.py, including text, CSV, JSON, and markitdown-based readers.
- Text is split by the chunkers in packages/graphrag-chunking/graphrag_chunking/chunker_factory.py, with token and sentence strategies.
- The graph itself is born in the extract workflow at packages/graphrag/graphrag/index/workflows/extract_graph.py.
- Communities come from hierarchical clustering in packages/graphrag/graphrag/graphs/hierarchical_leiden.py, and embeddings from the workflow in packages/graphrag/graphrag/index/workflows/generate_text_embeddings.py.
- Artifacts are written through packages/graphrag-storage/graphrag_storage/storage_factory.py and packages/graphrag-vectors/graphrag_vectors/vector_store_factory.py.
- Questions are answered by the structured search family in packages/graphrag/graphrag/query/structured_search/base.py, with every model call flowing through the middleware-rich LLM package in packages/graphrag-llm/graphrag_llm/completion/completion_factory.py.
Why You Need This
The first reason is that GraphRAG answers the question vanilla RAG fumbles: questions whose answer is spread across many documents rather than sitting in one chunk. The global search engine in packages/graphrag/graphrag/query/structured_search/global_search/search.py works map-reduce style over community reports, asking a batch of communities what they know about your question and then distilling those partial answers into one response, with the community context assembled in packages/graphrag/graphrag/query/structured_search/global_search/community_context.py and an optional dynamic community selection that rates relevance before spending tokens. For questions like what are the major themes in this corpus, a map beats a haystack every time, and the map is exactly what the indexing side builds.
The second reason is the engineering honesty of the infrastructure layer. Every LLM call, whether it extracts entities, summarizes descriptions, or writes community reports, flows through graphrag-llm, where a completion factory wraps the model in a middleware chain from packages/graphrag-llm/graphrag_llm/middleware/with_cache.py for caching, with_retries for exponential backoff, with_rate_limiting backed by a sliding-window limiter, and with_metrics for token accounting. Cache backends come from graphrag-cache in JSON, SQLite, or memory flavors; text units and tables persist through graphrag-storage as parquet, CSV, Cosmos DB, or Azure blobs; and embeddings land in graphrag-vectors backed by LanceDB or Azure AI Search. Because each is a small library behind a factory, the lesson transfers: model-facing code should treat retries, caching, rate limits, and cost metrics as composable middleware, not as scattered try-except blocks.
The third reason is that the query package is a comparative study of retrieval strategies in one tree. Local search in packages/graphrag/graphrag/query/structured_search/local_search/search.py anchors on entities matched to the question and pulls their neighbors, relationships, and text units through a mixed-context builder; basic search is the humble fallback that just ranks text units; and DRIFT search in packages/graphrag/graphrag/query/structured_search/drift_search/search.py expands from local hits into follow-up questions guided by community data, using the primer, state, and action modules in the same folder. All four share the context builders in packages/graphrag/graphrag/query/context_builder/builders.py and the parquet loaders in packages/graphrag/graphrag/query/input/loaders/dfs.py, so you can compare how the same index supports four different reasoning styles without four different storage formats.
The detail view: CLI and API on top, the infrastructure packages below, the indexing workflows from input to finalize in the middle, and the four query engines with their context builders on the right.
Walk the detail diagram and the indexing story unfolds as an ordered list of workflows run by the workflow runner in packages/graphrag/graphrag/index/run, which the API entry point in packages/graphrag/graphrag/api drives. Documents are read by graphrag-input, chunked by graphrag-chunking, and then the extract workflow calls the graph extractor in packages/graphrag/graphrag/index/operations/extract_graph/graph_extractor.py, which runs an LLM extraction pass per chunk and keeps gleaning more entities and relationships until the model stops finding new ones. If you cannot afford LLM extraction, the NLP alternative in packages/graphrag/graphrag/index/workflows/extract_graph_nlp.py builds a noun-phrase graph deterministically through packages/graphrag/graphrag/index/operations/build_noun_graph/build_noun_graph.py with pluggable extractors, a budget-friendly sibling that many users do not know exists.
The graph then gets socialized. Duplicate entity descriptions are merged by summarize_descriptions in packages/graphrag/graphrag/index/operations/summarize_descriptions/summarize_descriptions.py, optional claims come from the covariates extractor in packages/graphrag/graphrag/index/operations/extract_covariates/extract_covariates.py, and the Louvain-style Leiden algorithm in packages/graphrag/graphrag/graphs/hierarchical_leiden.py clusters the graph into nested communities, which the create_community_reports workflow turns into layered narratives through the extractor in packages/graphrag/graphrag/index/operations/summarize_communities/community_reports_extractor.py. Finally create_final_text_units stitches chunks to entities and communities, generate_text_embeddings embeds them into the vector store, and finalize writes the parquet tables plus a GraphML snapshot via packages/graphrag/graphrag/index/operations/snapshot_graphml.py. Incremental updates get their own machinery in packages/graphrag/graphrag/index/update/incremental_index.py, so adding documents does not force a full rebuild, and the whole run reports progress through the callbacks package in packages/graphrag/graphrag/callbacks.
From Install to a Working App
The documented path is command-line first. Install the package, run graphrag init with a root directory to emit a settings file and starter prompts, point the input at a folder of text, and run the index command; the console fills with workflow progress as documents become chunks, chunks become a graph, the graph becomes communities, and communities become reports. Then ask questions with the query command, choosing local, global, drift, or basic as the method, each of which loads its tables from the output folder and builds context as the diagram showed. The init step is deliberately repeatable: the README advises rerunning graphrag init with force between minor version bumps to pick up config format changes, and prompt tuning gets its own command, which regenerates extraction and report prompts from your domain through the generator modules in packages/graphrag/graphrag/prompt_tune/generator.
Budget with your eyes open, because the README warns that indexing can be an expensive operation and that the code serves as a demonstration rather than a supported offering, with the project now largely in maintenance mode accepting mainly bug fixes. Extraction makes one or more LLM calls per chunk, description summarization and community reports add more, and the utility of community reports depends on your corpus being narrative rather than transactional. For small, evolving doc sets a plain vector RAG may remain the pragmatic default. But for the corpus-shaped questions that keep defeating chunk retrieval, where the answer is a theme, an actor network, or a storyline spanning hundreds of documents, GraphRAG remains the clearest published blueprint for building a map instead of shuffling pages, and its monorepo makes every layer of that blueprint a self-contained lesson. Enjoyed this post? Never miss out on future posts by following us