GraphRAG: Knowledge Graph-Enhanced Retrieval for Multi-Hop Reasoning

This article explains GraphRAG, which combines knowledge graphs with RAG to solve multi-hop reasoning and global query limitations of traditional RAG, detailing Microsoft's GraphRAG approach, Java implementation with Neo4j for entity extraction and graph storage, retrieval processes, community summarization, cost trade-offs, and when to choose GraphRAG over standard RAG.

Coder Trainee
Coder Trainee
Coder Trainee
GraphRAG: Knowledge Graph-Enhanced Retrieval for Multi-Hop Reasoning

Traditional RAG Limitations

Traditional RAG hits three ceilings: multi-hop reasoning (e.g., "Where did A's boss graduate?" requires linking across documents), global questions (e.g., "What are the main trends in this report?" cannot be answered from a single chunk), and relationship queries (e.g., "What is the relationship between A and B?" where relations are scattered). The root cause is that traditional RAG is flat — all chunks exist in a single vector space without structure.

GraphRAG Core Idea

GraphRAG = Knowledge Graph + RAG . Extract entities and relations from documents to build a knowledge graph. During retrieval, fetch both text chunks and associated graph information.

Traditional RAG vs GraphRAG
Traditional RAG:
Document → Chunk → Vector → Top-K Retrieval → LLM
(Flat structure, only similarity)

GraphRAG:
Document → Entity Extraction → Relation Extraction → Knowledge Graph
                      │
Query → Entity Recognition → Graph Traversal → Subgraph + Text Chunks → LLM
(Structured + text, supports multi-hop reasoning)

Microsoft GraphRAG Process

Microsoft's open-source GraphRAG (2024) follows these steps:

Entity Extraction : Extract entities (people, organizations, products) from each text chunk.

Relation Extraction : Extract relationships between entities.

Community Detection : Use Leiden algorithm to partition the graph into communities.

Community Summarization : Generate a summary for each community.

Query : Local query (specific entity) + Global query (community summaries).

Java Implementation: Knowledge Graph Construction

3.1 Entity and Relation Extraction

An LLM-based extractor uses a prompt to output JSON with entities and relations:

@Service
public class EntityExtractor {
    private final ChatClient chatClient;

    public GraphExtraction extract(String chunk) {
        String prompt = """
            Extract entities and relations from the following text.

            Text:
            %s

            Output JSON format:
            {
              "entities": [
                {"name": "Entity Name", "type": "Type (Person/Org/Product/Location)", "description": "Description"}
              ],
              "relations": [
                {"source": "Entity A", "target": "Entity B", "relation": "Relation Description"}
              ]
            }

            Only output JSON, nothing else.
            """.formatted(chunk);
        String result = chatClient.prompt(prompt).call().content();
        return parseGraph(result);
    }
}

Example output for a text about products and suppliers:

{
  "entities": [
    {"name": "Product A", "type": "Product", "description": "Company's main product"},
    {"name": "Supplier X", "type": "Organization", "description": "Supplier of Product A"},
    {"name": "Product B", "type": "Product", "description": "Another company product"}
  ],
  "relations": [
    {"source": "Product A", "target": "Supplier X", "relation": "supplied by"},
    {"source": "Product B", "target": "Supplier X", "relation": "supplied by"}
  ]
}

3.2 Graph Storage with Neo4j

Dependency: org.neo4j.driver:neo4j-java-driver:5.15.0.

@Service
public class GraphStore {
    private final Driver driver;

    public GraphStore(@Value("${neo4j.uri}") String uri,
                      @Value("${neo4j.username}") String username,
                      @Value("${neo4j.password}") String password) {
        this.driver = GraphDatabase.driver(uri, AuthTokens.basic(username, password));
    }

    public void saveEntity(String name, String type, String description) {
        try (Session session = driver.session()) {
            session.run("""
                MERGE (e:Entity {name: $name})
                SET e.type = $type, e.description = $description
                """, Map.of("name", name, "type", type, "description", description));
        }
    }

    public void saveRelation(String source, String target, String relation) {
        try (Session session = driver.session()) {
            session.run("""
                MATCH (a:Entity {name: $source})
                MATCH (b:Entity {name: $target})
                MERGE (a)-[r:RELATES {type: $relation}]->(b)
                """, Map.of("source", source, "target", target, "relation", relation));
        }
    }

    public List<Map<String, Object>> queryRelated(String entityName, int depth) {
        try (Session session = driver.session()) {
            return session.run("""
                MATCH path = (e:Entity {name: $name})-[*1..%d]-(related)
                RETURN related.name AS name, related.description AS description
                LIMIT 20
                """.formatted(depth), Map.of("name", entityName))
                .list(record -> Map.of(
                    "name", record.get("name").asString(),
                    "description", record.get("description").asString()));
        }
    }
}

3.3 Graph Building Pipeline

@Service
public class GraphBuilder {
    @Autowired
    private EntityExtractor extractor;

    @Autowired
    private GraphStore graphStore;

    @Async
    public void buildGraph(List<Chunk> chunks) {
        for (Chunk chunk : chunks) {
            GraphExtraction extraction = extractor.extract(chunk.getContent());

            // Save entities
            for (Entity entity : extraction.getEntities()) {
                graphStore.saveEntity(entity.getName(), entity.getType(), entity.getDescription());
            }

            // Save relations
            for (Relation relation : extraction.getRelations()) {
                graphStore.saveRelation(relation.getSource(), relation.getTarget(), relation.getRelation());
            }
        }
    }
}

GraphRAG Retrieval

4.1 Retrieval Flow

@Service
public class GraphRagRetriever {
    @Autowired
    private GraphStore graphStore;

    @Autowired
    private HybridRetriever hybridRetriever;

    public List<Document> retrieve(String query) {
        // 1. Identify entities from query
        List<String> entities = extractEntities(query);

        // 2. Graph retrieval: find related entities
        Set<String> relatedEntities = new HashSet<>();
        for (String entity : entities) {
            List<Map<String, Object>> related = graphStore.queryRelated(entity, 2);
            for (Map<String, Object> r : related) {
                relatedEntities.add((String) r.get("name"));
            }
        }

        // 3. Text retrieval: search by entities and original query
        List<Document> graphDocs = searchByEntities(relatedEntities);
        List<Document> vectorDocs = hybridRetriever.retrieve(query, 10);

        // 4. Merge and deduplicate
        return mergeAndDeduplicate(graphDocs, vectorDocs);
    }

    private List<Document> searchByEntities(Set<String> entities) {
        List<Document> results = new ArrayList<>();
        for (String entity : entities) {
            results.addAll(hybridRetriever.retrieve(entity, 3));
        }
        return results;
    }
}

4.2 Global Queries: Community Summaries

For global questions like "What are the main trends in this report?", GraphRAG uses community summaries:

@Service
public class CommunitySummarizer {
    public List<Community> detectCommunities() {
        // Use Leiden algorithm for community detection
        // Output: each community contains a set of entities
        return leidenAlgorithm.detect(graph);
    }

    public String summarizeCommunity(Community community) {
        // Collect all entity descriptions in the community
        String context = community.getEntities().stream()
                .map(Entity::getDescription)
                .collect(Collectors.joining("
"));

        // Generate community summary
        return chatClient.prompt("""
            Based on the following information, summarize the core theme of this community:

            %s
            """.formatted(context)).call().content();
    }
}

At query time: global question → retrieve community summaries → generate answer.

GraphRAG Costs

Comparison across dimensions:

Build Cost : Traditional RAG low; GraphRAG high (requires entity extraction).

Build Time : Traditional RAG minutes; GraphRAG hours.

Storage Cost : Traditional RAG vector store only; GraphRAG vector store + graph database.

Query Cost : Traditional RAG low; GraphRAG medium.

Suitable Scenarios : Traditional RAG for simple QA; GraphRAG for multi-hop reasoning, relationship queries, global questions.

Key insight : GraphRAG does not replace traditional RAG but supplements it in specific scenarios.

When to Use GraphRAG

Simple fact queries → Traditional RAG.

Multi-hop reasoning → GraphRAG.

Relationship queries → GraphRAG.

Global questions → GraphRAG.

Large-scale documents → Hybrid approach.

Recommendation : Start with traditional RAG; consider GraphRAG only when you hit the ceiling.

Next Episode Preview

Part 15: Combining RAG with Agents — Making the Knowledge Base "Alive". Topics include RAG-Agent collaboration modes, active vs. passive retrieval, and multi-step reasoning.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

JavaLLMRAGNeo4jKnowledge GraphCommunity DetectionEntity ExtractionGraphRAGMulti-hop Reasoning
Coder Trainee
Written by

Coder Trainee

Experienced in Java and Python, we share and learn together. For submissions or collaborations, DM us.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.