Evaluating Graph-RAG vs. Standard RAG: A Hallucination Benchmark on Fact-Dense Queries

In this article, you will learn how to benchmark a deterministic 3-Tiered Graph-RAG system against a standard vector RAG pipeline on fact-dense queries, and what the results reveal about matching prompt complexity to model capacity.

Topics we will cover include:

  • How to build a synthetic, fact-dense sports dataset that intentionally introduces noisy, contradictory text into a vector database while keeping ground truth in a graph store.
  • How to implement and compare two retrieval functions — one for standard vector RAG and one for the 3-Tiered Graph-RAG — using ChromaDB and a local Flan-T5 model.
  • Why a small-capacity local LLM can underperform with complex, conflict-resolution prompts, and what this means for production RAG system design.

Evaluating Graph-RAG vs. Standard RAG: A Hallucination Benchmark on Fact-Dense Queries

Introduction

In a previous article, a deterministic, 3-Tiered Graph-RAG system architecture was proposed to address “lossiness” in dense vector databases when dealing with atomic facts. Because semantic proximity vector search and retrieval rely on similarity, classical RAG pipelines often confuse numeric data, historical stats, or entity relationships when several similar entities coexist in a latent space.

The proposed architecture tackled this limitation through a layered approach consisting of:

  • A lightweight QuadStore knowledge graph based on a SPOC (Subject-Predicate-Object-Context) schema.
  • A standard ChromaDB vector database capable of managing fuzzy, long-tail contexts.

In production-oriented engineering, however, hard numbers are critical. Accordingly, we now illustrate a benchmark between standard Vector RAG and our 3-Tiered Graph-RAG to measure how much reduction in hallucinations can be expected when querying over a fact-dense dataset.

Prerequisites and Setup

To run the example benchmark in this article, you’ll need to install these libraries if you don’t have them in your IDE — if you try the example out in a Google Colab notebook, install them all, as follows:

Next, we import all the modules and functions we will need and set up a local LLM, in particular, the google/flan-t5-base model available with Hugging Face’s transformers library:

Next, we build a simple QuadStore, following the same principles and guidelines as the one shown in the previous article:

Likewise, we need to initialize a Vector DB. We do this by instantiating a local ChromaDB collection that will handle unstructured chunks of text.

Next, we get some fresh data — and for this example, it couldn’t be fresher: we synthetically create it from scratch. The following code generates a synthetic dataset with performance information for 50 basketball players. It is designed to inject accurate truths into the graph database while intentionally feeding messy, contradictory paragraphs into the vector DB, just as often happens in the real-world texts that RAG systems are built upon.

We are almost there. Before running the benchmarking pipeline, the two missing pieces are the retrieval functions — one for each system. The standard RAG uses a simple context prompt, while the 3-Tiered Graph-RAG employs a strict rule-based prompt that forces the LLM to give first priority to graph facts over vector fallbacks.

To run the benchmark, we loop through all 50 test queries in the dataset, check whether the text produced by the model contains the exact target number, and calculate overall accuracy.

Final results:

Summary of Results and Conclusion

These results carry a crucial lesson for production engineering of RAG systems: prompt complexity and model capacity must be matched. The google/flan-t5-base model is well-optimized but relatively small (250M parameters). In practice, this means it excels at straightforward reading comprehension in a standard RAG pipeline, but struggles with the complex instruction following required to handle conflicting data sources — exactly the kind found in the Graph-RAG prompt. The slight accuracy drop shows that enforcing graph-based rules via prompting requires a much larger reasoning model, such as Llama 3 or GPT-4. Forcing complex conflict-resolution prompts onto lighter, local LLMs can hurt performance in unexpected ways. These are not the results you might expect, but they shed light on the real-world trade-offs between different RAG approaches.

No comments yet.

Leave a Reply

Machine Learning Mastery is part of Guiding Tech Media, a leading digital media publisher focused on helping people figure out technology. Visit our corporate website to learn more about our mission and team.