In this article, you will learn how to add a lightweight temporal reasoning layer to a Graph-RAG system so that it can distinguish fresh facts from stale ones.
Topics we will cover include:
- How to extend standard subject-predicate-object triples into time-stamped quadruples stored in a simple temporal graph.
- How to calculate recency weights with exponential decay and use them to rank conflicting facts as of a given query date.
- How to tune the half-life parameter and integrate the temporal graph into a deterministic 3-tiered Graph-RAG retrieval pipeline.
![]()
Introduction and Motivation
In a previous article, Building a Deterministic 3-Tiered Graph-RAG System, we addressed the challenge of handling conflicting information in RAG (Retrieval-Augmented Generation) architectures. In particular, we built a hierarchy to manage conflicting information by giving “fresh” facts top priority over less-fresh ones.
Throughout that journey, a key question arose: how does our graph-based RAG system know exactly what is fresh? Standard knowledge graphs treat facts as “timeless,” context-independent SPO (subject-predicate-object) triples, such as (Company, HAS_CEO, Alice). This approach doesn’t quite fit the real world we live in, where things are messy and change constantly: what if Alice switched jobs and is no longer CEO? Feeding these facts to an LLM without temporal context is the perfect recipe for hallucinations, resulting in confusing or factually incorrect responses.
To tackle this challenge, this article shows the key steps to build a dedicated, lightweight temporal reasoning engine for Graph-RAG through a few simple Python functions. We also discuss its integration with the 3-tiered Graph-RAG system built previously. The key idea consists of upgrading standard triples into time-stamped quadruples and calculating recency weights that indicate degrees of “fact freshness.”
Ready? Let’s go!
A New, Temporal Journey, Step by Step
The first step is to extend standard SPO triples into “temporal quads,” where the fourth dimension introduces time, concretely a timestamp: (Subject, Predicate, Object, Timestamp).
The following Python class is defined to hold our new, extended knowledge, also called a temporal graph. If you are working in a notebook environment, simply paste this code into your first code cell:
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 |
import datetime import math class TemporalGraph: def __init__(self): # Data will be stored in this knowledge base as: Subject -> Predicate -> List of (Object, Date) self.knowledge_base = {} def add_fact(self, subject, predicate, obj, date_string): """Adds a time-stamped fact to the graph.""" # Converting string to a comparable date object fact_date = datetime.datetime.strptime(date_string, "%Y-%m-%d").date() if subject not in self.knowledge_base: self.knowledge_base[subject] = {} if predicate not in self.knowledge_base[subject]: self.knowledge_base[subject][predicate] = [] self.knowledge_base[subject][predicate].append((obj, fact_date)) print(f"Added: {subject} {predicate} {obj} (as of {date_string})") # Initializing our graph tg = TemporalGraph() |
Now it’s time to populate our newly created temporal graph, tg, following a real-world scenario where facts change at light speed … well, maybe not that fast, but still rapidly! If we were tracking the leadership roles in a tech company during a chaotic week full of changes, we could have something like:
|
1 2 3 4 5 6 7 8 |
# A timeline of shifting facts tg.add_fact("TechCorp", "HAS_CEO", "Alice", "2021-01-15") tg.add_fact("TechCorp", "HAS_CEO", "Bob", "2023-11-17") tg.add_fact("TechCorp", "HAS_CEO", "Charlie", "2023-11-19") tg.add_fact("TechCorp", "HAS_CEO", "Bob", "2023-11-21") # Bob came back! # Adding also a static fact for a bit of contrast tg.add_fact("TechCorp", "FOUNDED_IN", "San Francisco", "2010-05-01") |
Output:
|
1 2 3 4 5 |
Added: TechCorp HAS_CEO Alice (as of 2021-01-15) Added: TechCorp HAS_CEO Bob (as of 2023-11-17) Added: TechCorp HAS_CEO Charlie (as of 2023-11-19) Added: TechCorp HAS_CEO Bob (as of 2023-11-21) Added: TechCorp FOUNDED_IN San Francisco (as of 2010-05-01) |
Bear in mind that in a standard RAG system, a search like “Who acts as the CEO of TechCorp?” would likely retrieve Alice, Bob, and Charlie, all at once! Thus, we need a mechanism to assign truthfulness weights to facts, and it’s simpler than you might think.
Calculating recency weights is the key to resolving possible conflicts mathematically. We just want a mechanism that says: “hey, this fact is newer than that one, so it’s more likely to constitute today’s truth.” A smart approach to do this is based on exponential decay, which consists of assigning a half-life time window to facts. For instance, if the half-life is set to one year (365 days), then a fact that is one year old will carry a weight of 0.5. Meanwhile, a fact asserted today would carry a weight of 1.0.
These two functions are designed to introduce the aforementioned weight scoring logic to our temporal graph:
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 |
def calculate_recency_weight(fact_date, query_date, half_life_days=365): """ Calculates a score between 0 and 1 based on how old the fact is. Using exponential decay: weight = (0.5) ^ (age_in_days / half_life) """ age_in_days = (query_date - fact_date).days # If the fact is from the future relative to our query, it is capped at 1.0 if age_in_days < 0: return 1.0 weight = 0.5 ** (age_in_days / half_life_days) return round(weight, 4) def query_temporal_graph(graph, subject, predicate, as_of_date_str, half_life_days=365): """Queries the graph and ranks answers by their temporal weight.""" query_date = datetime.datetime.strptime(as_of_date_str, "%Y-%m-%d").date() try: facts = graph.knowledge_base[subject][predicate] except KeyError: return f"No information found for {subject} -> {predicate}" scored_results = [] for obj, fact_date in facts: # We only consider facts that happened ON or BEFORE our query date if fact_date <= query_date: weight = calculate_recency_weight(fact_date, query_date, half_life_days) scored_results.append({ "answer": obj, "date": fact_date.strftime("%Y-%m-%d"), "weight": weight }) # Sorting by weight (highest/freshest first) scored_results.sort(key=lambda x: x['weight'], reverse=True) return scored_results |
Finally, we are in a position to see it all in action. We will finish by showing an example that queries our graph. Temporal reasoning acts as a kind of “time travel” at execution time: if we added code to persist our facts and then asked who the CEO was a few days later, the mechanism we implemented would simply adjust its weights on the fly:
|
1 2 3 4 5 6 7 8 9 |
print("--- Query 1: Who is the CEO as of Nov 18, 2023? ---") results_past = query_temporal_graph(tg, "TechCorp", "HAS_CEO", "2023-11-18") for res in results_past: print(f"Candidate: {res['answer']} | Fact Date: {res['date']} | Confidence Weight: {res['weight']}") print("\n--- Query 2: Who is the CEO as of Dec 01, 2023? ---") results_present = query_temporal_graph(tg, "TechCorp", "HAS_CEO", "2023-12-01") for res in results_present: print(f"Candidate: {res['answer']} | Fact Date: {res['date']} | Confidence Weight: {res['weight']}") |
Results:
|
1 2 3 4 5 6 7 8 9 |
--- Query 1: Who is the CEO as of Nov 18, 2023? --- Candidate: Bob | Fact Date: 2023-11-17 | Confidence Weight: 0.9981 Candidate: Alice | Fact Date: 2021-01-15 | Confidence Weight: 0.1396 --- Query 2: Who is the CEO as of Dec 01, 2023? --- Candidate: Bob | Fact Date: 2023-11-21 | Confidence Weight: 0.9812 Candidate: Charlie | Fact Date: 2023-11-19 | Confidence Weight: 0.9775 Candidate: Bob | Fact Date: 2023-11-17 | Confidence Weight: 0.9738 Candidate: Alice | Fact Date: 2021-01-15 | Confidence Weight: 0.1362 |
As one might expect, running the first query gives us Bob as the top answer with an almost full weight: Charlie doesn’t even appear, as he hadn’t been appointed at that point! Meanwhile, running the second query reveals a caveat: perhaps the 365-day half-life window is too long, since it takes a whole year for facts to lose 50% of their relevance. Thus, in a frenetic week full of organizational changes, we can see that even though Bob is again the top answer, he is very closely followed by Charlie and even by Bob’s own prior appointment. The quick fix consists of adjusting the half_life_days parameter, for instance, by changing it from 365 to 7. Try it yourself and enjoy the new results!
Wrapping Up
Now that we have built this mechanism to take temporal graph information into account, how could it be integrated into the deterministic 3-tiered architecture built in the previous, related article? When the user sends a prompt to the LLM in the RAG system, you’ll want a retriever that no longer only fetches texts: instead, it should run the query against the temporal graph, sort facts by their confidence weight, and pass only the top-weighted one (or, at most, a small ranked list) into the prompt’s context. This has the potential to remove the LLM’s need to guess which fact is the most current one: that issue is sorted out even before the final prompt reaches the model.






No comments yet.