When Relevance Isn't Enough: Temporal Validity in Retrieval
A retrieved passage can be about the right thing and still be wrong for the time a question asks about. Which times matter, and what that means for designing retrieval.
A relevant answer at the wrong time
A retrieval system is asked a historical question about a patient: What was the patient taking at tq? The record holds three short notes. They are hypothetical and deliberately simple: A and B stand for two treatments, and the times are ordered t1 < tq < t2 < t3 < t4.
- Note 1, written at t1: “Drug A started.”
- Note 2, written at t2: “Switched from Drug A to Drug B.”
- Note 3, written at t3: “Currently taking Drug B since switching from Drug A; review planned.”
Which one should the system hand to the model? All three concern the patient's treatment, and all three mention A. A common heuristic breaks the tie: when several passages look relevant, prefer the newest. That promotes note 3. It is the most recent note, it speaks about the patient's medication in the present tense, and it names a drug. The answer it suggests is B.
The answer is A. The note that starts A shows it began at t1. The switch note shows it ended at t2. The question's time, tq, falls between them, so A held at tq. Neither note settles this alone; together they bound the period during which A held.
Note 3 is accurate, but the state it describes, B, only begins after tq. It is about the right thing at the wrong time. Used as evidence for this question, it is wrong-time evidence.
Latest first gave a reasonable answer to a different question, what is the patient taking now?, and the wrong answer to the one that was asked. Fig. 1 sets the three notes against the timeline: all three are relevant, and only the start note and the switch note together establish what held at tq.
The distinction the example turns on is not new, but it is easy to lose inside a retrieval pipeline. Semantic relevance asks whether evidence is about the right thing. Temporal validity asks whether what it describes holds at the time the question intends. The rest of this article is about why the two come apart, and what that changes in design. The short version: recency is not validity.
Why semantic relevance does not settle temporal validity
Dense retrieval is built for the relevance judgment. In Dense Passage Retrieval (Karpukhin et al., 2020 ), a question encoder and a passage encoder map text to vectors, and the similarity between a question and a passage is the dot product of their vectors. The authors frame training as metric learning: each question comes with one relevant passage and several irrelevant ones, and the loss rewards the relevant pair for scoring higher than the negatives.
That objective is about relevance between questions and passages. Nothing in it explicitly requires a passage to be valid at the time the question intends. A training pair says that this passage should score well for this question. It does not say which time either of them refers to. By that measure, “Currently taking Drug B” can be a perfectly good match for a question about what the patient was taking.
This does not make dense retrievers blind to time. Dates are text, and embeddings can encode them. A model can pick up temporal cues indirectly when its training data happens to reward them. The narrower point is that temporal validity is not the property a standard relevance objective directly optimizes. Any temporal information it learns is incidental to an objective that directly optimizes relevance, not temporal validity.
Some work on time-sensitive question answering treats the two judgments separately. MRAG (Zhang et al., 2025 ) splits a question into its main content and a temporal constraint, then scores semantic relevance and temporal relevance separately for each piece of evidence. In a controlled analysis, the authors also found that the retrieval methods they tested prioritized matching exact dates, which they describe as a shortcut for temporal reasoning. Matching a date is useful temporal evidence, but it is not the same as establishing whether evidence is valid for the time a question asks about.
So the two properties need to be kept apart. Similarity directly targets the first; neither similarity nor recency, by itself, guarantees the second. Either can still surface valid evidence. Neither is designed to establish that it is valid.
Which time are we talking about?
“The note's time” sounds like one thing. Look at what note 3 actually says: “Currently taking Drug B since switching from Drug A; review planned.” It was written at t3. It describes a switch that had already happened, at t2. And it mentions a review that has not happened yet, planned for t4. That is three times in one short note, and the question brings a fourth. They need to stay separate.
- Document time, or document creation time (DCT), is when the note was written.
- Event time is when something the note describes happened, or is planned to happen. One note can refer to several.
- A validity interval is the period during which a state holds. Taking A is a state; the switch at t2 is an event, and it ends that state.
- Query time is the time the question intends. Here, tq.
Note 3 shows why these cannot be collapsed. Its document time is tD = t3. It refers back to the switch at t2, which is why A appears in it only in the past tense, and forward to the planned review at t4, after the note was written. One note, one document time, and event times on both sides of it. Fig. 2 puts those four notions on a single line.
The query time has not moved. It is still tq, before the switch, and at tq B does not yet hold. Note 3 has the latest document time in the record, and document time is exactly what latest first ranks by. Temporal validity is a different question: whether the relevant state holds at the question's time. Latest first can be useful for questions about the present, but recency alone does not establish temporal validity; for historical questions, it may answer a different question entirely.
This is also where freshness and validity part ways. Work on freshness, such as FreshLLMs (Vu et al., 2024 ), targets questions that need up-to-date world knowledge. That is a real objective, and a different one from answering correctly about a past time. A historical question can need older evidence, and a recently written note can describe earlier history. Treating the newest document as the most valid one conflates the two.
None of these distinctions are new to temporal NLP. Relating events to temporal expressions in clinical narratives is an established task: the 2012 i2b2 challenge evaluated systems on extracting clinical events, temporal expressions and the temporal relations between them from discharge summaries (Sun, Rumshisky & Uzuner, 2013 ). A retrieval design that collapses these distinctions into a single timestamp can lose information the query depends on.
Validity is often an interval, not a timestamp
Back to the main example. “Drug A started” opens A's interval at t1. On its own, it says when A began and nothing about when A stopped. A start is not a duration: from that note alone, a system can conclude that A held at tq only by assuming it continued.
The switch note closes the interval at t2. It was written after tq, yet it is part of the evidence for the answer, because it shows that A was still in place up to the switch, and therefore at tq. Establishing a validity interval can take start, stop, continuation or resolution evidence from several notes, some of it written after the time in question.
The same view changes how to read notes that seem to disagree. “Drug A started” implies A; “Currently taking Drug B” implies B. Neither is an error. A held, and was later superseded by B. The conflict goes away once each state is placed on its own interval. This is the kind of reasoning interval representations exist for: Allen's interval algebra (Allen, 1983 ) was motivated partly by the observation that when temporal knowledge is imprecise or relative, representations built on dates or time instants fall short, and it reasons about how intervals relate to one another instead.
A natural objection: why not filter the evidence so that only notes written by the query time are used, document_time <= query_time? That looks like a clean way to stop the system from seeing the future.
Take a separate hypothetical. It is not an extension of the three-note record: here, a retrospective note 2′ replaces the switch note. The timeline is t1 < ts < tq < t2′.
- Note 1, written at t1: “Drug A started.”
- A actually stops at ts, before tq.
- Note 2′, written later at t2′, after tq: “Drug A had been stopped at ts.”
At tq, A does not hold, and only the later note says so. The filter removes note 2′, because its document time is later than tq.
The filter leaves “Drug A started” as the available evidence. It says when A began, not whether A was still being taken at tq. That can lead the system to A only by assuming the state persisted, while later evidence needed to confirm or contradict that persistence has been discarded. Fig. 3 shows which note survives the filter and which one it drops. What it cannot show is why the drop matters: the filter constrains when notes were written, while the question concerns when things held. Note 2′ is written late about something early. Judged by its document time, it looks like evidence from the future. In this record, it is the only evidence that A stopped.
Hard temporal filters are not wrong. They are useful when the temporal constraint is explicit and the timestamps are reliable. They become brittle when document time is only a proxy for event time, or when evidence is relative ("two weeks after starting"), uncertain, or recorded retrospectively. Keeping those anchors, with the precision each one actually has, is an annotation problem in its own right; the temporal annotation case study covers that side: temporal anchors, granularity, and marking underspecification explicitly instead of filling it in.
What this changes when designing retrieval
Whether semantic relevance is enough for time-sensitive retrieval is the question that motivated the retrieval work in my State Engineer thesis. It first came up while I was writing my Master's thesis, a structured review that stepped back to survey temporal retrieval-augmented generation as a whole: its methods, evaluation and open challenges. As design reasoning, a few consequences follow from the examples above.
Document time is useful context, but insufficient on its own. It is often the most dependable timestamp a note has, and it should be kept, but it records when a note was written, not when the things it describes happened. Note 3 was written at t3 and describes a switch at t2; indexed only by its date, it looks like evidence about t3. The event anchors inside a note deserve to be kept too, along with their uncertainty and granularity. "In 2024" does not identify a day, and storing it as one invents precision the source never had. Relative and underspecified anchors should stay marked as such.
With those anchors in place, time can act as a second signal beside semantic relevance. A passage still has to be about the right thing, so time should complement relevance rather than blindly replace it. Hard temporal filters fit constraints that are explicit and reliable. Elsewhere, using temporal information as a ranking signal can preserve evidence that a hard cutoff would remove: a strict document-date cutoff at tq would have dropped note 2′, the one note that says A had already stopped.
The remaining changes sit on either side of retrieval. Before it, the question's temporal scope needs interpreting. “What was the patient taking at tq?”, “What is the patient taking now?” and “Has the patient ever taken A?” share most of their content words and need different evidence: the start and switch notes together for the first, the newest note for the second, and any note that mentions A for the third. After it, retrieval should be evaluated separately from generation. When only final answers are scored, wrong-time evidence can stay hidden behind an answer that happens to look right; scoring the retrieved evidence directly makes wrong-time retrieval something you can observe.
Return to the question at the start: what was the patient taking at tq? All three notes are relevant, and note 3 is the newest. The answer is decided by A's validity interval, opened by “Drug A started” and closed by the switch note, and where tq falls relative to it. Which note is newest does not decide it.
Two questions stay open for the field. How should retrieval represent validity when the end of an interval is unknown, or anchored only relative to another event? And how should retrieval be evaluated when correctness itself depends on time, so that the same passage is valid evidence for one question and wrong-time evidence for another?
References
- Allen, J. F. (1983). Maintaining Knowledge about Temporal Intervals. Communications of the ACM, 26(11), 832–843. doi:10.1145/182.358434
- Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W. (2020). Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of EMNLP 2020, 6769–6781. doi:10.18653/v1/2020.emnlp-main.550
- Sun, W., Rumshisky, A., & Uzuner, O. (2013). Evaluating temporal relations in clinical text: 2012 i2b2 Challenge. Journal of the American Medical Informatics Association, 20(5), 806–813. doi:10.1136/amiajnl-2013-001628
- Vu, T., Iyyer, M., Wang, X., Constant, N., Wei, J., Wei, J., Tar, C., Sung, Y.-H., Zhou, D., Le, Q., & Luong, T. (2024). FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation. In Findings of ACL 2024, 13697–13720. doi:10.18653/v1/2024.findings-acl.813
- Zhang, S., Xue, Y., Zhang, Y., Wu, X., Luu, A. T., & Zhao, C. (2025). MRAG: A Modular Retrieval Framework for Time-Sensitive Question Answering. In Findings of EMNLP 2025. doi:10.18653/v1/2025.findings-emnlp.167