Skip to content
Chamel Nadir Bouacha

Temporal RAG for Healthcare

A retrieval-first architecture that combines semantic relevance with explicit temporal constraints over longitudinal healthcare information, addressing temporal granularity, event-time anchoring, provenance, and retrieval-time consistency.

Status
State Engineer thesis, defended 22 June 2026.
Role
Research Intern
Institution
Télécom SudParis
Period
Sept. 2025 – June 2026
(a) “Is the patient currently taking [medication]?”A · “[medication] started”B · “[medication] stopped” · valid at tqt1t2tq(b) Similarity alone: A first. Relevant, but no longer true at tq.(c) With temporal constraints: B first, valid at tq.(a) “Is the patient currently taking [medication]?”A · “[medication] started”B · “[medication] stopped”valid at tqt1t2tq(b) Similarity alone: A first. Relevant, but no longer true at tq.(c) With temporal constraints: B first, valid at tq.(a) “Is the patient currentlytaking [medication]?”A · “started”B · “stopped” · valid at tqt1t2tq(b) Similarity alone: A first.Relevant, no longer true at tq.(c) With temporal constraints:B first, valid at tq.
Fig. 1 Semantic vs temporal validity. Both records are relevant to the question; only one is valid when it is asked. Ranks are order only; no scores. Illustrative example; synthetic patient data, never real patient records.

Question

“How can temporal information be extracted, represented, and integrated into a Retrieval-Augmented Generation framework in order to support more temporally aware question answering over healthcare documents?”

State Engineer thesis, General Introduction.

Semantic relevance is insufficient for time-sensitive healthcare retrieval: this is the work’s first hypothesis. Evidence that is medically related to a query is not necessarily valid for the time the query intends, so ranking by similarity alone can surface temporally inappropriate passages.

Temporal compatibility should complement semantic retrieval rather than replace it. Time is a second signal for choosing evidence, not a filter that works on its own.

Context

Longitudinal records describe a condition as it changes: medications start and stop, symptoms resolve, examinations are planned and later completed. A retrieval system has to preserve relative dates interpreted against the note’s own date, event order, validity periods, whether an event is planned or completed, granularity, and duration.

Granularity is a constraint in its own right. A note that says a diagnosis was made “in 2024” does not support a conclusion about the exact day, and retrieval should not add precision the evidence lacks.

The work uses synthetic patient timelines, never real patient records. Access to real records is limited by ethical, legal and institutional requirements, and synthetic timelines keep the evaluation controlled and auditable.

My contribution

Mine
Implemented the Temporal RAG system end to end, designed and evaluated the temporal-encoder component, designed the evaluation methodology, and integrated the retrieval and temporal components into the working system.
Shared
The overall retrieval architecture was conceptually co-designed with the Télécom SudParis supervisor.
Not mine
The original Java TIMEX stage pre-existed this work; Chamel migrated and adapted it to Python for English/French processing. GPU-dependent implementation or testing carried out by collaborators is not attributed to Chamel.

Approach

Retrieval runs as a patient-scoped pipeline with two scoring branches:

Restrict to the patient. The candidate set is limited to chunks from the question’s patient, so evidence from other patients cannot enter by construction.

Score semantic relevance. A semantic backbone scores each candidate for topical relevance to the question.

Score temporal compatibility. A DCT-aware temporal embedding scores the same candidates. Each chunk is encoded together with its document creation time; the question is encoded from its text.

Normalize per query. Both scores are brought to a common range by min-max normalization within the question’s candidate set.

Combine and rank. A linear blend of the two normalized scores gives a hybrid score, and the top-ranked candidates are selected as evidence.

The system is evaluated at the retrieval level; end-to-end answer generation is outside the scope of this work.

Temporalquestionfor one patientPatient-scopedcandidatesonly thatpatient’s chunksSemantic backbonesemantic scoreDCT-aware temporalembeddingtemporal scoreHybrid scorenormalized per query,then combinedRankedtop-k evidenceTemporal questionfor one patientPatient-scoped candidatesonly that patient’s chunksSemanticbackbonesemantic scoreDCT-awaretemporalembeddingtemporal scoreHybrid scorenormalized per query, then combinedRanked top-k evidence
Fig. 2 Patient-scoped temporal retrieval. Each candidate chunk is scored for semantic relevance and, with its document creation time, for temporal compatibility; the two scores are normalized per query and combined to rank evidence. Retrieval level only; answer generation is not shown. Simplified redraw from the State Engineer thesis, combining the retrieval sequence in §6.4.5 with the two-branch organisation of the evaluation setup (Fig. 16). No scores.

Evidence

The full text documents the work behind this system: the temporally rich corpus and the temporal annotation pipeline, the temporal representation model, and retrieval-level evaluations on public temporal benchmarks and on TempMedBench.

Cite this thesis

Chamel Nadir Bouacha. Building a Temporal Retrieval-Augmented Generation System for Healthcare Applications. State Engineer thesis, ESI Algiers and Télécom SudParis, 2026. DOI: 10.13140/RG.2.2.24676.33924.