Skip to content
Chamel Nadir Bouacha

Research

Relevant is not always valid.

A retrieval system can return evidence that matches a question and still be wrong about when it was true: a medication that was stopped, a symptom that resolved, a procedure that was only planned. My research asks how retrieval can account for when evidence holds, not only what it is about.

(a) “Is the patient currently taking [medication]?”A · “[medication] started”B · “[medication] stopped” · valid at tqt1t2tq(b) Similarity alone: A first. Relevant, but no longer true at tq.(c) With temporal constraints: B first, valid at tq.(a) “Is the patient currentlytaking [medication]?”A · “started”B · “stopped” · valid at tqt1t2tq(b) Similarity alone: A first.Relevant, no longer true at tq.(c) With temporal constraints:B first, valid at tq.
Fig. 1 Semantic vs temporal validity. Both records are relevant to the question; only one is valid when it is asked. Ranks are order only; no scores. Illustrative example; synthetic patient data, never real patient records.

How this came together

  1. Sept. 2025

    Télécom SudParis: Research internship on temporal retrieval-augmented generation for healthcare.

  2. Undated

    Master's thesis: A structured state-of-the-art review of temporal retrieval-augmented generation, covering methods, evaluation, and open challenges. The temporal-validity research gap was identified while writing it.

  3. 22 June 2026

    State Engineer thesis defended: “Building a Temporal Retrieval-Augmented Generation System for Healthcare Applications”.

  4. 2026 · in development

    Developing TempMedBench, a synthetic patient-timeline retrieval benchmark and generation pipeline.

Alongside the thesis: the temporal annotation pipeline. Its annotation output contributed data used for the work that became WaTE.

Current and selected work

Measuring it

  • Benchmark · in development · 2026

    TempMedBench

    A synthetic patient-timeline retrieval benchmark and generation pipeline for temporally faithful healthcare retrieval and reasoning. Structured longitudinal patient timelines are built from synthetic patient records, then used to generate temporally dependent questions, gold evidence relevance judgements, and model-ready note chunks against which retrieval strategies are evaluated.

    The benchmark idea and core contribution are Chamel's. TempMedBench is separate from WaTE and from the temporal annotation pipeline; it grew out of the broader Temporal RAG thesis work.

Retrieving under temporal constraints

  • Thesis research · Télécom SudParis · Sept. 2025 – June 2026

    Temporal RAG for Healthcare

    A retrieval-first architecture that combines semantic relevance with explicit temporal constraints over longitudinal healthcare information, addressing temporal granularity, event-time anchoring, provenance, and retrieval-time consistency.

    Patient-scopedcandidatesSemantic backboneDCT-aware temporalembeddingHybridscoreRankedevidencePatient-scoped candidatesSemanticbackboneDCT-awaretemporalembeddingHybrid scoreRanked evidence
    Fig. 2 Excerpt of the patient-scoped retrieval pipeline: semantic and temporal scores are combined to rank evidence. Simplified from the State Engineer thesis; the full figure is in the case study. No scores.

The data underneath

  • Research software · Télécom SudParis · Sept. 2025 – June 2026

    Temporal annotation pipeline

    A temporal annotation pipeline extended from temporal-expression detection and normalization to event extraction, event-to-time linking, and document-creation-time anchoring. The annotator combines rules, spaCy, and a small language model.

    Held primary responsibility for the pipeline. The TIMEX stage existed previously in Java; migrated it to Python and extended the pipeline. Independently proposed the synthetic temporal-document direction and built the generation pipeline.

Representing time

  • Paper · 2026

    WaTE: Continuous Distributional Temporal Representations for Text Encoders

    Co-author. Primary responsibility for temporal annotation and data construction; contributed tokenizer-related ideas, benchmark selection (TempEval-3, TimeBank-Dense, TS-RETRIEVER), evaluation methodology, ablation studies and controls, and manuscript preparation. Training strategy, tokenizer design, and ablation methodology were joint research decisions; the final pretraining experiments and reported results were carried out by collaborators.

Papers & theses

  • Paper · 2026

    WaTE: Continuous Distributional Temporal Representations for Text Encoders

    Accepted short paper, 38th IEEE International Conference on Tools with Artificial Intelligence (ICTAI 2026).

  • State Engineer thesis · 2026

    Building a Temporal Retrieval-Augmented Generation System for Healthcare Applications

    ESI Algiers and Télécom SudParis, defended 22 June 2026.

  • Master's thesis · 2026

    Exploring Temporal Retrieval-Augmented Generation: Methods, Evaluation, and Open Challenges

    ESI Algiers, defended 22 June 2026.