Skip to content
Chamel Nadir Bouacha

Temporal annotation pipeline

A temporal annotation pipeline extended from temporal-expression detection and normalization to event extraction, event-to-time linking, and document-creation-time anchoring.

Type
Research software
Role
Research Intern
Institution
Télécom SudParis
Period
Sept. 2025 – June 2026
Existed in Java;migrated to PythonExtension: event extraction,event-to-time linking, DCT anchoringDocumentand DCTtext with itsreference timeTIMEXextraction andnormalizationrule-based taggerEVENTextractionspaCy featuresEvent-timeanchoringlinks to TIMEXor to the DCTCandidatetemporalrelationsordering cuesStructuredtemporal recordfor representationlearningDocument and DCTtext with its reference timeExisted in Java;migrated to PythonTIMEX extraction and normalizationrule-based taggerExtension: event extraction,event-to-time linking, DCT anchoringEVENT extractionspaCy featuresEvent-time anchoringlinks to TIMEX or to the DCTCandidate temporal relationsordering cuesStructured temporal recordfor representation learning
Fig. 1 The temporal annotation pipeline. Each document, with its document creation time (DCT), is enriched stage by stage into a structured temporal record. Brackets mark only attribution explicitly established in the public record; unbracketed stages are not ownership claims. Responsibility for the pipeline as a whole is stated under My contribution. Redrawn from the State Engineer thesis, ch. 3. No scores.

Problem

Finding the dates in a text is not enough to place its evidence in time. A relative expression such as “yesterday” or “next month” cannot be interpreted without a reference time. A date says nothing about what happened unless the events around it are identified. And the link between an event and its time has to be made explicit before a model can learn from it, instead of being left for the model to rediscover from surface tokens.

The pipeline turns documents into structured temporal data for the representation-learning and retrieval work that follows: normalized temporal expressions, events, event-time anchors, and how reliable each of them is.

Context

The corpus combines source families with complementary temporal strengths, in English and French: news, web text, encyclopedic articles, biomedical abstracts and legal documents.

Their document dates are not equally reliable. Publication-grounded sources usually give a strong document creation time (DCT); web and encyclopedic sources often give a weaker one, or none. The pipeline keeps that difference visible: a weak DCT is still usable for normalizing absolute expressions, but anchors that depend on the DCT are treated with caution.

My contribution

Held primary responsibility for the pipeline. The TIMEX stage existed previously in Java; migrated it to Python and extended the pipeline. Independently proposed the synthetic temporal-document direction and built the generation pipeline. The supervising team executed the final large-scale run, which applied the pipeline to approximately 45 million documents, producing more than 1.2 billion TIMEX and event annotations.

Approach

The annotator combines rules, spaCy, and a small language model. The stages in Fig. 1 run in order:

Normalize temporal expressions. A HeidelTime-compatible rule-based temporal tagger detects dates, times, durations and recurring expressions and normalizes each to a TIMEX3 / ISO-8601-style value that keeps the source granularity: a year stays a year and a month stays a month. Relative expressions are resolved against the DCT; vague ones keep markers such as PAST_REF rather than gaining false precision. This TIMEX stage existed previously in Java and was migrated to Python.

Extract events. spaCy linguistic features (part of speech, lemmas, dependency relations) and heuristic trigger filters identify event mentions: mainly verbal predicates and, more conservatively, eventive nominalizations. Each event keeps its negation, modality and aspect, because a negated, a possible and an asserted event make different temporal commitments.

Anchor events in time. A deterministic priority cascade attaches each event to the strongest evidence available: a normalized date linked through the dependency structure, a directional cue such as “before” or “after”, a label relative to the DCT, or an explicit UNDERSPECIFIED when nothing reliable supports an anchor. The method that produced the anchor is stored with it.

Keep candidate relations. TLINK-like relations derived from syntactic patterns preserve further ordering cues between events, temporal expressions and the DCT. They are structural evidence, not validated temporal links, and they upgrade a vague anchor only when they carry stronger evidence.

Export structured records. Each document becomes a record of its normalized expressions, events, anchors and relation cues, with provenance and reliability kept alongside, ready for representation learning.

Evidence

Chapter 3 of the State Engineer thesis documents the pipeline: its design requirements, DCT handling, each annotation stage and the exported record format. The annotator code is not public.