Temporal annotation pipeline
A temporal annotation pipeline extended from temporal-expression detection and normalization to event extraction, event-to-time linking, and document-creation-time anchoring.
Problem
Finding the dates in a text is not enough to place its evidence in time. A relative expression such as “yesterday” or “next month” cannot be interpreted without a reference time. A date says nothing about what happened unless the events around it are identified. And the link between an event and its time has to be made explicit before a model can learn from it, instead of being left for the model to rediscover from surface tokens.
The pipeline turns documents into structured temporal data for the representation-learning and retrieval work that follows: normalized temporal expressions, events, event-time anchors, and how reliable each of them is.
Context
The corpus combines source families with complementary temporal strengths, in English and French: news, web text, encyclopedic articles, biomedical abstracts and legal documents.
Their document dates are not equally reliable. Publication-grounded sources usually give a strong document creation time (DCT); web and encyclopedic sources often give a weaker one, or none. The pipeline keeps that difference visible: a weak DCT is still usable for normalizing absolute expressions, but anchors that depend on the DCT are treated with caution.
My contribution
Held primary responsibility for the pipeline. The TIMEX stage existed previously in Java; migrated it to Python and extended the pipeline. Independently proposed the synthetic temporal-document direction and built the generation pipeline. The supervising team executed the final large-scale run, which applied the pipeline to approximately 45 million documents, producing more than 1.2 billion TIMEX and event annotations.
Approach
The annotator combines rules, spaCy, and a small language model. The stages in Fig. 1 run in order:
Normalize temporal expressions. A HeidelTime-compatible rule-based temporal tagger detects dates, times, durations and recurring expressions and normalizes each to a TIMEX3 / ISO-8601-style value that keeps the source granularity: a year stays a year and a month stays a month. Relative expressions are resolved against the DCT; vague ones keep markers such as PAST_REF rather than gaining false precision. This TIMEX stage existed previously in Java and was migrated to Python.
Extract events. spaCy linguistic features (part of speech, lemmas, dependency relations) and heuristic trigger filters identify event mentions: mainly verbal predicates and, more conservatively, eventive nominalizations. Each event keeps its negation, modality and aspect, because a negated, a possible and an asserted event make different temporal commitments.
Anchor events in time. A deterministic priority cascade attaches each event to the strongest evidence available: a normalized date linked through the dependency structure, a directional cue such as “before” or “after”, a label relative to the DCT, or an explicit UNDERSPECIFIED when nothing reliable supports an anchor. The method that produced the anchor is stored with it.
Keep candidate relations. TLINK-like relations derived from syntactic patterns preserve further ordering cues between events, temporal expressions and the DCT. They are structural evidence, not validated temporal links, and they upgrade a vague anchor only when they carry stronger evidence.
Export structured records. Each document becomes a record of its normalized expressions, events, anchors and relation cues, with provenance and reliability kept alongside, ready for representation learning.
Evidence
Chapter 3 of the State Engineer thesis documents the pipeline: its design requirements, DCT handling, each annotation stage and the exported record format. The annotator code is not public.
Related
- Temporal RAG for Healthcare · retrieving under temporal constraints
- WaTE · representing time
- TempMedBench · measuring it
- All research · the research direction and papers & theses