Triple

T1521043
Position Surface form Disambiguated ID Type / Status
Subject Tetum E32226 entity
Predicate macrolanguageOf P23734 FINISHED
Object Tetun Terik E148730 NE FINISHED

How this triple was built (2 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Tetun Terik | Statement: [Tetum, macrolanguageOf, Tetun Terik]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Tetun Terik
Context triple: [Tetum, macrolanguageOf, Tetun Terik]
  • A. Tetun language chosen
    The Tetun language is an Austronesian language spoken primarily in Timor-Leste, where it serves as one of the main national and official languages.
  • B. Blablanga
    Blablanga is an Austronesian language spoken in the Solomon Islands, belonging to the Meso-Melanesian subgroup.
  • C. Kapingamarangi language
    The Kapingamarangi language is a Polynesian outlier language spoken primarily on Kapingamarangi Atoll in the Federated States of Micronesia.
  • D. Tontemboan language
    The Tontemboan language is an Austronesian language spoken by the Tontemboan people of North Sulawesi, Indonesia, and is one of the traditional Minahasan languages of the region.
  • E. Lamaholot language
    The Lamaholot language is an Austronesian language spoken primarily in eastern Flores and nearby islands in Indonesia, known for its numerous dialects and complex verbal morphology.
  • F. None of above.
  • G. Unsure - the case is ambiguous/there is not enough information to decide.

Provenance (3 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69a885e9b0ac819093a9806ad0efc82c completed March 4, 2026, 7:20 p.m.
NER Named-entity recognition batch_69a907f071848190a5fb8fa1b97ef4de completed March 5, 2026, 4:34 a.m.
NED1 Entity disambiguation (via context triple) batch_69ad51a8b8b88190a320071964846518 completed March 8, 2026, 10:38 a.m.
Created at: March 4, 2026, 7:26 p.m.