Triple
T5361609
| Position | Surface form | Disambiguated ID | Type / Status |
|---|---|---|---|
| Subject | Northeastern Turkic |
E103033
|
entity |
| Predicate | hasMember |
P10
|
FINISHED |
| Object |
Tofa language
Tofa is a critically endangered Turkic language traditionally spoken by the Tofa (Tofalars) people of south-central Siberia in Russia.
|
E514630
|
NE FINISHED |
How this triple was built (4 steps)
Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.
NER
Named-entity recognition
gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Tofa language | Statement: [Northeastern Turkic, hasMember, Tofa language]
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: Tofa language Context triple: [Northeastern Turkic, hasMember, Tofa language]
-
A.
Towa language
Towa is a Native American language spoken by the Towa (Jemez) people of New Mexico and is part of the Puebloan language family.
-
B.
Tobian language
The Tobian language is a Micronesian language spoken primarily on Tobi Island in Palau, known for its small speaker population and close relation to other Carolinean languages.
-
C.
Defaka language
The Defaka language is a highly endangered Niger-Congo language spoken by a small community in Nigeria’s Niger Delta region.
-
D.
Tokodede language
Tokodede is an Austronesian language spoken primarily in the Liquiçá region of northwestern East Timor.
-
E.
Gofa language
The Gofa language is an Omotic language of southwestern Ethiopia spoken by the Gofa people and closely related to Wolaytta.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg
Description generation
gpt-5.1
Instruction
Generate a one-sentence description of the target entity. You are given a context triple in the form (subject, predicate, object), where the object is the target entity. # Instructions Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. Avoid repeating the information from the triple, unless really essential. # Response Format Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: Tofa language Triple: [Northeastern Turkic, hasMember, Tofa language]
Generated description
Tofa is a critically endangered Turkic language traditionally spoken by the Tofa (Tofalars) people of south-central Siberia in Russia.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: Tofa language Target entity description: Tofa is a critically endangered Turkic language traditionally spoken by the Tofa (Tofalars) people of south-central Siberia in Russia.
-
A.
Towa language
Towa is a Native American language spoken by the Towa (Jemez) people of New Mexico and is part of the Puebloan language family.
-
B.
Tobian language
The Tobian language is a Micronesian language spoken primarily on Tobi Island in Palau, known for its small speaker population and close relation to other Carolinean languages.
-
C.
Defaka language
The Defaka language is a highly endangered Niger-Congo language spoken by a small community in Nigeria’s Niger Delta region.
-
D.
Tokodede language
Tokodede is an Austronesian language spoken primarily in the Liquiçá region of northwestern East Timor.
-
E.
Gofa language
The Gofa language is an Omotic language of southwestern Ethiopia spoken by the Gofa people and closely related to Wolaytta.
- F. None of above. chosen
Provenance (5 batches)
The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.
| Step | Stage | Batch ID | Status | When |
|---|---|---|---|---|
| creating | Elicitation | batch_69bd43daa3e4819090b59d127db70e57 |
completed | March 20, 2026, 12:55 p.m. |
| NER | Named-entity recognition | batch_69bd865a0bb081909579cfe7c7974075 |
completed | March 20, 2026, 5:39 p.m. |
| NED1 | Entity disambiguation (via context triple) | batch_69bf21ef7acc8190aa7ff990a04ff09c |
completed | March 21, 2026, 10:55 p.m. |
| NEDg | Description generation | batch_69bf25467ad481909a24fb4cb2953af6 |
completed | March 21, 2026, 11:09 p.m. |
| NED2 | Entity disambiguation (via description) | batch_69bf25c28ad48190ad7f0e5cd9487a08 |
completed | March 21, 2026, 11:12 p.m. |
Created at: March 20, 2026, 2:02 p.m.