Triple

T10880233
Position Surface form Disambiguated ID Type / Status
Subject Tupian languages E256900 entity
Predicate includesLanguage P2177 FINISHED
Object Tembe language
The Tembe language is an indigenous Tupi-Guarani language spoken by the Tembé people of northern Brazil.
E890399 NE FINISHED

How this triple was built (4 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Tembe language | Statement: [Tupian languages, includesLanguage, Tembe language]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Tembe language
Context triple: [Tupian languages, includesLanguage, Tembe language]
  • A. Kalanga language
    The Kalanga language is a Bantu language spoken primarily by the Kalanga people in parts of Botswana and southwestern Zimbabwe.
  • B. Sanglechi language
    The Sanglechi language is an Eastern Iranian language spoken by a small community in the Sanglech Valley region of Afghanistan and Tajikistan.
  • C. Murle language
    The Murle language is an Eastern Sudanic language spoken primarily by the Murle people of South Sudan.
  • D. Marakwet language
    The Marakwet language is a Southern Nilotic language spoken by the Marakwet people of Kenya and is closely related to other Kalenjin languages such as Kipsigis.
  • E. Teke-Ngungwel language
    The Teke-Ngungwel language is a Bantu language spoken by the Teke people in Central Africa, particularly in parts of the Republic of the Congo and neighboring regions.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: Tembe language
Triple: [Tupian languages, includesLanguage, Tembe language]
Generated description
The Tembe language is an indigenous Tupi-Guarani language spoken by the Tembé people of northern Brazil.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: Tembe language
Target entity description: The Tembe language is an indigenous Tupi-Guarani language spoken by the Tembé people of northern Brazil.
  • A. Kalanga language
    The Kalanga language is a Bantu language spoken primarily by the Kalanga people in parts of Botswana and southwestern Zimbabwe.
  • B. Sanglechi language
    The Sanglechi language is an Eastern Iranian language spoken by a small community in the Sanglech Valley region of Afghanistan and Tajikistan.
  • C. Murle language
    The Murle language is an Eastern Sudanic language spoken primarily by the Murle people of South Sudan.
  • D. Marakwet language
    The Marakwet language is a Southern Nilotic language spoken by the Marakwet people of Kenya and is closely related to other Kalenjin languages such as Kipsigis.
  • E. Teke-Ngungwel language
    The Teke-Ngungwel language is a Bantu language spoken by the Teke people in Central Africa, particularly in parts of the Republic of the Congo and neighboring regions.
  • F. None of above. chosen

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69d6aa848804819081b2713ca0bedf06 completed April 8, 2026, 7:20 p.m.
NER Named-entity recognition batch_69d751b031a88190b1182dfc1f520264 completed April 9, 2026, 7:13 a.m.
NED1 Entity disambiguation (via context triple) batch_69dff7e2322c8190a55605237ae6ce95 completed April 15, 2026, 8:41 p.m.
NEDg Description generation batch_69e002709d38819099c4402d30824612 completed April 15, 2026, 9:26 p.m.
NED2 Entity disambiguation (via description) batch_69e005873ba48190b8c24c77611562fa completed April 15, 2026, 9:39 p.m.
Created at: April 8, 2026, 9:21 p.m.