Triple

T212287
Position Surface form Disambiguated ID Type / Status
Subject Perso-Arabic script E4744 entity
Predicate usedFor P98 FINISHED
Object Uyghur language (historically)
The historical Uyghur language refers to earlier Turkic varieties used by the Uyghur people, which were written in scripts such as Perso-Arabic and played a key role in the literary and cultural history of Central Asia.
E27175 NE FINISHED

How this triple was built (4 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Uyghur language (historically) | Statement: [Perso-Arabic script, usedFor, Uyghur language (historically)]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Uyghur language (historically)
Context triple: [Perso-Arabic script, usedFor, Uyghur language (historically)]
  • A. Tocharian languages
    The Tocharian languages were an extinct branch of the Indo-European family once spoken in the Tarim Basin of Central Asia, known from early medieval manuscripts and notable for their archaic linguistic features.
  • B. Turkmen language
    The Turkmen language is a Turkic language spoken primarily in Turkmenistan and surrounding regions, closely related to Turkish and other Oghuz languages.
  • C. Southwestern Turkic
    Southwestern Turkic is a major branch of the Turkic language family that includes languages such as Turkish, Azerbaijani, and Turkmen, primarily spoken across Anatolia, the Caucasus, and parts of Central and Western Asia.
  • D. Kazakh language
    The Kazakh language is a Turkic language spoken primarily in Kazakhstan and surrounding regions, written today mainly in Cyrillic but also in Latin and Arabic scripts.
  • E. Utian languages
    The Utian languages are a small group of Native American languages once spoken in central California, traditionally including the Miwok and Costanoan (Ohlone) language branches.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: Uyghur language (historically)
Triple: [Perso-Arabic script, usedFor, Uyghur language (historically)]
Generated description
The historical Uyghur language refers to earlier Turkic varieties used by the Uyghur people, which were written in scripts such as Perso-Arabic and played a key role in the literary and cultural history of Central Asia.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: Uyghur language (historically)
Target entity description: The historical Uyghur language refers to earlier Turkic varieties used by the Uyghur people, which were written in scripts such as Perso-Arabic and played a key role in the literary and cultural history of Central Asia.
  • A. Tocharian languages
    The Tocharian languages were an extinct branch of the Indo-European family once spoken in the Tarim Basin of Central Asia, known from early medieval manuscripts and notable for their archaic linguistic features.
  • B. Turkmen language
    The Turkmen language is a Turkic language spoken primarily in Turkmenistan and surrounding regions, closely related to Turkish and other Oghuz languages.
  • C. Southwestern Turkic
    Southwestern Turkic is a major branch of the Turkic language family that includes languages such as Turkish, Azerbaijani, and Turkmen, primarily spoken across Anatolia, the Caucasus, and parts of Central and Western Asia.
  • D. Kazakh language
    The Kazakh language is a Turkic language spoken primarily in Kazakhstan and surrounding regions, written today mainly in Cyrillic but also in Latin and Arabic scripts.
  • E. Utian languages
    The Utian languages are a small group of Native American languages once spoken in central California, traditionally including the Miwok and Costanoan (Ohlone) language branches.
  • F. None of above. chosen

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69a2575cb1dc8190a01ad332426dc339 completed Feb. 28, 2026, 2:47 a.m.
NER Named-entity recognition batch_69a25c313d108190a65d3e939f961bef completed Feb. 28, 2026, 3:08 a.m.
NED1 Entity disambiguation (via context triple) batch_69a338e8432081909c1b924250751022 completed Feb. 28, 2026, 6:50 p.m.
NEDg Description generation batch_69a339955c1081909ab38da591b836cf completed Feb. 28, 2026, 6:53 p.m.
NED2 Entity disambiguation (via description) batch_69a33a1d394c8190a9e1d686e20f1ade completed Feb. 28, 2026, 6:55 p.m.
Created at: Feb. 28, 2026, 2:52 a.m.