Triple

T2959337
Position Surface form Disambiguated ID Type / Status
Subject Khetrani E80006 entity
Predicate codeSwitchingWith P41297 FINISHED
Object Urdu E6054 NE FINISHED

How this triple was built (3 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Urdu | Statement: [Khetrani, codeSwitchingWith, Urdu]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Urdu
Context triple: [Khetrani, codeSwitchingWith, Urdu]
  • A. Urdu language chosen
    Urdu is a major South Asian language, written in a Perso-Arabic script and widely used in Pakistan and parts of India in literature, media, and everyday communication.
  • B. Sindhi
    Sindhi is an Indo-Aryan language spoken primarily in Pakistan and India, known for its rich literary tradition and distinct script variants.
  • C. Saraiki
    Saraiki is an Indo-Aryan language spoken primarily in central and southern Pakistan, especially in the southern Punjab region.
  • D. Balochi
    Balochi is an Iranian language spoken primarily by the Baloch people across Pakistan, Iran, and Afghanistan, with several dialects and a rich oral literary tradition.
  • E. Punjabi
    Punjabi refers to the ethnolinguistic group native to the Punjab region of South Asia, known for its distinct language, culture, and traditions shared across parts of India and Pakistan.
  • F. None of above.
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
PD Predicate disambiguation gpt-5-mini-2025-08-07
Target predicate: codeSwitchingWith
Context triple: [Khetrani, codeSwitchingWith, Urdu]
  • A. languageShift
    Indicates a change in the primary language used by an entity, such as switching from one language to another over time or in a given context.
  • B. languagePair
    Indicates a relationship that associates two specific languages as a paired combination, typically for translation, comparison, or mapping between them.
  • C. coexistsWithLanguage chosen
    Indicates that one entity exists or functions alongside a particular language at the same time, without excluding or replacing it.
  • D. languageBranch
    Indicates that one language belongs to, or is classified under, a broader linguistic branch or subgroup.
  • E. usedAsContactLanguageBetween
    Indicates that a language functions as the medium of communication between two or more distinct language communities.
  • F. None of above.

Provenance (4 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69ad8b1341848190bd19dbf46892887d completed March 8, 2026, 2:43 p.m.
NER Named-entity recognition batch_69ad992c4c7c819084b5bef299255181 completed March 8, 2026, 3:43 p.m.
NED1 Entity disambiguation (via context triple) batch_69b108de7e3c81908b1fee310515e4c1 completed March 11, 2026, 6:17 a.m.
PD Predicate disambiguation batch_69ad960c5c8881909d679912bd7d78f3 completed March 8, 2026, 3:30 p.m.
Created at: March 8, 2026, 2:57 p.m.