Triple

T21901600
Position Surface form Disambiguated ID Type / Status
Subject Betsileo Malagasy E540820 entity
Predicate belongsTo P35 FINISHED
Object Malagasy dialect continuum NE NERFINISHED

How this triple was built (3 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Malagasy dialect continuum | Statement: [Betsileo Malagasy, belongsTo, Malagasy dialect continuum]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Malagasy dialect continuum
Context triple: [Betsileo Malagasy, belongsTo, Malagasy dialect continuum]
  • A. Mono language continuum
    The Mono language continuum is a group of closely related Native American languages traditionally spoken by the Mono people of central California, encompassing varieties such as Western Mono.
  • B. Tat dialect continuum
    The Tat dialect continuum is a group of closely related Southwestern Iranian dialects spoken primarily in the eastern Caucasus region, notably in parts of Azerbaijan and Russia.
  • C. Southwest Indian Ocean linguistic area
    The Southwest Indian Ocean linguistic area is a region encompassing parts of Madagascar and nearby islands where diverse languages, including Bantu and Austronesian varieties, have converged and influenced one another through prolonged contact.
  • D. Tabasaran dialect continuum
    The Tabasaran dialect continuum is a group of closely related Northeast Caucasian speech varieties of the Tabasaran language, spoken primarily in southern Dagestan, Russia.
  • E. Fala dialect continuum
    The Fala dialect continuum is a group of closely related Romance varieties spoken in a small area of Extremadura in western Spain, often considered transitional between Galician-Portuguese and Spanish.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: Malagasy dialect continuum
Target entity description: The Malagasy dialect continuum is a group of closely related Austronesian dialects spoken across Madagascar that vary regionally yet remain largely mutually intelligible.
  • A. Mono language continuum
    The Mono language continuum is a group of closely related Native American languages traditionally spoken by the Mono people of central California, encompassing varieties such as Western Mono.
  • B. Tat dialect continuum
    The Tat dialect continuum is a group of closely related Southwestern Iranian dialects spoken primarily in the eastern Caucasus region, notably in parts of Azerbaijan and Russia.
  • C. Southwest Indian Ocean linguistic area
    The Southwest Indian Ocean linguistic area is a region encompassing parts of Madagascar and nearby islands where diverse languages, including Bantu and Austronesian varieties, have converged and influenced one another through prolonged contact.
  • D. Tabasaran dialect continuum
    The Tabasaran dialect continuum is a group of closely related Northeast Caucasian speech varieties of the Tabasaran language, spoken primarily in southern Dagestan, Russia.
  • E. Fala dialect continuum
    The Fala dialect continuum is a group of closely related Romance varieties spoken in a small area of Extremadura in western Spain, often considered transitional between Galician-Portuguese and Spanish.
  • F. None of above. chosen

Provenance (2 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69e0c47b4e8c81908c8076eaa4c8e4f2 completed April 16, 2026, 11:14 a.m.
NER Named-entity recognition batch_69f11fcb18748190a21071c122b7e6d5 completed April 28, 2026, 8:59 p.m.
Created at: April 16, 2026, 7:16 p.m.