Triple

T32614790
Position Surface form Disambiguated ID Type / Status
Subject Hakka grammars E833753 entity
Predicate focusesOn P31 FINISHED
Object Mainland China Hakka
Mainland China Hakka is a group of Hakka Chinese dialects spoken across various regions of mainland China, known for their distinct phonology, vocabulary, and grammar within the broader Sinitic language family.
E2018887 NE FINISHED

How this triple was built (2 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Mainland China Hakka | Statement: [Hakka grammars, focusesOn, Mainland China Hakka]
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: Mainland China Hakka
Triple: [Hakka grammars, focusesOn, Mainland China Hakka]
Generated description
Mainland China Hakka is a group of Hakka Chinese dialects spoken across various regions of mainland China, known for their distinct phonology, vocabulary, and grammar within the broader Sinitic language family.

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69f3492bfa648190b6ae472074634e29 completed April 30, 2026, 12:21 p.m.
NER Named-entity recognition batch_69f6c6e9d5d48190a3d85678d08ada40 completed May 3, 2026, 3:54 a.m.
NED1 Entity disambiguation (via context triple) batch_6a349eaae3e88190975e0e53906008ce completed June 19, 2026, 1:43 a.m.
NEDg Description generation batch_6a349f7f7c1c81908cf908084632c815 completed June 19, 2026, 1:46 a.m.
NED2 Entity disambiguation (via description) batch_6a34a0220314819091450074e783e841 completed June 19, 2026, 1:49 a.m.
Created at: May 1, 2026, 1:06 a.m.