Triple

T15217866
Position Surface form Disambiguated ID Type / Status
Subject Modified National Institute of Standards and Technology database E363685 entity
Predicate dataSplit P117562 FINISHED
Object training set LITERAL FINISHED

How this triple was built (2 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: training set | Statement: [Modified National Institute of Standards and Technology database, dataSplit, training set]
PD Predicate disambiguation gpt-5-mini-2025-08-07
Target predicate: dataSplit
Context triple: [Modified National Institute of Standards and Technology database, dataSplit, training set]
  • A. developmentSplitWith
    Indicates that a development effort, project, or process is divided or shared between multiple parties or components.
  • B. dividedBetween
    Indicates that something is partitioned or shared among two or more distinct entities or groups.
  • C. data
    Indicates that one entity serves as informational content, measurements, or recorded values associated with another entity.
  • D. trainingDataType
    Indicates the type or category of data used for training a model, system, or process.
  • E. trainsetComposition
    Indicates the relationship specifying how a trainset is composed from its constituent vehicles or units.
  • F. None of above. chosen

Provenance (4 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69d85a0ce24c81909c4d3b6475548c95 completed April 10, 2026, 2:01 a.m.
NER Named-entity recognition batch_69e0076f90c481909989befe031a2cae completed April 15, 2026, 9:47 p.m.
PD Predicate disambiguation batch_69deca8479188190b2e5d3bc708d7d07 completed April 14, 2026, 11:15 p.m.
PDg Predicate description generation batch_69decf2ca6148190967c319728ec3661 completed April 14, 2026, 11:35 p.m.
Created at: April 10, 2026, 3:11 a.m.