Triple

T214297
Position Surface form Disambiguated ID Type / Status
Subject MARC standards E4784 entity
Predicate hasComponent P35 FINISHED
Object UNIMARC
UNIMARC is an international variant of the MARC bibliographic format designed to facilitate the exchange and interpretation of library cataloging data across different countries and systems.
E4784 NE FINISHED

How this triple was built (4 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: UNIMARC | Statement: [MARC standards, hasComponent, UNIMARC]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: UNIMARC
Context triple: [MARC standards, hasComponent, UNIMARC]
  • A. MARC
    MARC is a commuter rail service in Maryland that connects Washington, D.C. with Baltimore and other regional destinations.
  • B. MARC standards
    MARC standards are a set of bibliographic data formats used worldwide to structure and exchange library catalog information in a consistent, machine-readable way.
  • C. BIBFRAME
    BIBFRAME (Bibliographic Framework) is a linked data model and standard developed by the Library of Congress to replace MARC for describing and sharing bibliographic information on the web.
  • D. Z39.50
    Z39.50 is a client-server protocol used primarily by libraries and information services to search and retrieve bibliographic and related data from remote databases in a standardized way.
  • E. Library of Congress Classification
    Library of Congress Classification is a comprehensive alphanumeric library classification system used primarily by academic and research libraries to organize and arrange their collections by subject.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: UNIMARC
Triple: [MARC standards, hasComponent, UNIMARC]
Generated description
UNIMARC is an international variant of the MARC bibliographic format designed to facilitate the exchange and interpretation of library cataloging data across different countries and systems.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: UNIMARC
Target entity description: UNIMARC is an international variant of the MARC bibliographic format designed to facilitate the exchange and interpretation of library cataloging data across different countries and systems.
  • A. MARC
    MARC is a commuter rail service in Maryland that connects Washington, D.C. with Baltimore and other regional destinations.
  • B. MARC standards chosen
    MARC standards are a set of bibliographic data formats used worldwide to structure and exchange library catalog information in a consistent, machine-readable way.
  • C. BIBFRAME
    BIBFRAME (Bibliographic Framework) is a linked data model and standard developed by the Library of Congress to replace MARC for describing and sharing bibliographic information on the web.
  • D. Z39.50
    Z39.50 is a client-server protocol used primarily by libraries and information services to search and retrieve bibliographic and related data from remote databases in a standardized way.
  • E. Library of Congress Classification
    Library of Congress Classification is a comprehensive alphanumeric library classification system used primarily by academic and research libraries to organize and arrange their collections by subject.
  • F. None of above.

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69a2575cb1dc8190a01ad332426dc339 completed Feb. 28, 2026, 2:47 a.m.
NER Named-entity recognition batch_69a25c32ae208190a03d504ef43ea659 completed Feb. 28, 2026, 3:08 a.m.
NED1 Entity disambiguation (via context triple) batch_69a34766506c8190a4c661410e0d8b33 completed Feb. 28, 2026, 7:52 p.m.
NEDg Description generation batch_69a34c38fc2881908afe5d3ef34db98e completed Feb. 28, 2026, 8:12 p.m.
NED2 Entity disambiguation (via description) batch_69a34ca4faec8190bd09fa0e87fe0bd5 completed Feb. 28, 2026, 8:14 p.m.
Created at: Feb. 28, 2026, 2:52 a.m.