Triple

T35915733
Position Surface form Disambiguated ID Type / Status
Subject Matei Zaharia E1038743 entity
Predicate notableWork P4 FINISHED
Object Resilient Distributed Datasets (RDDs)
Resilient Distributed Datasets (RDDs) are Apache Spark’s fundamental fault-tolerant, distributed data abstraction that enables efficient in-memory cluster computing for large-scale data processing.
E2162386 NE FINISHED

How this triple was built (2 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Resilient Distributed Datasets (RDDs) | Statement: [Matei Zaharia, notableWork, Resilient Distributed Datasets (RDDs)]
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: Resilient Distributed Datasets (RDDs)
Triple: [Matei Zaharia, notableWork, Resilient Distributed Datasets (RDDs)]
Generated description
Resilient Distributed Datasets (RDDs) are Apache Spark’s fundamental fault-tolerant, distributed data abstraction that enables efficient in-memory cluster computing for large-scale data processing.

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69f76e2320748190b7f5c4750d0cd0d3 completed May 3, 2026, 3:47 p.m.
NER Named-entity recognition batch_69f7aaa516b0819094cb3f025f128eb9 completed May 3, 2026, 8:05 p.m.
NED1 Entity disambiguation (via context triple) batch_6a38ae3168d081908a7224fb590d66f5 completed June 22, 2026, 3:38 a.m.
NEDg Description generation batch_6a38b1b72c2081909885dd9f39656fb5 completed June 22, 2026, 3:53 a.m.
NED2 Entity disambiguation (via description) batch_6a38b24aab1881909d587cf945d0390b completed June 22, 2026, 3:55 a.m.
Created at: May 3, 2026, 4:07 p.m.