Triple

T4651255
Position Surface form Disambiguated ID Type / Status
Subject Common Crawl E102298 entity
Predicate distributionPlatform P1486 FINISHED
Object AWS Open Data Sponsorship Program E399063 NE FINISHED

How this triple was built (2 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: AWS Open Data Sponsorship Program | Statement: [Common Crawl, distributionPlatform, AWS Open Data Sponsorship Program]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: AWS Open Data Sponsorship Program
Context triple: [Common Crawl, distributionPlatform, AWS Open Data Sponsorship Program]
  • A. Amazon Web Services Open Data chosen
    Amazon Web Services Open Data is a program that hosts and freely provides access to large, publicly available datasets—such as satellite imagery, genomics, and climate data—on AWS cloud infrastructure for research, innovation, and application development.
  • B. Open Data Lab
    Open Data Lab is a World Wide Web Foundation initiative that supports the use of open data to drive social impact, innovation, and better governance, particularly in developing countries.
  • C. Amazon SageMaker
    Amazon SageMaker is a fully managed cloud service that enables developers and data scientists to build, train, and deploy machine learning models at scale.
  • D. Amazon Athena
    Amazon Athena is a serverless, interactive query service from AWS that lets users analyze data directly in Amazon S3 using standard SQL.
  • E. Open Data Institute
    The Open Data Institute is a UK-based non-profit organization that advocates for and supports the use of open data to drive innovation, improve governance, and benefit society.
  • F. None of above.
  • G. Unsure - the case is ambiguous/there is not enough information to decide.

Provenance (3 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69bd43d71a308190afea7280841b0de8 completed March 20, 2026, 12:55 p.m.
NER Named-entity recognition batch_69bd630343f88190954d19fcd18a5864 completed March 20, 2026, 3:08 p.m.
NED1 Entity disambiguation (via context triple) batch_69bdfae7636881908244b86cba1c66b7 completed March 21, 2026, 1:56 a.m.
Created at: March 20, 2026, 1:14 p.m.