BookCorpus
E1312410
UNEXPLORED
BookCorpus is a large collection of freely available books commonly used as a pretraining dataset for natural language processing models.
All labels observed (3)
| Label | Occurrences |
|---|---|
| BookCorpus canonical | 1 |
| BookCorpus (for some variants) | 1 |
| BooksCorpus | 1 |
How this entity was disambiguated
This entity first appeared as the object of triple T18204243 — resolving that mention is where its identity was fixed. The disambiguator weighed these candidate entities and picked the highlighted one (or “None”, minting a new entity). This is how homonymy is resolved: the same surface form can point to different entities.
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: BookCorpus Context triple: [RoBERTa, trainingDataSource, BookCorpus]
-
A.
WebText dataset
The WebText dataset is a large-scale corpus of web pages curated by OpenAI to train language models like GPT-2 on diverse, high-quality internet text.
-
B.
Corpus
Corpus is a common shortened name for Corpus Christi College, one of the historic constituent colleges of the University of Cambridge.
-
C.
Collins Corpus
Collins Corpus is a large, computer-readable collection of real-world English texts used by Collins for corpus-based lexicography and language research.
-
D.
CORDE corpus
The CORDE corpus is a large historical Spanish language corpus compiled by the Royal Spanish Academy, used for studying the evolution and usage of Spanish over time.
-
E.
Waverley Novels corpus
The Waverley Novels corpus is the collective body of historical novels by Sir Walter Scott that helped establish the genre and shaped 19th-century historical fiction.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: BookCorpus Target entity description: BookCorpus is a large collection of freely available books commonly used as a pretraining dataset for natural language processing models.
-
A.
WebText dataset
The WebText dataset is a large-scale corpus of web pages curated by OpenAI to train language models like GPT-2 on diverse, high-quality internet text.
-
B.
Corpus
Corpus is a common shortened name for Corpus Christi College, one of the historic constituent colleges of the University of Cambridge.
-
C.
Collins Corpus
Collins Corpus is a large, computer-readable collection of real-world English texts used by Collins for corpus-based lexicography and language research.
-
D.
CORDE corpus
The CORDE corpus is a large historical Spanish language corpus compiled by the Royal Spanish Academy, used for studying the evolution and usage of Spanish over time.
-
E.
Waverley Novels corpus
The Waverley Novels corpus is the collective body of historical novels by Sir Walter Scott that helped establish the genre and shaped 19th-century historical fiction.
- F. None of above. chosen
Referenced by (3)
Full triples — surface form annotated when it differs from this entity's canonical label.
linked to: BookCorpus
linked to: BookCorpus