Colossal Clean Crawled Corpus
E1312424
UNEXPLORED
The Colossal Clean Crawled Corpus (C4) is a massive, cleaned web-text dataset widely used to train large language models and other state-of-the-art NLP systems.
All labels observed (1)
| Label | Occurrences |
|---|---|
| Colossal Clean Crawled Corpus canonical | 1 |
How this entity was disambiguated
This entity first appeared as the object of triple T18204408 — resolving that mention is where its identity was fixed. The disambiguator weighed these candidate entities and picked the highlighted one (or “None”, minting a new entity). This is how homonymy is resolved: the same surface form can point to different entities.
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: Colossal Clean Crawled Corpus Context triple: [T5, trainingData, Colossal Clean Crawled Corpus]
-
A.
Common Crawl
Common Crawl is a massive, publicly available web archive that regularly crawls and stores petabytes of web page data for use in research and large-scale data analysis.
-
B.
Collins Corpus
Collins Corpus is a large, computer-readable collection of real-world English texts used by Collins for corpus-based lexicography and language research.
-
C.
Corpus
Corpus is a common shortened name for Corpus Christi College, one of the historic constituent colleges of the University of Cambridge.
-
D.
COSMAS II corpus search system
COSMAS II corpus search system is a large-scale linguistic search platform for German language text corpora, maintained by the Institut für Deutsche Sprache for research and lexicographic analysis.
-
E.
WebText dataset
The WebText dataset is a large-scale corpus of web pages curated by OpenAI to train language models like GPT-2 on diverse, high-quality internet text.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: Colossal Clean Crawled Corpus Target entity description: The Colossal Clean Crawled Corpus (C4) is a massive, cleaned web-text dataset widely used to train large language models and other state-of-the-art NLP systems.
-
A.
Common Crawl
Common Crawl is a massive, publicly available web archive that regularly crawls and stores petabytes of web page data for use in research and large-scale data analysis.
-
B.
Collins Corpus
Collins Corpus is a large, computer-readable collection of real-world English texts used by Collins for corpus-based lexicography and language research.
-
C.
Corpus
Corpus is a common shortened name for Corpus Christi College, one of the historic constituent colleges of the University of Cambridge.
-
D.
COSMAS II corpus search system
COSMAS II corpus search system is a large-scale linguistic search platform for German language text corpora, maintained by the Institut für Deutsche Sprache for research and lexicographic analysis.
-
E.
WebText dataset
The WebText dataset is a large-scale corpus of web pages curated by OpenAI to train language models like GPT-2 on diverse, high-quality internet text.
- F. None of above. chosen
Referenced by (1)
Full triples — surface form annotated when it differs from this entity's canonical label.