COCO captioning challenge
E1300977
UNEXPLORED
The COCO captioning challenge is a computer vision and natural language processing competition where systems generate descriptive text captions for images from the COCO dataset.
All labels observed (3)
| Label | Occurrences |
|---|---|
| COCO Captioning Challenge | 1 |
| COCO captioning challenge canonical | 1 |
| MS COCO Captioning Challenge | 1 |
How this entity was disambiguated
This entity first appeared as the object of triple T18016053 — resolving that mention is where its identity was fixed. The disambiguator weighed these candidate entities and picked the highlighted one (or “None”, minting a new entity). This is how homonymy is resolved: the same surface form can point to different entities.
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: COCO captioning challenge Context triple: [COCO, associatedWith, COCO captioning challenge]
-
A.
CLIP
CLIP is an OpenAI model that learns joint representations of images and text, enabling tasks like zero-shot image classification and natural language-based image retrieval.
-
B.
Flickr30k
Flickr30k is a large-scale image dataset of 31,000 photographs each paired with multiple human-written captions, widely used for training and evaluating image captioning and vision-language models.
-
C.
Show and Tell: A Neural Image Caption Generator
"Show and Tell: A Neural Image Caption Generator" is a pioneering deep learning model that automatically generates natural-language descriptions for images by combining convolutional and recurrent neural networks.
-
D.
Images and Words
Images and Words is a landmark 1992 progressive metal album by Dream Theater, widely credited with bringing the band mainstream recognition and defining their signature sound.
-
E.
Flickr8k
Flickr8k is a benchmark image-captioning dataset consisting of 8,000 images each paired with multiple human-written descriptions, widely used for training and evaluating vision-language models.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: COCO captioning challenge Target entity description: The COCO captioning challenge is a computer vision and natural language processing competition where systems generate descriptive text captions for images from the COCO dataset.
-
A.
CLIP
CLIP is an OpenAI model that learns joint representations of images and text, enabling tasks like zero-shot image classification and natural language-based image retrieval.
-
B.
Flickr30k
Flickr30k is a large-scale image dataset of 31,000 photographs each paired with multiple human-written captions, widely used for training and evaluating image captioning and vision-language models.
-
C.
Show and Tell: A Neural Image Caption Generator
"Show and Tell: A Neural Image Caption Generator" is a pioneering deep learning model that automatically generates natural-language descriptions for images by combining convolutional and recurrent neural networks.
-
D.
Images and Words
Images and Words is a landmark 1992 progressive metal album by Dream Theater, widely credited with bringing the band mainstream recognition and defining their signature sound.
-
E.
Flickr8k
Flickr8k is a benchmark image-captioning dataset consisting of 8,000 images each paired with multiple human-written descriptions, widely used for training and evaluating vision-language models.
- F. None of above. chosen
Referenced by (3)
Full triples — surface form annotated when it differs from this entity's canonical label.