Flickr8k

E899058

Flickr8k is a benchmark image-captioning dataset consisting of 8,000 images each paired with multiple human-written descriptions, widely used for training and evaluating vision-language models.

All labels observed (2)

Label Occurrences
Flickr8k canonical 4
Flickr8k dataset 1

How this entity was disambiguated

Statements (46)

Predicate Object
instanceOf benchmark dataset
image-captioning dataset
vision-language dataset
hasAnnotationType image captions
natural language descriptions
hasApproximateNumberOfImages 8000
hasBenchmarkRole baseline dataset for image captioning
hasCaptionSource human annotators
hasCaptionsPerImage 5
hasDataModality images
text captions
hasDataType photographic images
hasDescriptionQuality human-written captions
hasDomain computer vision
natural language processing
hasEvaluationMetrics BLEU
CIDEr
METEOR
ROUGE
linked to: ROUGE-L
hasImageSource Flickr
hasLanguage English
hasLicenseType research use
hasName Flickr8k
hasNumberOfImages 8000
hasScale small-scale image-captioning dataset
hasSource Flickr
hasTask multimodal learning
vision-language grounding
hasTypicalSplit test set
training set
validation set
isBenchmarkFor automatic image description
multimodal representation learning
isConsidered standard benchmark in image captioning
isSmallerThan Flickr30k
MS COCO
linked to: COCO
isUsedIn vision-and-language research
isUsedTo compare image-captioning algorithms
evaluate caption generation quality
isWidelyUsedFor benchmarking captioning models
isWidelyUsedIn academic research
usedFor evaluating image-captioning models
evaluating vision-language models
image captioning
training image-captioning models
training vision-language models

How these facts were elicited

Referenced by (5)

Full triples — surface form annotated when it differs from this entity's canonical label.

Show and Tell trainedOn Flickr8k dataset
linked to: Flickr8k
Flickr8k hasName Flickr8k
Flickr30k isSuccessorOf Flickr8k