Language Models are Few-Shot Learners
E457860
"Language Models are Few-Shot Learners" is a landmark research paper that demonstrated large-scale transformer-based language models can perform diverse tasks from just a few examples without task-specific training.
All labels observed (4)
How this entity was disambiguated
This entity first appeared as the object of triple T4651147 — resolving that mention is where its identity was fixed. The disambiguator weighed these candidate entities and picked the highlighted one (or “None”, minting a new entity). This is how homonymy is resolved: the same surface form can point to different entities.
Target entity: Language Models are Few-Shot Learners Context triple: [Tom B. Brown, notableWork, Language Models are Few-Shot Learners]
-
A.
Language Models are Unsupervised Multitask Learners
"Language Models are Unsupervised Multitask Learners" is a 2019 OpenAI research paper that demonstrated how large-scale unsupervised language models like GPT-2 can perform a wide range of tasks without task-specific training.
-
B.
Exploring the Limits of Language Modeling
"Exploring the Limits of Language Modeling" is a research paper that investigates how far large-scale neural language models can be pushed in terms of performance, scalability, and generalization on natural language tasks.
-
C.
LLaMA
LLaMA is a family of large language models developed by Meta AI, designed for efficient training and inference across a range of natural language processing tasks.
-
D.
LLM
LLM is the ICAO airline designator assigned to Yamal Airlines, a Russian regional carrier.
-
E.
Longformer
Longformer is a transformer-based neural network architecture designed for efficient processing of very long sequences using sparse attention mechanisms.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
Target entity: Language Models are Few-Shot Learners Target entity description: "Language Models are Few-Shot Learners" is a landmark research paper that demonstrated large-scale transformer-based language models can perform diverse tasks from just a few examples without task-specific training.
-
A.
Language Models are Unsupervised Multitask Learners
"Language Models are Unsupervised Multitask Learners" is a 2019 OpenAI research paper that demonstrated how large-scale unsupervised language models like GPT-2 can perform a wide range of tasks without task-specific training.
-
B.
Exploring the Limits of Language Modeling
"Exploring the Limits of Language Modeling" is a research paper that investigates how far large-scale neural language models can be pushed in terms of performance, scalability, and generalization on natural language tasks.
-
C.
LLaMA
LLaMA is a family of large language models developed by Meta AI, designed for efficient training and inference across a range of natural language processing tasks.
-
D.
LLM
LLM is the ICAO airline designator assigned to Yamal Airlines, a Russian regional carrier.
-
E.
Longformer
Longformer is a transformer-based neural network architecture designed for efficient processing of very long sequences using sparse attention mechanisms.
- F. None of above. chosen
Statements (59)
| Predicate | Object |
|---|---|
| instanceOf |
research paper
ⓘ
scientific article ⓘ |
| alsoKnownAs |
GPT-3 paper
ⓘ
linked to:
GPT-3
|
| architecture | transformer ⓘ |
| author |
Aditya Ramesh
ⓘ
Alec Radford ⓘ Amanda Askell ⓘ Ariel Herbert-Voss ⓘ Arvind Neelakantan ⓘ Benjamin Chess ⓘ Benjamin Mann ⓘ Christopher Berner ⓘ Christopher Hesse ⓘ Clemens Winter ⓘ Daniel M. Ziegler ⓘ Dario Amodei ⓘ Eric Sigler ⓘ Girish Sastry ⓘ Gretchen Krueger ⓘ Ilya Sutskever ⓘ Jack Clark ⓘ Jared Kaplan ⓘ Jeffrey Wu ⓘ Mark Chen ⓘ Mateusz Litwin ⓘ Melanie Subbiah ⓘ Nick Ryder ⓘ Prafulla Dhariwal ⓘ Pranav Shyam ⓘ Rewon Child ⓘ Sam McCandlish ⓘ Sandhini Agarwal ⓘ Scott Gray ⓘ Tom B. Brown ⓘ Tom Henighan ⓘ |
| demonstrates |
few-shot learning capabilities of large language models
ⓘ
one-shot learning capabilities of large language models ⓘ zero-shot learning capabilities of large language models ⓘ |
| field |
artificial intelligence
ⓘ
machine learning ⓘ natural language processing ⓘ |
| impact | landmark paper in large-scale language modeling ⓘ |
| institution | OpenAI ⓘ |
| language | English ⓘ |
| mainSubject |
few-shot learning
ⓘ
large language models ⓘ transformer models ⓘ |
| modelParameterCount | 175 billion ⓘ |
| proposes | GPT-3 ⓘ |
| publicationYear | 2020 ⓘ |
| publishedIn |
Proceedings of the 34th Conference on Neural Information Processing Systems
ⓘ
linked to:
NeurIPS
|
| publisher |
NeurIPS 2020
ⓘ
linked to:
NeurIPS
|
| shows | performance scaling with model size across many NLP tasks ⓘ |
| taskTypesEvaluated |
cloze tasks
ⓘ
commonsense reasoning ⓘ question answering ⓘ reading comprehension ⓘ translation ⓘ |
| title | Language Models are Few-Shot Learners ⓘ |
How these facts were elicited
The pipeline generated the facts above by prompting gpt-5.1 with this entity's name + description and the instruction below.
You are a knowledge base construction expert. Given a subject entity and a description of it, return factual statements that you know for the subject as a JSON list of dictionaries(triples), where keys must be "subject", "predicate" and "object". The number of facts may be very high, between 25 to 50 or more, for very popular subjects. For less popular subjects, the number of facts can be very low, like 5 or 10. # Requirements - If you don't know the subject at all, return an empty list. - If the subject is not a named entity, return an empty list. - Include at least one triple where predicate is "instanceOf". - Do not get too wordy. - Separate several objects into multiple triples with one object.
Subject: Language Models are Few-Shot Learners Description of subject: "Language Models are Few-Shot Learners" is a landmark research paper that demonstrated large-scale transformer-based language models can perform diverse tasks from just a few examples without task-specific training.
Referenced by (12)
Full triples — surface form annotated when it differs from this entity's canonical label.