Big Bird: Transformers for Longer Sequences
E1312470
UNEXPLORED
"Big Bird: Transformers for Longer Sequences" is a research paper that introduces a sparse-attention Transformer architecture enabling efficient processing of much longer input sequences than standard Transformers while retaining strong performance.
All labels observed (1)
| Label | Occurrences |
|---|---|
| Big Bird: Transformers for Longer Sequences canonical | 1 |
How this entity was disambiguated
This entity first appeared as the object of triple T18204984 — resolving that mention is where its identity was fixed. The disambiguator weighed these candidate entities and picked the highlighted one (or “None”, minting a new entity). This is how homonymy is resolved: the same surface form can point to different entities.
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: Big Bird: Transformers for Longer Sequences Context triple: [BigBird, paperTitle, Big Bird: Transformers for Longer Sequences]
-
A.
Attention Is All You Need
"Attention Is All You Need" is the landmark 2017 research paper that introduced the Transformer architecture and revolutionized modern natural language processing and sequence modeling.
-
B.
Transformer-XL
Transformer-XL is a neural network architecture for language modeling that extends the Transformer with segment-level recurrence and relative positional encodings to better capture long-range dependencies.
-
C.
Sequence to Sequence Learning with Neural Networks
"Sequence to Sequence Learning with Neural Networks" is a seminal 2014 paper that introduced the sequence-to-sequence (seq2seq) neural network framework for tasks like machine translation, laying the groundwork for many modern NLP models.
-
D.
Neural Discrete Representation Learning
Neural Discrete Representation Learning is a machine learning framework that introduces Vector Quantized Variational Autoencoders (VQ-VAE) to learn discrete latent representations for high-dimensional data such as images, audio, and video.
-
E.
Neural Machine Translation in Linear Time
"Neural Machine Translation in Linear Time" is a research paper that introduces a more computationally efficient neural architecture for machine translation, reducing translation complexity to linear time with respect to input length.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: Big Bird: Transformers for Longer Sequences Target entity description: "Big Bird: Transformers for Longer Sequences" is a research paper that introduces a sparse-attention Transformer architecture enabling efficient processing of much longer input sequences than standard Transformers while retaining strong performance.
-
A.
Attention Is All You Need
"Attention Is All You Need" is the landmark 2017 research paper that introduced the Transformer architecture and revolutionized modern natural language processing and sequence modeling.
-
B.
Transformer-XL
Transformer-XL is a neural network architecture for language modeling that extends the Transformer with segment-level recurrence and relative positional encodings to better capture long-range dependencies.
-
C.
Sequence to Sequence Learning with Neural Networks
"Sequence to Sequence Learning with Neural Networks" is a seminal 2014 paper that introduced the sequence-to-sequence (seq2seq) neural network framework for tasks like machine translation, laying the groundwork for many modern NLP models.
-
D.
Neural Discrete Representation Learning
Neural Discrete Representation Learning is a machine learning framework that introduces Vector Quantized Variational Autoencoders (VQ-VAE) to learn discrete latent representations for high-dimensional data such as images, audio, and video.
-
E.
Neural Machine Translation in Linear Time
"Neural Machine Translation in Linear Time" is a research paper that introduces a more computationally efficient neural architecture for machine translation, reducing translation complexity to linear time with respect to input length.
- F. None of above. chosen
Referenced by (1)
Full triples — surface form annotated when it differs from this entity's canonical label.