On Layer Normalization in the Transformer Architecture
E1339919
UNEXPLORED
"On Layer Normalization in the Transformer Architecture" is a research paper that analyzes and refines how layer normalization is applied within Transformer neural networks to improve their stability and performance.
All labels observed (1)
| Label | Occurrences |
|---|---|
| On Layer Normalization in the Transformer Architecture canonical | 1 |
How this entity was disambiguated
This entity first appeared as the object of triple T18724118 — resolving that mention is where its identity was fixed. The disambiguator weighed these candidate entities and picked the highlighted one (or “None”, minting a new entity). This is how homonymy is resolved: the same surface form can point to different entities.
NED1
Entity disambiguation (via context triple)
gpt-5-mini-2025-08-07
Target entity: On Layer Normalization in the Transformer Architecture Context triple: [Noam Shazeer, coAuthorOf, On Layer Normalization in the Transformer Architecture]
-
A.
Reformer: The Efficient Transformer
Reformer: The Efficient Transformer is a research paper introducing a more memory- and computation-efficient Transformer architecture using techniques like locality-sensitive hashing attention and reversible layers.
-
B.
Layer Normalization
Layer Normalization is a neural network normalization technique that stabilizes and accelerates training by normalizing activations across features within each data sample, particularly useful in recurrent and transformer-based models.
-
C.
Attention Is All You Need
"Attention Is All You Need" is the landmark 2017 research paper that introduced the Transformer architecture and revolutionized modern natural language processing and sequence modeling.
-
D.
OPT: Open Pre-trained Transformer Language Models
OPT: Open Pre-trained Transformer Language Models is a family of openly released large-scale transformer-based language models developed by Meta AI to provide transparent, reproducible alternatives to proprietary models like GPT-3.
-
E.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
"Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer" is the seminal research paper that introduced the T5 model, framing all NLP tasks in a unified text-to-text format and demonstrating state-of-the-art transfer learning performance across diverse benchmarks.
- F. None of above. chosen
- G. Unsure - the case is ambiguous/there is not enough information to decide.
NED2
Entity disambiguation (via description)
gpt-5-mini-2025-08-07
Target entity: On Layer Normalization in the Transformer Architecture Target entity description: "On Layer Normalization in the Transformer Architecture" is a research paper that analyzes and refines how layer normalization is applied within Transformer neural networks to improve their stability and performance.
-
A.
Reformer: The Efficient Transformer
Reformer: The Efficient Transformer is a research paper introducing a more memory- and computation-efficient Transformer architecture using techniques like locality-sensitive hashing attention and reversible layers.
-
B.
Layer Normalization
Layer Normalization is a neural network normalization technique that stabilizes and accelerates training by normalizing activations across features within each data sample, particularly useful in recurrent and transformer-based models.
-
C.
Attention Is All You Need
"Attention Is All You Need" is the landmark 2017 research paper that introduced the Transformer architecture and revolutionized modern natural language processing and sequence modeling.
-
D.
OPT: Open Pre-trained Transformer Language Models
OPT: Open Pre-trained Transformer Language Models is a family of openly released large-scale transformer-based language models developed by Meta AI to provide transparent, reproducible alternatives to proprietary models like GPT-3.
-
E.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
"Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer" is the seminal research paper that introduced the T5 model, framing all NLP tasks in a unified text-to-text format and demonstrating state-of-the-art transfer learning performance across diverse benchmarks.
- F. None of above. chosen
Referenced by (1)
Full triples — surface form annotated when it differs from this entity's canonical label.