ALBERT

E435869

ALBERT is a lightweight, parameter-efficient variant of the BERT language model designed to achieve strong natural language understanding performance with reduced memory and computation costs.

All labels observed (1)

Label Occurrences
ALBERT canonical 3

How this entity was disambiguated

Statements (48)

Predicate Object
instanceOf BERT variant
language model
neural network model
transformer-based model
achieves state-of-the-art results on several benchmarks at release
acronymFor A Lite BERT
linked to: DistilBERT
aimsTo maintain strong performance
reduce computation cost
reduce memory usage
basedOn BERT
compatibleWith Hugging Face Transformers
describedInPaper ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
designedFor natural language understanding
evaluatedOn GLUE benchmark
RACE dataset
SQuAD
linked to: SQuAD 2.0
fullName A Lite BERT
linked to: DistilBERT
hasArchitecture Transformer
hasFeature cross-layer parameter sharing
smaller embedding size with projection
hasObjective sentence-order prediction
hasOpenSourceImplementation Yes
hasProperty computationally-efficient
lightweight
memory-efficient
parameter-efficient
hasVariant ALBERT-base
ALBERT-large
ALBERT-xlarge
ALBERT-xxlarge
linked to: ALBERT-xlarge
implementedIn TensorFlow
introducedBy Google Research
Toyota Technological Institute at Chicago
introducedIn 2019
language English (pretrained models)
paperAuthorsInclude Kevin Gimpel
Mingda Chen
Piyush Sharma
Radu Soricut
linked to: Mihai Surdeanu

Sebastian Goodman
Zhenzhong Lan
replacesObjective next sentence prediction
supports sentence-level tasks
token-level tasks
trainedWith masked language modeling
uses self-attention
usesTechnique factorized embedding parameterization
parameter sharing across layers

How these facts were elicited

Referenced by (3)

Full triples — surface form annotated when it differs from this entity's canonical label.