DeBERTa

E435870

DeBERTa is a transformer-based language model developed by Microsoft that improves upon BERT and RoBERTa using disentangled attention and enhanced mask decoder mechanisms for superior natural language understanding.

All labels observed (10)

How this entity was disambiguated

Statements (50)

Predicate Object
instanceOf pretrained language model
transformer-based language model
architectureComponent embedding layer
feed-forward network
layer normalization
multi-head self-attention
availableOn Hugging Face Transformers
basedOn Transformer architecture
benchmark GLUE
SQuAD
linked to: SQuAD 2.0

SuperGLUE
developer Microsoft
hasVersion DeBERTa-XL
linked to: DeBERTa

DeBERTa-base
linked to: DeBERTa

DeBERTa-large
linked to: DeBERTa

DeBERTa-v1
linked to: DeBERTa

DeBERTa-v2
linked to: DeBERTa

DeBERTa-v3
linked to: DeBERTa

DeBERTa-xlarge
linked to: DeBERTa

DeBERTa-xxlarge
linked to: DeBERTa
implementation PyTorch
improves context representation
word representation
improvesUpon BERT
RoBERTa
introducedBy Microsoft Research
language English
license MIT License (for official Microsoft implementation)
linked to: MIT License
optimization Adam optimizer (typical training setup)
outperforms BERT on GLUE
RoBERTa on GLUE (for larger variants)
paperTitle DeBERTa: Decoding-enhanced BERT with Disentangled Attention
linked to: DeBERTa
publicationType research paper
releasedYear 2020
supportsTask natural language inference
question answering
sentiment analysis
sequence labeling
text classification
token classification
task natural language understanding
trainingObjective masked language modeling
replaced token detection (for some versions)
uses absolute position embeddings (for some variants)
relative position embeddings
usesMechanism disentangled attention
enhanced mask decoder
usesPretrainingData BookCorpus (for some variants)
linked to: BookCorpus

Wikipedia
large-scale web text

How these facts were elicited

Referenced by (11)

Full triples — surface form annotated when it differs from this entity's canonical label.

DeBERTa hasVersion DeBERTa-v1
linked to: DeBERTa
DeBERTa hasVersion DeBERTa-v2
linked to: DeBERTa
DeBERTa hasVersion DeBERTa-v3
linked to: DeBERTa
DeBERTa hasVersion DeBERTa-XL
linked to: DeBERTa
DeBERTa hasVersion DeBERTa-large
linked to: DeBERTa
DeBERTa hasVersion DeBERTa-base
linked to: DeBERTa
DeBERTa hasVersion DeBERTa-xlarge
linked to: DeBERTa
DeBERTa hasVersion DeBERTa-xxlarge
linked to: DeBERTa
DeBERTa paperTitle DeBERTa: Decoding-enhanced BERT with Disentangled Attention
linked to: DeBERTa