HuBERT

E435884

HuBERT is a self-supervised speech representation learning model that learns powerful audio features from unlabeled speech for tasks like automatic speech recognition and audio classification.

All labels observed (5)

How this entity was disambiguated

Statements (51)

Predicate Object
instanceOf deep learning model
self-supervised speech representation learning model
speech foundation model
basedOn Transformer architecture
designedFor audio classification
automatic speech recognition
downstream speech tasks
self-supervised learning from speech
speech representation learning
spoken language understanding
developedBy Facebook AI Research
Meta AI
evaluationBenchmark Libri-light
LibriSpeech
TIMIT
hasAuthor Benjamin Bolte
James Glass
Kyunghyun Cho
Wei-Ning Hsu
Yao-Hung Hubert Tsai
hasVariant HuBERT Base
linked to: HuBERT

HuBERT Large
linked to: HuBERT

HuBERT X-Large
linked to: HuBERT
improvesOver wav2vec 2.0 on several speech benchmarks
inputType acoustic features such as log-mel filterbanks
raw audio waveforms
introducedIn 2021
introducedInPaper HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
linked to: HuBERT
language primarily English in original experiments
learningSignalSource discrete units obtained by clustering MFCC or filterbank features
maskingStrategy masking of contiguous time spans in the input sequence
openSourceImplementation Hugging Face Transformers
fairseq
linked to: Fairseq
outputType contextualized speech representations
frame-level audio embeddings
pretrainingStage masked region prediction of cluster assignments
offline clustering of acoustic features
publishedAt Interspeech 2021
linked to: INTERSPEECH
relatedTo Data2Vec
WavLM
wav2vec 2.0
linked to: Wav2Vec2
supportsTask audio event classification
automatic speech recognition fine-tuning
emotion recognition from speech
keyword spotting
phoneme recognition
speaker recognition
trainingDataType unlabeled speech audio
trainingParadigm self-supervised learning
usesObjective cluster-based prediction task
masked prediction of latent speech units

How these facts were elicited

Referenced by (6)

Full triples — surface form annotated when it differs from this entity's canonical label.

Wav2Vec2 inspired HuBERT
HuBERT introducedInPaper HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
linked to: HuBERT
HuBERT hasVariant HuBERT Base
linked to: HuBERT
HuBERT hasVariant HuBERT Large
linked to: HuBERT
HuBERT hasVariant HuBERT X-Large
linked to: HuBERT