Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

E1339918 UNEXPLORED

"Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity" is a research paper that introduces a sparsely activated mixture-of-experts transformer architecture enabling efficient training and inference of language models with up to a trillion parameters.

All labels observed (1)

How this entity was disambiguated

Referenced by (1)

Full triples — surface form annotated when it differs from this entity's canonical label.

Noam Shazeer coAuthorOf Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity