LayoutLM

E435880

LayoutLM is a transformer-based document understanding model that jointly leverages text, layout, and visual information to process and analyze scanned documents and forms.

All labels observed (4)

How this entity was disambiguated

Statements (45)

Predicate Object
instanceOf document understanding model
multimodal transformer model
pretrained language model
availableAs open-source model
basedOn Transformer architecture
category document AI model
vision-language model
designedFor document image classification
document understanding
form understanding
forms
information extraction
key information extraction
scanned documents
developer Microsoft Research Asia
hasAuthor Furu Wei
Lei Cui
Ming Zhou
Minghao Li
linked to: Yiheng Xu

Shaohan Huang
Yiheng Xu
hasVersion LayoutLMv2
linked to: LayoutLM

LayoutLMv3
linked to: LayoutLM
hostedOn Hugging Face Transformers
implementedIn PyTorch
inputModality image
layout
text
introducedAt KDD 2020
linked to: SIGKDD
introducedIn 2019
language English
leverages layout information
text information
visual information
optimizationObjective masked language modeling
multi-task learning for document understanding
paperTitle LayoutLM: Pre-training of Text and Layout for Document Image Understanding
linked to: LayoutLM
pretrainedOn large-scale document image datasets
supportsTask document question answering
invoice understanding
receipt understanding
uses 2D positional embeddings
bounding box coordinates
image region features
token-level text embeddings

How these facts were elicited

Referenced by (4)

Full triples — surface form annotated when it differs from this entity's canonical label.

LayoutLM hasVersion LayoutLMv2
linked to: LayoutLM
LayoutLM hasVersion LayoutLMv3
linked to: LayoutLM
LayoutLM paperTitle LayoutLM: Pre-training of Text and Layout for Document Image Understanding
linked to: LayoutLM