arxiv:2507.08606

DocPolarBERT: A Pre-trained Model for Document Understanding with Relative Polar Coordinate Encoding of Layout Structures

Published on Jul 11

Authors:

Abstract

DocPolarBERT, a layout-aware BERT model using relative polar coordinates for self-attention, achieves state-of-the-art document understanding with less pre-training data.

AI-generated summary

We introduce DocPolarBERT, a layout-aware BERT model for document understanding that eliminates the need for absolute 2D positional embeddings. We extend self-attention to take into account text block positions in relative polar coordinate system rather than the Cartesian one. Despite being pre-trained on a dataset more than six times smaller than the widely used IIT-CDIP corpus, DocPolarBERT achieves state-of-the-art results. These results demonstrate that a carefully designed attention mechanism can compensate for reduced pre-training data, offering an efficient and effective alternative for document understanding.

View arXiv page View PDF Add to collection