CREPE: Coordinate-Aware End-to-End Document Parser | allainews.com

May 2, 2024, 4:44 a.m. | Yamato Okamoto, Youngmin Baek, Geewook Kim, Ryota Nakao, DongHyun Kim, Moon Bin Yim, Seunghyun Park, Bado Lee

cs.CV updates on arXiv.org arxiv.org

arXiv:2405.00260v1 Announce Type: new
Abstract: In this study, we formulate an OCR-free sequence generation model for visual document understanding (VDU). Our model not only parses text from document images but also extracts the spatial coordinates of the text based on the multi-head architecture. Named as Coordinate-aware End-to-end Document Parser (CREPE), our method uniquely integrates these capabilities by introducing a special token for OCR text, and token-triggered coordinate decoding. We also proposed a weakly-supervised framework for cost-efficient training, requiring only parsing …

abstract architecture arxiv cs.cv document document understanding free head images multi-head ocr spatial study text type understanding visual

More from arxiv.org / cs.CV updates on arXiv.org

3D Human Pose Perception from Egocentric Stereo Videos 23 hours ago | arxiv.org

abstract arxiv compact cs.cv +11

SqueezeSAM: User friendly mobile interactive segmentation 23 hours ago | arxiv.org

abstract architecture arxiv computational +21

Polarimetric Light Transport Analysis for Specular Inter-reflection 23 hours ago | arxiv.org

abstract analysis arxiv cs.cv +13

Density-Guided Dense Pseudo Label Selection For Semi-supervised Oriented Object Detection 23 hours ago | arxiv.org

arxiv cs.cv detection object +4

CtxMIM: Context-Enhanced Masked Image Modeling for Remote Sensing Image Understanding 23 hours ago | arxiv.org

abstract arxiv clear context +19

nnSAM: Plug-and-play Segment Anything Model Improves nnUNet Performance 23 hours ago | arxiv.org

abstract arxiv clinical cs.cv +21

CoFiI2P: Coarse-to-Fine Correspondences for Image-to-Point Cloud Registration 23 hours ago | arxiv.org

arxiv cloud cs.ai cs.cv +5

Detail Reinforcement Diffusion Model: Augmentation Fine-Grained Visual Categorization in Few-Shot Conditions 23 hours ago | arxiv.org

abstract annotated data arxiv augmentation +18

Generative Image Dynamics 23 hours ago | arxiv.org

abstract arxiv clothes collection +15

Software Engineer for AI Training Data (School Specific)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Python)

@ G2i Inc | Remote

View on ai-jobs.net

Software Engineer for AI Training Data (Tier 2)

@ G2i Inc | Remote

View on ai-jobs.net

Data Engineer

@ Lemon.io | Remote: Europe, LATAM, Canada, UK, Asia, Oceania

View on ai-jobs.net

Artificial Intelligence – Bioinformatic Expert

@ University of Texas Medical Branch | Galveston, TX

View on ai-jobs.net

Lead Developer (AI)

@ Cere Network | San Francisco, US

View on ai-jobs.net