TY - GEN
T1 - Vision-driven Adaptive Decoding for Document Understanding
T2 - 9th International Conference on Vision, Image and Signal Processing, ICVISP 2025
AU - Shen, Ao
AU - Chen, Tongfei
AU - Sun, Yangyang
AU - Jiang, Feng
AU - Feng, Yaogong
AU - Hu, Kun
AU - Liu, Mengqi
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Real-world document understanding faces challenges such as layout diversity, low-quality imaging, and multilingual text, which hinder text recognition and table structure reconstruction. Traditional OCR-based pipelines achieve high character-level precision but lack global semantic understanding, while general-purpose vision-language models (VLMs) often struggle on text-heavy documents. To address these limitations, we propose VADoc (Vision-driven Adaptive Document Understanding), a dual-branch multimodal fusion framework that integrates an OCR-based branch for token-level grounding with Qwen2.5-VL for layout and semantic reasoning. A lightweight fusion module aligns both outputs into coherent text and structured HTML representations. Experiments on OmniDocBench show that VADoc substantially improves both text accuracy and table reconstruction quality over standalone VLMs and OCR systems. These results indicate that combining OCR precision with VLM contextual reasoning effectively bridges perception and understanding, and that VADoc supports robust document digitization and structured data extraction for downstream NLP and analytics tasks.
AB - Real-world document understanding faces challenges such as layout diversity, low-quality imaging, and multilingual text, which hinder text recognition and table structure reconstruction. Traditional OCR-based pipelines achieve high character-level precision but lack global semantic understanding, while general-purpose vision-language models (VLMs) often struggle on text-heavy documents. To address these limitations, we propose VADoc (Vision-driven Adaptive Document Understanding), a dual-branch multimodal fusion framework that integrates an OCR-based branch for token-level grounding with Qwen2.5-VL for layout and semantic reasoning. A lightweight fusion module aligns both outputs into coherent text and structured HTML representations. Experiments on OmniDocBench show that VADoc substantially improves both text accuracy and table reconstruction quality over standalone VLMs and OCR systems. These results indicate that combining OCR precision with VLM contextual reasoning effectively bridges perception and understanding, and that VADoc supports robust document digitization and structured data extraction for downstream NLP and analytics tasks.
KW - Document Understanding
KW - OCR-free Recognition
KW - Table Structure Analysis
KW - Vision-Language Models
UR - https://www.scopus.com/pages/publications/105036369349
U2 - 10.1109/ICVISP68610.2025.11451708
DO - 10.1109/ICVISP68610.2025.11451708
M3 - 会议稿件
AN - SCOPUS:105036369349
T3 - ICVISP 2025 Proceedings - 2025 9th International Conference on Vision, Image and Signal Processing
BT - ICVISP 2025 Proceedings - 2025 9th International Conference on Vision, Image and Signal Processing
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 28 November 2025 through 30 November 2025
ER -