跳到主要导航 跳到搜索 跳到主要内容

Vision-driven Adaptive Decoding for Document Understanding: A Dual-Path Visual-Language and OCR Framework with Prompt-aware Task Routing

  • Ao Shen
  • , Tongfei Chen
  • , Yangyang Sun
  • , Feng Jiang
  • , Yaogong Feng
  • , Kun Hu*
  • , Mengqi Liu
  • *此作品的通讯作者
  • Beihang University
  • Beijing Institute of Control and Electronics Technology
  • China Unicom (Hong Kong) Ltd.

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Real-world document understanding faces challenges such as layout diversity, low-quality imaging, and multilingual text, which hinder text recognition and table structure reconstruction. Traditional OCR-based pipelines achieve high character-level precision but lack global semantic understanding, while general-purpose vision-language models (VLMs) often struggle on text-heavy documents. To address these limitations, we propose VADoc (Vision-driven Adaptive Document Understanding), a dual-branch multimodal fusion framework that integrates an OCR-based branch for token-level grounding with Qwen2.5-VL for layout and semantic reasoning. A lightweight fusion module aligns both outputs into coherent text and structured HTML representations. Experiments on OmniDocBench show that VADoc substantially improves both text accuracy and table reconstruction quality over standalone VLMs and OCR systems. These results indicate that combining OCR precision with VLM contextual reasoning effectively bridges perception and understanding, and that VADoc supports robust document digitization and structured data extraction for downstream NLP and analytics tasks.

源语言英语
主期刊名ICVISP 2025 Proceedings - 2025 9th International Conference on Vision, Image and Signal Processing
出版商Institute of Electrical and Electronics Engineers Inc.
ISBN(电子版)9798331556822
DOI
出版状态已出版 - 2025
活动9th International Conference on Vision, Image and Signal Processing, ICVISP 2025 - Xi'an, 中国
期限: 28 11月 202530 11月 2025

出版系列

姓名ICVISP 2025 Proceedings - 2025 9th International Conference on Vision, Image and Signal Processing

会议

会议9th International Conference on Vision, Image and Signal Processing, ICVISP 2025
国家/地区中国
Xi'an
时期28/11/2530/11/25

学术指纹

探究 'Vision-driven Adaptive Decoding for Document Understanding: A Dual-Path Visual-Language and OCR Framework with Prompt-aware Task Routing' 的科研主题。它们共同构成独一无二的学术指纹。

引用此