Skip to main navigation Skip to search Skip to main content

Vision-driven Adaptive Decoding for Document Understanding: A Dual-Path Visual-Language and OCR Framework with Prompt-aware Task Routing

  • Ao Shen
  • , Tongfei Chen
  • , Yangyang Sun
  • , Feng Jiang
  • , Yaogong Feng
  • , Kun Hu*
  • , Mengqi Liu
  • *Corresponding author for this work
  • Beihang University
  • Beijing Institute of Control and Electronics Technology
  • China Unicom (Hong Kong) Ltd.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Real-world document understanding faces challenges such as layout diversity, low-quality imaging, and multilingual text, which hinder text recognition and table structure reconstruction. Traditional OCR-based pipelines achieve high character-level precision but lack global semantic understanding, while general-purpose vision-language models (VLMs) often struggle on text-heavy documents. To address these limitations, we propose VADoc (Vision-driven Adaptive Document Understanding), a dual-branch multimodal fusion framework that integrates an OCR-based branch for token-level grounding with Qwen2.5-VL for layout and semantic reasoning. A lightweight fusion module aligns both outputs into coherent text and structured HTML representations. Experiments on OmniDocBench show that VADoc substantially improves both text accuracy and table reconstruction quality over standalone VLMs and OCR systems. These results indicate that combining OCR precision with VLM contextual reasoning effectively bridges perception and understanding, and that VADoc supports robust document digitization and structured data extraction for downstream NLP and analytics tasks.

Original languageEnglish
Title of host publicationICVISP 2025 Proceedings - 2025 9th International Conference on Vision, Image and Signal Processing
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331556822
DOIs
StatePublished - 2025
Event9th International Conference on Vision, Image and Signal Processing, ICVISP 2025 - Xi'an, China
Duration: 28 Nov 202530 Nov 2025

Publication series

NameICVISP 2025 Proceedings - 2025 9th International Conference on Vision, Image and Signal Processing

Conference

Conference9th International Conference on Vision, Image and Signal Processing, ICVISP 2025
Country/TerritoryChina
CityXi'an
Period28/11/2530/11/25

Keywords

  • Document Understanding
  • OCR-free Recognition
  • Table Structure Analysis
  • Vision-Language Models

Fingerprint

Dive into the research topics of 'Vision-driven Adaptive Decoding for Document Understanding: A Dual-Path Visual-Language and OCR Framework with Prompt-aware Task Routing'. Together they form a unique fingerprint.

Cite this