TY - GEN
T1 - Breaking Size Barrier
T2 - 30th International Conference on Database Systems for Advanced Applications, DASFAA 2025
AU - Wu, Xianjie
AU - Liang, Di
AU - Yang, Jian
AU - Cheng, Xianfu
AU - Chai, Lin Zheng
AU - Li, Tongliang
AU - Yang, Liqun
AU - Li, Zhoujun
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Large language models (LLMs) significantly enhance their ability to process tabular data through chain-of-thought reasoning, particularly in table question answering tasks. However, LLMs encounter substantial challenges when dealing with large tables in real-world applications. Prompting LLMs with the entire table not only encounters context-length constraints but also significantly extends the reasoning path, heightening the risk of reasoning hallucination and information truncation. To address this, we construct a large-size table reasoning (LSTR) benchmark, featuring tables larger than those in existing benchmarks, to thoroughly investigate how table size affects the reasoning abilities of LLMs in answering table-related questions. Subsequently, we propose a size-adaptive-thought (SAT) approach that instructs the LLM utilizing refined metadata to employ Python commands for manipulating tables step by step, thereby facilitating efficient reasoning with tables of any size. Furthermore, we develop SAT-Llama, fine-tuned SAT on Llama3.1 (8B), which delivers performance comparable to large-size LLMs at a much lower cost, addressing the issue of inadequate code manipulation capabilities in small-size LLMs. Experimental results on the LSTR and WTQ datasets demonstrate that SAT achieves a new state-of-the-art in handling large-size tables, exhibiting significant performance advantages and high context token efficiency.
AB - Large language models (LLMs) significantly enhance their ability to process tabular data through chain-of-thought reasoning, particularly in table question answering tasks. However, LLMs encounter substantial challenges when dealing with large tables in real-world applications. Prompting LLMs with the entire table not only encounters context-length constraints but also significantly extends the reasoning path, heightening the risk of reasoning hallucination and information truncation. To address this, we construct a large-size table reasoning (LSTR) benchmark, featuring tables larger than those in existing benchmarks, to thoroughly investigate how table size affects the reasoning abilities of LLMs in answering table-related questions. Subsequently, we propose a size-adaptive-thought (SAT) approach that instructs the LLM utilizing refined metadata to employ Python commands for manipulating tables step by step, thereby facilitating efficient reasoning with tables of any size. Furthermore, we develop SAT-Llama, fine-tuned SAT on Llama3.1 (8B), which delivers performance comparable to large-size LLMs at a much lower cost, addressing the issue of inadequate code manipulation capabilities in small-size LLMs. Experimental results on the LSTR and WTQ datasets demonstrate that SAT achieves a new state-of-the-art in handling large-size tables, exhibiting significant performance advantages and high context token efficiency.
KW - Large Language Model
KW - Large-Size Table
KW - Table Question Answering
UR - https://www.scopus.com/pages/publications/105028314378
U2 - 10.1007/978-981-95-3830-0_16
DO - 10.1007/978-981-95-3830-0_16
M3 - 会议稿件
AN - SCOPUS:105028314378
SN - 9789819538294
T3 - Lecture Notes in Computer Science
SP - 241
EP - 256
BT - Database Systems for Advanced Applications - 30th International Conference, DASFAA 2025, Proceedings
A2 - Zhu, Feida
A2 - Lim, Ee-Peng
A2 - Yu, Philip S.
A2 - Nadamoto, Akiyo
A2 - Shim, Kyuseok
A2 - Ding, Wei
A2 - Zhang, Bingxue
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 26 May 2025 through 29 May 2025
ER -