TY - GEN
T1 - Fine-Tuning Multimodal Models for Multilingual Event and Opinion Extraction in Science and Technology Intelligence
AU - Hong, Sheng
AU - Wickramaratne, Thisura Bojitha
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Multimodal large language models (MLLMs) have demonstrated high performance in information extraction, but their ability to conduct multilingual, multimodal event extraction (EE) and opinion extraction (OE) for science and technology intelligence (STI) is yet to be proven. This study examines VideoLLaMA2 (VL2) and VideoLLaMA2.1 (VL2.1) on a manually annotated dataset of 5, 728 E E and 2,194 OE samples for STI across English, Chinese, Spanish, and Russian, using text, image, and video inputs. Zero-shot results show VL2 achieving 46.11% for OE, 28.40% for EE trigger (Tr), and 23.89% for EE argument (Arg), and VL2.1 achieving 47.39% for OE, 24.76% for EE trigger (Tr), and 20.88% for EE argument (Arg), improved by prompt techniques: Chain-of-Thought, Tree-of-Thought and 4 shots prompting. While fine-tuning delivers the highest gains with VL2.1 achieving 74.43% accuracy in opinion extraction (27.04% baseline improvement) and 65.48% for event trigger identification (40.72% baseline improvement). Fine-tuning outperformed prompt engineering, offering a robust framework for multimodal, multilingual STI extraction.
AB - Multimodal large language models (MLLMs) have demonstrated high performance in information extraction, but their ability to conduct multilingual, multimodal event extraction (EE) and opinion extraction (OE) for science and technology intelligence (STI) is yet to be proven. This study examines VideoLLaMA2 (VL2) and VideoLLaMA2.1 (VL2.1) on a manually annotated dataset of 5, 728 E E and 2,194 OE samples for STI across English, Chinese, Spanish, and Russian, using text, image, and video inputs. Zero-shot results show VL2 achieving 46.11% for OE, 28.40% for EE trigger (Tr), and 23.89% for EE argument (Arg), and VL2.1 achieving 47.39% for OE, 24.76% for EE trigger (Tr), and 20.88% for EE argument (Arg), improved by prompt techniques: Chain-of-Thought, Tree-of-Thought and 4 shots prompting. While fine-tuning delivers the highest gains with VL2.1 achieving 74.43% accuracy in opinion extraction (27.04% baseline improvement) and 65.48% for event trigger identification (40.72% baseline improvement). Fine-tuning outperformed prompt engineering, offering a robust framework for multimodal, multilingual STI extraction.
KW - CoT
KW - Event Extraction
KW - Multilingual NLP
KW - Multilingual Prompt Optimization for STI
KW - Multimodal LLMs
KW - Opinion Extraction
KW - QLoRA
KW - ToT
UR - https://www.scopus.com/pages/publications/105031168719
U2 - 10.1109/ICBDSE65491.2025.11220058
DO - 10.1109/ICBDSE65491.2025.11220058
M3 - 会议稿件
AN - SCOPUS:105031168719
T3 - Proceeding of 2025 IEEE 2nd International Conference on Big Data Science and Engineering, ICBDSE 2025
BT - Proceeding of 2025 IEEE 2nd International Conference on Big Data Science and Engineering, ICBDSE 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2nd IEEE International Conference on Big Data Science and Engineering, ICBDSE 2025
Y2 - 13 June 2025 through 15 June 2025
ER -