跳到主要导航 跳到搜索 跳到主要内容

OmniNova: A General Multimodal Multi-Agent Framework

  • Bingzhen Li
  • , Zihan Wang
  • , Pengfei Du*
  • , Yupeng Jiang
  • *此作品的通讯作者
  • Beihang University
  • National Computer Network Emergency Response Technical Team/Coordination Center of China
  • Shandong University of Political Science and Law
  • Hong Kong Research Institute of Technology

科研成果: 期刊稿件会议文章同行评审

摘要

The integration of Large Language Models (LLMs) with external tools enables intelligent automation far beyond text generation, yet coordinating multiple LLM-based agents remains difficult due to interaction overhead, resource waste, and fragile information flow. This paper introduces OmniNova, a modular and hierarchical multi-agent framework that unifies language models with capabilities for web search, browser automation, and code execution. OmniNova advances the state of the art through a hierarchical architecture that separates coordination, planning, supervision, and specialization; a dynamic routing mechanism that activates agents according to task complexity and state; and a multi-layered LLM integration strategy that allocates high-capability reasoning models only where they are cognitively necessary while assigning routine work to lighter models. Across 50 complex tasks in research, data analysis, and web interaction, OmniNova improves task completion (87% versus a 62% baseline), reduces token usage by 41%, and delivers higher human-rated quality (4.2/5 versus 3.1/5). The contribution includes both a principled system design and an open-source implementation intended to support research and practical deployment. Code is available at https://github.com/Superagentsys/OmniNoval.git.

指纹

探究 'OmniNova: A General Multimodal Multi-Agent Framework' 的科研主题。它们共同构成独一无二的指纹。

引用此