摘要
Air traffic controllers (ATCos) play a critical role in aviation safety through aircraft movement management and emergency response. Current ATCo state monitoring methods focus on discrete actions or simple metrics like blink frequency and gaze patterns. These approaches lack the depth needed for effective management decisions. This study presents a hierarchical collaborative framework that integrates lightweight deep learning models, vision-language models (VLM), and large language models (LLM) for comprehensive ATCo state analysis. The framework operates in three stages: lightweight models detect key temporal intervals from video streams; VLMs extract meaningful behavioral patterns from these intervals; LLMs generate interpretive reports for operational monitoring. We collected 153 video segments from 18 controllers in operational environments and developed a natural language annotation system with expert validation. Experimental results show our multi-feature fusion approach achieves 0.55 Mean IoU and 0.65 Label Coverage for interval detection, outperforming traditional methods. VLM evaluation reveals Gemini 2.5 Pro excels in micro-motion sensitivity (4.8/5.0) while Claude 4 Sonnet demonstrates superior consistency (4.8/5.0). LLM assessment shows GPT-4.1 and Claude 4 Sonnet achieve highest performance in evidence use and risk calibration respectively. This framework could be a useful tool for air traffic management, enhancing monitoring efficiency and enabling data-driven, proactive safety interventions to mitigate human-factor risks.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 103926 |
| 期刊 | International Journal of Industrial Ergonomics |
| 卷 | 113 |
| DOI | |
| 出版状态 | 已出版 - 5月 2026 |
联合国可持续发展目标
此成果有助于实现下列可持续发展目标:
-
可持续发展目标 3 良好健康与福祉
学术指纹
探究 'Integrating lightweight and large-scale models for state analysis of air traffic controllers' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver