TY - JOUR
T1 - CTISum
T2 - A new benchmark dataset for Cyber Threat Intelligence summarization
AU - Peng, Wei
AU - Ding, Junmei
AU - Wang, Wei
AU - Cui, Lei
AU - Cai, Wei
AU - Hao, Zhiyu
AU - Yun, Xiaochun
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/9
Y1 - 2026/9
N2 - Cyber Threat Intelligence (CTI) summarization involves generating concise and accurate highlights from web intelligence data with domain knowledge, which is critical to automatically summarize the knowledge and conclusion contained in CTI reports. Despite that, the development of efficient techniques for summarizing CTI reports, comprising facts, analytical insights, attack processes, and more, has been hindered by the lack of suitable datasets. To address this gap, we introduce CTISum, a new benchmark dataset designed for the CTI summarization task. Recognizing the significance of understanding attack processes, we also propose a novel fine-grained subtask: attack process summarization, which aims to help defenders assess risks, identify security gaps, and uncover vulnerabilities. Specifically, a multi-stage annotation pipeline is designed to collect and annotate CTI data from diverse web sources, alongside a comprehensive benchmarking of CTISum using both extractive, abstractive and LLMs-based summarization methods. Experimental results reveal that current state-of-the-art AI models (including GPT-4o) face significant challenges when applied to CTISum, highlighting that automatic summarization of CTI reports remains an open research problem. The code and example dataset can be made publicly available at https://github.com/pengwei-iie/CTISum.
AB - Cyber Threat Intelligence (CTI) summarization involves generating concise and accurate highlights from web intelligence data with domain knowledge, which is critical to automatically summarize the knowledge and conclusion contained in CTI reports. Despite that, the development of efficient techniques for summarizing CTI reports, comprising facts, analytical insights, attack processes, and more, has been hindered by the lack of suitable datasets. To address this gap, we introduce CTISum, a new benchmark dataset designed for the CTI summarization task. Recognizing the significance of understanding attack processes, we also propose a novel fine-grained subtask: attack process summarization, which aims to help defenders assess risks, identify security gaps, and uncover vulnerabilities. Specifically, a multi-stage annotation pipeline is designed to collect and annotate CTI data from diverse web sources, alongside a comprehensive benchmarking of CTISum using both extractive, abstractive and LLMs-based summarization methods. Experimental results reveal that current state-of-the-art AI models (including GPT-4o) face significant challenges when applied to CTISum, highlighting that automatic summarization of CTI reports remains an open research problem. The code and example dataset can be made publicly available at https://github.com/pengwei-iie/CTISum.
KW - Cyber threat intelligence
KW - Dataset and benchmark
KW - Information systems
KW - Summarization
UR - https://www.scopus.com/pages/publications/105037444221
U2 - 10.1016/j.cose.2026.104928
DO - 10.1016/j.cose.2026.104928
M3 - 文章
AN - SCOPUS:105037444221
SN - 0167-4048
VL - 168
JO - Computers and Security
JF - Computers and Security
M1 - 104928
ER -