TY - GEN
T1 - A Study on C Code Defect Detection with Fine-Tuned Large Language Models
AU - Wang, Yue
AU - Wang, Xu
AU - Yu, Hongwei
AU - Gao, Fei
AU - Liu, Xueshi
AU - Wang, Xiaoling
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Large Language Models(LLMs) have demonstrated excellent capabilities in many areas of software engineering(SE), including code completion, code generation, code understanding, code repair, etc., and the most prominent performer in this regard is ChatGPT. However, its cost of use makes the integration of ChatGPT into code defect detection techniques costly. In this paper, we focus on low-cost-of-use, fine-tunable, open-source large language models with less than 10B parameters, and study their capabilities of C code defect detection when fine-tuned with real-world data and improved with prompt engineering. We studied LLaMa3-8B, DeepSeek-Coder-7b and Qwen2-7B, as they are the typical models with prompt capabilities, whose performance in SE is close to ChatGPT, and they are open-source models. Experimental results show that our method can significantly improve the performance of LLMs within 10B parameters on code defect detection, and the output of the models can be applied to several downstream tasks, such as improving the report quality of static analysis tools.
AB - Large Language Models(LLMs) have demonstrated excellent capabilities in many areas of software engineering(SE), including code completion, code generation, code understanding, code repair, etc., and the most prominent performer in this regard is ChatGPT. However, its cost of use makes the integration of ChatGPT into code defect detection techniques costly. In this paper, we focus on low-cost-of-use, fine-tunable, open-source large language models with less than 10B parameters, and study their capabilities of C code defect detection when fine-tuned with real-world data and improved with prompt engineering. We studied LLaMa3-8B, DeepSeek-Coder-7b and Qwen2-7B, as they are the typical models with prompt capabilities, whose performance in SE is close to ChatGPT, and they are open-source models. Experimental results show that our method can significantly improve the performance of LLMs within 10B parameters on code defect detection, and the output of the models can be applied to several downstream tasks, such as improving the report quality of static analysis tools.
KW - Defect Detection
KW - Fine-tuning
KW - Large Language Models
KW - Prompt Engineering
UR - https://www.scopus.com/pages/publications/105004734288
U2 - 10.1109/APSEC65559.2024.00055
DO - 10.1109/APSEC65559.2024.00055
M3 - 会议稿件
AN - SCOPUS:105004734288
T3 - Proceedings - Asia-Pacific Software Engineering Conference, APSEC
SP - 437
EP - 441
BT - Proceedings - 2024 31st Asia-Pacific Software Engineering Conference, APSEC 2024
PB - IEEE Computer Society
T2 - 31st Asia-Pacific Software Engineering Conference, APSEC 2024
Y2 - 3 December 2024 through 6 December 2024
ER -