TY - JOUR
T1 - Identification of DNA-binding proteins by Kernel Sparse Representation via L2,1-matrix norm
AU - Ming, Yutong
AU - Liu, Hongzhi
AU - Cui, Yizhi
AU - Guo, Shaoyong
AU - Ding, Yijie
AU - Liu, Ruijun
N1 - Publisher Copyright:
© 2023 Elsevier Ltd
PY - 2023/6
Y1 - 2023/6
N2 - An understanding of DNA-binding proteins is helpful in exploring the role that proteins play in cell biology. Furthermore, the prediction of DNA-binding proteins is essential for the chemical modification and structural composition of DNA, and is of great importance in protein functional analysis and drug design. In recent years, DNA-binding protein prediction has typically used machine learning-based methods. The prediction accuracy of various classifiers has improved considerably, but researchers continue to spend time and effort on improving prediction performance. In this paper, we combine protein sequence evolutionary information with a classification method based on kernel sparse representation for the prediction of DNA-binding proteins, and based on the field of machine learning, a model for the identification of DNA-binding proteins by sequence information was finally proposed. Based on the confirmation of the final experimental results, we achieved good prediction accuracy on both the PDB1075 and PDB186 datasets. Our training result for cross-validation on PDB1075 was 81.37%, and our independent test result on PDB186 was 83.9%, both of which outperformed the other methods to some extent. Therefore, the proposed method in this paper is proven to be effective and feasible for predicting DNA-binding proteins.
AB - An understanding of DNA-binding proteins is helpful in exploring the role that proteins play in cell biology. Furthermore, the prediction of DNA-binding proteins is essential for the chemical modification and structural composition of DNA, and is of great importance in protein functional analysis and drug design. In recent years, DNA-binding protein prediction has typically used machine learning-based methods. The prediction accuracy of various classifiers has improved considerably, but researchers continue to spend time and effort on improving prediction performance. In this paper, we combine protein sequence evolutionary information with a classification method based on kernel sparse representation for the prediction of DNA-binding proteins, and based on the field of machine learning, a model for the identification of DNA-binding proteins by sequence information was finally proposed. Based on the confirmation of the final experimental results, we achieved good prediction accuracy on both the PDB1075 and PDB186 datasets. Our training result for cross-validation on PDB1075 was 81.37%, and our independent test result on PDB186 was 83.9%, both of which outperformed the other methods to some extent. Therefore, the proposed method in this paper is proven to be effective and feasible for predicting DNA-binding proteins.
KW - DNA-binding proteins
KW - Evolutionary information features
KW - Kernel sparse representation-based classification
KW - L-matrix norm
KW - Machine learning
UR - https://www.scopus.com/pages/publications/85152227393
U2 - 10.1016/j.compbiomed.2023.106849
DO - 10.1016/j.compbiomed.2023.106849
M3 - 文章
C2 - 37060772
AN - SCOPUS:85152227393
SN - 0010-4825
VL - 159
JO - Computers in Biology and Medicine
JF - Computers in Biology and Medicine
M1 - 106849
ER -