TY - JOUR
T1 - Knowledge graph modeling for data asset pricing
AU - Ren, Shiqi
AU - Wang, Wenhu
AU - Jia, Shengyin
AU - Wang, Shujie
AU - Han, Liyan
N1 - Publisher Copyright:
© 2026 Elsevier Ltd.
PY - 2026/7
Y1 - 2026/7
N2 - The valuation of data assets is a fundamental problem in data circulation and data trading systems, yet remains challenging due to heterogeneous descriptions, unclear usage contexts, and the lack of standardized comparability across products. To address this issue, we propose a knowledge-graph-based valuation framework that models data assets and their semantic and relational attributes in a structured graph space. Unstructured product descriptions collected from AWS Data Exchange are processed using a large language model (LLM) ensemble with strict-majority voting, which improves triple extraction stability and mitigates hallucination. The constructed knowledge graph encodes data products, providers, industries, and application concepts. To obtain machine-operable representations, we evaluate two embedding strategies: Node2Vec, which captures global semantic proximity, and GraphSAGE, which preserves localized neighborhood patterns. Experimental results show that Node2Vec achieves superior overall performance, especially in concept-rich sub-graph structures, while GraphSAGE performs better in industry-homogeneous scenarios. Using the learned embeddings, we perform clustering to identify comparable data assets and combine these features with supervised learning models to predict pricing levels. Extensive experiments demonstrate that our framework improves robustness, interpretability, and consistency in data asset valuation.
AB - The valuation of data assets is a fundamental problem in data circulation and data trading systems, yet remains challenging due to heterogeneous descriptions, unclear usage contexts, and the lack of standardized comparability across products. To address this issue, we propose a knowledge-graph-based valuation framework that models data assets and their semantic and relational attributes in a structured graph space. Unstructured product descriptions collected from AWS Data Exchange are processed using a large language model (LLM) ensemble with strict-majority voting, which improves triple extraction stability and mitigates hallucination. The constructed knowledge graph encodes data products, providers, industries, and application concepts. To obtain machine-operable representations, we evaluate two embedding strategies: Node2Vec, which captures global semantic proximity, and GraphSAGE, which preserves localized neighborhood patterns. Experimental results show that Node2Vec achieves superior overall performance, especially in concept-rich sub-graph structures, while GraphSAGE performs better in industry-homogeneous scenarios. Using the learned embeddings, we perform clustering to identify comparable data assets and combine these features with supervised learning models to predict pricing levels. Extensive experiments demonstrate that our framework improves robustness, interpretability, and consistency in data asset valuation.
KW - Data asset pricing
KW - Graph embedding
KW - Knowledge graph
KW - Market comparison method
KW - Prediction performance
KW - Two layer ontology
UR - https://www.scopus.com/pages/publications/105033210125
U2 - 10.1016/j.aei.2026.104591
DO - 10.1016/j.aei.2026.104591
M3 - 文章
AN - SCOPUS:105033210125
SN - 1474-0346
VL - 73
JO - Advanced Engineering Informatics
JF - Advanced Engineering Informatics
M1 - 104591
ER -