Abstract
The valuation of data assets is a fundamental problem in data circulation and data trading systems, yet remains challenging due to heterogeneous descriptions, unclear usage contexts, and the lack of standardized comparability across products. To address this issue, we propose a knowledge-graph-based valuation framework that models data assets and their semantic and relational attributes in a structured graph space. Unstructured product descriptions collected from AWS Data Exchange are processed using a large language model (LLM) ensemble with strict-majority voting, which improves triple extraction stability and mitigates hallucination. The constructed knowledge graph encodes data products, providers, industries, and application concepts. To obtain machine-operable representations, we evaluate two embedding strategies: Node2Vec, which captures global semantic proximity, and GraphSAGE, which preserves localized neighborhood patterns. Experimental results show that Node2Vec achieves superior overall performance, especially in concept-rich sub-graph structures, while GraphSAGE performs better in industry-homogeneous scenarios. Using the learned embeddings, we perform clustering to identify comparable data assets and combine these features with supervised learning models to predict pricing levels. Extensive experiments demonstrate that our framework improves robustness, interpretability, and consistency in data asset valuation.
| Original language | English |
|---|---|
| Article number | 104591 |
| Journal | Advanced Engineering Informatics |
| Volume | 73 |
| DOIs | |
| State | Published - Jul 2026 |
Keywords
- Data asset pricing
- Graph embedding
- Knowledge graph
- Market comparison method
- Prediction performance
- Two layer ontology
Fingerprint
Dive into the research topics of 'Knowledge graph modeling for data asset pricing'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver