Abstract:
As a storage and application system for knowledge graphs, graph databases excel in handling complex relationships and connectivity among data, enabling efficient search and analysis of entity associations. NebulaGraph, recognized as a typical representative of open-source domestic graph database, possesses a distributed architecture and linear scalability, making it adept at processing ultra-large datasets consisting of hundreds of billions of vertices and trillions of edges. However, when handling large-scale data imports, challenges arise due to independent, fragmented, and difficult-to-serialize operational processes. As a result of inefficient manual operation and poor coordination between various processes, the data import and index association processes become time-consuming, severely constraining application efficiency and user experience in big data scenarios. A kind of data standard applicable to the domestic graph database NebulaGraph is proposed. Based on this standard, the integration of graph space construction, ontology design, data import, index creation, and data association is achieved, significantly simplifying the data import process and substantially enhancing the efficiency of data import. This approach supports the automated creation of large-scale graphs effectively. Moreover, by automatically identifying entities, attributes, and relationships within the imported massive data, a reasonable graph structure can be dynamically generated. This method not only saves the time required for manual graph space construction but also ensures a high degree of alignment between the graph structure and the actual data, thereby enhancing the accuracy and effectiveness of graph analysis. Experiments demonstrate that, for the scenario of importing gigabyte-scale data into NebulaGraph, methods involving human-computer interaction fail to complete data import, while methods that import data sequentially via access to the graph database are excessively time-consuming. In contrast, the proposed method increases the import speed from 596 records per second to 14 331 records per second for 100 million-scale data, reducing the overall processing time by 95.8%. This significantly reduces the time cost of data import and enhances the application value of domestic graph databases in big data scenarios.