VOYAGE: A Large Collection of Vocabulary Usage in Open RDF Datasets
<p><strong>List of files:</strong></p> <ul> <li>odps.json: for each of the accessed ODPs, its name, URL, API type, API URL, and the IDs of RDF datasets collected from it <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - 'name' (string), 'URL' (string), 'API type' (string), 'API URL' (string), and 'collected datasets IDs' (list of integers)</p> </li> </ul> </li> <li>datasets.json: for each of the crawled RDF datasets, its ID, title, description, author, license, dump file URLs, and PLDs <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - 'ID' (integer), 'title' (string), 'description' (string), 'author' (string), 'license' (string), 'dump file URLs' (list of strings), and 'PLDs' (list of strings)</p> </li> </ul> </li> <li>deduplicated_datasets.json: the IDs of the deduplicated RDF datasets and whether they are in the LOD Cloud <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - 'ID' (integer) and 'in LOD Cloud' (boolean)</p> </li> </ul> </li> <li>terms.json: the extracted classes, properties, and the IDs of RDF datasets using each term <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - 'term' (string), 'is class' (boolean), 'is property' (boolean), and 'used in dataset IDs' (list of integers)</p> </li> </ul> </li> <li>vocabularies.json: the extracted vocabularies, the classes and properties in each vocabulary, and the IDs of RDF datasets using each vocabulary <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - 'vocabulary' (string), 'classes' (list of strings), 'properties' (list of strings), and 'used in dataset IDs' (list of integers).</p> </li> </ul> </li> <li>edps.json: the extracted distinct EDPs and the IDs of RDF datasets using each EDP <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - 'classes' (list of strings), 'forward properties' (list of strings), 'backward properties' (list of strings), and 'used in dataset IDs' (list of integers)</p> </li> </ul> </li> <li>clusters.json: the clusters of vocabularies generated by MV-ITCC and LDA <ul> <li> <p>JSON structure: {"LDA": {"vocabularies": {VOCABULARY_CLUSTER_ID_1: [LIST_OF_VOCABULARIES], VOCABULARY_CLUSTER_ID_2: [LIST_OF_VOCABULARIES], ...}}, "MV-ITCC": {"vocabularies": {VOCABULARY_CLUSTER_ID_1: [LIST_OF_VOCABULARIES], VOCABULARY_CLUSTER_ID_2: [LIST_OF_VOCABULARIES], ...}, "dataset IDs": {DATASET_CLUSTER_ID_1: [LIST_OF_DATASET_IDS], DATASET_CLUSTER_ID_2: [LIST_OF_DATASET_IDS], ...}}}</p> </li> </ul> </li> </ul>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0