Skip to main content
zenodoopen

VOYAGE: A Large Collection of Vocabulary Usage in Open RDF Datasets

<p><strong>List of files:</strong></p> <ul> <li>odps.json: for each of the accessed ODPs, its name, URL, API type, API URL, and the IDs of RDF datasets collected from it <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;name&#39; (string), &#39;URL&#39; (string), &#39;API type&#39; (string), &#39;API URL&#39; (string), and &#39;collected datasets IDs&#39; (list of integers)</p> </li> </ul> </li> <li>datasets.json: for each of the crawled RDF datasets, its ID, title, description, author, license, dump file URLs, and PLDs <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;ID&#39; (integer), &#39;title&#39; (string), &#39;description&#39; (string), &#39;author&#39; (string), &#39;license&#39; (string), &#39;dump file URLs&#39; (list of strings), and &#39;PLDs&#39; (list of strings)</p> </li> </ul> </li> <li>deduplicated_datasets.json: the IDs of the deduplicated RDF datasets and whether they are in the LOD Cloud <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;ID&#39; (integer) and &#39;in LOD Cloud&#39; (boolean)</p> </li> </ul> </li> <li>terms.json: the extracted classes, properties, and the IDs of RDF datasets using each term <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;term&#39; (string), &#39;is class&#39; (boolean), &#39;is property&#39; (boolean), and &#39;used in dataset IDs&#39; (list of integers)</p> </li> </ul> </li> <li>vocabularies.json: the extracted vocabularies, the classes and properties in each vocabulary, and the IDs of RDF datasets using each vocabulary <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;vocabulary&#39; (string), &#39;classes&#39; (list of strings), &#39;properties&#39; (list of strings), and &#39;used in dataset IDs&#39; (list of integers).</p> </li> </ul> </li> <li>edps.json: the extracted distinct EDPs and the IDs of RDF datasets using each EDP <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;classes&#39; (list of strings), &#39;forward properties&#39; (list of strings), &#39;backward properties&#39; (list of strings), and &#39;used in dataset IDs&#39; (list of integers)</p> </li> </ul> </li> <li>clusters.json: the clusters of vocabularies generated by MV-ITCC and LDA <ul> <li> <p>JSON&nbsp;structure:&nbsp;{&quot;LDA&quot;:&nbsp;{&quot;vocabularies&quot;:&nbsp;{VOCABULARY_CLUSTER_ID_1:&nbsp;[LIST_OF_VOCABULARIES],&nbsp;VOCABULARY_CLUSTER_ID_2:&nbsp;[LIST_OF_VOCABULARIES],&nbsp;...}},&nbsp;&quot;MV-ITCC&quot;:&nbsp;{&quot;vocabularies&quot;:&nbsp;{VOCABULARY_CLUSTER_ID_1:&nbsp;[LIST_OF_VOCABULARIES],&nbsp;VOCABULARY_CLUSTER_ID_2:&nbsp;[LIST_OF_VOCABULARIES],&nbsp;...},&nbsp;&quot;dataset&nbsp;IDs&quot;:&nbsp;{DATASET_CLUSTER_ID_1:&nbsp;[LIST_OF_DATASET_IDS],&nbsp;DATASET_CLUSTER_ID_2:&nbsp;[LIST_OF_DATASET_IDS],&nbsp;...}}}</p> </li> </ul> </li> </ul>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
20
Reuse readiness
8
Engagement
0