Skip to main content
zenodorestricted

Input JSON data for the pipeline of the CLARA Knowledge Graph

<p><span>/!\</span>&nbsp;This deposit is deprecated; a more complete version of the deposit can be found here: <a href="https://zenodo.org/records/8403142">8403142</a>. <span>/!\</span></p> <p><strong>CLARA</strong><br>This deposit is part of the <a href="https://project.inria.fr/clara/">CLARA project</a>. The CLARA project aims to empower teachers in the task of creating new educational resources. And in particular with the task of handling the licenses of reused educational resources.</p> <p>The present deposit contains&nbsp;the JSON files extracted from the <a href="https://www.x5gon.org/">X5GON</a> Postgresql database. The files&nbsp;are fed to the&nbsp;pipeline of the CLARA project for the creation of 4 different RDF graphs. This is achieved through the use of RDF mappings (<a href="https://rml.io/">RML</a>, <a href="https://ceur-ws.org/Vol-2980/paper374.pdf">RML-star</a>).<br>That pipeline can be found on <a href="https://gitlab.univ-nantes.fr/clara/pipeline">Gitlab</a>.</p> <p>The results of this pipeline can also be found on Zenodo, on those four different deposits:</p> <ul> <li><a href="https://doi.org/10.5281/zenodo.8108855">Standard reification</a></li> <li><a href="http://doi.org/10.5281/zenodo.8108962">Singleton properties</a></li> <li><a href="http://doi.org/10.5281/zenodo.8108947">Named graphs</a></li> <li><a href="http://doi.org/10.5281/zenodo.8108970">RDF-star</a></li> </ul> <p>&nbsp;</p> <p><strong>Content</strong></p> <p>The JSON files contain information on a total of 45K educational resources, linked to a total of 135K subjects (extracted from DBpedia). Each educational resource is&nbsp;linked to the&nbsp;subjects it talks about. Each of those links has two corresponding scores which represent the certainty of the given link. Those scores are <em>"norm_cosine" </em>and <em>"norm_pageRank"</em>.</p> <p>The dataset was cut into multiple JSON files in order to make its processing easier.&nbsp;<br>There are two type of json files in this deposit:</p> <ul> <li><strong>authors_[</strong>X<strong>].json</strong> - Which lists the authors names</li> <li><strong>ER_[</strong>X<strong>].json</strong>&nbsp;- Which lists the educational resources and their related information.<br>That information contains: <ul> <li>their <em>title.</em></li> <li>their <em>description.</em></li> <li>their <em>language</em> (and <em>language_detected</em>, only the first one is used in the pipeline here).</li> <li>their <em>license.</em></li> <li>their <em>mimetype.</em></li> <li>the&nbsp;<em>authors.</em></li> <li>the <em>date</em> of creation of the resource.</li> <li>a&nbsp;<em>url</em>&nbsp;linking to the resource itself.</li> <li>and finally the subjects (named&nbsp;<em>concepts</em>) associated to the resource. With the corresponding scores.</li> </ul> </li> </ul>

ShareScore

12/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
0
Reuse readiness
0
Engagement
0

Topics