Input JSON data for the pipeline of the CLARA Knowledge Graph
<p><span>/!\</span> This deposit is deprecated; a more complete version of the deposit can be found here: <a href="https://zenodo.org/records/8403142">8403142</a>. <span>/!\</span></p> <p><strong>CLARA</strong><br>This deposit is part of the <a href="https://project.inria.fr/clara/">CLARA project</a>. The CLARA project aims to empower teachers in the task of creating new educational resources. And in particular with the task of handling the licenses of reused educational resources.</p> <p>The present deposit contains the JSON files extracted from the <a href="https://www.x5gon.org/">X5GON</a> Postgresql database. The files are fed to the pipeline of the CLARA project for the creation of 4 different RDF graphs. This is achieved through the use of RDF mappings (<a href="https://rml.io/">RML</a>, <a href="https://ceur-ws.org/Vol-2980/paper374.pdf">RML-star</a>).<br>That pipeline can be found on <a href="https://gitlab.univ-nantes.fr/clara/pipeline">Gitlab</a>.</p> <p>The results of this pipeline can also be found on Zenodo, on those four different deposits:</p> <ul> <li><a href="https://doi.org/10.5281/zenodo.8108855">Standard reification</a></li> <li><a href="http://doi.org/10.5281/zenodo.8108962">Singleton properties</a></li> <li><a href="http://doi.org/10.5281/zenodo.8108947">Named graphs</a></li> <li><a href="http://doi.org/10.5281/zenodo.8108970">RDF-star</a></li> </ul> <p> </p> <p><strong>Content</strong></p> <p>The JSON files contain information on a total of 45K educational resources, linked to a total of 135K subjects (extracted from DBpedia). Each educational resource is linked to the subjects it talks about. Each of those links has two corresponding scores which represent the certainty of the given link. Those scores are <em>"norm_cosine" </em>and <em>"norm_pageRank"</em>.</p> <p>The dataset was cut into multiple JSON files in order to make its processing easier. <br>There are two type of json files in this deposit:</p> <ul> <li><strong>authors_[</strong>X<strong>].json</strong> - Which lists the authors names</li> <li><strong>ER_[</strong>X<strong>].json</strong> - Which lists the educational resources and their related information.<br>That information contains: <ul> <li>their <em>title.</em></li> <li>their <em>description.</em></li> <li>their <em>language</em> (and <em>language_detected</em>, only the first one is used in the pipeline here).</li> <li>their <em>license.</em></li> <li>their <em>mimetype.</em></li> <li>the <em>authors.</em></li> <li>the <em>date</em> of creation of the resource.</li> <li>a <em>url</em> linking to the resource itself.</li> <li>and finally the subjects (named <em>concepts</em>) associated to the resource. With the corresponding scores.</li> </ul> </li> </ul>
ShareScore
12/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 0
- Reuse readiness
- 0
- Engagement
- 0