Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
33
datasets available to search
ShareScore release 0.9.0
Dataset results
33 results for “research graph”
Expert judgements for evaluating deduplication of OpenAIRE Research Graph
<p>Expert judgments used to evaluate the deduplication algorithm used for constructing the OpenAIRE Research Graph.</p> <p>Each expert assigned each group with one of the following predetermined classes (that also indicate whether contained entities are equivalent or not):</p> <ul> <li>AMBIGUOUS: At least one DOIs is invalid (no metadata are available) (N/A)</li> <li>DELETED-DUPLICATES: DOIs once pointing to the same research object, currently deleted. (TRUE)</li> <li>MULTI-PUBLISHED: Article published in more than one locations (full or abstract) (TRUE)</li> <li>VERSIONS: Multiple versions of the same research object (e.g. pre-prints, post-prints etc). (TRUE)</li> <li>ERRONEOUS: Unrelated set of objects. (FALSE)</li> <li>PAPER-EXTENSIONS: Extended version of a conference paper in a journal. (FALSE)</li> <li>PART-OF-A-GROUP: Multiple parts of the same research object (e.g. multi-part publication, photos of the same collection etc). (FALSE)</li> <li>SUPPLEMENTARY: A publication and its supplementary material (including errata). (FALSE)</li> </ul> <p>The following are provided:</p> <ul> <li>OpenAIRE identifier</li> <li>Judgement</li> <li>Count of numbers in group</li> <li>DOIs in group</li> </ul>
OpenAIRE Graph: Dataset for research communities and initiatives
<p>This dataset contains metadata records of the OpenAIRE Graph relevant for the research communities and initiatives collaborating with OpenAIRE and with a public Community Gateway on <a href="https://connect.openaire.eu">OpenAIRE CONNECT</a> as of July 2025.</p> <p>Each file is a tar archive containing gzip files with one json per line. Each json is compliant to the schema available at <a href="https://doi.org/10.5281/zenodo.14891476" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.14891476</a></p> <table style="width: 100%; height: 822.939px;"> <tbody> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;"><strong>Research community name</strong></td> <td style="width: 38.9414%; height: 19.5938px;"><strong>File name</strong></td> <td style="width: 23.1067%; height: 19.5938px;"><strong>URL to the gateway</strong></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Argo France</td> <td style="width: 38.9414%; height: 19.5938px;">argo-france.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://argo-france.openaire.eu/">https://argo-france.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Aurora University Alliance</td> <td style="width: 38.9414%; height: 19.5938px;">aurora.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://aurora.openaire.eu">https://aurora.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Transport Research (EC projects BE OPEN and SciLake)</td> <td style="width: 38.9414%; height: 19.5938px;">beopen.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://beopen.openaire.eu">https://beopen.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">CIVICA Alliance</td> <td style="width: 38.9414%; height: 19.5938px;">civica.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://civica.openaire.eu" target="_blank" rel="noopener">https://civica.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">COVID 19</td> <td style="width: 38.9414%; height: 19.5938px;">covid-19.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://covid-19.openaire.eu">https://covid-19.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">DARIAH EU</td> <td style="width: 38.9414%; height: 19.5938px;">dariah.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://dariah.openaire.eu">https://dariah.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Digital Humanities and Cultural Heritage</td> <td style="width: 38.9414%; height: 19.5938px;">dh-ch.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://dh-ch.openaire.eu">https://dh-ch.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Virtual Twins in Health (EC project EDITH)</td> <td style="width: 38.9414%; height: 19.5938px;">dth.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://dth.openaire.eu">https://dth.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">European Digital Innovation Hubs Network ADRIA</td> <td style="width: 38.9414%; height: 19.5938px;">edih-adria.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://edih-adria.openaire.eu">https://edih-adria.openaire.eu</a></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 33.9453%; height: 39.1875px;">European Geothermal Research and Innovation Search Engine </td> <td style="width: 38.9414%; height: 39.1875px;">egrise.tar</td> <td style="width: 23.1067%; height: 39.1875px;"><a href="https://egrise.openaire.eu">https://egrise.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">ELIXIR Greece</td> <td style="width: 38.9414%; height: 19.5938px;">elixir-gr.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://elixir-gr.openaire.eu">https://elixir-gr.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Energy Planning (EC project SciLake)</td> <td style="width: 38.9414%; height: 19.5938px;">energy-planning_1.tar, energy_planning_2.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://energy-planning.openaire.eu/">https://energy-planning.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Energy Research (EC project Enermaps)</td> <td style="width: 38.9414%; height: 19.5938px;">enermaps.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://enermaps.openaire.eu">https://enermaps.openaire.eu</a></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 33.9453%; height: 39.1875px;">European University for Smart Urban Coastal Sustainability</td> <td style="width: 38.9414%; height: 39.1875px;">eu-conexus.tar</td> <td style="width: 23.1067%; height: 39.1875px;"><a href="https://eu-conexus.openaire.eu/">https://eu-conexus.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">European University of Technology+</td> <td style="width: 38.9414%; height: 19.5938px;">eut.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://eut.openaire.eu/">https://eut.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">EUTOPIA Alliance</td> <td style="width: 38.9414%; height: 19.5938px;">eutopia.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://eutopia.openaire.eu/">https://eutopia.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">FORTHEM Alliance</td> <td style="width: 38.9414%; height: 19.5938px;">forthem.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://forthem.openaire.eu">https://forthem.openaire.eu</a></td> </tr> <tr> <td style="width: 33.9453%;">[NEW] GoTriple </td> <td style="width: 38.9414%;">gotriple_1.tar, gotriple_2.tar</td> <td style="width: 23.1067%;"><a href="https://gotriple.openaire.eu/">https://gotriple.openaire.eu/</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Heritage Science (EC project IPERION HS)</td> <td style="width: 38.9414%; height: 19.5938px;">heritage-science.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://heritage-science.openaire.eu/">https://heritage-science.openaire.eu</a></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 33.9453%; height: 39.1875px;">Institut national de recherche en informatique et en automatique (EC project GraspOS)</td> <td style="width: 38.9414%; height: 39.1875px;">inria.tar</td> <td style="width: 23.1067%; height: 39.1875px;"><a href="https://inria.openaire.eu" target="_blank" rel="noopener">https://inria.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">IPERION HS</td> <td style="width: 38.9414%; height: 19.5938px;">iperionhs.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://iperionhs.openaire.eu">https://iperionhs.openaire.eu</a></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 33.9453%; height: 39.1875px;">Knowmad Institut</td> <td style="width: 38.9414%; height: 39.1875px;">knowmad_1.tar, knowmad_2.tar, knowmad_3.tar, knowmad_4.tar, knowmad_5.tar, knowmad_6.tar, knowmad_7.tar, knowmad_8.tar, knowmad_9.tar</td> <td style="width: 23.1067%; height: 39.1875px;"><a href="https://knowmad.openaire.eu/">https://knowmad.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">LifeWathc ERIC</td> <td style="width: 38.9414%; height: 19.5938px;">lifewatch-eric.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://lifewatch-eric.openaire.eu/">https://lifewatch-eric.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">European Marine Science</td> <td style="width: 38.9414%; height: 19.5938px;">mes.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://mes.openaire.eu">https://mes.openaire.eu</a></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 33.9453%; height: 39.1875px;">Atmospheric Research Community (EC project NEANIAS)</td> <td style="width: 38.9414%; height: 39.1875px;">neanias-atmospheric.tar</td> <td style="width: 23.1067%; height: 39.1875px;"><a href="https://neanias-atmospheric.openaire.eu/">https://neanias-atmospheric.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Space Research Community (EC project NEANIAS)</td> <td style="width: 38.9414%; height: 19.5938px;">neanias-space.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://neanias-space.openaire.eu/">https://neanias-space.openaire.eu</a></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 33.9453%; height: 39.1875px;">Underwater Research Community (EC project NEANIAS)</td> <td style="width: 38.9414%; height: 39.1875px;">neanias-underwater.tar</td> <td style="width: 23.1067%; height: 39.1875px;"><a href="https://neanias-underwater.openaire.eu/">https://neanias-underwater.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Netherlands Research Portal</td> <td style="width: 38.9414%; height: 19.5938px;">netherlands_1.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://netherlands.openaire.eu/">https://netherlands.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">[NEW] Neuroscience<br>(EC project SciLake - former Neuroinformatics community is now a subcommunity of Neuroscience)</td> <td style="width: 38.9414%; height: 19.5938px;">neuroscience_1.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://neuroscience.openaire.eu">https://neuroscience.openaire.eu</a></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 33.9453%; height: 39.1875px;">North American Studies</td> <td style="width: 38.9414%; height: 39.1875px;">north-american-studies.tar</td> <td style="width: 23.1067%; height: 39.1875px;"><a href="https://north-american-studies.openaire.eu">https://north-american-studies.openaire.eu</a></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 33.9453%; height: 39.1875px;">Rural Digital Europe (EC project DESIRA)</td> <td style="width: 38.9414%; height: 39.1875px;">rural-digital-europe.tar</td> <td style="width: 23.1067%; height: 39.1875px;"><a href="https://rural-digital-europe.openaire.eu/">https://rural-digital-europe.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Sustainable Development Solutions Network - Greece </td> <td style="width: 38.9414%; height: 19.5938px;">sdsn-gr.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://sdsn-gr.openaire.eu/">https://sdsn-gr.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">Technological University Network</td> <td style="width: 38.9414%; height: 19.5938px;">tunet.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://tunet.openaire.eu">https://tunet.openaire.eu</a></td> </tr> <tr> <td style="width: 33.9453%;">[NEW] UNITE! University Alliance</td> <td style="width: 38.9414%;">unite.tar</td> <td style="width: 23.1067%;"><a href="https://unite.openaire.eu">https://unite.openaire.eu</a></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 33.9453%; height: 19.5938px;">University of the Arctic (UArctic) </td> <td style="width: 38.9414%; height: 19.5938px;">uarctic_1.tar, uarctic_2.tar</td> <td style="width: 23.1067%; height: 19.5938px;"><a href="https://uarctic.openaire.eu/">https://uarctic.openaire.eu</a></td> </tr> </tbody> </table>
Supplementary data files for manuscript titled "From spreadsheet lab data templates to knowledge graphs: A FAIR data journey in the domain of AMR research"
<div>This data repository contains all the necessary supplementary files for the manuscript titled "<strong>From spreadsheet lab data templates to knowledge graphs: A FAIR data journey in the domain of AMR research.</strong>"</div> <div> </div> <div>The repository is a copy of the <a href="https://github.com/IMI-COMBINE/template2graphs">GitHub page</a> with the source code used to generate the graph and additional files required for the Lab Data Template.</div> <div> </div> <div>Below we provide a brief overview of the data files in the `additional folder` and their underlying purpose:</div> <div> <ul> <li>The <strong>Data Survey</strong> collects relevant project and data set information to set up a Data Management Plan. It can serve as an input for Lab Data Template development.</li> <li>The <strong>Lab Data Templates</strong> facilitate the collection of AMR research data (in vivo and in vitro) in several sub-tables. The Excel format is compatible with upload procedures into the data repository 'grit' and serves as input for a knowledge graph workflow.</li> <li>The <strong>Data dictionary</strong> is connected to the Lab Data Templates and ensures harmonized data entries. In addition, the dictionaries collect metadata beyond the content of the Lab Data Template (e.g. bacterial strain information or compound information) and link to ontologies where possible.</li> <li>The <strong>FAIR assessments</strong> have been used as a primer for improving the template. This report is generated using the FAIR-DSM model.</li> </ul> </div> <div>The templates have been used during the IMI2 GNA NOW project to collect information and have been improved according to FAIR standards in collaboration with the IMI FAIRplus project ("post FAIRification").</div>
NeSy4VRD: A Multifaceted Resource for Neurosymbolic AI Research using Knowledge Graphs in Visual Relationship Detection
<p><strong>NeSy4VRD</strong></p> <p>NeSy4VRD is a multifaceted, multipurpose resource designed to foster neurosymbolic AI (NeSy) research, particularly NeSy research using Semantic Web technologies such as OWL ontologies, OWL-based knowledge graphs and OWL-based reasoning as symbolic components. The NeSy4VRD research resource pertains to the <em>computer vision</em> field of AI and, within that field, to the application tasks of <em>visual relationship detection (VRD) and scene graph generation</em>.</p> <p>Whilst the core motivation of the NeSy4VRD research resource is to foster computer vision-based NeSy research using Semantic Web technologies such as OWL ontologies and OWL-based knowledge graphs, AI researchers can readily use NeSy4VRD to either: 1) pursue computer vision-based NeSy research without involving Semantic Web technologies as symbolic components, or 2) pursue computer vision research without NeSy (i.e. pursue research that focuses purely on deep learning alone, without involving symbolic components of any kind). This is the sense in which we describe NeSy4VRD as being <em>multipurpose</em>: it can readily be used by diverse groups of computer vision-based AI researchers with diverse interests and objectives.</p> <p>The NeSy4VRD research resource in its entirety is distributed across two locations: Zenodo and GitHub.</p> <p> </p> <p><strong>NeSy4VRD on Zenodo: the NeSy4VRD dataset package</strong></p> <p>This entry on Zenodo hosts the <em>NeSy4VRD dataset package</em>, which includes the <em>NeSy4VRD dataset</em> and its companion <em>NeSy4VRD ontology</em>, an OWL ontology called VRD-World.</p> <p>The <em>NeSy4VRD dataset</em> consists of an image dataset with associated visual relationship annotations. The images of the <em>NeSy4VRD dataset</em> are the same as those that were once publicly available as part of the <a href="https://cs.stanford.edu/people/ranjaykrishna/vrd/">VRD</a> dataset. The NeSy4VRD visual relationship annotations are a highly customised and quality-improved version of the original VRD visual relationship annotations. The <em>NeSy4VRD dataset</em> is designed for computer vision-based research that involves detecting objects in images and predicting relationships between ordered pairs of those objects. A visual relationship for an image of the <em>NeSy4VRD dataset</em> has the form <'subject', 'predicate', 'object'>, where the 'subject' and 'object' are two objects in the image, and the 'predicate' describes some relation between them. Both the 'subject' and 'object' objects are specified in terms of bounding boxes and object classes. For example, representative annotated visual relationships are <'person', 'ride', 'horse'>, <'hat', 'on', 'teddy bear'> and <'cat', 'under', 'pillow'>.</p> <p>Visual relationship detection is pursued as a computer vision application task in its own right, and as a building block capability for the broader application task of scene graph generation. Scene graph generation, in turn, is commonly used as a precursor to a variety of enriched, downstream visual understanding and reasoning application tasks, such as image captioning, visual question answering, image retrieval, image generation and multimedia event processing.</p> <p>The <em>NeSy4VRD ontology</em>, VRD-World, is a rich, well-aligned, companion OWL ontology engineered specifically for use with the <em>NeSy4VRD dataset.</em> It directly describes the domain of the <em>NeSy4VRD dataset</em>, as reflected in the NeSy4VRD visual relationship annotations. More specifically, all of the object classes that feature in the NeSy4VRD visual relationship annotations have corresponding classes within the VRD-World OWL class hierarchy, and all of the predicates that feature in the NeSy4VRD visual relationship annotations have corresponding properties within the VRD-World OWL object property hierarchy. The rich structure of the VRD-World class hierarchy and the rich characteristics and relationships of the VRD-World object properties together give the VRD-World OWL ontology rich inference semantics. These provide ample opportunity for OWL reasoning to be meaningfully exercised and exploited in NeSy research that uses OWL ontologies and OWL-based knowledge graphs as symbolic components. There is also ample potential for NeSy researchers to explore supplementing the OWL reasoning capabilities afforded by the VRD-World ontology with Datalog rules and reasoning.</p> <p>Use of the <em>NeSy4VRD ontology</em>, VRD-World, in conjunction with the <em>NeSy4VRD dataset </em>is, of course, purely optional, however. Computer vision AI researchers who have no interest in NeSy, or NeSy researchers who have no interest in OWL ontologies and OWL-based knowledge graphs, can ignore the <em>NeSy4VRD ontology</em> and use the <em>NeSy4VRD dataset </em>by itself.</p> <p>All computer vision-based AI research user groups can, if they wish, also avail themselves of the other components of the NeSy4VRD research resource available on GitHub.</p> <p> </p> <p><strong>NeSy4VRD on GitHub: open source infrastructure supporting extensibility, and sample code</strong></p> <p>The NeSy4VRD research resource incorporates additional components that are companions to the <em>NeSy4VRD dataset package</em> here on Zenodo. These companion components are available at <a href="https://github.com/djherron/NeSy4VRD/">NeSy4VRD on GitHub</a>. These companion components consist of:</p> <ul> <li>comprehensive open source Python-based infrastructure supporting the extensibility of the NeSy4VRD visual relationship annotations (and, thereby, the extensibility of the <em>NeSy4VRD ontology</em>, VRD-World, as well)</li> <li>open source Python sample code showing how one can work with the NeSy4VRD visual relationship annotations in conjunction with the <em>NeSy4VRD ontology</em>, VRD-World, and RDF knowledge graphs.</li> </ul> <p>The NeSy4VRD infrastructure supporting extensibility consists of:</p> <ul> <li>open source Python code for conducting deep and comprehensive analyses of the <em>NeSy4VRD dataset</em> (the VRD images and their associated NeSy4VRD visual relationship annotations)</li> <li>an open source, custom-designed <em>NeSy4VRD protocol</em> for specifying visual relationship annotation customisation instructions declaratively, in text files</li> <li>an open source, custom-designed <em>NeSy4VRD workflow, </em>implemented using Python scripts and modules, for applying small or large volumes of customisations or extensions to the NeSy4VRD visual relationship annotations in a configurable, managed, automated and repeatable process.</li> </ul> <p>The purpose behind providing comprehensive infrastructure to support extensibility of the NeSy4VRD visual relationship annotations is to make it easy for researchers to take the <em>NeSy4VRD dataset</em> in new directions, by further enriching the annotations, or by tailoring them to introduce new or more data conditions that better suit their particular research needs and interests. The option to use the NeSy4VRD extensibility infrastructure in this way applies equally well to each of the diverse potential NeSy4VRD user groups already mentioned.</p> <p>The NeSy4VRD extensibility infrastructure, however, may be of particular interest to NeSy researchers interested in using the <em>NeSy4VRD ontology</em>, VRD-World, in conjunction with the <em>NeSy4VRD dataset. </em>These researchers can of course tailor the VRD-World ontology if they wish without needing to modify or extend the NeSy4VRD visual relationship annotations in any way. But their degrees of freedom for doing so will be limited by the need to maintain alignment with the NeSy4VRD visual relationship annotations and the particular set of object classes and predicates to which they refer. If NeSy researchers want full freedom to tailor the VRD-World ontology, they may well need to tailor the NeSy4VRD visual relationship annotations first, in order that alignment be maintained.</p> <p>To illustrate our point, and to illustrate our vision of how the NeSy4VRD extensibility infrastructure can be used, let us consider a simple example. It is common in computer vision to distinguish between <em>thing</em> objects (that have well-defined shapes) and <em>stuff</em> objects (that are amorphous). Suppose a researcher wishes to have a greater number of <em>stuff</em> object classes with which to work. Water is such a <em>stuff</em> object. Many VRD images contain water but it is not currently one of the annotated object classes and hence is never referenced in any visual relationship annotations. So adding a <em>Water</em> class to the class hierarchy of the VRD-World ontology would be pointless because it would never acquire any instances (because an object detector would never detect any). However, our hypothetical researcher could choose to do the following:</p> <ul> <li>use the analysis functionality of the NeSy4VRD extensibility infrastructure to find images containing water (by, say, searching for images whose visual relationships refer to object classes such as 'boat', 'surfboard', 'sand', 'umbrella', etc.);</li> <li>use free image analysis software (such as GIMP, at gimp.org) to get bounding boxes for instances of water in these images;</li> <li>use the <em>NeSy4VRD protocol</em> to specify new visual relationships for these images that refer to the new 'water' objects (e.g. <'boat', 'on', 'water'>);</li> <li>use the <em>NeSy4VRD workflow</em> to introduce the new object class 'water' and to apply the specified new visual relationships to the sets of annotations for the affected images;</li> <li>introduce class Water to the class hierarchy of the VRD-World ontology (using, say, the free Protege ontology editor);</li> <li>continue experimenting, now with the added benefit of the additional <em>stuff</em> object class 'water';</li> <li>contribute the enriched set of NeSy4VRD visual relationship annotations, and the enriched companion VRD-World ontology, to research communities.</li> </ul> <p> </p> <p><strong>Information pertaining to the VRD dataset</strong></p> <p>Information about the original VRD dataset is available <a href="https://cs.stanford.edu/people/ranjaykrishna/vrd/">here</a>. </p> <p>Public availability of the VRD images (via information accessible from that location) ceased sometime in the latter part of 2021. We thank Dr. Ranjay Krishna, one of the principals associated with the VRD dataset, for granting us permission to re-establish the public availability of the VRD images as part of NeSy4VRD.</p> <p>The original VRD visual relationship annotations are still publicly available from that location. But our deep analysis of those annotations, driven by our desire to design a robust companion ontology, revealed them to be highly problematic in many ways that made credible ontology modelling infeasible. They were also found to be replete with all manner of errors. The NeSy4VRD visual relationship annotations are far superior and we recommend them over the original VRD annotations to anyone contemplating conducting research using the VRD images. The NeSy4VRD annotations also have the added benefit of the rich, well-aligned companion <em>NeSy4VRD ontology</em>, VRD-World, for those whose research requires such a companion ontology.</p> <p>Researchers wishing to use the original VRD dataset may still do so. They can access the VRD images here, from within the <em>NeSy4VRD dataset</em> on Zenodo, and access the VRD visual relationship annotations from the location in the link.</p> <p><em>A note of caution</em>: the <em>NeSy4VRD ontology</em>, VRD-World, is <em>not</em><strong> </strong>compatible with the original VRD visual relationship annotations and cannot be used in conjunction with them. The VRD-World ontology has been engineered in relation to the highly customised and quality-improved NeSy4VRD visual relationship annotations. The customisations that were applied include ones that introduced many new object classes, merged some of the existing object classes, introduced one new predicate, and changed several predicate names.</p> <p>However, researchers can, if they wish, use the NeSy4VRD extensibility infrastructure (described above) to undertake their own customisation and quality-improvement exercise with respect to the original VRD visual relationship annotations. This is precisely how the NeSy4VRD visual relationship annotations were created in the first place. The primary intended use case of NeSy4VRD's extensibility infrastructure, however, is for researchers to use the NeSy4VRD visual relationship annotations as their starting point, and to take these annotations forward with onward customisations and extensions, as illustrated in the example use case given above.</p> <p> </p> <p> </p>
Medieval manuscripts and their migrations: Using SPARQL to investigate the research potential of an aggregated Knowledge Graph
<p>This dataset contains the <strong>SPARQL queries</strong> presented and discussed in our article published in <em>Digital Medievalist</em> 2022 (as a PDF file), together with the <strong>results of those queries</strong> as CSV files. The query and step numbering follows that given in the article.</p> <p>The queries can be run against the SPARQL endpoint for the <strong>Mapping Manuscript Migrations</strong> project: <a href="https://ldf.fi/mmm/sparql">https://ldf.fi/mmm/sparql</a></p> <p>The full <strong>Mapping Manuscript Migrations dataset </strong>can also be downloaded from the Zenodo repository and installed in your own triple store: <a href="https://zenodo.org/record/4440464">https://zenodo.org/record/4440464</a></p> <p>When copying and pasting these SPARQL queries into a SPARQL client like <a href="https://yasgui.triply.cc/">YASGUI</a>, please check that the line numbering has been copied over correctly. Copying from a PDF file can sometimes break a single long line into multiple separate lines, which will cause a SPARQL validation error.</p> <p>The CSV files contain the results of the queries when run against the Mapping Manuscript Migrations SPARQL endpoint as of 17 December 2021. Please note that Query 2, Step 2, produces no results, so a CSV file has not been provided.</p> <p>The<strong> Mapping Manuscript Migrations portal </strong>can be found at <a href="https://mappingmanuscriptmigrations.org/en/">https://mappingmanuscriptmigrations.org/en/ </a></p> <p>SPARQL tutorials are included in the project's <strong>GitHub documentation</strong>: <a href="https://mapping-manuscript-migrations.github.io/">https://mapping-manuscript-migrations.github.io/</a></p>
Evaluation Set - Contributions Similarity in the Open Research Knowledge Graph
<p>This evaluation set has been created for evaluating a content-based recommender system in the context of the Open Research Knowledge Graph (ORKG). The recommender system accepts structured ORKG contribution as input and recommends existing contributions in the ORKG semantically relevant to the given one.</p> <p> </p> <p>The evaluation set is manually annotated based on the <a href="https://www.orkg.org/orkg/featured-comparisons">featured comparisons</a> in the ORKG. In the course of this, it has been distinguished between homogeneous (those who are dissimilar in 2-3 properties) and heterogeneous (otherwise) instances. Multiple annotations have been obtained for the former and exactly one for the latter.</p> <p> </p> <p>It has been also distinguished between "with_response" and "without_response" instances (50 instances for each). The former are those contributions for them the initial version of the contributions similarity service has found similarities and the latter are the opposite case.</p> <p> </p> <p>This evaluation set has been created and applied on a modified version of the contributions similarity service in the context of <a href="https://doi.org/10.15488/11834">this master's thesis</a>. The modified version of the service has simplified the document representation of contributions that are stored in an ElasticSearch index by omitting redundant terms.</p> <p>The evaluation set has the following schema:</p> <pre><code class="language-json">{ "with_response": [ { "contribution_id": "some_id", "comparison_id": "some_id", "comparison_label": "some_label", "contribution_label": "some_label", "paper": "some_id", "research_field": "some_id", "research_problems": [ "some_id" ], "annotations": [ "some_id of a similar contribution", ... ] }, ... ], "without_response": [ ... ] }</code></pre> <p> </p>
MIRA-KG: A Knowledge Graph of Hypotheses and Findings for Social Demography Research
<p>A shift in scientific publishing from paper-based to knowledge-based practices promotes reproducibility, machine actionability and knowledge discovery. This is important for disciplines like social science, as study indicators are often social constructs such as race or education; hypothesis tests are challenging to compare in demographic research due to their limited temporal and spatial coverage; and natural language in research papers is often imprecise and ambiguous. Therefore, we present the MIRA-KG, consisting of: (1) an ontology for capturing social demography research, which links hypotheses and findings to evidence, (2) annotations of papers on health inequality in terms of the ontology, gathered by (i) prompting a Large Language Model to annotate paper abstracts using the ontology, (ii) mapping concepts to terms from NCBO BioPortal ontologies and GeoNames, and (iii) refining the final graph by a set of SHACL constraints, developed according to data quality criteria. The utility of the resource lies in its use for formally representing social demography research hypotheses, discovering research biases, discovery of knowledge, and the derivation of novel questions.<br><br>This dataset was generated using the code available on Github at <a href="https://w3id.org/mira/">https://w3id.org/mira/</a> at version v1.0. It uses the following ontology: <a href="https://w3id.org/mira/ontology/">https://w3id.org/mira/ontology/</a>. </p>
OpenAire Research Graph linked with OpenAlex
<p>This package contains linked datasets of OpenAire Research Graph and OpenAlex. </p> <p>Files descriptions:</p> <p>- author_to_publication_dic.json contains a mapping of authors to their publications</p> <p>- downloads_views_dic.json contains mappings of the publication id to the number of its downloads and views</p> <p>- id_doi_dic.json contains a mapping of the publication id to its doi</p> <p>- merged1..5.json contain all publication data from the OARG dataset</p> <p>- necessary_fields_dic.json contains extracted publications’ fields necessary for the work</p> <p>- oarg_ref_rel_dic.json contains mapping of publication id to referenced and related work present in OpenAlex dataset</p> <p>- openalex_found_publications5_4.json contains all data on found publications from the OpenAlex</p> <p>- publication_to_author_dic.json contains a mapping of publications to their authors</p>
OpenAIRE Graph: dataset for research community in Virtual Human Twins
<p>This dataset contains metadata records of publications, research data, software and projects relevant for the research community in Virtual Twins in health.<br>The dump contains the records available in the <a href="https://dth.openaire.eu/" target="_blank" rel="noopener">OpenAIRE Gateway on Digital Twins in Health</a> of the <a href="https://www.edith-csa.eu/" target="_blank" rel="noopener">EDITH CSA project </a>of the European Commission (grant agreement n. 101083771).</p> <p>Records are identified via full-text mining and inference techniques applied to the <a href="https://graph.openaire.eu/">OpenAIRE Graph</a>.<br>The OpenAIRE Graph is one of the largest Open Access collections of metadata records and links between publications, datasets, software, projects, funders, and organizations, aggregating thousands of scholarly data sources world-wide.</p> <p>The dump consists of a tar archive containing gzip files with one json per line.<br>Each json is compliant to the schema available at <a href="https://doi.org/10.5281/zenodo.10519297">https://doi.org/10.5281/zenodo.10519297</a>.</p>
Books from the OpenAIRE Research Graph
<p>This dataset is the subset of the OpenAIRE Research Graph about research products of type "Book".</p> <p>The tar archive contains gz files, each with one json per line. Each json compliant to the schema available at <a href="http://doi.org/10.5281/zenodo.5799514">http://doi.org/10.5281/zenodo.5799514</a>. </p>
Combat-TB-NeoDB: fostering Tuberculosis research through integrative analysis using graph database technologies.
<p>NeoDB is a free and open source integrated M.tuberculosis ‘omics’ knowledge-base. NeoDB is based on Neo4j and enables researchers to execute complex federated queries by linking well-known, curated and widely used biological data resources, and supplementary TB variants data from published literature.</p> <p>Documentation can be found at https://combat-tb-db.readthedocs.io</p>
Description of the Features of the Open Research Knowledge Graph as a Crowdsourcing Platform based on the 4 Pillars of Crowdsourcing
<p>This dataset provides the description of the features of the <a href="https://www.orkg.org/orkg/">Open Research Knowledge Graph</a> (ORKG) based on the 4 pillars of crowdsourcing according to the reference model for crowdsourcing by Hosseini et al. [1]. This overview represents the features of the current implementation status of ORKG as a crowdsourcing platform.</p> <p>[1] M. Hosseini, K. Phalp, J. Taylor, and R. Ali, "<a href="https://ieeexplore.ieee.org/abstract/document/6861072?casa_token=zWTHNBHeH6kAAAAA:1n3EJjajqgSkSk154g4DlNFAmJs_7o3KY7LnobvP_W7AUIr-FvM5OyGP1FaRn68zUX-2oYHh7A">The Four Pillars of Crowdsourcing: A Reference Model</a>", in 2014 IEEE 8th International Conference on Research Challenges in Information Science (RCIS). IEEE, 2014, pp. 1–12.</p>
LLM-Based Knowledge Graph Construction from Materials Research Scientific Literature
<p>This dataset was constructed by creating a benchmark of 349 manually annotated triples, which were extracted from four different research articles in the field of materials science.</p>
A traffic graph dataset for autonomous driving research: Commonroad-Nuplan-Dataset
<p>We provide a standardized graph dataset for traffic based on the large-scale <a href="https://www.nuscenes.org/nuplan">NuPlan </a>v1.1 dataset, converted to a <a href="https://pytorch-geometric.readthedocs.io/en/latest/">PyTorch-Geometric </a>dataset using our <a href="https://commonroad.in.tum.de/tools/commonroad-geometric">CommonRoad-Geometric </a>tool. The dataset is collected from 3 megacities: Singapore, Boston and Pittsburgh.</p>
Workflow for structured literature reviews using the Open Research Knowledge Graph (ORKG)
<p>Figure showing a workflow of making a structured literature review using the core features of the Open Research Knowledge Graph (ORKG). </p>
Co-authoring graphs of research teams in a laboratory in computer science
<p>Our aim is to study inter-organisational collaborations initiated by researchers in their research activity. We considered the co-authoring graph involving at least researchers from LORIA (<a href="https://www.loria.fr/fr/">https://www.loria.fr/fr/</a>), a French laboratory in computer science.</p> <p>The dataset is collected from the open French archive HAL (<a href="https://data.archives-ouvertes.fr/">https://data.archives-ouvertes.fr/</a>).</p> <p>Each file encodes (in <a href="http://www.graphviz.org/about/">DOT</a>) the co-authoring graph of a team of LORIA. A node represents a researcher, two nodes are linked only if the corresponding researchers published together over the three considered years 2017, 2018 and 2019. An affiliation attribute is attached to each considered node.</p> <p>The name of researchers and the teams as well as the HAL:id are anonymised. Only affiliations remain the same.</p>
Dataset - Templates Recommendation in the Open Research Knowledge Graph
<p>This dataset has been created for implementing a content-based recommender system in the context of the Open Research Knowledge Graph (ORKG). The recommender system accepts research paper's title and abstracts as input and recommends existing templates in the ORKG semantically relevant to the given paper.</p> <p> </p> <p>Two approaches have been trained on this dataset in the context of <a href="https://doi.org/10.15488/11834">this master's thesis</a>, namely a Natural Language Inference (NLI) approach based on SciBERT embeddings and an unsupervised approach based on ElasticSearch.</p> <p> </p> <p>This publication consists therefore of one general dataset, two training sets for each approach, validation set for the supervised approach and a test set for both approaches.</p> <p> </p> <p><strong>dataset.json</strong></p> <p>The main JSON object consists of a list of templates and a list of neutral papers.</p> <p>Each template object has an ID, label, list of research fields, list of properties and list of papers using that template, whereas each paper object has ID, label, DOI, research field and abstract.</p> <p>Each neutral paper object has the same schema of a paper object using that template.</p> <p>See an example instance below.</p> <p> </p> <pre><code class="language-json">{ "templates": [ { "id": "R138668", "label": "Psychiatric Disorders AI Overview", "research_fields": [ { "id": "http://orkg.org/orkg/resource/R133", "label": "Artificial Intelligence" } ... ], "properties": [ "Study cohort", ... ], "papers": [ { "id": "R138698", "label": "Application of Autoencoder in Depression Diagnosis", "doi": "10.12783/dtcse/csma2017/17335", "research_field": { "id": "R104", "label": "Bioinformatics" }, "abstract": "Major depressive disorder (MDD) is a mental disorder characterized by at least two weeks of low mood which is present across most situations. Diagnosis of MDD using rest-state functional magnetic resonance imaging (fMRI) data faces many challenges due to the high dimensionality, small samples, noisy and individual variability. No method can automatically extract discriminative features from the origin time series in fMRI images for MDD diagnosis. In this study, we proposed a new method for feature extraction and a workflow which can make an automatic feature extraction and classification without a prior knowledge. An autoencoder was used to learn pre-training parameters of a dimensionality reduction process using 3-D convolution network. Through comparison with the other three feature extraction methods, our method achieved the best classification performance. This method can be used not only in MDD diagnosis, but also other similar disorders." }, ... }, ... ] "neutral_papers": [ { "id": "R109377", "label": "Structural basis of SARS-CoV-2 3CLpro and anti-COVID-19 drug discovery from medicinal plants", "doi": "10.1016/j.jpha.2020.03.009", "research_field": { "id": "R104", "label": "Bioinformatics" }, "abstract": "Abstract The recent outbreak of coronavirus disease 2019 (COVID-19) caused by SARS-CoV-2 in December 2019 raised global health concerns. The viral 3-chymotrypsin-like cysteine protease (3CLpro) enzyme controls coronavirus replication and is essential for its life cycle. 3CLpro is a proven drug discovery target in the case of severe acute respiratory syndrome coronavirus (SARS-CoV) and middle east respiratory syndrome coronavirus (MERS-CoV). Recent studies revealed that the genome sequence of SARS-CoV-2 is very similar to that of SARS-CoV. Therefore, herein, we analysed the 3CLpro sequence, constructed its 3D homology model, and screened it against a medicinal plant library containing 32,297 potential anti-viral phytochemicals/traditional Chinese medicinal compounds. Our analyses revealed that the top nine hits might serve as potential anti- SARS-CoV-2 lead molecules for further optimisation and drug development process to combat COVID-19." }, ... ] }</code></pre> <p> </p> <p><strong>All other files</strong></p> <p>The main JSON object consists of a list of entailments, a list of contradiction and a list of neutrals.</p> <p>Each object of the above mentioned lists has the same schema. An instance_id created by concatenating the template_id (when exists) with the paper_id, a template_id, a paper_id, premise (representing the paper's title), hypthesis (representing the paper's abstract), their concatenation in sequence and the target class.</p> <p>See an example instance below.</p> <p> </p> <pre><code class="language-json">{ "entailments": [ { "instance_id": "R138668xR138698", "template_id": "R138668", "paper_id": "R138698", "premise": "psychiatric disorders ai overview study cohort outcome assessment aims performance findings used models data", "hypothesis": "application of autoencoder in depression diagnosis major depressive disorder (mdd) is a mental disorder characterized by at least two weeks of low mood which is present across most situations diagnosis of mdd using rest state functional magnetic resonance imaging (fmri) data faces many challenges due to the high dimensionality, small samples, noisy and individual variability no method can automatically extract discriminative features from the origin time series in fmri images for mdd diagnosis in this study, we proposed a new method for feature extraction and a workflow which can make an automatic feature extraction and classification without a prior knowledge an autoencoder was used to learn pre training parameters of a dimensionality reduction process using 3 d convolution network through comparison with the other three feature extraction methods, our method achieved the best classification performance this method can be used not only in mdd diagnosis, but also other similar disorders", "sequence": "[CLS] psychiatric disorders ai overview study cohort outcome assessment aims performance findings used models data [SEP] application of autoencoder in depression diagnosis major depressive disorder (mdd) is a mental disorder characterized by at least two weeks of low mood which is present across most situations diagnosis of mdd using rest state functional magnetic resonance imaging (fmri) data faces many challenges due to the high dimensionality, small samples, noisy and individual variability no method can automatically extract discriminative features from the origin time series in fmri images for mdd diagnosis in this study, we proposed a new method for feature extraction and a workflow which can make an automatic feature extraction and classification without a prior knowledge an autoencoder was used to learn pre training parameters of a dimensionality reduction process using 3 d convolution network through comparison with the other three feature extraction methods, our method achieved the best classification performance this method can be used not only in mdd diagnosis, but also other similar disorders [SEP]", "target": "entailment" }, ... ], "contradictions": [ ... ], "neutrals": [ ... ] } </code></pre> <p> </p> <p><strong>Statistics</strong></p> <table align="center"> <tbody> <tr> <td>-</td> <td><strong>Training (supervised)</strong></td> <td><strong>Validation (supervised)</strong></td> <td><strong>Training (unsupervised)</strong></td> <td><strong>Test</strong></td> </tr> <tr> <td>Entailment</td> <td>180</td> <td>20</td> <td>200</td> <td>52</td> </tr> <tr> <td>Neutral</td> <td>180</td> <td>20</td> <td>200</td> <td>64</td> </tr> <tr> <td>Contradictrion</td> <td>736</td> <td>84</td> <td>0</td> <td>0</td> </tr> <tr> <td>Total</td> <td>1096</td> <td>124</td> <td>400</td> <td>116</td> </tr> </tbody> </table> <p> </p>
Interesting Scientific Idea Generation Using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders
<p>Dataset for the Knowledge graph used in the paper "<a href="https://arxiv.org/abs/2405.17044">Interesting Scientific Idea Generation Using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders</a>" by Xueme Gu and Mario Krenn. </p> <p>Nodes represent scientific concepts extracted from 2.44 million paper titles and abstracts and edges are formed when two concepts co-occur in titles or abstracts of over 58 million papers from OpenAlex, augmented with citation information. </p>
Submission to Special Issue of Quantitative Science Studies "Scientific Knowledge Graphs and Research Impact Assessment"
<p>To be added</p>
Research on key generic technology prediction based on graph neural networks under the perspective of patent citation - An example from the field of genetic engineering
<p>In this research, we adopted graph neural network models for key generic prediction based on cited patent data. Through the construction of the patent citation network and the design of a key generic evaluation system, 20879 relevant patents and 51,610 irrelevant patents were screened out. Further, we utilized the LDA topic model to interpret technical topics at a finer granularity. Finally, to test the effectiveness of this method, we took the field of genetic engineering as an example for key generic technology prediction, with an accuracy rate of 95%.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.