Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
209
datasets available to search
ShareScore release 0.9.0
Dataset results
209 results for βKGβ
PheKnowLator Human Disease KG Benchmarks: Class-Standard Relations-OWL (v2.1.0 - September 2021)
<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds (v2.1.0)</strong></p><p><strong>Build Type: </strong><i>Class-Standard Relations-OWL</i></p><p><strong>Build Date: </strong>September 01, 2021</p><p> </p><h3><strong>Important Build Information</strong></h3><p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed an archive for the builds on Zenodo. While the original GCP resources contained all associated files, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the associated logs.</p><p>Details on each of the files generated by the build process can be found in the file associated with this directory (<a href="https://zenodo.org/records/10065431/files/PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx?download=1">PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</a>).</p><p> </p><p>π¨ <strong>AVAILABLE FILES </strong>π¨ </p><ul><li>Available KG benchmark files are zipped and listed below.</li><li>For additional details on what each file contains, please see the associated Wiki page π <a href="https://github.com/callahantiff/PheKnowLator/wiki/September-01%2C-2021">here</a>.</li></ul>
PheKnowLator Human Disease KG Benchmarks: Instance-Standard Relations-OWLNETS (v2.1.0 - August 2021)
<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds (v2.1.0)</strong></p><p><strong>Build Type: </strong><i>Instance-Standard Relations-OWLNETS</i></p><p><strong>Build Date: </strong>August 01, 2021</p><p> </p><h3><strong>Important Build Information</strong></h3><p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed an archive for the builds on Zenodo. While the original GCP resources contained all associated files, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the associated logs.</p><p>Details on each of the files generated by the build process can be found in the file associated with this directory (<a href="https://zenodo.org/records/10065431/files/PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx?download=1">PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</a>).</p><p> </p><p>π¨ <strong>AVAILABLE FILES </strong>π¨ </p><ul><li>Available KG benchmark files are zipped and listed below.</li><li>For additional details on what each file contains, please see the associated Wiki page π <a href="https://github.com/callahantiff/PheKnowLator/wiki/August-01%2C-2021">here</a>.</li></ul>
PheKnowLator Human Disease KG Benchmarks: Instance-Inverse Relations-OWL (v2.1.0 - August 2021)
<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds (v2.1.0)</strong></p><p><strong>Build Type: </strong><i>Instance-Inverse Relations-OWL</i></p><p><strong>Build Date: </strong>August 01, 2021</p><p> </p><h3><strong>Important Build Information</strong></h3><p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed an archive for the builds on Zenodo. While the original GCP resources contained all associated files, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the associated logs.</p><p>Details on each of the files generated by the build process can be found in the file associated with this directory (<a href="https://zenodo.org/records/10065431/files/PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx?download=1">PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</a>).</p><p> </p><p>π¨ <strong>AVAILABLE FILES </strong>π¨ </p><ul><li>Available KG benchmark files are zipped and listed below.</li><li>For additional details on what each file contains, please see the associated Wiki page π <a href="https://github.com/callahantiff/PheKnowLator/wiki/August-01%2C-2021">here</a>.</li></ul>
PheKnowLator Human Disease KG Benchmarks: Instance-Inverse Relations-OWLNETS (v2.1.0 - August 2021)
<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds (v2.1.0)</strong></p><p><strong>Build Type: </strong><i>Instance-Inverse Relations-OWLNETS</i></p><p><strong>Build Date: </strong>August 01, 2021</p><p> </p><h3><strong>Important Build Information</strong></h3><p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed an archive for the builds on Zenodo. While the original GCP resources contained all associated files, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the associated logs.</p><p>Details on each of the files generated by the build process can be found in the file associated with this directory (<a href="https://zenodo.org/records/10065431/files/PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx?download=1">PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</a>).</p><p> </p><p>π¨ <strong>AVAILABLE FILES </strong>π¨ </p><ul><li>Available KG benchmark files are zipped and listed below.</li><li>For additional details on what each file contains, please see the associated Wiki page π <a href="https://github.com/callahantiff/PheKnowLator/wiki/August-01%2C-2021">here</a>.</li></ul>
PheKnowLator Human Disease KG Benchmarks: Class-Standard Relations-OWLNETS (v2.1.0 - September 2021)
<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds (v2.1.0)</strong></p><p><strong>Build Type: </strong><i>Class-Standard Relations-OWLNETS</i></p><p><strong>Build Date: </strong>September 01, 2021</p><p> </p><h3><strong>Important Build Information</strong></h3><p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed an archive for the builds on Zenodo. While the original GCP resources contained all associated files, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the associated logs.</p><p>Details on each of the files generated by the build process can be found in the file associated with this directory (<a href="https://zenodo.org/records/10065431/files/PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx?download=1">PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</a>).</p><p> </p><p>π¨ <strong>AVAILABLE FILES </strong>π¨ </p><ul><li>Available KG benchmark files are zipped and listed below.</li><li>For additional details on what each file contains, please see the associated Wiki page π <a href="https://github.com/callahantiff/PheKnowLator/wiki/September-01%2C-2021">here</a>.</li></ul>
ISPRA Linked Open Dataset part of the WHOW-KG
<p>This linked open dataset contains Italian data for:</p> <ul> <li>soil consumption (soilc);</li> <li>spread in water of Ostreopsis ovata (ostreopsis);</li> <li>environmental protection interventions (dissesto);</li> <li>Italian national tide gauge network (rmn);</li> <li>Italian National Wave Network (ron);</li> <li>bathing waters (bathw).</li> </ul>
DDB-KG v0.1
<p><strong>DDB-KG </strong>is a knowledge graph (KG) containing a selection of resources from the German Digital Library or <a href="https://www.deutsche-digitale-bibliothek.de"><em>Deutsche Digitale Bibliothek</em></a> (DDB). The resources stored in this KG were provided by libraries and represented using FaBiO. For more information, please refer to the <a href="https://ise-fizkarlsruhe.github.io/ddbkg/docs/">documentation</a>. The list of publications pertaining to this dataset is available at <a href="https://ise-fizkarlsruhe.github.io/ddbkg/publication/">https://ise-fizkarlsruhe.github.io/ddbkg/publication/</a>.</p> <p>This dataset is also published as a SPARQL endpoint at <a href="https://ddbkg.fiz-karlsruhe.de">https://ddbkg.fiz-karlsruhe.de</a>.</p>
MLSea-KG
<p>MLSea-KG is a declaratively constructed and regularly updated KG with more than 1.44 billion RDF triples of ML experiments, regarding datasets used in ML experiments, tasks, implementations and related hyper-parameters, experiment executions, their configuration settings and evaluation results, code notebooks and repositories, algorithms, publications, models, scientists and practitioners. It integrates data from OpenML, Kaggle and Papers with Code.<br><br>Latest update: May 23rd 2024</p>
FLIP-KG: Enriching Automated Knowledge Graph Extraction with Tacit Knowledge. The Case of Lyrical Implicatures within Poems
<p>FLIP-KG: Enriching Automated Knowledge Graph Extraction with Tacit Knowledge. The Case of Lyrical Implicatures within Poems. Dataset for the task force "Vulcan" from ISWS 2024 led by Aldo Gangemi and Andrea Poltronieri</p>
Supplementary Table S1. Combined analysis of variance containing the degrees of freedom (DF), mean squares (MS), P value (P val.), mean, coefficient of experimental variation (CEV%) and selective accuracy (SA) for the traits of luminosity (L*), chromaticity a* (a*), chromaticity b* (b*), grain length (length, mm), grain width (width, mm), grain thickness (thickness, mm), mass of 100 grains (Mass, g), normal grains (Ng, %), water absorption (absorption, %), cooking time (Ct, min:s), and concentrations of potassium (K, g kg-1 dry matter - DM), phosphorus (P, g kg-1 DM), calcium (Ca, g kg-1 DM), magnesium (Mg, g kg-1 DM), iron (Fe, mg kg-1 DM), zinc (Zn, mg kg-1 DM), and copper (Cu, mg kg-1 DM) obtained in 25 common bean cultivars evaluated in four experiments carried out from 2019 to 2021
<p><strong><span>Table S1.</span></strong><span> Combined analysis of variance.</span></p> <p><strong><span>Indirect selection for multiple technological and nutritional traits in common bean cultivars under different degrees of multicollinearity</span></strong></p> <p><strong><span>Bragantia, 2024.</span></strong></p>
Inductive Freebase and Wikidata for KG Completion
<p>UPD 2.0: Regenerated datasets free of potential test set leakages</p> <p>This repository contains 10 inductive link prediction datasets (graphs only) published in "Inductive Logical Query Answering in Knowledge Graphs" (NeurIPS 2022). 9 datasets (106-550) were created from FB15k-237, the wikikg dataset was created from OGB WikiKG 2 graph. In the datasets, all inference graphs extend training graphs and include new nodes and edges. Dataset numbers indicate a relative size of the inference graph compared to the training graph, e.g., in 175, the number of nodes in the inference graph is 175% compared to the number of nodes in the training graph. The higher the ratio, the more new unseen nodes appear at inference time, the more complex the task is. The Wikikg split has a fixed 133% ratio.</p> <p>Each dataset is a zip archive containing 5 files:</p> <ul> <li>train_graph.txt (pt for wikikg) - original training graph</li> <li>val_inference.txt (pt) - inference graph (validation split), new nodes in validation are disjoint with the test inference graph</li> <li>val_predict.txt (pt) - missing edges in the validation inference graph to be predicted. </li> <li>test_intference.txt (pt) - inference graph (test splits), new nodes in test are disjoint with the validation inference graph</li> <li>test_predict.txt (pt) - missing edges in the test inference graph to be predicted;</li> </ul> <p>This is a light-weight version of the full datasets for inductive query answering published here: https://zenodo.org/record/7231344 </p> <p>Here, we only provide graph data for training inductive link prediction models.</p> <p>Paper pre-print: https://arxiv.org/abs/2210.08008</p> <p>The full source code of training/inference models is available at https://github.com/DeepGraphLearning/InductiveQE</p>
Experimental results and the user questionnaires for KG profiling tool (ABSTAT vs ProtΓ©gΓ©) evaluation
<p>Experimental results and the user questionnaires for KG profiling tool (ABSTAT vs Protégé) evaluation</p>
Semantic Web resources and Machine Learning systems - Knowledge Graph (SWeMLS-KG)
<p>This resource is part of our submission to ESWC 2023 resource track, which includes:</p> <p>Datasets:<br> - Folder "pattern" - a set of SWeMLS patterns represented based on OPMW and P-Plan ontology,<br> - Folder "shapes" - a set of SHACL constraints to check the conformance of SWeML Systems against SWeMLS patterns as well as a set of SHACL-AF rules to generate links between system components,<br> - File "swemls-ontology.ttl" - an ontology to represent Semantic Web resources and Machine Learning systems (SWeMLS),<br> - File "swemls-instances.ttl" - a set of triples representing the extracted metadata from 476 SWeML systems and papers,<br> - File "swemls-kg.ttl" - an integrated and validated KG containing all above files, including enrichment from SHACL-AF rules using "swemls-toolkit" [2].</p> <p>These resources are produced based on the result of the Systematic Mapping Study (SMS) reported in [1]. The latest SNAPSHOT-version of the resource can be accessed through our resource landing page: <a href="https://w3id.org/semsys/sites/swemls-kg/">https://w3id.org/semsys/sites/swemls-kg/</a></p> <p>[1] Breit, A., Waltersdorfer, L., Ekaputra, J.F., Sabou, M., Ekelhart, A., Iana, A., Paulheim, H., Portisch, J., Revenko, A., Ten Teije, A., van Harmelen, F.: Combining Machine Learning and Semantic Web -A Systematic Mapping Study (under review). ACM CSUR (2022)<br> [2] Source code of swemls-toolkit is available at: https://github.com/semanticsystems/swemls-toolkit</p>
SURE-KG dataset
<p>The SURE-KG RDF dataset provides a knowledge graph built from a real dataset to represent Real Estate and Uncertain Spatial Data from Advertisements. It relies on natural language processing and machine learning methods for information extraction, and semantic Web frameworks for representation and integration. It describes more than 100K real estate ads and 6K place-names extracted from French Real Estate advertisements from various online advertiser and located in the French Riviera. It can be exploited by real estate search engines, real estate professionals, or geographers willing to analyze local place-names</p> <p>Homepage: <a href="https://github.com/Wimmics/sure">https://github.com/Wimmics/sure</a> </p>
PheKnowLator Human Disease KG Benchmarks: Instance-Inverse Relations-OWL (v2.1.0 - June 2021)
<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds (<code>v2.1.0</code>)</strong></p> <p><strong>Build Type: </strong><em>Instance-Inverse Relations-OWL</em></p> <p><strong>Build Date: </strong><code>June 01, 2021</code></p> <blockquote> <p>Please note that all resources linked below redirect to a publicly Google Cloud Storage bucket where all data are publicly accessible. Routing users from this wiki page is perfectly safe and allows us to avoid requiring users to have a Google account and login to download data. If you have any questions or concerns, please email the project maintainer at <a href="https://github.com/callahantiff/PheKnowLator/wiki/callahantiff@gmail.com">callahantiff@gmail.com</a>.</p> </blockquote> <p>If you have a Google account you can access the data directly via π <a href="https://console.cloud.google.com/storage/browser/pheknowlator">here</a></p> <p> </p> <p>π For additional information on the builds please see the following <a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">README</a><br> π For additional information on the KG file types please see the following <a href="https://github.com/callahantiff/PheKnowLator/wiki/KG-Construction#table-knowledge-graph-build-output">Wiki page</a></p> <p> </p> <p> </p> <p>π¨ <strong>AVAILABLE FILES </strong>π¨ </p> <p>Available KG benchmark files are zipped and listed below. For additional details on what each file contains, please see the associated Wiki page π <a href="https://github.com/callahantiff/PheKnowLator/wiki/June-01%2C-2021">here</a>.</p>
Results and log of LLM-KG-Bench runs described in article "Benchmarking the Abilities of Large Language Models for RDF Knowledge Graph Creation and Comprehension: How Well Do LLMs Speak Turtle?", Frey et al. 2023
<p>Results and log of LLM-KG-Bench runs described in article ""Benchmarking the Abilities of Large Language Models for RDF Knowledge Graph Creation and Comprehension: How Well Do LLMs Speak Turtle?", Frey et al. 2023, to appear in proceedings for workshop DL4KG@ISWC 2023.</p> <p>For data on task FactExtractStatic please contact authors.</p>
Efficacy, Safety and Pharmacokinetics of Artemether-lumefantrine Dispersible Tablet in the Treatment of Malaria in Infants < 5 kg
ClinicalTrials.gov study NCT01619878. IPD Sharing: Not stated. Countries: 5. Publications: 1.
SoSEN-KG: Knowledge Graph Dump v0 for SoSEn: Software Search Engine.
<p>SoSEn is a semantic search engine for scientific software. We index scientific software in a knowledge graph, which is used for search and understanding of the software. The graph `graph.ttl` contains information about the software, and `keywords.ttl` contains keyword information about the software. `graph.ttl` can be used alone, or it can be combined with `keywords.ttl` to facilitate tf-idf based keyword search.</p>
Supplementary data β OSH knowledge graph (OSH-KG)
<div> <div> <div> <p>Additional materials for the paper "Obtaining Occupational Safety and Health Insights with a Knowledge Base" that introduces the Occupational Safety and Health Knowledge Graph (OSH-KG).</p> <h2>Content</h2> <ul> <li><code>ontology.tar.gz</code> – the OSH-KG ontology. The archive contains both the individual ontology files and a merged version.</li> <li><code>queries.tar.gz</code> – SPARQL queries used in the paper, along with their results.</li> <li><code>rules.tar.gz</code> – SHACL-AF reasoning rules used in the paper.</li> <li><code>assist.tar.gz</code> – vocabulary, reasoning rules, and queries for integrating OSH-KG with data from the <a href="https://assist-iot.eu/">ASSIST-IoT project</a>.</li> <li><code>*.nq.gz</code> / <code>*.nt.gz</code> – individual named graphs of OSH-KG in the N-Quads format.</li> <li><code>*.jelly.gz</code> – individual named graphs of OSH-KG in the <a href="https://w3id.org/jelly">Jelly format</a>.</li> <li><code>code.tar.gz</code> – code in Python (Jupyter notebooks) used to generate the vocabularies and RDF translations of Eurostat data.</li> </ul> <h2>License</h2> <p>Please see LICENSE.md for details. Different resources in this repository have different licenses, but are generally freely licensed.</p> </div> </div> </div>
PheKnowLator Human Disease KG Benchmarks: Class-Inverse Relations-OWLNETS (v2.1.0 - May 2021)
<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds (v2.1.0)</strong></p><p><strong>Build Type: </strong><i>Class-Inverse Relations-OWLNETS</i></p><p><strong>Build Date: </strong>May 01, 2021</p><p> </p><h3><strong>Important Build Information</strong></h3><p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed an archive for the builds on Zenodo. While the original GCP resources contained all associated files, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the associated logs.</p><p>Details on each of the files generated by the build process can be found in the file associated with this directory (<a href="https://zenodo.org/records/10065431/files/PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx?download=1">PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</a>).</p><p> </p><p>π¨ <strong>AVAILABLE FILES </strong>π¨ </p><ul><li>Available KG benchmark files are zipped and listed below.</li><li>For additional details on what each file contains, please see the associated Wiki page π <a href="https://github.com/callahantiff/PheKnowLator/wiki/May-01%2C-2021">here</a>.</li></ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.