Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
219
datasets available to search
ShareScore release 0.7.1
Dataset results
219 results for “Knowledge Graph”
TaxGraph - A Knowledge Graph of Multi-National Companies and their Relationships
<p>The taxation of multi-national companies is a complex field, since it is influenced by the legislation of several states. Laws in different states may have unforeseen interaction effects, which can be exploited by allowing multi-national companies to minimize and avoid taxes.</p> <p>Thus, we created a knowledge graph of multi-national companies and their relationships. Many commonly known tax avoidance strategies can be formulated as subgraph queries to this graph, which allows for identifying companies using certain strategies. Moreover, we can identify anomalies in the graph which hint at potential tax avoidance strategies.</p>
SoSEN-KG: Knowledge Graph Dump v0 for SoSEn: Software Search Engine.
<p>SoSEn is a semantic search engine for scientific software. We index scientific software in a knowledge graph, which is used for search and understanding of the software. The graph `graph.ttl` contains information about the software, and `keywords.ttl` contains keyword information about the software. `graph.ttl` can be used alone, or it can be combined with `keywords.ttl` to facilitate tf-idf based keyword search.</p>
Predicting Phenotype from Multi-Scale Genomic and Environment Data using Neural Networks and Knowledge Graphs
<p><strong>Background: To mitigate the effects of climate change on public health and conservation, we need to better understand the dynamic interplay between biological processes and environmental effects. Machine learning (ML) methods in general, and Deep Learning (DL) methods in particular, are a potential way forward because they are able to cope with the nonlinearity of natural systems. However, there are several barriers that exist, including the absence of ML-ready data. We propose to develop a machine learning framework capable of predicting phenotypes based on multi-scale data about genes and environments. A critical part of this framework are data transformation methods that map the heterogeneous input data into formats that are consumable by the ML techniques. The central hypothesis of this research is that deep learning algorithms and biological knowledge graphs will predict phenotypes more accurately across more taxa and more ecosystems than do current numerical and traditional statistical modeling methods. Our long term goal is to develop predictive analytics for organismal response to environmental perturbations using innovative data science approaches. This pilot project on predicting emergent properties of complex systems and multidimensional interactions is funded by the NSF (Award # 1939945, 1940059, 1940062, 1940330). </strong></p> <p> </p> <p><strong>Results: We have established shared project governance, communication channels, project timeline, and data and computing environment across four universities. We have successfully reached out to three other projects for broader collaboration.</strong></p>
Beyond 2022 Knowledge Graph Sample Data
<p>This dataset contains a CSV file and an RDF Turtle file. Both files contain information on a few people mentioned in the Irish Exchequer Payments 1270-1326, a book written by Connolly, P and published by the Irish Manuscripts Commission in 1998. A historian transcribed those people in a CSV file, subsequently transformed into RDF using an R2RML mapping. This dataset contains the records and the output of a handful of people transcribed in this way. This dataset illustrates how the Beyond 2022 project avails of CIDOC-CRM to populate its knowledge graph.</p> <p>Beyond 2022 is funded by the Government of Ireland, through the Department of Culture, Heritage and the Gaeltacht, under the Project Ireland 2040 framework. The project is also partially supported by the ADAPT Centre for Digital Content Technology under the SFI Research Centres Programme (Grant 13/RC/2106).</p>
CSKG: The CommonSense Knowledge Graph
<pre>Associated data for the ESWC'21 submission "CSKG: The CommonSense Knowledge Graph". Contents: 1. `cskg.tsv.gz` - the CSKG graph 2. `cskg_star.tsv.gz` - version of the CSKG graph before merging identical nodes</pre> <p>The embedding files can be found on google drive.</p>
Mapping Manuscript Migrations Knowledge Graph
<p>The Mapping Manuscript Migrations (MMM) project transformed three separate datasets into a unified knowledge graph. The source databases include</p> <p>- <a href="https://sdbm.library.upenn.edu">Schoenberg Database of Manuscripts</a> from the Schoenberg Institute for Manuscript Studies,<br> - <a href="http://bibale.irht.cnrs.fr">Bibale database</a> from the Institute for Research and History of Texts, and<br> - <a href="https://medieval.bodleian.ox.ac.uk">Medieval Manuscripts in Oxford Libraries</a>.</p> <p>The Knowledge Graph has been created using the <a href="https://github.com/mapping-manuscript-migrations/mmm-data-conversion">MMM data transformation pipeline</a>. This dataset is available on a public SPARQL endpoint (<em>http://ldf.fi/mmm/sparql</em>) and it can be directly deployed onto a SPARQL endpoint using <a href="https://github.com/mapping-manuscript-migrations/mmm-fuseki">this docker recipe</a>.</p> <p>To test and demonstrate its usefulness, the MMM Knowledge Graph is in use in the <a href="https://mappingmanuscriptmigrations.org/">MMM Semantic Portal</a>.</p> <p><strong>Version History</strong></p> <ul> <li>1.0.0, January 2020.</li> <li>1.1.0, February 2020, Updated Bibale and SDBM source datasets, data processing improvements.</li> <li>2.0.0, May 2020, Updated SDBM source dataset, data processing fixes affecting some generated URIs.</li> <li>2.1.0, September 2020, Updated Bibale and SDBM source datasets, minor data processing bug fixes.</li> <li>2.2.0, January 2021, Updated Oxford source dataset, minor data processing improvements.</li> </ul> <p> </p> <p> </p>
Supplementary data – OSH knowledge graph (OSH-KG)
<div> <div> <div> <p>Additional materials for the paper "Obtaining Occupational Safety and Health Insights with a Knowledge Base" that introduces the Occupational Safety and Health Knowledge Graph (OSH-KG).</p> <h2>Content</h2> <ul> <li><code>ontology.tar.gz</code> – the OSH-KG ontology. The archive contains both the individual ontology files and a merged version.</li> <li><code>queries.tar.gz</code> – SPARQL queries used in the paper, along with their results.</li> <li><code>rules.tar.gz</code> – SHACL-AF reasoning rules used in the paper.</li> <li><code>assist.tar.gz</code> – vocabulary, reasoning rules, and queries for integrating OSH-KG with data from the <a href="https://assist-iot.eu/">ASSIST-IoT project</a>.</li> <li><code>*.nq.gz</code> / <code>*.nt.gz</code> – individual named graphs of OSH-KG in the N-Quads format.</li> <li><code>*.jelly.gz</code> – individual named graphs of OSH-KG in the <a href="https://w3id.org/jelly">Jelly format</a>.</li> <li><code>code.tar.gz</code> – code in Python (Jupyter notebooks) used to generate the vocabularies and RDF translations of Eurostat data.</li> </ul> <h2>License</h2> <p>Please see LICENSE.md for details. Different resources in this repository have different licenses, but are generally freely licensed.</p> </div> </div> </div>
Artifact for "Efficient Construction of Practical Python Call Graphs with Entity Knowledge Base"
<p>This is the artifact for the paper entitled "Efficient Construction of Practical Python Call Graphs with Entity Knowledge Base"</p>
nuScenes Knowledge Graph
<p>Ontologies and Knowledge Graphs of the ICCV 2023 workshop paper "nuScenes Knowledge Graph - A comprehensive semantic representation of traffic scenes for trajectory prediction".</p><h2>Content</h2><ul><li><strong>nSKG</strong> (nuScenes Knowledge Graph): knowledge graph for the <a href="https://www.nuscenes.org/nuscenes">nuScenes dataset</a>, that models all scene participants and road elements, as well as their semantic and spatial relationships</li><li><strong>nSTP</strong> (nuScenes Trajectory Prediction Graph): heterogeneous graph of the nuScenes dataset for trajectory prediction in <a href="https://pytorch-geometric.readthedocs.io/en/latest/">PyTorch Geometric (PyG)</a> format. It extends nSKG for example by transformation into agents' local coordinate systems, relevant agent extraction, semantic relationships between agents.</li><li><strong>nuScenes_agent_onto</strong>: ontology for the traffic participants (agents)</li><li><strong>nuScenes_map_onto</strong>: ontology for the extended map</li><li><strong>stardog_rules</strong>: <a href="https://www.w3.org/TR/rdf-sparql-query/">SPARQL</a> rules for mapping nuScenes concepts to the agent ontology</li></ul><h2>How to use</h2><p><strong>nSKG</strong> represents the KG created on the basis of the nuScenes ontologies nuScenes_agent_onto.ttl and nuScenes_map_onto.ttl and materializing the <a href="https://www.nuscenes.org/nuscenes#data-annotation">nuScenes annotation dataset</a>. It can be used for applications where relational information between entities are important. The ontologies are in <a href="https://www.w3.org/TR/turtle/">Turtle</a> format and can be viewed by ontology editors such as <a href="https://protege.standord.edu/">Protege</a>.</p><p><strong>nSTP</strong> represents the extended version of nSKG, where agents are represented in local coordinate systems to enforce shift- and rotation-invariance. It also includes semantic relationships between agents, e.g. whether agents are on neighboring lanes, the same lane or might intersect. This is done based on the semantic scene graph describe in <a href="https://arxiv.org/abs/2111.10196">(Towards Traffic Scene Description: The Semantic Scene Graph)</a>. The data is provided in <a href="https://pyg.org/">PyTorch Geometric</a> format and directly be used to train graph neural network for trajectory prediction.</p><h2>Acknowledgements</h2><p>Special thanks to Motional for the permission to distribute this modified version of the nuScenes dataset.</p><h2>Citation</h2><p>If you use this work please cite</p><blockquote><p>@inproceedings{ <br>title={nuScenes Knowledge Graph - A comprehensive semantic representation of traffic scenes for trajectory prediction}, <br>author={Leon Mlodzian and Zhigang Sun and Hendrik Berkemeyer and Sebastian Monka and Zixu Wang and Stefan Dietze and Lavdim Halilaj and Juergen Luettin},<br>booktitle={International Conference on Computer Vision (ICCV), Workhsop on Scene Graphs and Graph Representation Learning (SG2RL)}, <br>year={2023} <br>}</p></blockquote>
RDF Knowledge Graph SemOpenAlex-SemanticWeb
<p>SemOpenAlex-SemanticWeb is a subset of the RDF knowledge graph <a href="(https://semopenalex.org/">SemOpenAlex</a> (version from 2023-04-24) modeling the Semantic Web Community.</p><p>SemOpenAlex-SemanticWeb is a suitable database for creating heterogeneous graph machine learning datasets (e.g., used for GNN-based recommendation and similar tasks).</p><p>The RDF knowledge graph consists of 21,978,026 semantic RDF triples including the following entities: works (95,575 entities), semantic web authors (19,970 entities), concepts (38,050 entities), sources (10,739 entities), institutions (5,846 entities) and publishers (786 entities).</p><p>Besides the RDF knowledge graph, we provide TransE embeddings for entities and relations (see embeddings.zip).</p><p>More information can be found at <a href="https://github.com/davidlamprecht/semopenalex-semanticweb">https://github.com/davidlamprecht/semopenalex-semanticweb</a>.</p>
WeKG-MF: Weather Knowledge Graph of Météo France Meteorological Observations
<p>WeKG-MF, the <strong>Weather RDF Knowledge Graph of Météo France </strong>represents the meteorological observations made by 62 Météo-France weather stations located in different regions in metropolitan France and overseas departments, from 2012 to 2022. This work is supported by the French National Research Agency under grant ANR-18-CE23-0017 (project <a href="https://d2kab.mystrikingly.com/">D2KAB</a>).</p> <p>The raw data was obtained from <a href="https://donneespubliques.meteofrance.fr/?fond=produit&id_produit=90&id_rubrique=32.">Météo France's website</a>.</p> <p>Code and description: https://github.com/Wimmics/weather-kg/</p> <p> </p>
Linked Papers With Code: Knowledge Graph Embeddings
<p>We provide <strong>knowledge graph embeddings</strong> for the <strong>Linked Papers With Code</strong> Knowledge Graph. More information at <a href="https://linkedpaperswithcode.com/">https://linkedpaperswithcode.com/</a></p>
Temporal Event Knowledge Graphs transformed from Object-Centric Event Logs
<p>In the paper"Transforming Object-Centric Event Logs to Temporal Event Knowledge Graphs", we introduced and formalized temporal Event Knowledge Graphs (tEKGs) and presented an algorithm to transform Object-Centric Event Logs (OCEL) 2.0 into tEKGs. Data sets are the results of transforming OCEL 2.0 log files. The source of the generated dump files are as follows:</p> <ol> <li>ContainerLogistics.neo4j.dum : <a href="../records/8428084">Link to the source data</a></li> <li>OrderManagement.neo4j.dump: <a href="../records/8428112">Link to the source data</a></li> <li>Procure-To-Payment.neo4j.dump: <a href="../records/8412920">Link to the source data</a></li> </ol> <p>The version of the dump files is <strong>5.12.0</strong>. Additionally, for restoring dump files inside Neo4j, you need to enter the username and password, which is indicated below:</p> <p><strong>Username</strong>: neo4j</p> <p><strong>Password</strong>: 12345678</p> <p> </p>
Results of Multi-Perspective Concept Drift Detection on BPIC17 data using Event Knowledge Graphs
<p>Extensive results of multi-perspective concept drift detection performed on the BPI Challenge 2017 dataset retrieved using the related implementation from Github: <a href="https://github.com/multi-dimensional-process-mining/ekg-bpic17-concept-drift-detection-multi-perspective">https://github.com/multi-dimensional-process-mining/ekg-bpic17-concept-drift-detection-multi-perspective</a></p> <p> </p>
SemTab 2024 - Table Metadata to Knowledge Graph Track Datasets
<p>Data Sets from the ISWC 2024 Semantic Web Challenge on Tabular Data to Knowledge Graph Matching, Table Metadata to Knowledge Graph Track.</p> <p>See round1/README.md and round2/README.md for details on how to use the dataset.</p> <p>Links to other datasets can be found on the challenge website: https://sem-tab-challenge.github.io/2024/ as well as the proceedings of the challenge published on CEUR.</p> <p>For details about the challenge, see: http://www.cs.ox.ac.uk/isg/challenges/sem-tab/</p> <p>For 2024 edition, see: https://sem-tab-challenge.github.io/2024/</p> <p>Note on License: This data includes data from the following sources. Refer to each source for license details:<br>- Wikidata https://www.wikidata.org/<br>- DBpedia https://www.dbpedia.org/<br>- Kaggle https://www.kaggle.com</p> <p>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>
A Framework for Automated Construction of Heterogeneous Large-Scale Biomedical Knowledge Graphs (Recorded Talk)
<p>This entry contains the recording of the presentation that was presented at the 2020 Intelligent Systems for Molecular Biology as part of the Bio-Ontologies COSI (https://www.iscb.org/ismb2020).</p>
Combining biomedical knowledge graphs and text to improve predictions for drug-target interactions and drug-indications.
<p>The datasets used in the publications titled "Combining biomedical knowledge graphs and text to improve predictions for drug-target interactions and drug-indications"</p>
Dataset - Templates Recommendation in the Open Research Knowledge Graph
<p>This dataset has been created for implementing a content-based recommender system in the context of the Open Research Knowledge Graph (ORKG). The recommender system accepts research paper's title and abstracts as input and recommends existing templates in the ORKG semantically relevant to the given paper.</p> <p> </p> <p>Two approaches have been trained on this dataset in the context of <a href="https://doi.org/10.15488/11834">this master's thesis</a>, namely a Natural Language Inference (NLI) approach based on SciBERT embeddings and an unsupervised approach based on ElasticSearch.</p> <p> </p> <p>This publication consists therefore of one general dataset, two training sets for each approach, validation set for the supervised approach and a test set for both approaches.</p> <p> </p> <p><strong>dataset.json</strong></p> <p>The main JSON object consists of a list of templates and a list of neutral papers.</p> <p>Each template object has an ID, label, list of research fields, list of properties and list of papers using that template, whereas each paper object has ID, label, DOI, research field and abstract.</p> <p>Each neutral paper object has the same schema of a paper object using that template.</p> <p>See an example instance below.</p> <p> </p> <pre><code class="language-json">{ "templates": [ { "id": "R138668", "label": "Psychiatric Disorders AI Overview", "research_fields": [ { "id": "http://orkg.org/orkg/resource/R133", "label": "Artificial Intelligence" } ... ], "properties": [ "Study cohort", ... ], "papers": [ { "id": "R138698", "label": "Application of Autoencoder in Depression Diagnosis", "doi": "10.12783/dtcse/csma2017/17335", "research_field": { "id": "R104", "label": "Bioinformatics" }, "abstract": "Major depressive disorder (MDD) is a mental disorder characterized by at least two weeks of low mood which is present across most situations. Diagnosis of MDD using rest-state functional magnetic resonance imaging (fMRI) data faces many challenges due to the high dimensionality, small samples, noisy and individual variability. No method can automatically extract discriminative features from the origin time series in fMRI images for MDD diagnosis. In this study, we proposed a new method for feature extraction and a workflow which can make an automatic feature extraction and classification without a prior knowledge. An autoencoder was used to learn pre-training parameters of a dimensionality reduction process using 3-D convolution network. Through comparison with the other three feature extraction methods, our method achieved the best classification performance. This method can be used not only in MDD diagnosis, but also other similar disorders." }, ... }, ... ] "neutral_papers": [ { "id": "R109377", "label": "Structural basis of SARS-CoV-2 3CLpro and anti-COVID-19 drug discovery from medicinal plants", "doi": "10.1016/j.jpha.2020.03.009", "research_field": { "id": "R104", "label": "Bioinformatics" }, "abstract": "Abstract The recent outbreak of coronavirus disease 2019 (COVID-19) caused by SARS-CoV-2 in December 2019 raised global health concerns. The viral 3-chymotrypsin-like cysteine protease (3CLpro) enzyme controls coronavirus replication and is essential for its life cycle. 3CLpro is a proven drug discovery target in the case of severe acute respiratory syndrome coronavirus (SARS-CoV) and middle east respiratory syndrome coronavirus (MERS-CoV). Recent studies revealed that the genome sequence of SARS-CoV-2 is very similar to that of SARS-CoV. Therefore, herein, we analysed the 3CLpro sequence, constructed its 3D homology model, and screened it against a medicinal plant library containing 32,297 potential anti-viral phytochemicals/traditional Chinese medicinal compounds. Our analyses revealed that the top nine hits might serve as potential anti- SARS-CoV-2 lead molecules for further optimisation and drug development process to combat COVID-19." }, ... ] }</code></pre> <p> </p> <p><strong>All other files</strong></p> <p>The main JSON object consists of a list of entailments, a list of contradiction and a list of neutrals.</p> <p>Each object of the above mentioned lists has the same schema. An instance_id created by concatenating the template_id (when exists) with the paper_id, a template_id, a paper_id, premise (representing the paper's title), hypthesis (representing the paper's abstract), their concatenation in sequence and the target class.</p> <p>See an example instance below.</p> <p> </p> <pre><code class="language-json">{ "entailments": [ { "instance_id": "R138668xR138698", "template_id": "R138668", "paper_id": "R138698", "premise": "psychiatric disorders ai overview study cohort outcome assessment aims performance findings used models data", "hypothesis": "application of autoencoder in depression diagnosis major depressive disorder (mdd) is a mental disorder characterized by at least two weeks of low mood which is present across most situations diagnosis of mdd using rest state functional magnetic resonance imaging (fmri) data faces many challenges due to the high dimensionality, small samples, noisy and individual variability no method can automatically extract discriminative features from the origin time series in fmri images for mdd diagnosis in this study, we proposed a new method for feature extraction and a workflow which can make an automatic feature extraction and classification without a prior knowledge an autoencoder was used to learn pre training parameters of a dimensionality reduction process using 3 d convolution network through comparison with the other three feature extraction methods, our method achieved the best classification performance this method can be used not only in mdd diagnosis, but also other similar disorders", "sequence": "[CLS] psychiatric disorders ai overview study cohort outcome assessment aims performance findings used models data [SEP] application of autoencoder in depression diagnosis major depressive disorder (mdd) is a mental disorder characterized by at least two weeks of low mood which is present across most situations diagnosis of mdd using rest state functional magnetic resonance imaging (fmri) data faces many challenges due to the high dimensionality, small samples, noisy and individual variability no method can automatically extract discriminative features from the origin time series in fmri images for mdd diagnosis in this study, we proposed a new method for feature extraction and a workflow which can make an automatic feature extraction and classification without a prior knowledge an autoencoder was used to learn pre training parameters of a dimensionality reduction process using 3 d convolution network through comparison with the other three feature extraction methods, our method achieved the best classification performance this method can be used not only in mdd diagnosis, but also other similar disorders [SEP]", "target": "entailment" }, ... ], "contradictions": [ ... ], "neutrals": [ ... ] } </code></pre> <p> </p> <p><strong>Statistics</strong></p> <table align="center"> <tbody> <tr> <td>-</td> <td><strong>Training (supervised)</strong></td> <td><strong>Validation (supervised)</strong></td> <td><strong>Training (unsupervised)</strong></td> <td><strong>Test</strong></td> </tr> <tr> <td>Entailment</td> <td>180</td> <td>20</td> <td>200</td> <td>52</td> </tr> <tr> <td>Neutral</td> <td>180</td> <td>20</td> <td>200</td> <td>64</td> </tr> <tr> <td>Contradictrion</td> <td>736</td> <td>84</td> <td>0</td> <td>0</td> </tr> <tr> <td>Total</td> <td>1096</td> <td>124</td> <td>400</td> <td>116</td> </tr> </tbody> </table> <p> </p>
Spatial domains identification in spatial transcriptomics by domain knowledge-aware and subspace-enhanced graph contrastive learning
<p>We propose a graph contrastive learning framework, GRAS4T, which combines contrastive learning and subspace module to accurately distinguish different spatial domains by capturing tissue microenvironment through self-expressiveness of spots within the same domain. To uncover the pertinent features for spatial domain identification, GRAS4T employs a graph augmentation based on histological images prior, preserving information crucial for the clustering task. Experimental results on 8 ST datasets from 5 different platforms show that GRAS4T outperforms five state-of-the-art competing methods in spatial domain identification. Significantly, GRAS4T excels at separating distinct tissue structures and unveiling more detailed spatial domains. GRAS4T combines the advantages of subspace analysis and graph representation learning with extensibility, making it an ideal framework for ST domain identification.</p>
Benchmark datasets for biomedical knowledge graphs with negative statements
<p>We present a collection of datasets for three relation prediction tasks - protein-protein interaction prediction, gene-disease association prediction and disease prediction - that aim at circumventing the difficulties in building benchmarks for knowledge graphs with negative statements. These datasets include data from two successful biomedical ontologies, Gene Ontology and Human Phenotype Ontology, enriched with negative statements. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.