Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

190

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

190 results for “Knowledge Base”

Learn how ShareScore rates datasets ↗
zenodo40/100

The CrossCult Knowledge Base

<p>The Crosscult Knowledge Base (CCKB) is a comprehensive structure of semantic definitions and formalisms, developed for facilitating interoperable connections between the cultural heritage datasets&nbsp;contributing to Crosscult. It is written in OWL2 (the standard ontology language for the Semantic Web) and enables augmentation, semantic linking, semantic-based reasoning and retrieval across disparate data resources</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

Data associated with "A collaborative filtering based approach to biomedical knowledge discovery"

<p>This is the data set associated with the publication: &quot;A collaborative filtering based approach to biomedical knowledge discovery&quot; published in Bioinformatics.</p> <p>The data are sets of cooccurrences of biomedical terms extracted from published abstracts and full text articles. The cooccurrences are then represented in sparse matrix form. There are three different splits of this data denoted by the prefix number on the files.</p> <p>1. All - All cooccurrences combined in a single file</p> <p>2. Training/Validation - All cooccurrences in publications before 2010 in training, all novel cooccurrences in publication in 2010 go in validation</p> <p>3. Training+Validation/Test - All cooccurrences in publication upto and including 2010 in training+validation. All novel cooccurrences after 2010 in year by year increments and also all combined together</p> <p>&nbsp;</p> <p>Furthermore there are subset files which are used in some experiments to deal with the computational cost of evaluating the full set. The associated cuids.txt file containing a link between the row/column in the matrix with the UMLS Metathesaurus CUIDs. Hence the first row of cuids.txt matches up to the 0th row/column in the matrix. Note that the matrix is square and symmetric. This work was done with UMLS Metathesaurus 2016AB.</p>

opencc-by-4.0Apr 2018View details →
zenodo40/100

SUBSOL Knowledge Base data

<p>This dataset includes the content of the SUBSOL Knowledge Base and Marketplace, excluding all informations related to personal data. The same content is provided as:</p> <ul> <li>MySQL Database dump and</li> <li>Excel files exported from the Database</li> </ul>

opencc-by-4.0Sep 2018View details →
zenodo40/100

HiDy: A Large-scale Hierarchical Dynamic Financial Knowledge Base

<p>HiDy is a hierarchical, dynamic, robust, diverse, and large-scale financial KB that aims to provide various valuable financial knowledge as critical benchmarking data for fair model testing in different financial tasks. Specifically, HiDy currently contains 34 relation types, more than 506,444 relations, 17 entity types, and more than 51,095 entities. The scale of HiDy is steadily growing due to its continuous updates. To make HiDy easily accessible and retrieved, HiDy is organized in a well-formed financial hierarchy with four branches, <em>Macro</em>, <em>Meso</em>,<em> Micro</em>, and<em> Others</em>.</p> <p>We then give explanations on the various csv files as follows.</p> <ul> <li>"hidy.nodes.entity_type.csv" includes a mapping dictionary&nbsp;of&nbsp;entities with a specific&nbsp;entity type. The meta data is ID, name, (code), Label. For example, "hidy.nodes.company.csv" includes "0,东诚药业,002675.SZ,company", "1,大庆华科,000985.SZ,company".</li> <li>"hidy.relationships.relation_type.csv" includes quadruple knowledge with a specific relation type. The meta data is START_ID, END_ID, TYPE, time. For example, "hidy.relationships.cooperate.csv" includes "2197,245,cooperate,2019/9/17 12:00, "1165,756,cooperate,2020/8/2 17:10".</li> </ul> <p>For details of HiDy, please refer to our <a href="https://github.com/K-Quant/HiDy">GitHub</a>.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

SPARC Connectivity Knowledge base of the Autonomic Nervous System

<p>The SPARC Knowledge base of the Autonomic Nervous System (SCKAN) is an integrated graph database composed of three parts: the SPARC dataset metadata graph, ApiNATOMY and NPO models of connectivity, and the larger ontology used&nbsp; by SPARC which is a combination of the NIF-Ontology and community ontologies.</p> <p>The fastest way to get querying is to follow the instructions in the <a href="https://github.com/SciCrunch/sparc-curation/blob/master/docs/sckan/README.org#getting-started">SCKAN readme file</a>.</p> <p>For background information please see <a href="https://scicrunch.org/sawg/about/SCKAN">https://scicrunch.org/sawg/about/SCKAN</a> and <a href="https://sparc.science/resources/6eg3VpJbwQR4B84CjrvmyD">the SPARC portal resource page about SCKAN.</a></p> <p>This release contains the raw and compiled data for SCKAN. The release-*.zip contains raw data inputs along with the Blazegraph journal file, the sparc-sckan-graph-*.zip contains the SciGraph database, and sckan-data-*.tar.gz is a Docker image that contains the Blazegraph journal file and the SciGraph database along with the configuration files for running each of the servers. The image is intended&nbsp; to be used as a data volume with another Docker container that runs the SciGraph and Blazegraph server software.</p> <p>The Docker image containing this data is available live and is likely easier to use than the archived image included in this release. See the <a href="https://github.com/SciCrunch/sparc-curation/blob/master/docs/sckan/README.org#getting-started">SCKAN readme file</a> for the most up-to-date instructions.</p> <p>We would like to thank the members of the SAWG (SPARC Anatomy Working Group, RRID:SCR_018709) for their work on the various connectivity models included in this release.</p> <p>This work was funded by the NIH Common Fund under 3OT2OD030541-01S1.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Dataset of "Knowledge-based Sense Disambiguation of Multiword Expressions in Requirements Documents"

<p>This is the dataset used in the paper &quot;Knowledge-based Sense Disambiguation of Multiword Expressions in Requirements Documents&quot; at AIRE&#39;21</p> <p>&nbsp;</p> <p>In this paper, we explore the use of a multiword expression detection in combination with a knowledge-based word sense disambiguation to disambiguate expressions in requirements documents.</p> <p>The dataset comprises a gold standard for multiword expression detection and sense disambiguation for Wikipedia and WordNet 3.1.</p> <p>It covers 18 projects: CM1, EBT and GANTT as well as the 15 projects of the NFR dataset.</p> <p>&nbsp;</p> <p><strong>File format</strong></p> <p>We use a tab-separated version of the DiMSUM file format and extended it with sense information.</p> <p>The nine original DiMSUM tab-separated columns:</p> <p>1. token offset</p> <p>2. word</p> <p>3. lowercase lemma</p> <p>4. POS</p> <p>5. MWE tag</p> <p>6. offset of parent token (i.e. previous token in the same MWE), if applicable</p> <p>7. strength level encoded in the tag, if applicable. Currently not used</p> <p>8. supersense label, Currently not used</p> <p>9. sentence ID</p> <p>&nbsp;</p> <p>and the two further columns for sense information:</p> <p>10. Wikipedia article name</p> <p>11. WordNet 3.1 synset</p> <p>&nbsp;</p> <p>The last two columns might end with .1 or .0 indicating that the sense is a fully applicable or partial sense of a multiword expression.</p> <p><strong>Attribution (of datasets used)</strong></p> <p>The NFR Dataset can be attributed to Jane Cleland-Huang.<br> Jane Cleland-Huang, Sepideh Mazrouee, Huang Liguo, &amp; Dan Port. (2007). nfr [Data set]. Zenodo. Available:&nbsp;<a href="http://doi.org/10.5281/zenodo.268542">http://doi.org/10.5281/zenodo.268542</a><br> &nbsp;</p> <p>The CM1, EBT and GANTT datasets were retrieved from the Center of Excellence for Software &amp; Systems Traceability (CoEST)&nbsp;<a href="https://doi.org/10.5281/zenodo.3309669">http://coest.org/</a></p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Description of the Features of the Open Research Knowledge Graph as a Crowdsourcing Platform based on the 4 Pillars of Crowdsourcing

<p>This dataset provides the description of the features of the <a href="https://www.orkg.org/orkg/">Open Research Knowledge Graph</a> (ORKG) based on the 4 pillars of crowdsourcing according to the reference model for crowdsourcing by Hosseini et al. [1]. This overview represents the features of the current implementation status of&nbsp; ORKG as a crowdsourcing platform.</p> <p>[1] M. Hosseini, K. Phalp, J. Taylor, and R. Ali, &quot;<a href="https://ieeexplore.ieee.org/abstract/document/6861072?casa_token=zWTHNBHeH6kAAAAA:1n3EJjajqgSkSk154g4DlNFAmJs_7o3KY7LnobvP_W7AUIr-FvM5OyGP1FaRn68zUX-2oYHh7A">The Four Pillars of Crowdsourcing: A Reference Model</a>&quot;, in 2014 IEEE 8th International Conference on Research Challenges in Information Science (RCIS). IEEE, 2014, pp. 1&ndash;12.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

CORE: Gene Expression-Cancer Knowledge Base

<p>This repository contains the Gene Expression-Cancer&nbsp;Knowledge Base&nbsp;generated by the CORE system.<br> The &#39;schema.owl&#39; file contains the KB schema, whereas the &#39;data.ttl&#39; file contains the actual data.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

FastKGQA: A modified knowledge base of the MoviesQA dataset for prototyping

<p><strong>Full Changelog</strong>: <a href="https://github.com/d1egoprog/FastKGQA/commits/1.0">https://github.com/d1egoprog/FastKGQA/commits/1.0</a></p>

openmit-licenseFeb 2023View details →
zenodo36/100

Teacher Characteristics, Knowledge and Use of Evidence-Based Practices in Autism Education in Ireland_Dataset_V1

<p>Dataset to accompany the paper &quot;<strong>Teacher Characteristics, Knowledge and Use of Evidence-Based Practices in Autism Education in Ireland&quot;.&nbsp;</strong>The purpose of this paper was to identify differences in teachers knowledge and use of evidence-based practices across different teacher characteristics.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
dryad36/100

Data from: Incorporating herders' knowledge and ecosystem-based adaptation strategies in local decision-making

<p class="CxSpFirst"> 1.    Ecosystem-based adaptation (EbA) relies upon the capacity of ecosystems to buffer communities against the adverse impacts of climate change. Maintaining ecosystems that deliver critical services to communities can also provide co-benefits beyond adaptation, such as climate mitigation and protection of biological diversity and livelihoods. EbA has to a limited extent drawn upon indigenous-and local knowledge (ILK) for defining critical services and for implementing EbA in decision-making. This is a paradox given that the primary focus of EbA is to enable communities to adapt to climate change.<br> 2.      The purpose of this study was to elucidate EbA strategies that take into account the knowledge of Sámi reindeer herders about pastures in tundra regions. We first examined what constitutes critical services as perceived by Sámi reindeer herders through a synthesis of data and literature. We thereafter used content analysis of 91 land use cases from 2010-2018 to investigate to what extent the herders' knowledge and maps over seasonal pastures and migratory routes are used in local decision-making. Finally, we propose EbA strategies of relevance to Sámi communities and pastoral communities elsewhere.<br> 3.      Our analysis revealed that reindeer herders and organizations representing their interests perceived threats from green energy development, tourism, recreation, public road construction and powerlines. These threats included the loss of key habitats and the loss of connectivity for migration between seasonal pastures. Herders' knowledge is incorporated through participatory tools to protect the ecosystems and services crucial for herders, but multiple competing land uses result in incremental loss of pastures regardless. <br> 4.      Synthesis and application. Pastoralists need access to diverse resources on seasonal pastures and the ability to move between pastures when snow, ice, rainfall and the timing of critical service supplies changes. Drawing on herders' knowledge to elicit EbA strategies is vital for buffering the adverse effects of climate change. EbA that incorporates indigenous perspectives cannot purely rely on co-benefit approaches. Fundamental trade-offs exist between adaptation needs and other land uses, such as infrastructure, tourism and green energy development</p>

opencc-zeroDec 2019View details →
zenodo36/100

KLIFS: A Knowledge-Based Structural Database To Navigate Kinase–Ligand Interaction Space

<p>The Kinase-Ligand Interaction Fingerprints and Structure database (KLIFS) contains a consistent structural alignment and deconstruction of the kinase domains from&nbsp;over 1734 PDB structures covering 190 different human kinases.&nbsp;</p> <p>Every crystal structure was structurally aligned in a consistent manner, subsequently broken down from the full complex into separate structural parts: the protein, the orthosteric ligand-binding pocket (85 aligned residues covering the catalytic cleft), orthosteric and allosteric ligand(s), ions, organometallics, cofactors, and waters. By combining the pocket with the orthosteric ligand all interactions are annotated using Interactions FingerPrints (IFPs) for systematic comparison.</p>

opencc-zeroAug 2013View details →
zenodo36/100

Learning stochastic process-based models of dynamical systems from knowledge and data - Libraries, incomplete models and data

<p>The archive contains all libraries of domain knowledge, the incomplete models and the data used in the experiments described in the manuscript titled &quot;Learning stochastic process-based models of dynamical systems from knowledge and data&quot; pubilshed in BMC Systems Biology</p>

openbsd-3-clauseNov 2015View details →
zenodo36/100

RE-DWELL Knowledge base dataset (RDF)

<p>This RDF dataset represents a collection of concepts and case studies focused on urban and housing projects, capturing a wide range of details such as project descriptions, types, locations, timeframes, construction systems, and associated Sustainable Development Goals (SDGs).</p> <p>It includes metadata about the authors of each study, related concepts, and the geographic contexts of the projects.&nbsp;</p> <p>The dataset is designed to support analysis and exploration of trends, relationships, and insights in urban development, housing strategies, and their alignment with broader social and environmental goals.&nbsp;</p> <p>It provides a rich foundation for research in architecture, urban planning, and sustainability studies.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Dataset for "Personalised Learning Environments Based on Knowledge Graphs and the Zone of Proximal Development"

<p>The dataset&nbsp;accompanying our paper &quot;Personalised Learning Environments Based on Knowledge Graphs and<br> the Zone of Proximal Development&quot; published in proceedings of CSEDU 2022.</p> <p>In the dataset you will find the raw csv results from both the explorative survey and the evaluation survey.<br> <br> The surveys were made in Google Forms and the results have also been exported as pdf files that are included as well.<br> <br> Finally, screenshots of the application, grouped by module can&nbsp;be found inside the screenshots.zip archive.</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Dataset for paper: " Knowledge Graph Embeddings based Approach for Author Name Disambiguation using Literals"

<p>This dataset consists in two distinct scholarly&nbsp;knowledge graph created from two publicly available bibliographic datasets: 1) a triplestore covering information about the journal <em>Scientometrics</em>&nbsp;provided by&nbsp;<em>OpenCitations</em> (available <a href="https://doi.org/10.5281/zenodo.5151264">here</a>), and 2) the <em>AMiner </em>AND benchmark from 2018&nbsp;available <a href="https://static.aminer.cn/misc/na-data-kdd18.zip">here</a>. This KG was extracted&nbsp;for a research project on knowledge graph embeddings (KGEs)&nbsp;for author disambiguation. Structural triples of the knowledge graphs are split into training, testing and validation for applying representation learning methods. Textual literals and numeric literals were stored separately in order to implement multimodal approaches for KGEs (see&nbsp;<a href="https://arxiv.org/abs/1802.00934">arXiv:1802.00934</a>). For the same reason, textual literals and numeric literals are already stored into sentence embeddings and a&nbsp;numeric matrix&nbsp;respectively in the files&nbsp;<em>textual_literals.npy&nbsp;</em>and&nbsp;<em>numeric_literals.npy </em>in order to simplify the representation learning task. The file <em>and_eval.json</em> of each KG&nbsp;contains the evaluation dataset used for evaluating our AND architecture. For the script used to gather this dataset see <a href="https://github.com/sntcristian/and-kge/tree/main/src/AMiner-534K">https://github.com/sntcristian/and-kge/tree/main/src/AMiner-534K</a> and&nbsp;<a href="https://github.com/sntcristian/and-kge/tree/main/src/OC-782K">https://github.com/sntcristian/and-kge/tree/main/src/OC-782K</a>.</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Cloud computing is one of the most popular and sophisticated technologies adopted by organizations worldwide. Some world-leading organizations enhance their efficiency and effectiveness by using cloud computing technology. Working from home (WFH) has been a popular trend among organizations during the coronavirus (COVID-19) pandemic. The COVID-19 saw a breakthrough in work cultures and environments where working from home was a remarkable success in remote working environments, despite being a rare phenomenon in Sri Lanka. Yet, it is argued that the deployment of work from home has not been effective among Sri Lankan business organizations due to a lack of IT infrastructure, facilities, and knowledge. The purpose of the study is to investigate the impact of cloud computing, embracing the service models (Infrastructure as a Service, Platform as a Service, and Software as a Service) as theoretical lenses and testing the COVID-19 as the moderator. The study has been conducted based on a deductive approach and adopted a stratified random sampling method. The sample consisted of 384 IT employees among those who had experienced working from home. The study utilized multiple regression and found that cloud computing service models significantly impact work from home with the moderating effect of COVID-19.

<p>Cloud computing is one of the most popular and sophisticated technologies adopted by organizations worldwide. Some world-leading organizations enhance their efficiency and effectiveness by using cloud computing technology. Working from home (WFH) has been a popular trend among organizations during the coronavirus (COVID-19) pandemic. The COVID-19 saw a breakthrough in work cultures and environments where working from home was a remarkable success in remote working environments, despite being a rare phenomenon in Sri Lanka. Yet, it is argued that the deployment of work from home has not been effective among Sri Lankan business organizations due to a lack of IT infrastructure, facilities, and knowledge. The purpose of the study is to investigate the impact of cloud computing, embracing the service models (Infrastructure as a Service, Platform as a Service, and Software as a Service) as theoretical lenses and testing the COVID-19 as the moderator. The study has been conducted based on a deductive approach and adopted a stratified random sampling method. The sample consisted of 384 IT employees &nbsp;among those who had experienced working from home. The study utilized multiple regression and found that cloud computing service models significantly impact work from home with the moderating effect of COVID-19.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Dataset: Impact of HIV knowledge and stigma on the uptake of HIV testing - results from a community-based participatory research survey among migrants from sub-Saharan Africa in Germany

<p>Dataset for: Kuehne A, Koschollek C, Santos-H&ouml;vener C, Thorlie A, M&uuml;llersch&ouml;n J, Tshibadi CM, Mayamba P, Batemona-Abeke H, Amoah S, Greiner VW, Bursi TD, Bremer V: Impact of HIV knowledge and stigma on the uptake of HIV testing - results from a community-based participatory research survey among migrants from sub-Saharan Africa in Germany.</p> <p>This dataset has been described in a PLoS One paper and contains all data necessary to replicate the results presented within this paper (10.1371/journal.pone.0194244). Please cite both the paper as well as the DOI of this dataset if you make use of the data.</p>

opencc-by-nc-4.0Mar 2018View details →
zenodo36/100

FLABASE: A Flamenco Knowledge Base

<p>FlaBase (Flamenco Knowledge Base) is the acronym of a new knowledge base of flamenco music. Its ultimate aim is to gather all available online editorial, biographical and musicological information related to flamenco music. A first version is just being released. Its content is the result of the curation and extraction processes. FlaBase is stored in JSON format, and it is freely available for download. This first release of FlaBase contains information about 1,102 artists, 74 palos (flamenco genres), 2,860 albums, 13,311 tracks, and 771 Andalusian locations.</p> <p><strong>Data sources</strong></p> <p>Data was compiled and curated from different sources: Wikipedia, DBpedia, Andalucia.org, elartedevivirelflamenco.com, MusicBrainz, flun.cica.es/index.php/grabaciones/base-datos-grabaciones and juntadeandalucia.es/institutodeestadisticaycartografia/sima</p> <p><strong>Using this dataset</strong></p> <p>For more details on how these files were generated, we refer to the following scientific publication. We would highly appreciate if scientific publications of works partly based on the FlaBase dataset quote the following publication:</p> <blockquote> <p>Oramas, S.,&nbsp;G&oacute;mez F.,&nbsp;G&oacute;mez E., &amp;&nbsp;Mora J.&nbsp;(2015).&nbsp;&nbsp;<a href="http://mtg.upf.edu/node/3319">FlaBase: Towards the Creation of a Flamenco Music Knowledge Base</a>.&nbsp;16th International Society for Music Information Retrieval Conference.&nbsp;</p> </blockquote> <p>&nbsp;We are interested in knowing if you find our datasets useful! If you use our dataset please email us at <a href="mailto:mtg-info@upf.edu">mtg-info@upf.edu</a> and tell us about your research.</p> <p>&nbsp;</p> <p><a href="https://www.upf.edu/web/mtg/flabase">https://www.upf.edu/web/mtg/flabase</a></p>

opencc-by-4.0Sep 2015View details →
zenodo36/100

FreeCiv games for the experiment on comparing Knowledge-Based Reinforcement Learning and Neural Networks in Strategic Games

<p>Dataset provides played&nbsp;FreeCiv games. The Tournament&nbsp;subset of them was played fully by two Artificial Intelligence (AI) agents against each other and one more computer player. The Human games subset was played by humans for demonstration/teaching purposes. The rest of the games were played 120 turns and stopped.</p>

opencc-by-4.0Jan 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record