Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

190

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

190 results for “knowledge base”

Learn how ShareScore rates datasets ↗
zenodo40/100

Fig. 13 in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages

Fig. 13. Male representatives of three subfamilies. A, C, E. Frontal view. B, D, F. Lateral view. — A–B. Tatuidris tatusia, Agroecomyrmecinae (Panama, CASENT0178870, E. Prado). C–D. Paraponera clavata, Paraponerinae (Guyana, CASENT0902407, R. Perry). E–F. Pseudoponera stigma, Ponerinae (Paraguay, CASENT0178182, A. Nobile). Scale bars: A = 0.1 mm, B = 0.5 mm, C, F = 1.0 mm, D = 2.0 mm, E = 0.2 mm.

opencc-by-4.0Apr 2015View details →
zenodo40/100

Fig. 7. A–B. Forewing. A. Ventral view, gyne. B. Dorsal view, male. C–D in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages

Fig. 7. A–B. Forewing. A. Ventral view, gyne. B. Dorsal view, male. C–D. Mesosoma, lateral view, male. E–F. Paramere, lateral view, male. — A. Aneuretus simoni (Sri Lanka, CASENT0172259, A. Nobile). B. Nylanderia vividula (U.S.A., CASENT0058918, A. Nobile). C. Apomyrma stygia (Central African Republic, CASENT0086073, E. Prado). D. Anochetus boltoni (Madagascar, CASENT0063847, A. Nobile). E. Dolichoderus validus (Costa Rica, INB0003662427, B. Boudinot). F. Formica pacifica (U.S.A., JTLC000006350, B. Boudinot). Scale bars: A–C, E–F = 0.5 mm, D = 1.0 mm.

opencc-by-4.0Apr 2015View details →
zenodo40/100

SAQI: An Ontology based Knowledge GraphPlatform for Social Air Quality Index

<p>This dataset consists of all contributions made by Social AQI (SAQI)&nbsp;project. The description of dataset is as below -</p> <p>Local Sensor Data (hyperlocal-air-quality-sensor-data) - contains all sensors values recorded through local neighbourhood sensors throught the length of the project&nbsp;<br> Locations for all these sensors are as below - In Najafgarh, Delhi, India : Jharoda Kalan, Nangli Dairy and&nbsp;DTC Bus terminal.<br> In Okhla : Sanjay Colony, Tekhand, Shaheen Bagh.</p> <p>Data from <a href="https://cpcb.nic.in/">Central Pollution Control Board</a>&nbsp;(central-air-quality-sensor-data) - Najafgarh_CPCB.csv,&nbsp;Okhla_CPCB.csv : Contains data provided by CPCB from Najafgarh,Delhi&nbsp;and Oklha, Delhi</p> <p><br> PollutionODP.owl : Ontology Design Pattern for pollution -&nbsp;http://ontologydesignpatterns.org/wiki/Submissions:Pollution.<br> <br> Ontology&nbsp;: SAQI ontology as triples (ttl), xml (rdf) and json-ld (json) serialization format<br> Ontology documentation : ontology/diagram contains figures describing ontology, ontology/documentation/saqi.html contains LODE documentation for the ontology</p> <p><br> ethnographic-survey-data - anonymized survey responses for initial pollution perception and literacy survey as well as SAQI app feedback survey.</p> <p>SHACL-shapes - for validating against SAQI ontology.</p> <p>&nbsp;sparql-queries - sample queries to run on our ontology.</p> <p>setup-rdf-store-script - script to setup rdf store with given data using rml mapper.</p> <p>&nbsp;</p> <p><br> &nbsp;&nbsp; &nbsp;</p>

openapache2.0Dec 2022View details →
zenodo40/100

Multilingual MigrationsKB: A Mulitlingual Knowledge Base of Migration related annotated Tweets

<p><strong>Multilingual MigrationskB (MGKB) </strong>is a mulitlingual extended version of English <a href="https://zenodo.org/record/5206820#.YRqF1nUza0o">MGKB</a>. The tweets geotagged with Geo location from 32 European Countries (<em><strong>Austria, Belgium, Bulgaria,&nbsp;Croatia, Cyprus, Czech, Denmark, Estonia, Finland, France, Germany, Greece, Hungary, Ireland, Italy, Latvia, Lithuania, Luxembourg, Malta, Netherlands, Poland, Portugal, Romania, Slovakia, Slovenia, Spain, Sweden, Iceland, Liechtenstein, Norway, Switzerland, the United Kingdom</strong></em>)&nbsp;&nbsp;are extracted and filtered by 11 languages (<em><strong>English, French, Finnish, German, Greek, Dutch, Hungarian, Italian, Polish, Spain, Swedish</strong></em>). Metadata information about the tweets, such as <strong>Geo information (place name, coordinates, country code)</strong> are included. <strong>MGKB</strong>&nbsp; contains <strong>sentiments, offensive and hate speeches, topics, hashtags, user mentions</strong> in RDF format. The schema of <strong>MGKB</strong> is an extension of TweetsKB for migration related information. Moreover, to associate and represent the potential economic and social factors driving the migration flows, the data from <a href="https://ec.europa.eu/eurostat/web/main/home">Eurostat</a>&nbsp; and <a href="https://spec.edmcouncil.org/fibo/ontology/">FIBO</a> ontology was used. To represent multilinguality, the<a href="https://www.cidoc-crm.org/"> CIDOC Conceptual Reference Model (CIDOC-CRM)</a>&nbsp;is used. The extracted economic indicators, i.e., GDP Growth Rate, Total Unemployment Rate, Youth Unemployment Rate, Long-term Unemployment Rate and Income per househould, are connected with each tweet in RDF using geographical and temporal dimensions.&nbsp;</p> <p>For this version, the Multilingual MGKB is delivered separated by year. The extracted topic words are also published.</p> <p>Code:&nbsp;<a href="https://github.com/migrationsKB/MRL">https://github.com/migrationsKB/MRL</a></p> <p>Please contact Yiyi Chen (yiyi.chen@partner.kit.edu) for pretrained models (Sentiment analysis/hate speech detection/ETM) if necessary.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

YAGO knowledge base

<p>YAGO is a knowledge base, i.e., a database with knowledge about the real world. YAGO contains both entities (such as movies, people, cities, countries, etc.) and relations between these entities (who played in which movie, which city is located in which country, etc.). All in all, YAGO contains more than 50 million entities and 2 billion facts.</p> <p>YAGO arranges its entities into classes: Elvis Presley belongs to the class of people, Paris belongs to the class of cities, and so on. These classes are arranged in a taxonomy: The class of cities is a subclass of the class of populated places, this class is a subclass of geographical locations, etc.</p> <p>YAGO also defines which relations can hold between which entities: birthPlace, e.g., is a relation that can hold between a person and a place. The definition of these relations, together with the taxonomy is called the ontology.</p> <p>For more information, see https://yago-knowledge.org</p> <p>&nbsp;</p>

opencc-by-sa-4.0Feb 2020View details →
zenodo40/100

Topical Classification of Food Safety Publications with a Knowledge Base – materials

<p>Supplementary material for the article titled &quot;<strong><a href="https://doi.org/10.1007/978-981-19-4364-5_48">Topical Classification of Food Safety Publications with a Knowledge Base</a></strong>&quot;.</p> <p><strong>Citation</strong></p> <p>If you use this data in research work, please cite this paper:</p> <p>Sowinski, P., Wasielewska-Michniewska, K., Ganzha, M., &amp; Paprzycki, M. (2022). Topical Classification of Food Safety Publications with a Knowledge Base. In&nbsp;<em>Sustainable Technology and Advanced Computing in Electrical Engineering</em>&nbsp;(pp. 673-693). Springer, Singapore.</p> <p>BibTeX:</p> <pre><code>@incollection{sowinski2022topical, title={Topical Classification of Food Safety Publications with a Knowledge Base}, author={Sowinski, Piotr and Wasielewska-Michniewska, Katarzyna and Ganzha, Maria and Paprzycki, Marcin}, booktitle={Sustainable Technology and Advanced Computing in Electrical Engineering}, pages={673--693}, year={2022}, publisher={Springer}, doi={10.1007/978-981-19-4364-5_48} }</code></pre> <p>&nbsp;</p>

openmit-licenseFeb 2022View details →
zenodo40/100

Conflict Event Knowledge Graph based on the ongoing Ukraine-Russia Conflict

<p><strong>Conflict Event Knowledge Graph</strong> is a <strong>Knowledge Graph</strong> that links the tweets and the current events&nbsp;of <strong>Russia-Ukraine Conflict</strong>&nbsp;portrayed in&nbsp;<strong>English Wikipedia</strong> from <strong>24th&nbsp;February 2022 to 4th March 2022</strong>.&nbsp;</p> <p><strong>Abstract</strong>: In the current situation of Russia-Ukraine Conflict, numerous contents are posted daily to engage in the discourse about different events in this conflict. The goal of this study is to provide a framework to enable analysis of these events and the Twitter data, utilizing entity linking. The relevant tweets and events are integrated into a Knowledge Graph <strong>ConflictEventKG</strong>, and the resources are made publicly available.</p> <p><strong>Homepage</strong>:&nbsp;<a href="https://siebeniris.github.io/ConflictEventKG/">https://siebeniris.github.io/ConflictEventKG/</a></p> <p><strong>Code</strong>:&nbsp;<a href="https://github.com/siebeniris/ConflictEventKG">https://github.com/siebeniris/ConflictEventKG</a></p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

MigrationsKB: A Knowledge Base of Migration related annotated Tweets

<p><strong>MigrationsKB(MGKB)</strong>&nbsp;is a public Knowledge Base of anonymized&nbsp;<strong>Migration</strong>&nbsp;related&nbsp;<strong>annotated</strong>&nbsp;tweets. The MGKB currently contains over&nbsp;<strong>200 thousand</strong>&nbsp;tweets, spanning over 9&nbsp;years (January 2013 to July 2021), filtered with 11 European countries of&nbsp;<em>the United Kingdom, Germany, Spain, Poland, France, Sweden, Austria, Hungary, Switzerland, Netherlands and Italy</em>.&nbsp;<strong>Metadata</strong>&nbsp;information about the tweets, such as Geo information (<strong>place name</strong>,&nbsp;<strong>coordinates</strong>,&nbsp;<strong>country code</strong>).&nbsp;<strong>MGKB</strong>&nbsp;contains&nbsp;<strong>entities</strong>,&nbsp;<strong>sentiments</strong>,&nbsp;<strong>hate speeches</strong>,&nbsp;<strong>topics</strong>,&nbsp;<strong>hashtags</strong>,&nbsp;<em>encrypted user mentions</em>&nbsp;in RDF format. The schema of&nbsp;<strong>MGKB</strong>&nbsp;is an extension of TweetsKB for migrations related information. Moreover, to associate and represent the potential economic and social factors driving the migration flows such as&nbsp;<a href="https://ec.europa.eu/eurostat/web/main/home"><strong>eurostat</strong></a>,&nbsp;<a href="https://www.statista.com/"><strong>statista</strong></a>, etc. FIBO ontology was used. The extracted&nbsp;<strong>economic indicators</strong>, such as GDP Growth Rate, are connected with each Tweet in RDF using geographical and temporal dimensions. The user IDs and the tweet texts are encrypted for privacy purposes, while the tweet IDs are preserved.</p> <p>For this version, the <strong>MGKB</strong> is delivered as a whole and separately by year. The extracted entities and topic words are also published.</p> <p>Online SPARQL endpoint&nbsp;&nbsp;<a href="https://mgkb.fiz-karlsruhe.de/sparql/">https://mgkb.fiz-karlsruhe.de/sparql/</a></p> <p>More information please refer to the website&nbsp;<a href="https://migrationskb.github.io/MGKB/">https://migrationskb.github.io/MGKB/</a>.</p> <p>Please contact Yiyi Chen (yiyi.chen@partner.kit.edu) for pretrained models (sentiment analysis/hate speech detection/ETM) if necessary.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

MineDojo Internet Knowledge Base (Reddit)

<p><strong>Project website:</strong>&nbsp;<a href="https://minedojo.org">minedojo.org</a></p> <p><strong>Paper:</strong>&nbsp;<a href="https://arxiv.org/abs/2206.08853">arxiv.org/abs/2206.08853</a></p> <p><strong>GitHub:</strong>&nbsp;<a href="https://github.com/MineDojo/MineDojo">github.com/MineDojo/MineDojo</a></p> <p><strong>We collect 340K+ Reddit posts along with 6.6M comments under the &ldquo;<a href="https://www.reddit.com/r/minecraft">r/Minecraft</a>&rdquo; subreddit.</strong>&nbsp;These posts ask questions on how to solve certain tasks, showcase cool architectures and achievements in image/video snippets, and discuss general tips and tricks for players of all expertise levels. Large language models can be finetuned on our Reddit corpus to internalize Minecraft-specific concepts and develop sophisticated strategies.</p> <p>Check out our&nbsp;paper!</p> <p>&nbsp;</p> <pre><code class="language-markdown">@article{fan2022minedojo, title = {MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge}, author = {Linxi Fan and Guanzhi Wang and Yunfan Jiang and Ajay Mandlekar and Yuncong Yang and Haoyi Zhu and Andrew Tang and De-An Huang and Yuke Zhu and Anima Anandkumar}, year = {2022}, journal = {arXiv preprint arXiv: Arxiv-2206.08853} }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

MineDojo Internet Knowledge Base (Wiki)

<p><strong>Project website:</strong>&nbsp;<a href="https://minedojo.org">minedojo.org</a></p> <p><strong>Paper:</strong>&nbsp;<a href="https://arxiv.org/abs/2206.08853">arxiv.org/abs/2206.08853</a></p> <p><strong>GitHub:</strong>&nbsp;<a href="https://github.com/MineDojo/MineDojo">github.com/MineDojo/MineDojo</a></p> <p>The Minecraft Wiki pages cover almost every aspect of the game mechanics, and supply a rich source of unstructured knowledge in multimodal tables, recipes, illustrations, and step-by-step tutorials.&nbsp;<strong>We scrape 6,735 pages that interleave text, images, tables, and diagrams.</strong>&nbsp;To preserve the layout information, we also save the screenshots of entire pages and extract bounding boxes of the visual elements.</p> <p>There are two files in&nbsp;our Wiki knowledge base.</p> <ul> <li><strong>wiki_samples.zip:</strong> A sample version of the full knowledge base&nbsp;(10 pages).&nbsp;</li> <li><strong>wiki_full.zip:</strong> The full knowledge base (6,735 pages).&nbsp;</li> </ul> <p>Check out our&nbsp;paper!</p> <pre><code class="language-markdown">@article{fan2022minedojo, title = {MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge}, author = {Linxi Fan and Guanzhi Wang and Yunfan Jiang and Ajay Mandlekar and Yuncong Yang and Haoyi Zhu and Andrew Tang and De-An Huang and Yuke Zhu and Anima Anandkumar}, year = {2022}, journal = {arXiv preprint arXiv: Arxiv-2206.08853} }</code></pre>

opencc-by-nc-sa-3.0Jun 2022View details →
zenodo40/100

C-RIDGE: Indoor CO2 Data Collection System for Large Venues Based on Prior Knowledge

<p>This CO2 of C-RIDGE system dataset contains the high spatial and temporal resolution of the CO2<br> measures with the corresponding timestamp &nbsp;of static wireless sensors.&nbsp;<br> 45 sensors are densely deployed on the stand in the venue. The id and relative positions of sensors&nbsp;<br> are shown in device message. The sampling rate is adaptively adjusted according to the competition schedule.&nbsp;<br> The sampling interval is 5 minutes during the competition and 15 minutes otherwise. Please refer to&nbsp;<br> the description of the experimental setup in the data descriptor paper.</p> <p>In the process of data collection, the data cleaning process is performed &nbsp;to remove&nbsp;<br> and calibrate outliers and abnormal trend data. The script of data cleaning algorithm is&nbsp;<br> provided in this repository. For details about the data cleaning process, please refer to the script in&nbsp;<br> this repository and data descriptor paper.</p> <p>The dataset in this repository is processed version. The raw dataset is not included in this repository.</p> <p>Data is stored as CSV file. Each device is numbered in order of placement. There are 45 sensors in total,<br> 1 to 45 in the csv file are sensor numbers, timestamp as the China Standard Time (GMT+8), Each timestamp&nbsp;<br> corresponds to 45 CO2 concentration data from different sensors.</p> <p>To access the dataset, any programming language that can access the CSV file is appropriate. Users can&nbsp;<br> also directly open the CSV file. To successfully execute script files, Pycharm with&nbsp;Python 3.0&nbsp;is required.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Figure 4. CBR Implementation-Intelligent Flowcharting Developmental Approach to Legal Knowledge Based System

<p>The development of the case based reasoning module is in done in java net-beans. Proper<br> verification and validation of this module was done by the legal experts. The cases related to<br> Transfer of property act were collected from different legal databases and compiled. The necessary<br> keywords were framed, which were used in searching for the related cases. The following Fig 2.0<br> gives the screen shot of the CBR module develoed in Java Net beans.</p>

opencc-by-4.0Jan 2018View details →
zenodo40/100

Fig 3 : VisiRule Implementation Module-1-Intelligent Flowcharting Developmental Approach to Legal Knowledge Based System

<p>The code of this flowchart is developed by the VisiRule in FLEX/ Prolog. As the source<br> code very huge it has not be incorporated in the paper.</p>

opencc-by-4.0Jun 2011View details →
zenodo40/100

Figure 1 - VisiRule architecture-Intelligent Flowcharting Developmental Approach to Legal Knowledge Based System

<p>In the development of RBR we used the intelligent flowcharting approach. VisiRule is a tool<br> for creating decision support software purely by drawing flowcharts. The end result is Flex or<br> Prolog code which is automatically generated, compiled and ready to run, but which can also be<br> copied and used in a separate program. Not only can VisiRule be used by people with minimal<br> programming skills. VisiRule also enhances productivity by considerably reducing the time it takes<br> to produce a decision support system. VisiRule is an intelligent flowcharting tool in two senses.<br> Firstly, it is used to create knowledge-based systems and, secondly, it intelligently guides the<br> construction process by constraining what you can and can&#39;t do on the basis of the semantic content<br> of the emerging program. VisiRule provides the automatic construction of menu dialogues from<br> questions. These are populated by items inferred from expression boxes throughout the flowchart<br> tree which have a path to the question.</p>

opencc-by-4.0Jun 2011View details →
zenodo40/100

Figure 5. Multi-agent system information and knowledge scheme.-Design and Development a Control and Monitoring System for Greenhouse Conditions Based-On Multi Agent System

<p>The negotiations between agents are subject to optimization based on &ldquo;knowledge&rdquo; that is<br> derived from complete production models, yield models or even sparse models as expressed in<br> fuzzy expert rules or practical rules of thumb. In addition, pest control and plant disease models<br> provide additional information useful to the design of a successful strategy for optimal management<br> [15] (illustrated in Figure 5).</p>

opencc-by-4.0Jun 2011View details →
zenodo40/100

Figure 3. Traditional Approach and Process Management System Approach – Effort Percentage Comparison-Business Process Management – A Traditional Approach versus a Knowledge Based Approach

<p>Comparing the results obtained in using the two approaches (Figure 3), it is possible to note<br> a significant reduction in terms of both effort and working hours in correspondence of design and<br> development phases.</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

BRAIN Journal-Right-Linear Languages Generated in Systems of Knowledge Representation based on LSG-Right-Figure 3. Representation of the grammar G1 in the labelled graph G0 1

<p>If we take the labeled graph G0 1 given in Figure 3 and construct the stratified graph structure over (99) such that (100) we obtain (101), (102).&nbsp;&nbsp;</p> <p>In this paper, we proposed a new system for formal language generation by means of stratified graphs structures. This mechanism can generate languages of the first type and of the second type. More precisely, we propose a new system for formal language generation by means of a system of knowledge based on stratified graphs. We exemplified that, using an interpretation&nbsp; system specially defined for stratified graphs representations, a particular formal language can be obtained by means of the resulted accepted structured paths.</p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

BRAIN Journal-Right-Linear Languages Generated in Systems of Knowledge Representation based on LSG-Right-Figure 2. The representation of the rule

<p>In order to model these derivations in the stratified graph G, each production of the grammar will be represented in the labeled graph G0 by a direct arc of the form given in Figure 2.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2017View details →
zenodo40/100

BRAIN Journal-Right-Linear Languages Generated in Systems of Knowledge Representation based on LSG-Right-Figure 1. The graphical representation of the morphism

<p>A morphism of partial algebras such that (30) and if (31), then (32) (see Figure 1). We obtain f(L) = T which means that &ldquo;for every element of L the associated element of T is computed by the morphism f&rdquo; (Ţăndăreanu, 2000).</p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

Crosscult Knowledge Base (CCKB)

<p>&ldquo;The Crosscult Knowledge Base (CCKB) is a comprehensive structure of semantic definitions and formalisms, developed for facilitating interoperable connections between the cultural heritage dataset of contributing to Crosscult . &nbsp;It is written in OWL2 (the standard ontology language for the Semantic Web) and enables augmentation, semantic linking, semantic-based reasoning and retrieval across disparate data resources&rdquo;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record