Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

483

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

483 results for “semantics”

Learn how ShareScore rates datasets ↗
OpenNeuro52/100

FeedBES - FeedBack signals from Episodic and Semantic memories.

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
OpenNeuro52/100

Shared neural codes for visual and semantic information about familiar faces in a common representational space

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
zenodo52/100

A Benchmark dataset on Semantic Change in Scholarly Publications on Disability

<p>This is a benchmark dataset for semantic shift detection in disability-related corpora, including collected title and abstract text from PubMed and ArXiv, annotation sets based on domain experts and LLMs, and extracted KGs (Wikidata entity claims). The corpus from PubMed covers the period from the 1900s to 2023, while the corpus from ArXiv covers the period from the 1990s to 2023. The corpus was filtered based on 16 disability-related target words. In the annotation sets, '1' indicates that a semantic shift occurred for a target word, while '0' indicates the opposite. In particular, the LLM-based annotation sets include their generated text, and we used the Llama2 and GPT-4 models. '7b' refers to the parameter size of the Llama2 model. Graph_data.zip contains Wikidata entity claims.</p>

opencc-by-4.0Apr 2024View details →
OpenNeuro48/100

Decoding of multisensory semantics and memories in low-level visual cortex

Open the record for dataset details and reuse information.

openCC0Jan 2019View details →
zenodo48/100

Data for "On the Practice of Semantic Versioning for Ansible Galaxy Roles: An Empirical Study and a Change Classification Model"

<p>This dataset accompanies a replication package provided for a study on Semantic Versioning for Ansible Galaxy roles.</p> <p>The replication package is available at https://github.com/ROpdebee/ansible_semver_ext_replication</p>

opencc-by-4.0Mar 2021View details →
zenodo48/100

Catalogues of Semantic Artefacts - Maturity Dimensions and Features

<p>This dataset contains, in three different formats (XSLX, CSV, and PDF), a description of twelve maturity dimensions identified from the literature that can be used to measure the maturity of the semantic artefacts catalogues (SAC). For each dimension, a number from 2 to 6 features has been added, for a total of 43 features overall.</p>

opencc-zeroMay 2023View details →
zenodo48/100

The Semantic Turkey metadata registry ontology

<p>An application profile of DCAT combining it with other metadata vocabularies (e.g. VoID, DCTERMS, LIME)&nbsp;to meet requirements elicited in various use cases of the Semantic Web platform Semantic Turkey</p>

opencc-by-4.0May 2022View details →
zenodo48/100

Supplementary Material for Embodied Emotions in Ancient Neo-Assyrian Texts Revealed by Bodily Mapping of Emotional Semantics

<p>This dataset accompanies the article "Embodied Emotions in Ancient Neo-Assyrian Texts Revealed by Bodily Mapping of Emotional Semantics" (Lahnakoski &amp; Bennett et al., submitted).&nbsp;</p> <p>It includes the Neo-Assyrian text corpus that is the basis for the word embeddings, a list of the Akkadian emotion and body words of interest for this study, and the scripts, toolboxes, and data used to generate the heat maps of the body.</p> <p>There is an additional folder containing the high resolution figures included in the article.</p> <p>A detailed ReadMe (README.txt) provides an overview of the folders.</p>

opencc-by-4.0May 2024View details →
zenodo48/100

S1S2-Water: A global dataset for semantic segmentation of water bodies from Sentinel-1 and Sentinel-2 satellite images

<p>The S1S2-Water dataset is a global reference dataset for training, validation and testing of convolutional neural networks for semantic segmentation of surface water bodies in publicly available Sentinel-1 and Sentinel-2 satellite images. The dataset consists of 65 triplets of Sentinel-1 and Sentinel-2 images with quality checked binary water mask. Samples are drawn globally on the basis of the Sentinel-2 tile-grid (100 x 100 km) under consideration of pre-dominant landcover and availability of water bodies. Each sample is complemented with metadata and Digital Elevation Model (DEM) raster from the Copernicus DEM.</p><p>This work was supported by the German Federal Ministry of Education and Research (BMBF) through the project "Künstliche Intelligenz zur Analyse von Erdbeobachtungs- und Internetdaten zur Entscheidungsunterstützung im Katastrophenfall" (AIFER) under Grant 13N15525, and by the Helmholtz Artificial Intelligence Cooperation Unit through the project "AI for Near Real Time Satellite-based Flood Response" (AI4FLOOD) under Grant ZT-IPF-5-39.&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo48/100

FrameNet Semantic Frame Disambiguation with CrowdTruth

<p>This repository contains a ground truth corpus for semantic frame disambiguation, acquired with crowdsourcing and processed with <strong><a href="http://crowdtruth.org/">CrowdTruth</a></strong> metrics that capture ambiguity in annotations by measuring inter-annotator disagreement.</p> <p>The dataset contains annotations for 433 sentence-word pairs from the <a href="https://framenet.icsi.berkeley.edu/">FrameNet corpus v.1.7</a>, with each sentence-word pair annotated for frame disambiguation by 15 workers. The crowdsourced data was collected from <a href="https://www.mturk.com/">Amazon Mechanical Turk</a>.</p> <p>The corpus has been referenced in the following paper:</p> <ul> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="https://arxiv.org/abs/1805.00270">Capturing and Interpreting Ambiguity in Crowdsourcing Frame Disambiguation</a></strong>. <a href="https://www.humancomputation.com/2018/">HCOMP 2018</a>.</li> </ul> <p>To replicate the data processing from the paper, use the Jupyter Notebook file <code>CrowdTruth metrics.ipynb</code>. It requires the installation of the <a href="https://github.com/CrowdTruth/CrowdTruth-core">CrowdTruth metrics</a> Python package (v &gt;= 2.0).</p> <p>The data aggregated with CrowdTruth metrics is available in folder <code>data/output/</code></p> <p>The raw crowdsourcing data is available in folder <code>data/input/</code></p> <p>If you find this data useful in your research, please consider citing:</p> <pre><code>@inproceedings{dumitrache2018frames, Author = {Anca Dumitrache and Lora Aroyo and Chris Welty}, Title = {Capturing Ambiguity in Crowdsourcing Frame Disambiguation}, Booktitle = {The sixth AAAI Conference on Human Computation and Crowdsourcing}, Year = {2018} } </code></pre>

opencc-by-sa-4.0Oct 2018View details →
zenodo48/100

MiRoR11 - P2 - Annotated corpus for semantic similarity of clinical trial outcomes

<p>Outcome similarity corpus</p> <p>This dataset contains annotations of semantic similarity for pairs of primary and reported outcomes.<br> Tab-separated format is used. The files contain the following columns:<br> filename, sentence pair ID, sentence pair text, primary outcome, primary outcome start position, primary outcome end position, reported outcome, reported outcome start position, reported outcome end position, label</p> <p>The folder out_relations_split contains the dataset splits for 10-fold cross-validation.</p>

opencc-by-4.0May 2019View details →
zenodo48/100

Uncovering the Semantics of Wikipedia Categories - Axioms and Assertions

<p>Resulting axioms and assertions from applying the Cat2Ax approach to the DBpedia knowledge graph.<br> The methodology is described in the conference publication &quot;N. Heist, H. Paulheim: Uncovering the Semantics of Wikipedia Categories, International Semantic Web Conference, 2019&quot;.</p>

openmit-licenseOct 2019View details →
zenodo48/100

Benchmark for the Evaluation of Lexical Semantic Change Detection for Ancient Greek

<p>This repository contains a benchmark of Ancient Greek lemmas which underwent semantic change. It is meant as a support for the evaluation of methods detecting lexical semantic change in Ancient Greek. It was created at the University of Groningen, The Netherlands.&nbsp;A publication will follow soon.</p> <p>&nbsp;</p> <p><strong>1. Overview of the repository</strong></p> <p>This benchmark was created by retrieving and selecting from existing scholarship cases of lexemes which underwent semantic change. The evaluation items are 44 Ancient Greek lemmas, accompanied by the following information (see the column headers in the CSV file):</p> <ul> <li><strong>reference:</strong> the literature source of information about the change;</li> <li><strong>which_change: </strong>an explanation of the change in meaning. NB: the older meaning(s) do not necessarily disappear after the change, but it can happen that the new meaning(s) are added to the existing one(s), increasing the polysemy of the lemma;</li> <li><strong>when_changed:&nbsp;</strong>information about the work(s) or time period in which the change was first recorded; this kind of information was not always available or precise;</li> <li><strong>christian_change:</strong> whether the change is triggered by social, religious, or cultural changes related to the spread of Christianity, according to the scholarship.</li> </ul> <p>&nbsp;</p> <p><strong>2. References</strong></p> <p>The literature used to build this benchmark is the following:</p> <p>&nbsp; &nbsp; BUCK, Carl Darling. A dictionary of selected synonyms in the principal Indo-European languages. University of Chicago Press, 1949.</p> <p>&nbsp; &nbsp; FINKELBERG, Aryeh. "On the History of the Greek &Kappa;&Omicron;&Sigma;&Mu;&Omicron;&Sigma;." Harvard Studies in Classical Philology (1998): 103-136.</p> <p>&nbsp; &nbsp; GINGRICH, F. Wilbur. "The Greek New Testament as a landmark in the course of semantic change." <em>Journal of Biblical Literature</em> (1954): 189-196.</p> <p>&nbsp; &nbsp; HORKY, Phillip Sidney. "When did Kosmos become the Kosmos." <em>Cosmos in the Ancient World</em> (2019): 22-41.</p> <p>&nbsp; &nbsp; LURAGHI, Silvia. "The verb ar&eacute;skein in Ancient Greek: Constructions and semantic change." <em>Acta Linguistica Petropolitana. Труды института лингвистических исследований</em> 18-1 (2022): 226-245.</p> <p>&nbsp;</p> <p>These dictionaries of Ancient Greek were also used to double-check the instances of change:</p> <p>&nbsp; &nbsp; LIDDELL, Henry George, and Robert Scott. <em>A Greek-English Lexicon</em>. revised and augmented throughout by. Sir Henry Stuart Jones. with the assistance of. Roderick McKenzie. Oxford. Clarendon Press. 1940.</p> <p>&nbsp; &nbsp; ROCCI, Lorenzo.<em> Vocabolario greco-italiano</em>. Roma. Societ&agrave; editrice Dante Alighieri. 1939.</p> <p>&nbsp; &nbsp; SLUITER, Ineke, and Lucien van Beek, and Ton Kessels, and Albert Rijksbaron. <em>Woordenboek Grieks/Nederlands</em>. 2024. <a href="https://woordenboekgrieks.nl/" target="_blank" rel="noopener">https://woordenboekgrieks.nl/</a></p> <p>&nbsp;</p> <p><strong>3. Acknowledgements</strong></p> <div>This work was partially supported by the Young Academy Groningen through the PhD scholarship of Silvia Stopponi.<br>&nbsp;<br>We acknowledge the financial support of Anchoring Innovation. Anchoring Innovation is the Gravitation Grant research agenda of the Dutch National Research School in Classical Studies, OIKOS. It is financially supported by the Dutch ministry of Education, Culture and Science (NWO project number 024.003.012). For more information about the research programme and its results, see the website&nbsp;<a href="https://www.anchoringinnovation.nl/">www.anchoringinnovation.nl</a>.</div> <div> <p>&nbsp;</p> <p><strong>4. How to cite</strong></p> </div> <div>Until there is no publication about this benchmark, please cite the resource as:</div> <div>Silvia Stopponi, Saskia Peels-Matthey, Malvina Nissim (2024), <em>Benchmark for the Evaluation of Lexical Semantic Change Detection Measures in Ancient Greek</em>, DOI: 10.5281/zenodo.13364555.</div> <div>&nbsp;</div> <div>&nbsp;</div>

opencc-by-4.0Aug 2024View details →
zenodo48/100

Dataset for Semantic Segmentation of Fishing Trajectories

<p>This is the dataset that was manually labelled by the author during his research work for the paper &quot;Semantic Segmentation of AIS Trajectories for Detecting Complete Fishing Activities&quot; in MDM 2022.</p>

opencc-by-4.0Nov 2022View details →
zenodo48/100

The POPREBEL semantic social network data

<p>The <a href="https://populism-europe.com/poprebel/">POPREBEL project</a> explores the phenomenon of populism in Europe. As part of it, a team of ethnographers generated and coded the corpus contained in this dataset. It consists of coded interviews, realized between spring 2021 and spring 2022, to Internet users in Czechia, Germany and Poland, who used social media to gather information about the COVID-19 pandemic. The dataset is pseudonymized. POPREBEL is supported by the European Union&#39;s Horizon 2020 programme, grant n. 822682.</p> <ul> <li><a href="https://zenodo.org/record/7494327">Final ethnographic report.</a> Section 1.2 contains a detailed description of how and why data were collected.</li> <li><a href="https://wellbeing.edgeryders.eu">Funnel website</a> of the project.</li> <li><a href="https://hal.archives-ouvertes.fr/hal-02478720/document">About semantic social networks</a>.</li> <li><a href="https://edgeryders.eu/t/long-term-ssna-data-storage-documentation-manual/12786">Data export and documentation process</a> (contains links to the code used to export the data)</li> </ul>

opencc-by-4.0Dec 2022View details →
zenodo48/100

The TREASURE semantic social network data on the circular economy aspect of automotive manufacturing

<p>The <a href="https://www.treasureproject.eu/">TREASURE</a> project looks at industrial innovation to address the problem of making onboard electronics in the automotive industry easier to recycle, increasing the industry&#39;s contribution to the circular economy. As part of it, a team of ethnographers generated and coded the corpus contained in this dataset. interviews conducted between January 2022 and June 2023 with car owners and enthusiasts at car industry events. The interviews focus on experiences with car electronics and perspectives on sustainability and the circular economy. The dataset is pseudonymized. TREASURE is supported by the European Union&#39;s Horizon 2020 programme, grant n. 101003587.</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Swedish Test Data for SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection

<p>This data collection contains the Swedish test data for <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection:</a></p> <p>- a Swedish text corpus pair (`corpus1/`, `corpus2/`)<br> - 31 lemmas which have been annotated for their lexical semantic change between the two corpora (`targets.txt`)<br> - the annotated binary change scores of the targets for subtask 1, and their annotated graded change scores for subtask 2 (`truth/`)</p> <p>We sample from the KubHist2 corpus, digitized by the National Library of Sweden, and available through the Spr&aring;kbanken corpus infrastructure Korp (<a href="https://www.researchgate.net/profile/Markus_Forsberg/publication/266352576_Korp_-_the_corpus_infrastructure_of_Sprakbanken/links/55bf1ee008aed621de121ba3/Korp-the-corpus-infrastructure-of-Sprakbanken.pdf">Borin et al., 2012</a>). The full corpus is available through a CC BY (attribution) license. Each word for which the lemmatizer in the Korp pipelien has found a lemma is replaced with the lemma. In cases where the lemmatizer cannot find a lemma, we leave the word as is (i.e., unlemmatized, no lower-casing). KubHist contains very frequent OCR errors, especially for the older data.More detail about the properties and quality of the Kubhist corpus can be found in (<a href="https://www.diva-portal.org/smash/get/diva2:1358014/FULLTEXT01.pdf#page=28">Adesam et al., 2019</a>).</p> <p>Lars Borin, Markus Forsberg, and Johan Roxendal. &quot;Korp-the corpus infrastructure of Spr&aring;kbanken.&quot; <em>LREC</em>. 2012.</p> <p>Adesam, Yvonne, Dana Dann&eacute;lls, and Nina Tahmasebi. &quot;Exploring the Quality of the Digital Historical Newspaper Archive KubHist.&quot; <em>DHN</em>. 2019.</p> <p>__Corpus 1__</p> <p>- based on: <a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1790-1830<br> - size: ~71 million tokens<br> - format: lemmatized, sentence length &gt; 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains frequent OCR errors</p> <p>__Corpus 2__</p> <p>- based on:&nbsp;<a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1895-1903<br> - size: ~111 million tokens<br> - format: lemmatized, sentence length &gt; 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains OCR errors</p> <p>Besides the official lemma version of the corpora for SemEval-2020 Task 1 we also provide the raw token version (`corpus1/token/`, `corpus2/token/`). It contains the raw sentences in the same order as in the lemma version. Find more information on the data and SemEval-2020 Task 1 in the paper referenced below.</p> <p>&nbsp;</p> <p>Reference:</p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi.<a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. To appear in SemEval@COLING2020.</p>

opencc-by-2.0Feb 2020View details →
zenodo44/100

MESINESP: Medical Semantic Indexing in Spanish - Development dataset

<p><em><strong>Please use the <a href="https://doi.org/10.5281/zenodo.4612274">MESINESP2 corpus (the second edition of the shared-task)</a> since it has a higher level of curation, quality and is organized by document type (scientific articles, patents and clinical trials).</strong></em></p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Introduction</strong></p> <p>The Mesinesp (Spanish BioASQ track, see https://temu.bsc.es/mesinesp) development set has a total of 750 records indexed manually by seven experienced medical literature indexers. Indexing is done using <em>DeCS codes, a sort of Spanish equivalent to MeSH terms</em>. Records were distributed in a way that each article was annotated, at least, by two different human indexers.</p> <p>The data annotation process consisted in two steps:</p> <ol> <li>Manual indexing step. DeCS codes were manually assigned to each record following the DeCS manual indexing guidelines.</li> <li>Manual validation and consensus. The joined set of manually indexed DeCS codes generated by both indexers were manually revised and corrections were done.</li> </ol> <p>These annotations were analyzed, resulting in an agreement using the Jaccard index.</p> <p>Records consisted basically in medical literature abstracts and titles from the IBECS and LILACS databases.</p> <p><strong>Zip structure</strong><br> The zip file contains two different development sets:</p> <ul> <li><em>Official development set</em>, which has the union of the annotations, with an agreement of macro = 0.6568 and micro = 0.6819. This set is composed by all the different (unique) DeCS codes that have been added by any annotator for each document; and</li> <li><em>Core-descriptors development set</em>, which has the intersection of the annotations, with an agreement of macro = 1.0 and micro = 1.0. This set is composed of the common DeCS codes that have been added by two or more annotators for each document.</li> </ul> <p><strong>Corpus format</strong></p> <p>Each dataset is a JSON object with one single key named &quot;articles&quot;, which contains a list of documents. So, the raw format of the file is one line per document plus two additional lines (the first and the last) to enclose that list of documents and the expected type of data is as follows:</p> <pre><code class="language-json">{"articles":[ {"abstractText":str,"db":str,"decsCodes":list,"id":str,"journal":str,"title":str,"year":int}, ... ]}</code></pre> <p>To clarify, the order of appearance of the fields in each document is as follows (note that this example it is pretty printed for readability purposes):</p> <pre><code class="language-json">{ "articles": [ { "abstractText": "Content of the abstract", "db": "Name of the source database", "decsCodes": [ "code1", "code2", "code3" ], "id": "Id of the document", "journal": "Name of the journal", "title": "Title of the document", "year": 2019 } ] }</code></pre> <p>Note: The fields &quot;db&quot;, &quot;journal&quot; and &quot;year&quot; might&nbsp;be null.</p> <p>Copyright (c) 2020 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

COVID-19@Semantics

<p>A semantic software framework in the context of the COVID-19 cases worldwide, named the COVID-19@Semantics, in which we have developed a specific ontology, COVID-Ont, an open and linked dataset and different services in the mentioned scope.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Semantic Annotation for Tabular Data with DBpedia: Adapted SemTab 2019 with DBpedia 2016-10

<p>Semantic Annotation for Tabular Data with DBpedia: Adapted SemTab 2019 with DBpedia 2016-10</p> <p>Github:&nbsp;https://github.com/phucty/mtab4dbpedia<br> ---------------------------------------------------------------------------------------------------------------------------------------</p> <p>CEA:&nbsp;</p> <ul> <li> <p>Keep only valid entities in DBpedia 2016-10</p> </li> <li> <p>Resolve percentage encoding</p> </li> <li> <p>Add missing redirect entities</p> </li> </ul> <p>CTA:&nbsp;</p> <ul> <li> <p>Keep only valid types</p> </li> <li> <p>Resolve transitive types (parents and equivalent types of the specific type) with DBpedia ontology 2016-10</p> </li> </ul> <p>CPA:</p> <ul> <li> <p>Add equivalent properties</p> </li> </ul> <p>Statistic of Adapted Tabular data SemTab 2019</p> <pre><code>| | CEA | | | CPA | | | CTA | | | |---------|:--------:|:-------:|:------:|:--------:|:-------:|:------:|:--------:|---------|--------| | | Orginal | Adapted | Change | Orginal | Adapted | Change | Orginal | Adapted | Change | | Round 1 | 8418 | 8406 | -0.14% | 116 | 116 | 0.00% | 120 | 120 | 0.00% | | Round 2 | 463796 | 457567 | -1.34% | 6762 | 6762 | 0.00% | 14780 | 14333 | -3.02% | | Round 3 | 406827 | 406820 | 0.00% | 7575 | 7575 | 0.00% | 5762 | 5673 | -1.54% | | Round 4 | 107352 | 107351 | 0.00% | 2747 | 2747 | 0.00% | 1732 | 1717 | -0.87% |</code></pre> <p>&nbsp;</p> <p>---------------------------------------------------------------------------------------------------------------------------------------<br> DBpedia 2016-10 extra resources: (Original dataset http://downloads.dbpedia.org/2016-10/)</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_classes_2016-10.csv</p> <p>Information: DBpedia classes and parents: (We remove the abstract types: Agent, Thing)</p> <p>Total: 759 classes</p> <p>Structure: [class, parents (separate with space)] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: &quot;City&quot;,&quot;Location Place PopulatedPlace Settlement&quot;</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_properties_2016-10.csv</p> <p>Information: DBpedia properties and these equivalents</p> <p>Total: 2865 properties</p> <p>Structure: [property, it&rsquo;s equivalent properties] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: &quot;restingDate&quot;,&quot;deathDate&quot;</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_domains_2016-10.csv</p> <p>Information: DBpedia properties and these domain types</p> <p>Total: 2421 properties (have types as their domain)</p> <p>Structure: [property, type (domain)] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: &quot;deathDate&quot;,&quot;Person&quot;</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_entities_2016-10.jsonl.bz2&nbsp;</p> <p>Information: DBpedia entity dump</p> <p>Format: json list bz2 (bz2 Compressed json list)</p> <p>Source: DBpedia dump 2016-10 core</p> <p>Total: 5,289,577 entities (No disambiguation entities)</p> <p>Structure:</p> <p>An entity: for example &ldquo;Tokyo&rdquo;: (datatype: dictionary),</p> <p>{</p> <p>&#39;wd&#39;: &#39;Q1322032&#39;, (Wikidata ID, datatype: string)</p> <p>&#39;wp&#39;: &#39;Tokyo&#39;, (Wikipedia ID, add prefix <a href="https://en.wikipedia.org/wiki/">https://en.wikipedia.org/wiki/</a> + wp to get the Wikipedia URL, datatype: string)</p> <p>&#39;dp&#39;: &#39;Tokyo&#39;, (DBpedia ID, add prefix <a href="http://dbpedia.org/resource/">http://dbpedia.org/resource/</a> + dp to get the DBpedia URL, datatype: string)</p> <p>&#39;label&#39;: &#39;Tokyo&#39;, (Entity label, datatype: string)</p> <p>&#39;aliases&#39;: [&#39;To-kyo&#39;, &#39;T&ocirc;ky&ocirc; Prefecture&#39;, ..], (Other entity names, datatype: list)&nbsp;</p> <p>&#39;aliases_multilingual&#39;: [&#39;东京小子&#39;, &#39;طوكيو&#39;, ...], (Other entity names in multilingual, datatype: list)</p> <p>&#39;types_specific&#39;: &#39;City&#39;, (Entity direct type, datatype: string)&nbsp;</p> <p>&#39;types_transitive&#39;: [&#39;Human settlement&#39;, &#39;City&#39;, &#39;PopulatedPlace&#39;, &#39;Location&#39;, &#39;Place&#39;, &#39;Settlement&#39;], (Entity transitive types, datatype: list)</p> <p>&#39;claims_entity&#39;: { (entity statements, datatype: dictionary. Keys: properties, Values: list of tail entities)</p> <p>&#39;governingBody&#39;: [&#39;Tokyo Metropolitan Government&#39;],&nbsp;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; &#39;subdivision&#39;: [&#39;Honshu&#39;, &#39;Kantō region&#39;],</p> <p>...</p> <p>},</p> <p>&#39;claims_literal&#39;: {</p> <p>&#39;string&#39;: { (String literal: datatype: dictionary. Keys: properties, Values: list of values</p> <p>&#39;postalCode&#39;: [&#39;JP-13&#39;],&nbsp;</p> <p>&#39;utcOffset&#39;: [&#39;+09:00&#39;, &#39;+9&#39;],</p> <p>&hellip;</p> <p>}</p> <p>&#39;time&#39;: { (Time literal: datatype: dictionary. Keys: properties, Values: list of date time</p> <p>&#39;populationAsOf&#39;: [&#39;2016-07-31&#39;],&nbsp;</p> <p>...</p> <p>}),&nbsp;</p> <p>&#39;quantity&#39;: { (Numerical literal: datatype: dictionary. Keys: properties, Values: list of values</p> <p>populationDesity: [6224.66, 6349.0],&nbsp;</p> <p>&#39;maximumElevation&#39;: [2017],&nbsp;</p> <p>...</p> <p>},</p> <p>&#39;pagerank&#39;: 2.2167366040153352e-06 (Entity page rank score calculated on DBpedia Graph)</p> <p>}</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>THIS DATA IS PROVIDED &quot;AS IS&quot;, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>

opencc-by-4.0Jun 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record