Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

483

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

483 results for “SEMANTICS”

Learn how ShareScore rates datasets ↗
zenodo36/100

tBiodiv: Semantic Table Annotations Benchmark for Biodiversity Domain

<p><strong>tBiodiv </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has&nbsp;two types of tables.&nbsp;On the one hand, <strong>Horizontal Relational Tables</strong>&nbsp;are where&nbsp;each table&nbsp;represents a collection of entities. On the other&nbsp;hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p><strong>tBiodiv </strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using two levels of a recursive hierarchy of related concepts in Wikidata.</p> <p>We updated this repository&nbsp; with full verion of the dataset, we will update it again with the test ground truth (gt) data in the future.</p> <p>The supported tasks for semantic table annotations are:&nbsp;</p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> <li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li> </ol>

opencc-by-4.0Dec 2023View details →
zenodo36/100

A dataset for semantic segmentation of typical oceanic and atmospheric phenomena from Sentinel-1 images

<p>We have constructed a SAR (Synthetic Aperture Radar) image semantic segmentation dataset that includes 12 oceanic and atmospheric phenomena: Atmospheric Front (AF), Oceanic Front (OF), Rainfall (RF), Iceberg (IC), Sea Ice (SI), Pure Ocean Wave (POW), Wind Streak (WS), Low Wind Area (LWA), Biological Slick (BS), Micro Convective Cells (MCC), Internal Wave (IW), and Eddy.</p> <p>This dataset is built using Sentinel-1 IW and WV mode images. For WV mode data, we referenced TenGeoP-SARwv and SAR_WV_SemanticSegmentation and selected 2,383 images for semantic segmentation and annotation. For IW mode images, we incorporated some images from Tao et al.'s internal wave detection dataset. We selected 484 Sentinel-1 IW mode images obtained from 2015 to 2022 and divided them into 2,628 sub-images.</p> <p>The dataset contains a total of 5,011 image slices, with approximately 400 images for each phenomenon. All images are 16-bit .tiff files with a resolution of 100m and a size of 256x256 pixels. The images were manually annotated using the Labelme software, generating corresponding JSON files, which were then used to create the related annotation .png files.</p> <p>The updated version(V2) provides geographic information for each image.</p> <p>Thank you for your interest in our dataset. Here are the meanings of each label:</p> <p>1. BG: The unlabelled parts in JSON files are "BG" (Background)<br>2. AF: Atmospheric Front<br>3. BS: Biological Slick<br>4. I: &ldquo;I&rdquo; is equivalent to &ldquo;IB&rdquo;, representing icebergs<br>5. LWA: Low Wind Area<br>6. MCC: Micro Convective Cells<br>7. OF: Oceanic Front<br>8. POW: Pure Ocean Wave<br>9. RC: &ldquo;RC&rdquo; (Rain Cells) is equivalent to &ldquo;RF&rdquo; (Rainfall), both representing the&nbsp; rainfall phenomenon in the SAR image.&nbsp;<br>10. SI: Sea Ice<br>11. WS: Wind Streak<br>12. Eddy<br>13. IW: Internal Wave<br><em>14. HM: Represents the artificial objects appearing in the image, such as ships, aquaculture floating rafts, wind power facilities, etc.</em><br><em>15. OS: Unlike &ldquo;BS&rdquo;,&ldquo;OS&rdquo; represents mineral oil spills appearing in the SAR image (currently, there is insufficient data available for training, which will be supplemented in the future).</em></p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Semantic audio-visual congruence modulates visual sensitivity to biological motion across awareness levels

<p>This repository contains all experimental scripts, analysis and data files for the manuscript "Semantic audio-visual congruence modulates visual sensitivity to biological motion across awareness levels".</p>

opencc-by-4.0May 2024View details →
dryad36/100

Decrypting cryptic crosswords: Semantically complex wordplay puzzles as a target for NLP

<p>Cryptic crosswords, the dominant crossword variety in the UK, are a promising target for advancing NLP systems that seek to process semantically complex, highly compositional language. Cryptic clues read like fluent natural language but are adversarially composed of two parts: a definition and a wordplay cipher requiring character-level manipulations. Expert humans use creative intelligence to solve cryptics, flexibly combining linguistic, world, and domain knowledge. In this paper, we make two main contributions. First, we present a dataset of cryptic clues as a challenging new benchmark for NLP systems that seek to process compositional language in more creative, human-like ways. After showing that three non-neural approaches and T5, a state-of-the-art neural language model, do not achieve good performance, we make our second main contribution: a novel curriculum approach, in which the model is first fine-tuned on related tasks such as unscrambling words. We also introduce a challenging data split, examine the meta-linguistic capabilities of subword-tokenized models, and investigate model systematicity by perturbing the wordplay part of clues, showing that T5 exhibits behavior partially consistent with human solving strategies. Although our curricular approach considerably improves on the T5 baseline, our best-performing model still fails to generalize to the extent that humans can. Thus, cryptic crosswords remain an unsolved challenge for NLP systems and a potential source of future innovation.</p>

opencc-zeroNov 2021View details →
zenodo36/100

Evaluation Dataset: SemOI2 – Building Adaptive And Cost-Effective Recognition Applications With Semantic Augmentation

<p>Evaluation dataset for the paper &quot;SemOI2 &ndash; Building Adaptive And Cost-Effective Recognition Applications With Semantic Augmentation&quot;, presented at the AAAI-Make conference 2022, to be published in: <em>A. Martin, K. Hinkelmann, H.-G. Fill, A. Gerber, D. Lenat, R. Stolle, F. van Harmelen (Eds.), Proceedings of the AAAI 2022 Spring Symposium on Machine Learning and Knowledge Engineering for Hybrid Intelligence (AAAI-MAKE 2022), Stanford University, Palo Alto, California, USA, March 21&ndash;23, 2022.</em></p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

The Semantic PASCAL-Part Dataset

<p><strong>The Semantic PASCAL-Part dataset</strong></p> <p>The Semantic PASCAL-Part dataset is the RDF version of the famous PASCAL-Part dataset used for object detection in Computer Vision. Each image is annotated with <strong>bounding boxes</strong> containing a single object. Couples of bounding boxes are annotated with the part-whole relationship. For example, the bounding box of a car has the part-whole annotation with the bounding boxes of its wheels.</p> <p>This original release joins Computer Vision with Semantic Web as the objects in the dataset are aligned with concepts from:</p> <ul> <li>the provided supporting ontology;</li> <li>the <a href="https://wordnet.princeton.edu/">WordNet</a> database through its synstes;</li> <li>the <a href="https://yago-knowledge.org/">Yago</a> ontology.</li> </ul> <p>The provided Python 3 code (see the <a href="https://github.com/ivanDonadello/semantic-PASCAL-Part">GitHub repo</a>) is able to browse the dataset and convert it in RDF knowledge graph format. This new format easily allows the fostering of research in both Semantic Web and Machine Learning fields.</p> <p><strong>Structure of the semantic PASCAL-Part Dataset</strong></p> <p>This is the folder structure of the dataset:</p> <ul> <li><code>semanticPascalPart</code>: it contains the refined images and annotations (e.g., small specific parts are merged into bigger parts) of the PASCAL-Part dataset in Pascal-voc style. <ul> <li><code>Annotations_set</code>: the test set annotations in <code>.xml</code> format. For further information See the PASCAL VOC format <a href="http://host.robots.ox.ac.uk/pascal/VOC/index.html">here</a>.</li> <li><code>Annotations_trainval</code>: the train and validation set annotations in <code>.xml</code> format. For further information See the PASCAL VOC format <a href="http://host.robots.ox.ac.uk/pascal/VOC/index.html">here</a>.</li> <li><code>JPEGImages_test</code>: the test set images in <code>.jpg</code> format.</li> <li><code>JPEGImages_trainval</code>: the train and validation set images in <code>.jpg</code> format.</li> <li><code>test.txt</code>: the 2416 image filenames in the test set.</li> <li><code>trainval.txt</code>: the 7687 image filenames in the train and validation set.</li> </ul> </li> </ul> <p><strong>The PASCAL-Part Ontology</strong></p> <p>The PASCAL-Part OWL ontology formalizes, through logical axioms, the part-of relationship between whole objects (22 classes) and their parts (39 classes). The ontology contains 85 logical axiomns in Description Logic in (for example) the following form:</p> <pre><code>Every potted_plant has exactly 1 plant AND has exactly 1 pot </code></pre> <p>We provide two versions of the ontology: with and without cardinality constraints in order to allow users to experiment with or without them. The WordNet alignment is encoded in the ontology as annotations. We further provide the <code>WordNet_Yago_alignment.csv</code> file with both WordNet and Yago alignments.</p> <p>The ontology can be browsed with many Semantic Web tools such as:</p> <ul> <li><a href="https://protege.stanford.edu/">Prot&eacute;g&eacute;</a>: a graphical tool for ongology modelling;</li> <li><a href="http://owlapi.sourceforge.net/">OWLAPI</a>: Java API for manipulating OWL ontologies;</li> <li><a href="https://rdflib.readthedocs.io/en/stable/">rdflib</a>: Python API for working with the RDF format.</li> <li>RDF stores: databases for storing and semantically retrieve RDF triples. See <a href="https://www.w3.org/wiki/LargeTripleStores">here</a> for some examples.</li> </ul> <p><strong>Citing semantic PASCAL-Part</strong></p> <p>If you use semantic PASCAL-Part in your research, please use the following BibTeX entry</p> <pre><code>@article{DBLP:journals/ia/DonadelloS16, author = {Ivan Donadello and Luciano Serafini}, title = {Integration of numeric and symbolic information for semantic image interpretation}, journal = {Intelligenza Artificiale}, volume = {10}, number = {1}, pages = {33--47}, year = {2016} } </code></pre>

opencc-by-4.0Jan 2022View details →
zenodo36/100

Identified journal descriptor (JD), semantic type (ST), and MAUI keywords for PubMed/MEDLINE articles

<p>The Journal descriptor (JD) and Semantic type (ST) of PubMed articles were identified using a tool, called&nbsp;Journal descriptor indexing (JDI).</p> <p>MAUI keywords are identified by the MAUI tool and they can be used as a complement to MESH terms, as MeSH keywords are not always available in all PubMed articles.</p> <p>The methods of building the datasets can be found in&nbsp;the two articles: &quot;Author name disambiguation in MEDLINE based on journal descriptors and semantic types&quot; and &quot; Exploring author name disambiguation on PubMed-scale&quot;</p> <p>The PubMed database used is the 2019 baseline version, the number of articles in the two datasets are as follows:</p> <p>$ wc -l pubmed-paper-jd-st.tsv&nbsp;<br> 29796281 pubmed-paper-jd-st.tsv</p> <p>$ wc -l pubmed-paper-maui-keywords.tsv&nbsp;<br> 26509210 pubmed-paper-maui-keywords.tsv<br> &nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Tango Spacecraft Dataset for Region of Interest Estimation and Semantic Segmentation

<p><strong>Reference Paper:&nbsp;</strong></p> <p><a href="https://doi.org/10.1016/j.actaastro.2023.01.012"><strong>M. Bechini, M. Lavagna, P. Lunghi, Dataset generation and validation for spacecraft pose estimation via monocular images processing, Acta Astronautica 204 (2023) 358&ndash;369</strong></a></p> <p><a href="https://www.researchgate.net/publication/361924362_Spacecraft_Pose_Estimation_via_Monocular_Image_Processing_Dataset_Generation_and_Validation">M. Bechini, P. Lunghi, M. Lavagna. &quot;Spacecraft Pose Estimation via Monocular Image Processing: Dataset Generation and Validation&quot;. In 9th European Conference for Aeronautics and Aerospace Sciences (EUCASS)</a></p> <p><strong>General Description:</strong></p> <p>The &quot;<em>Tango Spacecraft Dataset for Region of Interest Estimation and Semantic Segmentation</em>&quot; dataset here published should be used for Region of Interest (ROI) and/or semantic segmentation tasks. It is split into 30002 train images and 3002 test images representing the Tango spacecraft from Prisma mission, being the largest publicly available dataset of synthetic space-borne noise-free images tailored to ROI extraction and Semantic Segmentation tasks (up to our knowledge). The label of each image gives, for the Bounding Box annotations, the filename of the image, the ROI top-left corner (minimum x, minimum y) in pixels, the ROI bottom-right corner (maximum x, maximum y) in pixels,&nbsp;and the center point of the ROI in pixels. The annotation are taken in image reference frame with the origin located at the top-left corner of the image, positive x rightward and positive y downward. Concerning the Semantic Segmentation, RGB masks are provided. Each RGB mask correspond to a single image in both train and test dataset. The RGB images are such that the R channel corresponds to the spacecraft, the G channel corresponds to the Earth (if present), and the B channel corresponds to the background (deep space). Per each channel the pixels have non-zero value only in correspondence of the object that they represent (Tango, Earth, Deep Space).&nbsp;More information on the dataset split and on the label format are reported below.&nbsp;</p> <p><strong>Images Information:</strong></p> <p>The dataset comprises 30002 synthetic grayscale images of Tango spacecraft from Prisma mission that serves as train set, while the test set is formed by 3002 synthetic grayscale images of Tango spacecraft from Prisma mission in PNG format.&nbsp;About 1/6 of the images both in the train and in the test set have a non-black background, obtained by rendering an Earth-like model in the raytracing process used to define the images reported.&nbsp;The images are noise-free to increase the flexibility of the dataset. The illumination direction of the spacecraft in the scene is uniformly distributed in the 3D space in agreement with the Sun position constraints.</p> <p><br> <strong>Labels Information:</strong></p> <p>Labels for the bounding box extraction are here provided in separated JSON files. The files are formatted per each image as in the following example:</p> <ul> <li>&nbsp; &nbsp; filename &nbsp; &nbsp;: tango_img_1 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;# name of the image to which the data are referred</li> <li>&nbsp;&nbsp; &nbsp;rol_tl&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; : [x, y] &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # ROI top-left corner (minimum x, minimum y) in pixels</li> <li>&nbsp; &nbsp;&nbsp;roi_br &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;: [x, y] &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;# ROI bottom-right corner (maximum x, maximum y) in pixels</li> <li>&nbsp; &nbsp; roi_cc &nbsp; &nbsp; &nbsp; &nbsp; : [x, y] &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; #&nbsp;center point of the ROI in pixels</li> </ul> <p>Notice that the annotation are taken in image reference frame with the origin located at the top-left corner of the image, positive x rightward and positive y downward.To make&nbsp;the usage of the dataset easier, both the training set and the test set are split in two folders containing the images with earth as background and without background.</p> <p>Concerning the Semantic Segmentation Labels, they are provided as RGB masks named as &quot;filename_mask.png&quot; where &quot;filename&quot; is the filename of the image of the training set or the test set to which a specific mask is referred.&nbsp;The RGB images are such that the R channel corresponds to the spacecraft, the G channel corresponds to the Earth (if present), and the B channel corresponds to the background (deep space). Per each channel the pixels have non-zero value only in correspondence of the object that they represent (Tango, Earth, Deep Space).&nbsp;</p> <p><strong>VERSION CONTROL</strong></p> <ul> <li>v1.0: This version contains&nbsp;the dataset (both train and test) of full scale images with ROI annotations and RGB masks for Semantic Segmentation tasks. These images have width=height=1024 pixels. The position of tango with respect to the camera is randomly selected from a uniform distribution, but it is ensured the full visibility in all the images.&nbsp;</li> </ul> <p>Note: this dataset contains the same images of the&nbsp;<em>&quot;Tango Spacecraft Wireframe Dataset Model for Line Segments Detection&quot;</em>&nbsp;v2.0 full-scale&nbsp;(DOI:&nbsp;<a href="https://doi.org/10.5281/zenodo.6372848">https://doi.org/10.5281/zenodo.6372848</a>) and also &quot;<em>Tango Spacecraft Dataset for Monocular Pose Estimation</em>&quot; v1.0 (DOI: <a href="https://doi.org/10.5281/zenodo.6499007">https://doi.org/10.5281/zenodo.6499007</a>)&nbsp;and they can be used&nbsp;together by combining the annotations of the relative pose and the ones of the reprojected wireframe model of Tango, with also the ones of the ROI. <strong>These three datasets give the most comprehensive dataset of space borne synthetic images ever published</strong> (up to our knowledge).</p>

opencc-by-nc-4.0Apr 2022View details →
zenodo36/100

GlossReader at LSCDiscovery: Train to Select a Proper Gloss in English -- Discover Lexical Semantic Change in Spanish

<pre>Precomputed vectors for the GlossReader system. LSCDiscovery Competition: https://codalab.lisn.upsaclay.fr/competitions/2243. </pre>

opencc-by-4.0May 2022View details →
zenodo36/100

Strawberry dataset for Semantic Segmentation

<p>The dataset was annotated using the labelme tool, and it was trained using the <a href="https://github.com/ayoolaolafenwa/PixelLib">pixellib</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Collaborative annotation and semantic enrichment of 3D media: Demo of a new FOSS toolchain

<p>A suite of tools for&nbsp;collaborative annotation and semantic enrichment of 3D cultural artefacts is being developed as part of the&nbsp;<a href="https://nfdi4culture.de/">NFDI4Culture</a>&nbsp;project across several partner organisations (led by the&nbsp;<a href="https://www.tib.eu/de/forschung-entwicklung/forschungsgruppen-und-labs/open-science">Open Science Lab at TIB, Hannover</a>). Operating within Task area 1: Data capture and enrichment, the proposed toolchain focuses on the annotation of 3D data within an open&nbsp;knowledge graph environment, so that 3D objects&rsquo; metadata&nbsp;and&nbsp;related annotations can be linked to various resources, part of the semantic web and all data is searchable via a public SPARQL endpoint.&nbsp;</p> <p>This short video presents the core steps in the media file upload and annotation workflow facilitated by the toolchain.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Semantically tagged Finnish Wikipedia 2017

<p><strong>Description of FI Wikipedia 2017 tagging</strong></p> <p><strong>Kimmo Kettunen</strong></p> <p><strong>University of Eastern Finl</strong><strong>and</strong></p> <p>The tagged data contains the texts of the Finnish Wikipedia of 2017. It has been first tagged syntactically in the Language Bank of Finland using the available UD2 tagger version of the Mylly service (https://mylly.rahtiapp.fi/home).</p> <p>Semantic tags to the UD2 parse have been added using a lexical semantic tagger FiST (Kettunen, 2019, <a href="https://aclanthology.org/W19-0306/">https://aclanthology.org/W19-0306/</a>).</p> <p>This published version has been condensed to a format where each analysed word contains the</p> <p>1. original running word form,</p> <p>2. lemma of the word form from UD2 parse,</p> <p>3. part-of-speech of the word from FiST</p> <p>4. semantic tag(s) for the word from FiST, and</p> <p>5. syntactic function of the word from UD2 parse.</p> <p>Semantic tags used are explained in this UCREL Semantic Analysis System (USAS) document: <a href="https://ucrel.lancs.ac.uk/usas/USASSemanticTagset.pdf">https://ucrel.lancs.ac.uk/usas/USASSemanticTagset.pdf</a></p> <p>Tagging includes all the semantic tags available for the word, as FiST does not perform disambiguation. Unknown words for the tagger are marked with tag Z99. Punctuation is tagged with PUNCT and numbers with NUMB. Lines beginning with # are output of UD2 and contain document, paragraph and sentence information.</p> <p>The output contains 6&nbsp;415&nbsp;027 sentences and 98.81 million lines. Lexical coverage of the semantic tagging is 76.59 %</p> <p><strong>Examples of output</strong></p> <p># newdoc</p> <p># newpar</p> <p># sent_id = 1</p> <p># text = Amsterdam</p> <p>Amsterdam#Amsterdam#Proper#Z2 root</p> <p># newpar</p> <p># sent_id = 2</p> <p># text = Amsterdam on Alankomaiden p&auml;&auml;kaupunki.</p> <p>Amsterdam#Amsterdam#Proper#Z2 nsubj:cop</p> <p>on#olla#Verb#A3+ A1.1.1 M6 Z5 cop</p> <p>Alankomaiden#Alankomaat#Proper#Z2 nmod:poss</p> <p>p&auml;&auml;kaupunki#p&auml;&auml;kaupunki#Noun#M7 root</p> <p>. PUNCT</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

OLIVES Dataset: Ophthalmic Labels for Investigating Visual Eye Semantics

<p>Clinical diagnosis of the eye is performed over multifarious data modalities including scalar clinical labels, vectorized biomarkers, two-dimensional fundus images, and three-dimensional Optical Coherence Tomography (OCT) scans. While the clinical labels, fundus images and OCT scans are instrumental measurements, the vectorized biomarkers are interpreted attributes from the other measurements. Clinical practitioners use all these data modalities for diagnosing and treating eye diseases like Diabetic Retinopathy (DR) or Diabetic Macular Edema (DME). Enabling usage of machine learning algorithms within the ophthalmic medical domain requires research into the relationships and interactions between these relevant data modalities. Existing datasets are limited in that: (i) they view the problem as disease prediction without assessing biomarkers, and (ii) they do not consider the explicit relationship among all four data modalities over the treatment period. In this paper, we introduce the Ophthalmic Labels for Investigating Visual Eye Semantics (OLIVES) dataset that addresses the above limitations. This is the first OCT and fundus dataset that includes clinical labels, biomarker labels, and time-series patient treatment information from associated clinical trials. The dataset consists of $1268$ fundus eye images each with 49&nbsp;OCT scans, and 16&nbsp;biomarkers, along with 3&nbsp;clinical labels and a disease diagnosis of DR or DME. In total, there are 96&nbsp;eyes&#39; data averaged over a period of at least two years with each eye treated for an average of 66&nbsp;weeks and 7&nbsp;injections. OLIVES dataset has advantages in other fields of machine learning research including self-supervised learning as it provides alternate augmentation schemes that are medically grounded.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Time-generalized multivariate analysis of EEG responses reveals a cascading architecture of semantic mismatch processing

<p>This entry includes the required data to conduct the analysis from&nbsp;our paper. It includes recordings, a look-up csv for target word onsets and montage file. Github page for analysis code:&nbsp;https://github.com/heikele/GAT_n4-p6</p> <p>&nbsp;</p> <p>Abstract from submitted paper:</p> <p>Event-related brain potentials have a strong impact on neurocognitive models, as they inform about the temporal sequence of cognitive processes. Nevertheless, their value for deciding among alternative cognitive architectures is partly limited by component overlap and the possibility of ambiguity regarding component identity. Here, we apply temporally-generalized multivariate pattern analysis &ndash; a recently-proposed machine learning method capable of tracking the evolution of neurocognitive processes over time &ndash; to constrain possible alternative architectures underlying the processing of semantic incongruency in sentences. In a spoken sentence paradigm, we replicate established N400/P600 correlates of semantic mismatch. Time-generalized decoding&nbsp;indicatesthat early vs. late mismatch-sensitive processes are (i) distinct in their neural substrate, arguing against recurrent or latency-shifted single process architectures, and (ii) partially overlapping in time,&nbsp;inconsistent withpredictions of strictly serial models. These results&nbsp;are in accordance withan incremental-cascading neurocognitive organization of semantic mismatch processing. We propose time-generalized multivariate decoding as a valuable tool for neurocognitive language studies.</p> <p>&nbsp;</p> <p>Keywords: EEG; ERP; semantic mismatch; N400; P600; multivariate pattern analysis; generalization across time decoding</p>

opencc-by-4.0Feb 2018View details →
zenodo36/100

SAS: Semantic Artist Similarity Dataset

<p>The Semantic Artist Similarity dataset consists of two datasets of artists entities with their corresponding biography texts, and the list of top-10 most similar artists within the datasets used as ground truth. The dataset is composed by a corpus of 268 artists and a slightly larger one of 2,336 artists, both gathered from Last.fm in March 2015. The former is mapped to the MIREX Audio and Music Similarity evaluation dataset, so that its similarity judgments can be used as ground truth. For the latter corpus we use the similarity between artists as provided by the Last.fm API. For every artist there is a list with the top-10 most related artists. In the MIREX dataset there are 188 artists with at least 10 similar artists, the other 80 artists have less than 10 similar artists. In the Last.fm API dataset all artists have a list of 10 similar artists.&nbsp;</p> <p>There are 4 files in the dataset.</p> <p><strong>mirex_gold_top10.txt</strong>&nbsp;and&nbsp;<strong>lastfmapi_gold_top10.txt</strong>&nbsp;have the top-10 lists of artists for every artist of both datasets. Artists are identified by MusicBrainz ID. The format of the file is one line per artist, with the artist mbid separated by a tab with the list of top-10 related artists identified by their mbid separated by spaces.</p> <p>artist_mbid \t artist_mbid_top10_list_separated_by_spaces \n</p> <p><strong>mb2uri_mirex</strong>&nbsp;and&nbsp;<strong>mb2uri_lastfmapi.txt</strong>&nbsp;have the list of artists. In each line there are three fields separated by tabs. First field is the MusicBrainz ID, second field is the last.fm name of the artist, and third field is the DBpedia uri.</p> <p>artist_mbid \t lastfm_name \t dbpedia_uri \n</p> <p>There are also 2 folders in the dataset with the biography texts of each dataset. Each .txt file in the biography folders is named with the MusicBrainz ID of the biographied artist. Biographies were gathered from the Last.fm wiki page of every artist.</p> <p><strong>Using this dataset</strong></p> <p>We would highly appreciate if scientific publications of works partly based on the Semantic Artist Similarity dataset quote the following publication:</p> <blockquote> <p>Oramas, S.,&nbsp;Sordo M.,&nbsp;Espinosa-Anke L., &amp;&nbsp;Serra X.&nbsp;(In Press).&nbsp;&nbsp;<a href="http://mtg.upf.edu/node/3316">A Semantic-based Approach for Artist Similarity</a>.&nbsp;16th International Society for Music Information Retrieval Conference.</p> </blockquote> <p>We are interested in knowing if you find our datasets useful! If you use our dataset please email us at <a href="mailto:mtg-info@upf.edu">mtg-info@upf.edu</a> and tell us about your research.</p> <p>&nbsp;</p> <p>https://www.upf.edu/web/mtg/semantic-similarity</p>

opencc-by-4.0Oct 2015View details →
zenodo36/100

DeConf: De-conflated Semantic Representations

<p>Sense representations generated using the algorithm introduced in:</p> <ul> <li>M. T. Pilehvar and N. Collier,&nbsp;<a href="http://www.pilevar.com/taher/pubs/Pilehvar_Collier_EMNLP2016.pdf">De-Conflated Semantic Representations</a>. EMNLP 2016, Austin, TX.</li> </ul>

opencc-by-4.0Aug 2018View details →
zenodo36/100

Masked datasets from an fMRI experiment on the impact of semantic priming on the perception of ambivalent (male versus female) faces

<p>Twenty-four female native Dutch speakers participated in the fMRI experiment and gained monetary compensation for their participation. Only female participants were recruited for the study, in order to avoid gender-related confounding factors. The study was approved by the local ethics committee (CMO Arnhem-Nijmegen, Radboud University Medical Center, ethical approval for studies on healthy human subjects at the Donders Centre for Cognitive Neuroimaging, no ECG 2012-0910-058) and conducted in accordance with their guidelines. All participants signed informed consent forms before the experiment. The data from seven subjects were excluded from the analysis: 3 subjects failed to finish the task and 4&nbsp;subjects exhibited head motion that exceeded the maximum acceptance rate of&nbsp;2 [mm]. The remaining 17 subjects (females, age 18-29&nbsp;years) reported no neurological diseases, and had normal or corrected-to-normal vision.&nbsp;</p> <p>A set of realistic 3D faces was morphed across gender (from extremely female to extremely male) using FaceGen Modeller 3.5 (Singular Inversions, www.facegen.com). The morphing procedure started from 40 distinct faces. For each face, we gradually modulated gender features in 5 steps with the same amount of feature transformation in each step. The face stimuli were presented frontally and cropped around the oval of the face. We controlled for luminance using SHINE toolbox for MATLAB. The perceptual boundary within gender continuum of faces was established in a separate behavioral experiment.</p> <p>Each trial started with priming: presentation of a gender-related word &#39;man&#39; or &#39;vrouw&#39; for 0.2 [s]. Then, after the fixation cross 0.25 [s]), a face was presented (0.5 [s]), followed by an inter-trial period of a randomized length of 5-7 [s]. Participants were asked to perform a matching task: respond &#39;yes&#39; if a word and subsequent picture corresponded in gender, and &#39;no&#39; otherwise. The experiment was carried out in Dutch. The buttons were counterbalanced across subjects. The experiment was divided into 6 blocks in order to avoid fatigue. Each block consisted of 50 trials. The order of stimuli was randomized across blocks and participants. We used Presentation software (version 17.1, www.neurobs.com) in order to screen the stimuli during the experiment.</p> <p>Functional images were acquired using 3T Skyra MRI system (Siemens Magnetom), T2* weighted echo-planar images (gradient-echo, repetition-time&nbsp;TR = 1760 [ms], echo-time&nbsp;TE = 32 [ms],&nbsp; 0.7 [ms] echo spacing, 1626 hz/Px bandwidth, generalized auto-calibrating partially parallel acquisition (GRAPPA), acceleration factor&nbsp;3, 32&nbsp;channel brain receiver coil). In total,&nbsp;78 axial slices were acquired (2.0 [mm] thickness,&nbsp;2.0*2.0 [mm] in plane resolution,&nbsp; 212 [mm] field of view (FOV) whole brain, anterior-to-posterior phase-encoding direction).</p> <p>The data reprocessing was performed using SPM12 (Welcome Trust Center for Neuroimaging, University College London, UK). Functional scans were realigned to the first scan of the first run with further realignment to the mean scan. We performed slice-time correction on realigned images to account for differences in image acquisition between slices. Motion-related components were removed from the data using a data-driven ICA-AROMA. Denoised functional scans were spatially normalized to the Montreal Neurological Institute (MNI) space without changing the voxel size. Normalized data were smoothed spatially with a Gaussian kernel of&nbsp;6 [mm] full-width at half-maximum.</p> <p>We extracted region-of-interest (ROI) mask using Anatomical Automatic Labeling atlas (AAL). According to our a priori hypothesis, we preselected the bilateral SPL (4288 voxels) and the bilateral IPL (3792 voxels).</p>

opencc-by-4.0Nov 2018View details →
zenodo36/100

Data, code, models for "Weakly Supervised Semantic Segmentation for Joint Key Local Structure Localization and Classification of Aurora Image"

<p>Data, code and models for https://ieeexplore.ieee.org/document/8410588/</p>

opencc-by-4.0Jul 2018View details →
zenodo36/100

SeSaMe: A Data Set of Semantically Similar Java Methods

<p>This is the data set presented in the paper</p> <p>Kamp, M., Kreutzer P., Philippsen M.: SeSaMe: A Data Set of Semantically<br> Similar Java Methods. 16th International Conference on Mining Software<br> Repositories (MSR 2019), Montreal, QC, Canada. 2019</p>

opencc-by-4.0Feb 2019View details →
zenodo36/100

Amazon Rainforest dataset for semantic segmentation

<p>This database contains images used for the semantic segmentation of forest and non-forest areas from a fully convolutional neural network U-Net .</p> <p>1. <strong>Training dataset: </strong>it contains 30 GeoTIFF images with 512x512 pixels and associated PNG masks (forest indicated in white and non-forest in black color).</p> <p>2. <strong>Validation dataset</strong>: it contains 15&nbsp;GeoTIFF images with 512x512 pixels and associated PNG masks used for U-Net validation step.</p> <p>3. <strong>Test dataset:&nbsp;</strong>it contains 15&nbsp;GeoTIFF images 512x512 pixels for testing.</p>

opencc-by-4.0May 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record