Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,139

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

2,139 results for “recognition”

Learn how ShareScore rates datasets ↗
zenodo40/100

Dvoice : An open source dataset for Automatic Speech Recognition on Moroccan dialectal Arabic

<p>Dialectal Voice is a community project initiated by AIOX Labs to facilitate voice recognition by Intelligent Systems. Today, the need for AI systems capable of recognizing the human voice is increasingly expressed within communities. However, we note that for some languages such as Darija, there are not enough voice technology solutions. To meet this need, we then proposed to establish this program of iterative and interactive construction of a dialectal database open to all in order to help improve models of voice recognition and generation.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

TWristAR - wristband activity recognition

<p>TWristAR is a small three subject dataset recorded using an <a href="https://www.empatica.com/research/e4/">e4 wristband</a>. &nbsp; Each subject performed&nbsp;six scripted activities: upstairs/downstairs, walk/jog, and sit/stand.&nbsp; Each activity except stairs was performed&nbsp;for one minute a total of&nbsp;three times alternating between the pairs.&nbsp; Subjects 1 &amp; 2 also&nbsp;completed a walking sequence of approximately 10 minutes.&nbsp; The dataset contains motion (accelerometer) data, temperature, electrodermal activity, and heart rate data.&nbsp; &nbsp;The .csv file associated with each datafile contains timing and labeling information and was built using the provided Excel files.</p> <p>Each two activity session was recorded using a downward facing action camera.&nbsp; &nbsp;This video was used to generate the labels&nbsp;and is provided to investigate&nbsp;any data anomalies, especially for the free-form long walk.&nbsp; For privacy reasons only the sub1_stairs video contains audio.&nbsp;&nbsp;</p> <p>The Jupyter notebook&nbsp;processes the acceleration data and performs hold-one-subject out evaluation of a 1D-CNN.&nbsp; Example results from a run performed on a google colab GPU instance (w/o GPU the training time increases to about 90 seconds per pass):</p> <table align="center"> <caption>Hold-one-subject-out results</caption> <thead> <tr> <th scope="col">Train Sub</th> <th scope="col">Test Sub</th> <th scope="col">Accuracy</th> <th scope="col">Training Time (HH:MM:SS)</th> </tr> </thead> <tbody> <tr> <td>[1,2]</td> <td>[3]</td> <td>0.757</td> <td>00:00:12</td> </tr> <tr> <td>[2,3]</td> <td>[1]</td> <td>0.849</td> <td>00:00:14</td> </tr> <tr> <td>[1,3]</td> <td>[2]</td> <td>0.800</td> <td>00:00:11</td> </tr> </tbody> </table> <p>This notebook can also be <a href="https://colab.research.google.com/drive/1oD2Hwz0wmyWZqj07RHslUoRmjg_N7jXo?usp=sharing">run in colab&nbsp;here</a>.&nbsp;&nbsp;This video describes the processing&nbsp;&nbsp;<a href="https://mediaflo.txstate.edu/Watch/e4_data_processing">https://mediaflo.txstate.edu/Watch/e4_data_processing</a>.</p> <p>We hope you find this a useful dataset with end-to-end code.&nbsp; &nbsp;We have several papers in process and would appreciate your citation of the dataset if you use it in your work.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Raw data for Facemasks and face recognition: Potential impact on synaptic plasticity

<p>Figure 2 legend &nbsp;Upper panel. In control condition, visual sensory inputs from in- dividual&rsquo;s face are encoded by the face recognition system. At system level (a), this process implies functional and structural modifications in multiple brain regions, whereas at cellular level (b), this promotes the induction of distinct forms of synaptic plasticity, such as long-term potentiation and long-term depression (LTP, LTD, respectively). Lower panel. Wearing face masks consis- tently reduces the amount of information, by excluding the lower part of the face, including nose and mouth. Thus, both at system and cellular level, such mismatch impairs long-term functional and structural plasticity. In particular, at synaptic level, LTP induction will be favored, whereas LTD will be impaired. The black traces indicate the excitatory postsynaptic potentials in control condition; the red traces represent the long-term changes in synaptic efficacy after the induction protocol.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

(VIDEOS) Glycosylation as a key for enhancing drug recognition into spike glycoprotein of SARS-CoV-2.

<ul> <li>S1_movie_1. Movie of MD1 trajectory showing the Interaction of ligand TCMDC-124223 (roto-translate phenomenon) on RBD in absence of glycosylations within 50ns of simulation.</li> <li>S2_movie_2. Movie of MD4 trajectory showing the Interaction of ligand TCMDC-133766 (induced fit phenomenon) on the cryptic pocket of NTD in presence of glycosylations within 50ns of simulation.</li> <li>S3_movie_3. Movie of MD2 trajectory showing the Interaction of ligand TCMDC-124223 on RBD in presence of glycosylations within 50ns of simulation.</li> <li>S4_movie_4. Movie of MD3 trajectory showing the Interaction of ligand TCMDC-133766 on the cryptic pocket of NTD in absence of glycosylations within 50ns of simulation.</li> <li>S1_movie_1_v2. Movie of MD1 trajectory showing the Interaction of ligand TCMDC-124223 (roto-translate phenomenon) on RBD in absence of glycosylations within 300ns of simulation.</li> <li>S2_movie_2_v2. Movie of MD4 trajectory showing the Interaction of ligand TCMDC-133766 (induced fit phenomenon) on the cryptic pocket of NTD in presence of glycosylations within 300ns of simulation.</li> <li>S3_movie_3_v2. Movie of MD2 trajectory showing the Interaction of ligand TCMDC-124223 on RBD in presence of glycosylations within 300ns of simulation.</li> <li>S4_movie_4_v2. Movie of MD3 trajectory showing the Interaction of ligand TCMDC-133766 on the cryptic pocket of NTD in absence of glycosylations within 300ns of simulation.</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Matching Network of Ontologies: a Pattern Recognition Approach

<p>Networks of Ontologies research deals with the need to combine several ontologies at the same time. In a world of integrated systems (system of systems), isolated systems are increasingly rare in the near future, and their integration creates opportunities to change, validate information and add more value to an information system. This system of systems can contain ontologies to support the corresponding knowledge model. Consequently, new integration requirements may have to deal with network alignment rather than single ontologies. This work delves into the area of network alignment and proposes new ways to approach a particular case of alignment of large ontologies. The contribution of the work is the use of algebraic operations on networks to eliminate candidates before alignment and to use a stochastic search method to discover the relevant nodes. These nodes should be retained as they increase accuracy and final alignment retrieval even though they are identical and removed by the algebraic operation. To find out the particular relevance of each node, we propose a random walk combined with a frequent itemsets approach that overcomes the force brute approaches in processing time, as the size of networks grows, and have close precision. The approach was validated using networks of ontologies created from the OAEI ontologies. The approach selected the entities to send to the matcher without losing significant preexisting alignments. Finally, two different matchers were used to get metrics and compare the results with the pairwise force brute approach.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Multi-site environmental field recordings for audio event recognition in Upstate NY

<p>The first soundscape is located in a lightly wooded suburban area north of Albany, New York, at about 100 meters above sea level.&nbsp;The surrounding vegetation consists of white pines, spruces, oaks, and maples. &nbsp;The acoustic data was originally collected using a Zoom H2n Handy Recorder in 4-channel mode,&nbsp;yet only data from a single microphone capsule is presented.. Recordings were&nbsp;originally made at 16 bits in Waveform Audio File Format (WAV) at a frequency of 44,100 Hz.&nbsp;This dataset&nbsp;was collected from August 2019 to August 2020.</p> <p><br> The second soundscape&nbsp;is located in a forested area in Lake George, New York, also at about 100 meters above sea level. The surrounding vegetation consists largely of white pines.&nbsp;The microphone array consisted of two H2n Zoom recorders mounted perpendicularly to simulate four-channel recording conditions, yet only data from a single microphone capsule is presented. Recordings were&nbsp;originally made at 256 kbps in MPEG-2 Audio Layer III (MP3) format.&nbsp;This dataset was collected from&nbsp;February 2020 until February 2021.</p> <p>Once all of the audio data was collected, it was segmented into 8 second-long clips, to ensure that both biological and anthropogenic sound sources with longer call lengths were provided with sufficient temporal context, while still achieving reasonable computational efficiency. Each segment was then downsampled to 16 kHz for processing efficiency, allowing each to be represented as a 128,000-sample vector.&nbsp;</p> <p>Sounds</p> <p>ECMK- Eastern chipmunk &quot;chuck&quot;</p> <p>DGBR- dog bark</p> <p>DOWO- Downy woodpecker drum</p> <p>RNFL- rainfall</p> <p>ENGE- engine</p> <p>RCCR- remote control car</p> <p>BUBP- back-up beeper (truck)</p> <p>SREN- siren</p> <p>FFCR- fall field cricket</p> <p>DDCC- dog-day cicada</p> <p>ECMC- Eastern chipmunk &quot;chirp&quot;</p> <p>EGSK- Eastern gray squirrel &quot;kuk&quot;</p> <p>AMRO- American robin</p> <p>AMCR- American crow</p> <p>BLJA- Bluejay</p> <p>NOCA- Northern cardinal</p> <p>BCCH- Black-capped chickadee</p> <p>EAPH- Eastern phoebe</p> <p>BAWW- Black-and-white warbler</p> <p>WBNU- White-breasted nuthatch</p> <p>RBNU- Red-breasted nuthatch</p> <p>**Sounds from site 1 have no specific&nbsp;label tag while&nbsp;sounds from site 2 are specifically labeled with &quot;-site2&quot;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

SROADEX: Dataset for binary recognition and semantic segmentation of road surface areas from high resolution Aerial Orthoimages Covering Approximately 8,650 km2 of the Spanish Territory Tagged with Road Information

<p>The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography representing the axes of the different types of roads (urban, interurban and rural). This cartography has been obtained from different Spanish official sources (National Geographic Institute and autonomic cartographic agencies) that we have revised and edited in a meticulous and systematic way to verify that the roads are represented on the cartography according to the orthoimages, available on January 1, 2021 in the download center of the National Center of Geographic Information (CNIG), on 16 rectangular areas (28,5 km * 18,5 km) of the Spanish territory (insular and peninsular).</p> <p>The dataset consists of &nbsp;777599&nbsp;images in png format of 256x256 pixels, organized in folders for the different trainings, separating those corresponding to training, testing and validation.</p> <p>The structure of the data is as follows:<br> 1-Road-Ortho and 1-Road-Mask contain the images and ground true for training the semantic segmentation networks.<br> 1-Road-Ortho and 2-NoRoad-Ortho contain aerial images containing or not containing vials, for the training of binary tessellation networks identifying tessellations with vials.<br> Moreover, in each folder the structure is the same: train, test, validation containing 90%, 5% and 5% of the total images and masks of each type.</p> <p>1-Road-Ortho</p> <p>&nbsp;&nbsp;&nbsp; |----Train</p> <p>&nbsp;&nbsp;&nbsp; |----Test</p> <p>&nbsp;&nbsp;&nbsp; -----Validation</p> <p>1-Road-Mask</p> <p>&nbsp;&nbsp;&nbsp; |----Train</p> <p>&nbsp;&nbsp;&nbsp; |----Test</p> <p>&nbsp;&nbsp;&nbsp; -----Validation</p> <p>2-NoRoad-Ortho</p> <p>&nbsp;&nbsp;&nbsp; |----Train</p> <p>&nbsp;&nbsp;&nbsp; |----Test</p> <p>&nbsp;&nbsp;&nbsp; -----Validation</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems

<p>This is the dataset created for the paper, &quot;EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems&quot; (https://arxiv.org/abs/2109.04919).</p> <p>EmoWOZ is based on MultiWOZ, a multi-domain task-oriented dialogue dataset (https://github.com/budzianowski/multiwoz). It contains more than 11K task-oriented dialogues with more than 83K emotion annotations of user utterances. In addition to Wizard-of-Oz dialogues from MultiWOZ, we collect human-machine dialogues within the same set of domains to sufficiently cover the space of various emotions that can happen during the lifetime of a data-driven dialogue system. There are 7 emotion labels, which are adapted from the OCC emotion models.</p> <p>For data format and label definition, please refer to README.md.&nbsp;</p>

opencc-by-nc-4.0Jan 2022View details →
dryad40/100

Asymmetric song recognition does not influence gene flow in an emergent songbird hybrid zone

<p>Hybrid zones can be used to examine the mechanisms affecting reproductive isolation and speciation, like song. Song has equivocal support as a driver of speciation; we did not find song to cause reproductive isolation. We examined an emerging secondary contact zone between White-crowned Sparrow subspecies <em>pugetensis </em>and <em>gambelii </em>by measuring song variation, song recognition, plumage, morphology and mtDNA. Plumage and morphological characters provided evidence of hybridization in the contact zone, with some birds possessing plumage and song characteristics intermediate between the subspecies. Playback experiments revealed asymmetric song recognition: male <em>pugetensis </em>displayed greater response to their own song than <em>gambelii </em>song, whereas <em>gambelii </em>did not discriminate significantly. If female choice operates similarly to male song discrimination, we predicted asymmetric gene flow, resulting in a greater number of hybrids with <em>gambelii </em>mitochondrial DNA (mtDNA). Contrary to our prediction, more <em>gambelii </em>and putative hybrids in the contact zone possessed <em>pugetensis </em>mtDNA haplotypes, possibly due to greater <em>pugetensis </em>abundance and female-biased dispersal.</p>

opencc-zeroMay 2022View details →
zenodo40/100

POPP Datasets : Datasets for handwriting recognition from French population census

<p><strong>POPP datasets</strong></p> <p>This repository contains 3 datasets created within the POPP project (<a href="https://popp.hypotheses.org/#ancre2">Project for the Oceration of the Paris Population Census</a>) for the task of handwriting text recognition. These datasets have been published in <a href="https://hal.science/hal-03675614/"><em>Recognition and information extraction in historical handwritten tables: toward understanding early 20th century Paris census</em> at DAS 2022.</a></p> <p>The 3 datasets are called &ldquo;Generic dataset&rdquo;, &ldquo;Belleville&rdquo;, and &ldquo;Chauss&eacute;e d&rsquo;Antin&rdquo; and contains lines made from the extracted rows of census tables from 1926. Each table in the Paris census contains 30 rows, thus each page in these datasets corresponds to 30 lines.</p> <p>The structure of each dataset is the following:</p> <ul> <li>double-pages : images of the double pages</li> <li>pages: <ul> <li>images: images of the pages</li> <li>xml: METS and ALTO files of each page containing the coordinates of the bounding boxes of each line</li> </ul> </li> <li>lines: contains the labels in the file <code>labels.json</code> and the line images splitted into the folders <em>train</em>, <em>valid</em> and <em>test</em>. The double pages were scanned at a resolution of 200dpi and saved as PNG images with 256 gray levels. The line and page images are shared in the TIFF format, also with 256 gray levels.</li> </ul> <p>Since the lines are extracted from table rows, we defined 4 special characters to describe the structure of the text:</p> <ul> <li>&curren; : indicates an empty cell</li> <li>/ : indicates the separation into columns</li> <li>? : indicates that the content of the cell following this symbol is written above the regular baseline</li> <li>! : indicates that the content of the cell following this symbol is written below the regular baseline</li> </ul> <p>We provide a script <code>format_dataset.py</code> to define which special character you want to use in the ground-truth.</p> <p>The split for the <em>Generic Dataset</em> and <em>Belleville</em> have been made at the double-page level so that each writer only appears in one subset among train, evaluation and test. The following table summarizes the splits and the number of writers for each dataset:</p> <table> <thead> <tr> <th>Dataset</th> <th>train - # of lines</th> <th>validation - # of lines</th> <th>test - # of lines</th> <th># of writers</th> </tr> </thead> <tbody> <tr> <td>Generic</td> <td>3840 (128 pages)</td> <td>480 (16 pages)</td> <td>480 (16 pages)</td> <td>80</td> </tr> <tr> <td>Belleville</td> <td>1140 (38 pages)</td> <td>150 (5 pages)</td> <td>180 (6 pages)</td> <td>1</td> </tr> <tr> <td>Chauss&eacute;e d&rsquo;Antin</td> <td>625</td> <td>78</td> <td>77</td> <td>10</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Generic dataset (or POPP dataset)</strong></p> <ul> <li>This dataset is made 4800 annotated lines extracted from 80 double pages of the 1926 Paris census.</li> <li>There is one double page for each of the 80 districts of Paris</li> <li>There is one writer per double page so the dataset contains 80 different writers.</li> </ul> <p>&nbsp;</p> <p><strong>Belleville dataset</strong></p> <p>This dataset is a mono-writer dataset made of 1470 lines (49 pages) from the <em>Belleville</em> district census of 1926.</p> <p>&nbsp;</p> <p><strong>Chauss&eacute;e d&rsquo;Antin dataset</strong></p> <p>This dataset is a multi-writer dataset made of 780 lines (26 pages) from the <em>Chauss&eacute;e d&rsquo;Antin</em> district census of 1926 and written by 10 different writers.</p> <p>&nbsp;</p> <p><strong>Error reporting</strong></p> <p>It is possible that errors persist in the ground truth, so any suggestions for correction are welcome. To do so, please make a merge request on the <a href="https://github.com/Shulk97/POPP-datasets">Github repository</a> and include the correction in both the labels.json file and in the XML file concerned.</p> <p>&nbsp;</p> <p><strong>Citation Request</strong></p> <p>If you publish material based on this database, we request you to include a reference to paper <a href="http://link.springer.com/chapter/10.1007/978-3-031-06555-2_10"><code>T. Constum, N. Kempf, T. Paquet, P. Tranouez, C. Chatelain, S. Br&eacute;e, and F. Merveille,Recognition and information extraction in historical handwritten tables: toward understanding early 20th century Paris census ,Document Analysis Systems (DAS), pp. 143- 157, La Rochelle, 2022.</code></a></p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Animal Recognition Using Methods Of Fine-Grained Visual Analysis - YOLOv5 Object Detection Dataset (Oxford-IIIT Pet)

<p>Preprocessed dataset for Oxford-IIIT Pet in YOLOv5 format..&nbsp; Ground truth labels for head bounding boxes, body bounding boxes (derived from segmentation mask).</p>

openmit-licenseJul 2022View details →
zenodo40/100

Animal Recognition Using Methods Of Fine-Grained Visual Analysis - YOLOv5 Breed Classification Dataset (Oxford-IIIT Pet)

<p>Oxford-IIIT Pet Dataset with ground truth labels for breeds&nbsp;(from https://public.roboflow.com/object-detection/oxford-pets).</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Animal Recognition Using Methods Of Fine-Grained Visual Analysis - Kashtanka Pets (All Dev and Test Images, Single Folder)

<p>Kashtanka Pets images, with all Dev and Test images (total 66639 images).&nbsp; In a single folder, with filenames indicating path of file in original dataset distribution.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Animal Recognition Using Methods Of Fine-Grained Visual Analysis - YOLOv5 Object Detection Dataset (Tsinghua Dogs)

<p>Preprocessed dataset for Tsinghua Dogs&nbsp;in YOLOv5 format.. &nbsp;Ground truth labels for head bounding boxes, body bounding boxes</p>

openmit-licenseJul 2022View details →
zenodo40/100

BembaSpeech: A Speech Recognition Corpus for the Bemba Language

<p>We present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting over 24 hours of read speech in the Bemba language, a written but low-resourced language spoken by over 30% of the population in Zambia. To assess its usefulness for training and testing ASR systems for Bemba, we explored different approaches; supervised pre-training (training from scratch), cross-lingual transfer learning from a monolingual English pre-trained model using DeepSpeech on the portion of the dataset and fine-tuning large scale self-supervised Wav2Vec2.0 based multilingual pre-trained models on the complete BembaSpeech corpus. From our experiments, the 1 billion XLS-R parameter model gives the best results. The model achieves a word error rate (WER) of 32.91%, results demonstrating that model capacity significantly improves performance and that multilingual pre-trained models transfers cross-lingual acoustic representation better than monolingual pre-trained English model on the BembaSpeech for the Bemba ASR. Lastly, results also show that the corpus can be used for building ASR systems for Bemba language</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words

<p>Hi,KIA dataset is a shared short Wakeup Word&nbsp;database focusing on perceived emotion in&nbsp;speech The dataset contains&nbsp;<strong>488 </strong>Wakeup Word&nbsp;speech.&nbsp;</p> <p>For more detailed information about the dataset, please refer to our paper:&nbsp;Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words</p> <p><strong>File Description</strong></p> <ul> <li><em><strong>wav/</strong></em>:&nbsp;wav files. <ul> <li>Filename f`{gender}_{pid}_{scene}_{trial}_{emotion}.wav`&nbsp;The first letter was used to express emotion.<br> &nbsp;</li> </ul> </li> <li><em><strong>annotation/</strong></em>:&nbsp;Information related to annotation and human validation of the entire speech</li> <li> <p><em><strong>split</strong></em>: 8fold data split with {train, valid, test}.csv&nbsp;</p> </li> <li> <p><em><strong>handcraft:</strong></em>&nbsp;Features used for data EDA and baseline performance</p> </li> <li> <p><em><strong>best_weights:</strong></em>&nbsp;wav2vec2.0 context network finetuning weights for re-implementation. Due to file size, we attach only fold M1, F5</p> </li> </ul> <p>&nbsp;</p> <p><strong>Reference</strong></p> <ul> </ul> <p>Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words [[ArXiv](https://arxiv.org/abs/2211.03371)]</p> <p>```<br> @inproceedings{kim2022hi,<br> &nbsp; title={Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words},<br> &nbsp; author={Taesu Kim, SeungHeon Doh, Gyunpyo Lee, Hyung seok Jun, Juhan Nam, Hyeon-Jeong Suk},<br> &nbsp; booktitle={Proceedings of the 14th Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)},<br> &nbsp; year={2022}<br> }<br> ```</p>

opencc-by-4.0Aug 2022View details →
dryad40/100

Data from: Domain-specific neural networks improve automated bird sound recognition already with small amount of local data

<p><span><span>An automatic bird sound recognition system is a useful tool for collecting data of different bird species for ecological analysis. Together with autonomous recording units (ARUs), such a system provides a possibility to collect bird observations on a scale that no human observer could ever match. During the last decades progress has been made in the field of automatic bird sound recognition, but recognizing bird species from untargeted soundscape recordings remains a challenge. <br></span></span></p> <p><span><span>In this article we demonstrate the workflow for building a global identification model and adjusting it to perform well on the data of autonomous recorders from a specific region. We show how data augmentation and a combination of global and local data can be used to train a convolutional neural network to classify vocalizations of 101 bird species. We construct a model and train it with a global data set to obtain a base model. The base model is then fine-tuned with local data from Southern Finland in order to adapt it to the sound environment of a specific location and tested with two data sets: one originating from the same Southern Finnish region and another originating from a different region in German Alps.<br></span></span></p> <p><span><span>Our results suggest that fine-tuning with local data significantly improves the network performance. Classification accuracy was improved for test recordings from the same area as the local training data (Southern Finland) but not for recordings from a different region (German Alps). Data augmentation enables training with a limited number of training data and even with few local data samples significant improvement over the base model can be achieved. Our model outperforms the current state-of-the-art tool for automatic bird sound classification.<br></span></span></p> <p><span><span>Using local data to adjust the recognition model for the target domain leads to improvement over general non-tailored solutions. The process introduced in this article can be applied to build a fine-tuned bird sound classification model for a specific environment.</span></span></p>

opencc-zeroSep 2022View details →
zenodo40/100

Computationally profiling peptide:MHC recognition by T-cell receptors and T-cell receptor-mimetic antibodies

<p>Supporting datasets for preprint version of &quot;Computationally profiling peptide:MHC recognition by T-cell receptors and T-cell receptor-mimetic antibodies&quot;.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Figure 9. Classification accuracy regardless the ethnic group (Total accuracy 75%)-Impact of Ethnic Group on Human Emotion Recognition Using Backpropagation Neural Network

<p>Our experiments show that, the impact of ethnic group on the accuracy of emotions<br> recognition is a positive where the accuracy of emotion recognition considering ethnic group is<br> 83.3% as shown in Figure 8, and we got 75% of accuracy regardless ethnic group as shown in<br> Figure 9.</p>

opencc-by-4.0Nov 2011View details →
zenodo40/100

Figure 8. Classification accuracy of emotions considering the ethnic group-Impact of Ethnic Group on Human Emotion Recognition Using Backpropagation Neural Network

<p>To study the accuracy of emotion recognition for our approach regardless the ethnic group<br> we used 108 images for the training representing six emotions of six persons. For testing, we used<br> 36 images representing six emotions of six persons.<br> On the other hand, to study the accuracy of emotion recognition for our approach<br> considering the ethnic group we used 36 images for training for each ethnic group representing six<br> emotions of six persons, and test the classifier by using 12 images representing six emotions of six<br> persons.</p>

opencc-by-4.0Nov 2011View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record