Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

32

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

32 results for “asr”

Learn how ShareScore rates datasets ↗
zenodo40/100

Allen Mouse CCF v3.1 ASR - Labels Modulo 65000

<p>If you use this dataset, please cite the original Allen Mouse CCF paper (<a href="https://doi.org/10.1016/j.cell.2020.04.007">https://doi.org/10.1016/j.cell.2020.04.007</a>):</p> <blockquote> <p>Quanxin Wang, Song-Lin Ding, Yang Li, Josh Royall, David Feng, Phil Lesnar, Nile Graddis, Maitham Naeemi, Benjamin Facer, Anh Ho, Tim Dolbeare, Brandon Blanchard, Nick Dee, Wayne Wakeman, Karla E. Hirokawa, Aaron Szafer, Susan M. Sunkin, Seung Wook Oh, Amy Bernard, John W. Phillips, Michael Hawrylycz, Christof Koch, Hongkui Zeng, Julie A. Harris, Lydia Ng,<br>The Allen Mouse Brain Common Coordinate Framework: A 3D Reference Atlas,<br>Cell,<br>Volume 181, Issue 4,<br>2020</p> </blockquote> <p>&nbsp;</p> <p>More information is available in</p> <p><a href="https://community.brain-map.org/t/allen-mouse-ccf-accessing-and-using-related-data-and-tools/359">https://community.brain-map.org/t/allen-mouse-ccf-accessing-and-using-related-data-and-tools/359</a></p> <p>This dataset is a 3D mutichannel image which is derived from the original data introduced by the Allen Mouse Common Coordinate Framework v3 (<a href="http://help.brain-map.org/download/attachments/2818171/MouseCCF.pdf">http://help.brain-map.org/download/attachments/2818171/MouseCCF.pdf</a>).</p> <p>The image file format is a xml/hdf5 file used by ImageJ/Fiji's bigdataviewer (<a href="https://docs.openmicroscopy.org/bio-formats/6.1.0/formats/big-data-viewer.html">https://docs.openmicroscopy.org/bio-formats/6.1.0/formats/big-data-viewer.html</a>).</p> <p>All channels are isotropically sampled with a voxel size of 10 micrometers.</p> <p>Channels:</p> <ul> <li>0: NISSL (<a href="http://download.alleninstitute.org/informatics-archive/current-release/mouse_ccf/ara_nissl/ara_nissl_10.nrrd">http://download.alleninstitute.org/informatics-archive/current-release/mouse_ccf/ara_nissl/ara_nissl_10.nrrd</a>)</li> <li>1: LABELS Border - Delimits the borders of the labels</li> <li>2: ARA</li> <li>3: LABELS Mod 65000 : Labels modulo 65000.</li> </ul> <p>The labels modulo 65000 channel is used because the xml/hdf5 file format handles only 16-bits images, while the reference atlas has (sparse) values above 65535. There is no overlap of labels after computing the label value mod 65000. Thus no loss of information in the mod labels channels.</p> <p>The related ontology from the allen brain (file 1.json) is also stored here.</p> <p>The only modification between this version and the previous one is to use an ASR convention for the positioning in 3D. It will match exactly the BrainGlobe allen_mouse atlases.</p> <p>The xml file which prevents the appearance of a fake extra timepoint in the dataset.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Voice of America: Ukrainian ASR Dataset of Broadcast Speech

<p>The dataset is based on public recordings of Voice of America (<a href="https://ukrainian.voanews.com">https://ukrainian.voanews.com</a>) extracted from their&nbsp;videos.</p> <p>The dataset contains&nbsp;398 hours of speech.</p> <p>The dataset is created by the ASR Corpus Creator (<a href="https://zenodo.org/record/7396705">https://zenodo.org/record/7396705</a>).</p> <p>The format of files: WAV with 16 kHz.</p> <p>&nbsp;</p>

openother-openDec 2022View details →
zenodo40/100

Untranscribed Radio Broadcast Audio Segments for Bemba ASR

<p>The audio collections segment for the radio broadcast speech dataset portion&nbsp;for Bemba on the Zambezi Voice dataset.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Griots Interviews, Bambara Language WAV, 30 hours, Recorded 2022, Cultural and ASR Training Resource

<p><strong>Source material to this project:</strong></p> <ul> <li> <p>Addition to 200,000 lines Bambara-French clean synchronized corpus</p> </li> <li> <p>Co-project with Google, recorded 30 hours video interviews with Griots</p> </li> <li> <p>30 hours manually transcribed and translated to French</p> </li> <li> <p>10 hours used in training ASR system and MT transformer</p> </li> <li> <p>100% Open Sourced</p> </li> <li> <p>Cultural/Technical Exhibition to be hosted online and in the National Museum of Mali</p> </li> <li> <p>Record, preserve, and share Malian culture with the world</p> </li> <li> <p>Contribute to the science of low-resource language NLP</p> </li> <li> <p>Reinforce the development of written Bambara</p> </li> <li> <p>Enable Bambara to reach status as a &ldquo;first-class internet language&rdquo;</p> </li> </ul> <p>Corresponding transcribed data can be found at the following&nbsp;<a href="https://github.com/robotsmali-ai/jeli-asr">Github repository</a></p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Air – Sea Rescue Craft (ASR-10)

This craft, known as ASR-10, is a rare example of a type once stationed in the North Sea and English Channel and played an important role during World War II. Its role was to provide emergency shelter for the crews of downed aircraft, and it contained vital equipment and supplies, including food, drinking water, bunks, towels, washing gear, books and playing cards. These comforts were more to reduce the shock of their ordeal than to prepare them for a long stay. Stranded men were able to radio for assistance ensuring that a fast rescue vessel would be sent out. ASR-10 was built by Carrier Engineering of Wembley in 1941. Its career after the war is not documented, but it may have ended its working life as a towed naval target on the Clyde. The Museum received ASR-10 after it lay derelict on the slipway at Battery Park, Gourock, for many years. It remains a representative of an ongoing challenge to save lives at sea and a reminder of the importance of the sea to Britain during wartime. Source: Objaverse 1.0 / Sketchfab

opencc-zeroMar 2020View details →
zenodo36/100

Waxholm Space atlas of the Sprague Dawley Rat Brain V4.2 ASR - BigDataViewer Xml/Hdf5 File Format

<p>This dataset is the <a href="https://www.nitrc.org/projects/whs-sd-atlas/">Waxholm Space atlas of the Sprague Dawley Rat Brain V4 atlas</a> converted to a Fiji BigDataViewer compatible file format.</p> <p>To cite this dataset, please follow the <a href="https://www.nitrc.org/citation/?group_id=1081">citation policy of the authors of this atlas.</a></p> <p>Channels:</p> <ul> <li>T2star_v1.01 masked by the label image</li> <li>T2star_v1.01 full data</li> <li>LABELS BORDER, a channel generated by delineating the original labels</li> <li>LABELS, the label image of the atlas</li> </ul> <p>ilf file : ontology structure</p> <p>Compared to the previous version, the first channel is the result of the second channel masked by the presence of labels (last channel). This allows to use more easily automated registration methods.</p> <p>The version 4.1 had an issue with the left/right definition (but not the 4.0)</p> <p>https://forum.image.sc/t/help-for-abba-in-fiji-rat-atlas-tags-atlas-regions-on-both-hemispheres-as-belonging-to-the-right-hemisphere/83020</p> <p>The only modification between this version and the previous one is to use an ASR convention for the positioning in 3D. It will match exactly the BrainGlobe atlases.</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Untranscribed Radio Broadcast Audio Segments for Silozi ASR

<p>Untranscribed audio collections of radio broadcast recordings segments in Silozi or Lozi from Zambia.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Untranscribed Radio Broadcast Audio Segments for Lunda ASR

<p>Untranscribed audio collections of radio broadcast recordings segments in Lunda or Chilunda from Zambia.</p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

ASR values of ferritic stainless steels

<p>Additional data to the article &quot;CORROSION BEHAVIOUR OF NITRIDED FERRITIC STAINLESS STEELS FOR USE IN SOLID OXIDE FUEL CELL DEVICES&quot; comparing the results of nitrided and non nitrided steel substrate.</p>

opencc-by-4.0Dec 2019View details →
zenodo32/100

Dataset for ASR for EE publication

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo32/100

PB Hindi ASR dataset

<p>Hindi ASR dataset -&nbsp; released along with the paper "TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR".&nbsp; If you find this useful, please cite as&nbsp;</p> <pre>@article{ravi2024teles, title={TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR}, author={Ravi, Nagarathna, T, Thishyan Raj and Arora, Vipul}, journal={arXiv preprint arXiv:2401.03251}, year={2024} }</pre>

opencc-by-4.0May 2024View details →
zenodo32/100

Untranscribed Radio Broadcast Audio Segments for Tonga ASR

<p>Untranscribed audio collections of radio broadcast recording segments for Tonga.&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Untranscribed Radio Broadcast Audio Segments for Nyanja ASR

<p>The audio collection of radio broadcast recordings segments for Nyanja.</p>

opencc-by-4.0Jan 2023View details →
ClinicalTrials.gov32/100

Massachusetts General Hospital Evaluation of DePuy ASR Hip System

ClinicalTrials.gov study NCT01611233. IPD Sharing: Not stated. Countries: 5. Publications: 9.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

DSing ASR task: Resources and Baseline for an unaccompanied singing ASR.

<p>DSing ASR task: Resources and Baseline for an unaccompanied singing ASR.</p> <p>In this repository, you will find the scripts used to construct the DSing ASR-oriented dataset and the baseline system constructed on Kaldi.</p> <p>Cite:</p> <pre><code>@inproceedings{Roa_Dabike-Barker_2019, author = {Roa Dabike, Gerardo and Barker, Jon} title = {{Automatic Lyric Transcription from Karaoke Vocal Tracks: Resources and a Baseline System}}, year = 2019, booktitle = {Proceedings of the 20th Annual Conference of the International Speech Communication Association (INTERSPEECH 2019)} } </code></pre> <p>1- DSing dataset</p> <p>DSing is an ASR-oriented dataset constructed from the&nbsp;<a href="https://ccrma.stanford.edu/damp/">Smule Sing!300x30x2</a>&nbsp;dataset (<strong>Sing!</strong>). This repository provides the scripts to transform&nbsp;<strong>Sing!</strong>&nbsp;to the DSing ASR task.</p> <p>2- Initial steps</p> <p>The first step before running any of the scripts is to obtain access to&nbsp;<strong>Sing!</strong>&nbsp;dataset. For more details, go to&nbsp;<a href="https://ccrma.stanford.edu/damp/">DAMP repository</a>.</p> <p>3- Transform Sing! to DSing dataset</p> <p>The scripts to transform the&nbsp;<strong>Sing!</strong>&nbsp;dataset to DSing ASR task dataset is located in the&nbsp;<strong>[DSing Construction](DSing Construction/)</strong>&nbsp;directory. The process is based on a series of python tools that are summarised in the runme_sing2dsing.sh bash script.</p> <ol> <li>Define the variable&nbsp;<em>version</em>&nbsp;with the name of the DSing version you want to construct (DSing1, DSing3 or DSing30). Any other option will raise an error.</li> <li>Set the variable&nbsp;<em>DSing_dest</em>&nbsp;with the path where the DSing version will be saved.</li> <li>Set the variable&nbsp;<em>SmuleSing_path</em>&nbsp;with the path to your copy of Smule Sing!300x30x2.</li> <li>Run code until step&nbsp;<strong>K</strong></li> <li>.....</li> </ol> <p>4- Extract DSing dataset using pre-segmented data.</p> <p>If you want to do some analysis in the segmentation results or to use DSing for different porpoise than ASR. In directory&nbsp;<strong>[DSing preconstructed](DSing preconstructed)</strong>&nbsp;you can find a small script that allows recovering the transcriptions and utterance wav files. Just need to to set the output directory and the path of your version of&nbsp;<strong>Sing!</strong></p>

openother-openMar 2020View details →
zenodo28/100

Graph Contrastive Learning with Adversarial Structure Refinement (GCL-ASR)

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo28/100

VIVOS: Vietnamese Speech Corpus for ASR

<p><strong>VIVOS Corpus</strong></p> <p>VIVOS is a free Vietnamese speech corpus consisting of 15 hours of recording speech prepared for Automatic Speech Recognition task.</p> <p>The corpus was published by AILAB, a computer science lab of VNUHCM - University of Science, with&nbsp;<strong>Prof. Vu Hai Quan</strong>&nbsp;is the head of.</p> <p>We publish this corpus in hope to attract more scientists to solve Vietnamese speech recognition problems. The corpus should only be used for academic purposes.</p> <p><strong>License</strong></p> <p>Creative Commons Attribution NonCommercial ShareAlike v4.0 (CC BY-NC-SA 4.0) (<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">details</a>)</p> <p><strong>Associated paper</strong></p> <p>Please cite this paper when using VIVOS corpus for research</p> <p>&quot;A non-expert Kaldi recipe for Vietnamese Speech Recognition System&quot;, Hieu-Thi Luong and Hai-Quan Vu, in Proc. <em>WLSI/OIAF4HLT2016</em> (<a href="https://aclanthology.org/W16-5207/">paper</a>)</p> <p><strong>Contact</strong></p> <p><a href="mailto:ailab@hcmus.edu.vn">ailab@hcmus.edu.vn</a></p> <p>&nbsp;</p> <p><strong>Corpus properties</strong></p> <p>Speech was recorded in a quiet environment with high quality microphone, speakers were asked to read one sentence at a time.</p> <table> <thead> <tr> <th scope="col">&nbsp;</th> <th scope="col">Training</th> <th scope="col">Testing</th> </tr> </thead> <tbody> <tr> <td>Speakers</td> <td>46</td> <td>19</td> </tr> <tr> <td>Male</td> <td>22</td> <td>12</td> </tr> <tr> <td>Female</td> <td>24</td> <td>7</td> </tr> <tr> <td>Utterances</td> <td>11660</td> <td>760</td> </tr> <tr> <td>Duration</td> <td>14:55</td> <td>00:45</td> </tr> <tr> <td>Unique Syllables</td> <td>4617</td> <td>1692</td> </tr> </tbody> </table> <p><br> <strong>Evaluations</strong></p> <p>The corpus was evaluated using our non-expert recipe for Vietnamese Speech Recognition system which is described&nbsp;in the associated paper.</p> <table> <thead> <tr> <th scope="col">&nbsp;</th> <th scope="col">baseline</th> <th scope="col">+pitch</th> <th scope="col">+tone</th> </tr> </thead> <tbody> <tr> <td>mGMM</td> <td>19.66</td> <td>15.14</td> <td>14.91</td> </tr> <tr> <td>mGMM+MMI</td> <td>18.08</td> <td>14.96</td> <td>13.91</td> </tr> <tr> <td>mGMM+SAT</td> <td>15.79</td> <td>12.07</td> <td>12.13</td> </tr> <tr> <td>mDNN+SAT</td> <td>13.34</td> <td>9.54</td> <td>9.48</td> </tr> </tbody> </table> <p><strong>Notice</strong></p> <p>This is the official replacement for http://ailab.hcmus.edu.vn/vivos/</p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Dec 2016View details →
ClinicalTrials.gov28/100

Multi-Center Comparative Trial of the ASR™-XL Acetabular Cup System vs. the Pinnacle™ Metal- on- Metal Total Hip System

ClinicalTrials.gov study NCT00561600. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov28/100

Prevention/Reduction of ASRs and PTSD to Sustain Civilian Performance With Sublingual Cyclobenzaprine HCl (TNX-102 SL)

ClinicalTrials.gov study NCT06636786. IPD Sharing: YES. Countries: 1. Publications: 0.

controlledIPD-YESFeb 2026View details →
geo24/100

Transcriptome analysis to uncover potential roles of a barley ASR gene in stress tolerance

GEO Series GSE123128. Oryza sativa. 3 samples. Type: Expression profiling by array.

openGEO-OpenNov 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record