Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
32
datasets available to search
ShareScore release 0.7.1
Dataset results
32 results for “asr”
Allen Mouse CCF v3.1 ASR - Labels Modulo 65000
<p>If you use this dataset, please cite the original Allen Mouse CCF paper (<a href="https://doi.org/10.1016/j.cell.2020.04.007">https://doi.org/10.1016/j.cell.2020.04.007</a>):</p> <blockquote> <p>Quanxin Wang, Song-Lin Ding, Yang Li, Josh Royall, David Feng, Phil Lesnar, Nile Graddis, Maitham Naeemi, Benjamin Facer, Anh Ho, Tim Dolbeare, Brandon Blanchard, Nick Dee, Wayne Wakeman, Karla E. Hirokawa, Aaron Szafer, Susan M. Sunkin, Seung Wook Oh, Amy Bernard, John W. Phillips, Michael Hawrylycz, Christof Koch, Hongkui Zeng, Julie A. Harris, Lydia Ng,<br>The Allen Mouse Brain Common Coordinate Framework: A 3D Reference Atlas,<br>Cell,<br>Volume 181, Issue 4,<br>2020</p> </blockquote> <p> </p> <p>More information is available in</p> <p><a href="https://community.brain-map.org/t/allen-mouse-ccf-accessing-and-using-related-data-and-tools/359">https://community.brain-map.org/t/allen-mouse-ccf-accessing-and-using-related-data-and-tools/359</a></p> <p>This dataset is a 3D mutichannel image which is derived from the original data introduced by the Allen Mouse Common Coordinate Framework v3 (<a href="http://help.brain-map.org/download/attachments/2818171/MouseCCF.pdf">http://help.brain-map.org/download/attachments/2818171/MouseCCF.pdf</a>).</p> <p>The image file format is a xml/hdf5 file used by ImageJ/Fiji's bigdataviewer (<a href="https://docs.openmicroscopy.org/bio-formats/6.1.0/formats/big-data-viewer.html">https://docs.openmicroscopy.org/bio-formats/6.1.0/formats/big-data-viewer.html</a>).</p> <p>All channels are isotropically sampled with a voxel size of 10 micrometers.</p> <p>Channels:</p> <ul> <li>0: NISSL (<a href="http://download.alleninstitute.org/informatics-archive/current-release/mouse_ccf/ara_nissl/ara_nissl_10.nrrd">http://download.alleninstitute.org/informatics-archive/current-release/mouse_ccf/ara_nissl/ara_nissl_10.nrrd</a>)</li> <li>1: LABELS Border - Delimits the borders of the labels</li> <li>2: ARA</li> <li>3: LABELS Mod 65000 : Labels modulo 65000.</li> </ul> <p>The labels modulo 65000 channel is used because the xml/hdf5 file format handles only 16-bits images, while the reference atlas has (sparse) values above 65535. There is no overlap of labels after computing the label value mod 65000. Thus no loss of information in the mod labels channels.</p> <p>The related ontology from the allen brain (file 1.json) is also stored here.</p> <p>The only modification between this version and the previous one is to use an ASR convention for the positioning in 3D. It will match exactly the BrainGlobe allen_mouse atlases.</p> <p>The xml file which prevents the appearance of a fake extra timepoint in the dataset.</p>
Voice of America: Ukrainian ASR Dataset of Broadcast Speech
<p>The dataset is based on public recordings of Voice of America (<a href="https://ukrainian.voanews.com">https://ukrainian.voanews.com</a>) extracted from their videos.</p> <p>The dataset contains 398 hours of speech.</p> <p>The dataset is created by the ASR Corpus Creator (<a href="https://zenodo.org/record/7396705">https://zenodo.org/record/7396705</a>).</p> <p>The format of files: WAV with 16 kHz.</p> <p> </p>
Untranscribed Radio Broadcast Audio Segments for Bemba ASR
<p>The audio collections segment for the radio broadcast speech dataset portion for Bemba on the Zambezi Voice dataset.</p>
Griots Interviews, Bambara Language WAV, 30 hours, Recorded 2022, Cultural and ASR Training Resource
<p><strong>Source material to this project:</strong></p> <ul> <li> <p>Addition to 200,000 lines Bambara-French clean synchronized corpus</p> </li> <li> <p>Co-project with Google, recorded 30 hours video interviews with Griots</p> </li> <li> <p>30 hours manually transcribed and translated to French</p> </li> <li> <p>10 hours used in training ASR system and MT transformer</p> </li> <li> <p>100% Open Sourced</p> </li> <li> <p>Cultural/Technical Exhibition to be hosted online and in the National Museum of Mali</p> </li> <li> <p>Record, preserve, and share Malian culture with the world</p> </li> <li> <p>Contribute to the science of low-resource language NLP</p> </li> <li> <p>Reinforce the development of written Bambara</p> </li> <li> <p>Enable Bambara to reach status as a “first-class internet language”</p> </li> </ul> <p>Corresponding transcribed data can be found at the following <a href="https://github.com/robotsmali-ai/jeli-asr">Github repository</a></p>
Air – Sea Rescue Craft (ASR-10)
This craft, known as ASR-10, is a rare example of a type once stationed in the North Sea and English Channel and played an important role during World War II. Its role was to provide emergency shelter for the crews of downed aircraft, and it contained vital equipment and supplies, including food, drinking water, bunks, towels, washing gear, books and playing cards. These comforts were more to reduce the shock of their ordeal than to prepare them for a long stay. Stranded men were able to radio for assistance ensuring that a fast rescue vessel would be sent out. ASR-10 was built by Carrier Engineering of Wembley in 1941. Its career after the war is not documented, but it may have ended its working life as a towed naval target on the Clyde. The Museum received ASR-10 after it lay derelict on the slipway at Battery Park, Gourock, for many years. It remains a representative of an ongoing challenge to save lives at sea and a reminder of the importance of the sea to Britain during wartime. Source: Objaverse 1.0 / Sketchfab
Waxholm Space atlas of the Sprague Dawley Rat Brain V4.2 ASR - BigDataViewer Xml/Hdf5 File Format
<p>This dataset is the <a href="https://www.nitrc.org/projects/whs-sd-atlas/">Waxholm Space atlas of the Sprague Dawley Rat Brain V4 atlas</a> converted to a Fiji BigDataViewer compatible file format.</p> <p>To cite this dataset, please follow the <a href="https://www.nitrc.org/citation/?group_id=1081">citation policy of the authors of this atlas.</a></p> <p>Channels:</p> <ul> <li>T2star_v1.01 masked by the label image</li> <li>T2star_v1.01 full data</li> <li>LABELS BORDER, a channel generated by delineating the original labels</li> <li>LABELS, the label image of the atlas</li> </ul> <p>ilf file : ontology structure</p> <p>Compared to the previous version, the first channel is the result of the second channel masked by the presence of labels (last channel). This allows to use more easily automated registration methods.</p> <p>The version 4.1 had an issue with the left/right definition (but not the 4.0)</p> <p>https://forum.image.sc/t/help-for-abba-in-fiji-rat-atlas-tags-atlas-regions-on-both-hemispheres-as-belonging-to-the-right-hemisphere/83020</p> <p>The only modification between this version and the previous one is to use an ASR convention for the positioning in 3D. It will match exactly the BrainGlobe atlases.</p>
Untranscribed Radio Broadcast Audio Segments for Silozi ASR
<p>Untranscribed audio collections of radio broadcast recordings segments in Silozi or Lozi from Zambia.</p>
Untranscribed Radio Broadcast Audio Segments for Lunda ASR
<p>Untranscribed audio collections of radio broadcast recordings segments in Lunda or Chilunda from Zambia.</p>
ASR values of ferritic stainless steels
<p>Additional data to the article "CORROSION BEHAVIOUR OF NITRIDED FERRITIC STAINLESS STEELS FOR USE IN SOLID OXIDE FUEL CELL DEVICES" comparing the results of nitrided and non nitrided steel substrate.</p>
Dataset for ASR for EE publication
Open the record for dataset details and reuse information.
PB Hindi ASR dataset
<p>Hindi ASR dataset - released along with the paper "TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR". If you find this useful, please cite as </p> <pre>@article{ravi2024teles, title={TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR}, author={Ravi, Nagarathna, T, Thishyan Raj and Arora, Vipul}, journal={arXiv preprint arXiv:2401.03251}, year={2024} }</pre>
Untranscribed Radio Broadcast Audio Segments for Tonga ASR
<p>Untranscribed audio collections of radio broadcast recording segments for Tonga. </p>
Untranscribed Radio Broadcast Audio Segments for Nyanja ASR
<p>The audio collection of radio broadcast recordings segments for Nyanja.</p>
Massachusetts General Hospital Evaluation of DePuy ASR Hip System
ClinicalTrials.gov study NCT01611233. IPD Sharing: Not stated. Countries: 5. Publications: 9.
DSing ASR task: Resources and Baseline for an unaccompanied singing ASR.
<p>DSing ASR task: Resources and Baseline for an unaccompanied singing ASR.</p> <p>In this repository, you will find the scripts used to construct the DSing ASR-oriented dataset and the baseline system constructed on Kaldi.</p> <p>Cite:</p> <pre><code>@inproceedings{Roa_Dabike-Barker_2019, author = {Roa Dabike, Gerardo and Barker, Jon} title = {{Automatic Lyric Transcription from Karaoke Vocal Tracks: Resources and a Baseline System}}, year = 2019, booktitle = {Proceedings of the 20th Annual Conference of the International Speech Communication Association (INTERSPEECH 2019)} } </code></pre> <p>1- DSing dataset</p> <p>DSing is an ASR-oriented dataset constructed from the <a href="https://ccrma.stanford.edu/damp/">Smule Sing!300x30x2</a> dataset (<strong>Sing!</strong>). This repository provides the scripts to transform <strong>Sing!</strong> to the DSing ASR task.</p> <p>2- Initial steps</p> <p>The first step before running any of the scripts is to obtain access to <strong>Sing!</strong> dataset. For more details, go to <a href="https://ccrma.stanford.edu/damp/">DAMP repository</a>.</p> <p>3- Transform Sing! to DSing dataset</p> <p>The scripts to transform the <strong>Sing!</strong> dataset to DSing ASR task dataset is located in the <strong>[DSing Construction](DSing Construction/)</strong> directory. The process is based on a series of python tools that are summarised in the runme_sing2dsing.sh bash script.</p> <ol> <li>Define the variable <em>version</em> with the name of the DSing version you want to construct (DSing1, DSing3 or DSing30). Any other option will raise an error.</li> <li>Set the variable <em>DSing_dest</em> with the path where the DSing version will be saved.</li> <li>Set the variable <em>SmuleSing_path</em> with the path to your copy of Smule Sing!300x30x2.</li> <li>Run code until step <strong>K</strong></li> <li>.....</li> </ol> <p>4- Extract DSing dataset using pre-segmented data.</p> <p>If you want to do some analysis in the segmentation results or to use DSing for different porpoise than ASR. In directory <strong>[DSing preconstructed](DSing preconstructed)</strong> you can find a small script that allows recovering the transcriptions and utterance wav files. Just need to to set the output directory and the path of your version of <strong>Sing!</strong></p>
Graph Contrastive Learning with Adversarial Structure Refinement (GCL-ASR)
Open the record for dataset details and reuse information.
VIVOS: Vietnamese Speech Corpus for ASR
<p><strong>VIVOS Corpus</strong></p> <p>VIVOS is a free Vietnamese speech corpus consisting of 15 hours of recording speech prepared for Automatic Speech Recognition task.</p> <p>The corpus was published by AILAB, a computer science lab of VNUHCM - University of Science, with <strong>Prof. Vu Hai Quan</strong> is the head of.</p> <p>We publish this corpus in hope to attract more scientists to solve Vietnamese speech recognition problems. The corpus should only be used for academic purposes.</p> <p><strong>License</strong></p> <p>Creative Commons Attribution NonCommercial ShareAlike v4.0 (CC BY-NC-SA 4.0) (<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">details</a>)</p> <p><strong>Associated paper</strong></p> <p>Please cite this paper when using VIVOS corpus for research</p> <p>"A non-expert Kaldi recipe for Vietnamese Speech Recognition System", Hieu-Thi Luong and Hai-Quan Vu, in Proc. <em>WLSI/OIAF4HLT2016</em> (<a href="https://aclanthology.org/W16-5207/">paper</a>)</p> <p><strong>Contact</strong></p> <p><a href="mailto:ailab@hcmus.edu.vn">ailab@hcmus.edu.vn</a></p> <p> </p> <p><strong>Corpus properties</strong></p> <p>Speech was recorded in a quiet environment with high quality microphone, speakers were asked to read one sentence at a time.</p> <table> <thead> <tr> <th scope="col"> </th> <th scope="col">Training</th> <th scope="col">Testing</th> </tr> </thead> <tbody> <tr> <td>Speakers</td> <td>46</td> <td>19</td> </tr> <tr> <td>Male</td> <td>22</td> <td>12</td> </tr> <tr> <td>Female</td> <td>24</td> <td>7</td> </tr> <tr> <td>Utterances</td> <td>11660</td> <td>760</td> </tr> <tr> <td>Duration</td> <td>14:55</td> <td>00:45</td> </tr> <tr> <td>Unique Syllables</td> <td>4617</td> <td>1692</td> </tr> </tbody> </table> <p><br> <strong>Evaluations</strong></p> <p>The corpus was evaluated using our non-expert recipe for Vietnamese Speech Recognition system which is described in the associated paper.</p> <table> <thead> <tr> <th scope="col"> </th> <th scope="col">baseline</th> <th scope="col">+pitch</th> <th scope="col">+tone</th> </tr> </thead> <tbody> <tr> <td>mGMM</td> <td>19.66</td> <td>15.14</td> <td>14.91</td> </tr> <tr> <td>mGMM+MMI</td> <td>18.08</td> <td>14.96</td> <td>13.91</td> </tr> <tr> <td>mGMM+SAT</td> <td>15.79</td> <td>12.07</td> <td>12.13</td> </tr> <tr> <td>mDNN+SAT</td> <td>13.34</td> <td>9.54</td> <td>9.48</td> </tr> </tbody> </table> <p><strong>Notice</strong></p> <p>This is the official replacement for http://ailab.hcmus.edu.vn/vivos/</p> <p> </p>
Multi-Center Comparative Trial of the ASR™-XL Acetabular Cup System vs. the Pinnacle™ Metal- on- Metal Total Hip System
ClinicalTrials.gov study NCT00561600. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Prevention/Reduction of ASRs and PTSD to Sustain Civilian Performance With Sublingual Cyclobenzaprine HCl (TNX-102 SL)
ClinicalTrials.gov study NCT06636786. IPD Sharing: YES. Countries: 1. Publications: 0.
Transcriptome analysis to uncover potential roles of a barley ASR gene in stress tolerance
GEO Series GSE123128. Oryza sativa. 3 samples. Type: Expression profiling by array.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.