Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

241

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

241 results for “Transcriber”

Learn how ShareScore rates datasets ↗
zenodo48/100

Transcribing audio data: overview and transcripts of several automatic transcription tools

<p>Throughout institutions, audio recordings are being made regularly. To be able to further process these recordings, the audio often needs to be transcribed. In order to avoid having to transcribe the audio manually, there is a wealth of tools available for doing so automatically. In this record, we present an overview of several often-used tools to automatically transcribe pre-recorded audio data, including their features, costs, and security.</p> <p>To check the quality of the tool, we also recorded an audio fragment in Dutch that we ran through all tools in this overview in March of 2022. This original audio fragment (Test_interview_20220203.mp3), the cleaned-up transcription (Test_interview_cleaned_transcript.odt) and each tool&rsquo;s raw transcript of the audio fragment (Test_interview_[name-tool]_raw_[date-run]) are included in this record as well.&nbsp;The raw transcripts were downloaded as .docx or .txt files and the .docx files saved as .odt. No edits to the transcripts were made before saving them, except an incidental removal of a personal&nbsp;email address or hyperlink.</p> <p>The overview contains information and transcripts of following transcription tools:</p> <ul> <li>Amberscript</li> <li>HappyScribe</li> <li>Kaldi</li> <li>NVIVO transcription</li> <li>Sonix</li> <li>SpokenOnline</li> <li>Transcribe</li> <li>Trint</li> <li>Microsoft Word 365 Online</li> </ul> <p><strong>About</strong></p> <p>This overview was created through a collaboration between Utrecht University&rsquo;s Research Data Management (RDM) Support and the <a href="https://datahub.sites.uu.nl/">DataHub SSH</a> programme situated at the faculty of Humanities.</p> <p>The details in the overview have last been updated April 19, 2022. Please note that at the time you are downloading these files, the quality of the (Dutch) speech-to-text conversion may have been improved by the respective supplier.</p>

opencc-by-4.0Apr 2022View details →
zenodo48/100

ATEPP: A Dataset of Automatically Transcribed Expressive Piano Performance

<p>ATEPP is a dataset of expressive piano performances by virtuoso pianists. The dataset contains 11742&nbsp;11677 performances (~1000 hours) by 49 pianists and covers 1580 movements by 25 composers. All of the MIDI files in the dataset come from the piano transcription of existing audio recordings of piano performances. Scores in MusicXML format are also available for around half of the tracks. The dataset is organized and aligned by compositions and movements for comparative studies. For more details, please check <a href="https://github.com/BetsyTang/ATEPP">here</a>.&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo48/100

Decade-level Word2Vec models from automatically transcribed 19th-century newspapers digitised by the British Library (1800-1919)

<p>Word embeddings trained on a 4.2-billion-word corpus of 19th-century British newspapers using Word2Vec and the following parameters:</p> <pre><code>sg = True min_count = 5 window = 5 vector_size = 100 epochs = 5</code></pre> <p>The embeddings&nbsp;are divided into periods of ten years each. Unlike those in <a href="https://doi.org/10.5281/zenodo.7181682">this repository</a>, these were not aligned and OCR errors skimmed from the vocabulary.&nbsp;</p> <p>See related GitHub repository for the full documentation:&nbsp;<a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a></p> <p>Project website (Living with Machines):&nbsp;<a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p>

opencc-by-4.0May 2023View details →
zenodo44/100

Duhumbi Personal Narratives - Transcribed, parsed, glossed, translated text files

<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUK230512D1A] / CMT / The story of the former CM&rsquo;s death</li> <li>[CHUK230512C1A] / LHT / The history of Laphek village</li> <li>[CHUK230512B1] / THT / Hunting takin</li> <li>[CHUK260413A3A]/ ACK / Alcohol consumption</li> <li>[CHUK131014] / DTPK / Chasing the demons</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file &ldquo;Settings&rdquo;, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Duhumbi Religious Texts and Song - Transcribed, parsed, glossed, translated text files

<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUK110413A2A] / RELJ / Buddhist admonition</li> <li>&nbsp;[CHUK221212D2A] / JIK / Bonpo prediction text</li> <li>&nbsp;[CHUK260413A1] / MSK / Impromptu song</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file &ldquo;Settings&rdquo;, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Duhumbi Grammar - Sound Files, Toolbox and Transcriber File, PDFs of files

<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file &ldquo;Settings&rdquo;, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>This data set contains the files belonging to the sound files as mentioned in the pdf file &ldquo;Duhumbi Grammar All Files Upload 1&rdquo;. The S/N code corresponds to the code used in the Grammar to identify the text from which an example was taken. The name of the file refers to the name of the .wav, .trs, .txt and .pdf files in this upload. The subject is a short description of the topic of the text. The duration is the duration of the recording.</p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collector&nbsp;of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Duhumbi Procedural Texts - Transcribed, parsed, glossed, translated text files

<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUK210512I1] /&nbsp;PHPT / Hunting for porcupine&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</li> <li>[CHUK230512A1A] /&nbsp;SBDC / Making fermented soybean</li> <li>[CHUK220413A1] /&nbsp;CTTT / Catching frogs&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</li> <li>[CHUK220413B1] /&nbsp;CWTT / Collecting beeswax&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</li> <li>[CHUK220413C1] /&nbsp;CHTT / Collecting hornets</li> <li>[CHUK240413A1] /&nbsp;SNAP / Collecting stinging nettle&nbsp;</li> <li>[CHUK240413B1] /&nbsp;LCYT / Leather craft&nbsp; &nbsp; &nbsp;&nbsp;</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file &ldquo;Settings&rdquo;, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collector&nbsp;of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit properly and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Duhumbi Discussions - Transcribed, parsed, glossed, translated text files

<p>This dataset contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUKxxxx13A6] / LEL / Local elections (not included on speaker&#39;s request)</li> <li>[CHUK300412J2] / LGT / Planning a trip to Lagam</li> <li>[CHUK260413A2A] / NNK / Nicknames</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file &ldquo;Settings&rdquo;, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Duhumbi Grammar - Sound Files, Toolbox and Transcriber File, PDFs of files (Part 2)

<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file &ldquo;Settings&rdquo;, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>This data set contains the files belonging to the sound files as mentioned in the pdf file &ldquo;Duhumbi Grammar All Files Upload 2&rdquo;. The S/N code corresponds to the code used in the Grammar to identify the text from which an example was taken. The name of the file refers to the name of the .wav, .trs, .txt and .pdf files in this upload. The subject is a short description of the topic of the text. The duration is the duration of the recording.</p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collector&nbsp;of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Kallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal

<p>This data is transcribed speech data, in Wolof, Pulaar and Sereer.</p> <p>The recordings are about agriculture. The recorded consist of farmers, agricultural advisers, and agri-food business managers.&nbsp;Type of recordings comprise interactive radio programmes, focus groups, voice messages, push messages and interviews. Therefore, spontaneous speech is prevailing. Quality of audio may vary depending on the type of programme.</p> <p>Content description :</p> <ul> <li><strong>speech_dataset_wol.tar.gz:</strong> Wolof (ISO Code 639-2: wol) speech dataset contains 55 hours of transcribed speech, including almost 13 hours of validated content check by an expert. It also contains a XSAMPA lexicon (49,132 phonetised entries) and a text corpus (1,140,508 words).</li> <li><strong>speech_dataset_fuc.tar.gz:</strong> Pulaar (ISO Code 639-2: fuc) speech dataset contains nearly 32 hours of transcribed speech, including around 11 hours of validated content check by an expert. It also contains a text corpus (742,024 words).</li> <li><strong>speech_dataset_srr.tar.gz:</strong> Sereer (ISO Code 639-2: srr) speech dataset contains 38 hours of transcribed speech, including nearly 11 hours of validated content check by an expert.<br>In total, these resources provide 125 hours of transcribed speech in the 3 most widely spoken languages in Senegal, including 35 hours of checked transcriptions.</li> </ul> <p>This work is a result of the Kallaama project, funded by Lacuna Fund for 1 year, in 2023.&nbsp;</p> <p>See the <a title="Kallaama speech dataset" href="https://github.com/gauthelo/kallaama-speech-dataset" target="_blank" rel="noopener">GitHub repository</a> for more details about the dataset.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

An Automatically Transcribed Piano Performer Dataset

<p>Data-driven approaches for performers&#39; style analysis need large corpora of music performances to derive expressive performance parameters. One good reason for Deep Neural Networks (DNNs) not being used for performer identification is the lack of large-scale datasets with overlapping performances by different performers. Hence, to bridge this&nbsp;gap, we created a score-aligned&nbsp;automatically transcribed performer dataset.&nbsp;There are a total of 474 performances in the dataset, played by 6 pianists, spanning 35 movements by 2 composers&nbsp;for a total of 474 Western classical piano recordings in MIDI format.&nbsp;Every single MIDI file that&#39;s included in this collection was derived from a piano transcription of an already-existing audio recording of a piano performance. Score files in musicXML format and the alignment results are also included in the dataset.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Integrative modeling results of in-cell architecture of an actively transcribing-translating expressome

<p>Repository containing good-scoring models, input files, modeling and analysis protocols for the integrative modeling of the M. pneumoniae expressome from in-cell cryo-electron tomography and crosslinking mass spectrometry using IMP.</p>

openother-openMay 2020View details →
zenodo40/100

[Demo Input Data] for SCAFE: a software suite for analysis of transcribed cis-regulatory elements in single cells

<p>This archive (input.tar.gz) contains the demo data for&nbsp;SCAFE v1.0.0 (on <a href="https://doi.org/10.5281/zenodo.7023163">Zenodo</a> or <a href="https://github.com/chung-lab/SCAFE/releases/tag/v1.0.0">Github</a>)</p> <p><em>SCAFE</em>&nbsp;(Single Cell Analysis of Five-prime Ends) provides an end-to-end solution for processing of single cell 5&rsquo;end RNA-seq data. It takes a read alignment file (*.bam) from single-cell RNA-5&rsquo;end-sequencing (e.g. 10xGenomics Chromimum&reg;), precisely maps the cDNA 5&#39;ends (i.e. transcription start sites, TSS), filters for the artefacts and identifies genuine TSS clusters using logistic regression. Based on the TSS clusters, it defines transcribed cis-regulatory elements (tCRE) and annotated them to gene models. It then counts the UMI in tCRE in single cells and returns a tCRE UMI/cellbarcode matrix ready for downstream analyses, e.g. cell-type clustering, linking promoters to enhancers by co-activity&nbsp;<em>etc</em>.</p> <p>For details on installation, usage and test run on demo data,&nbsp;visit&nbsp;<a href="https://github.com/chung-lab/SCAFE">https://github.com/chung-lab/SCAFE</a></p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Fig. 3. Sequence variation among 4 internal transcribed spacer 2 in Population genetics of Oligonychus perseae (Acari: Tetranychidae) collected from avocados in Mexico and California

Fig. 3. Sequence variation among 4 internal transcribed spacer 2 (ITS2) genotypes identified from Oligonychus perseae populations in California,Mexico, and Costa Rica. Genotypes are named according to 3 genetic clusters identified from cytochrome oxidase subunit 1 (COI) haplotypes (see Fig. 2).

opencc-by-4.0Sep 2017View details →
zenodo40/100

Fig. 3 Species diagnostic internal transcribed spacer 2 in Anopheles (Anopheles) petragnani Del Vecchio 1939-a new mosquito species for Germany

Fig. 3 Species diagnostic internal transcribed spacer 2 (ITS2) fragments from all An. petragnani (367 bp) and some An. claviger s.s. (269 bp) individuals found during this study (M, Quantitas DNA Marker 100 bp–1 kb, Biozym; lanes 1–5, An. petragnani; lanes 6–10, An. claviger s.s.; − negative control)

opencc-by-4.0Mar 2016View details →
zenodo40/100

Fig. 2. Bayesian Inference tree constructed from Internal transcribed Spacer 2 in Ecological and geographical speciation in Lucilia bufonivora: The evolution of amphibian obligate parasitism

Fig. 2. Bayesian Inference tree constructed from Internal transcribed Spacer 2 (non-coding) sequence data. Each specimen is labelled with the species name and location abbreviation as indicated in Table 1. Green text corresponds to European samples of Lucilia bufonivora; red represents Lucilia elongata; purple represents Canadian L. bufonivora; orange represents Lucilia silvarum. Scale bar represents expected changes per site. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)

opencc-by-4.0Dec 2019View details →
zenodo40/100

A Compiled Archaeobiological Dataset for Central Asia's Chalcolithic through Bronze Age: macrobotanical and zooarchaeological data transcribed, standardized, and summarized from original publications

<p>This dataset contains archaeobotnaical and zooarchaeological data that have been compiled, transcribed, standardized, and summarized from original published data sources. Original data publications are given herein as a List of References (Microsoft Word file). These publications appeared between 1960-2022, presenting data in various formats, in various languages, and in scientific works that included journals, books, and conference proceedings - .</p> <p>Data have been compiled and are given in two Microsoft Excel files, one corresponding to archaeobotanical data and one to zooarchaeological data. On the first tab (worksheet) of each of these files, original publication sources are given in a summarized reference (Author, Year, Table/Figure Number) that corresponds to the full bibliographic reference in the accompanying List of References file. This first tab (worksheet) also summarizes additional information on the archaeological context of each dataset, collection and analysis methods (when reported), and the availabilty of other relevant and/or corresponding datasets. The remaining tabs (worksheets) in each file, organized alphabetically by author last name, tabularize the original data in a standardized format; these data are compiled and transcribed as necessary from the various formatting of original data sources, though they keep the original reported species names, table ordering, and numerical data. In the cases where totals were obviously erroneous or superceded by later or additional analyses, transcription notes have been offered directly in the worksheet.</p> <p>A fair number of these data sources are now out of print, have no digital distribution, and are otherwise difficult to access physically and/or linguistically. Accordingly, the sole aim in compiling these data here is to facilitate their widespread availablity, proper citation, and increased use within the international community of archaeological scientists working in Central Asia in the present day. Other scholars are encouraged to utilize these compiled datasets for foundational regional data and extended analyses, and to add to and expand these datasets going forward.</p>

opencc-by-4.0Aug 2024View details →
dryad40/100

Data from: Contrasting distributions and expression characteristics of transcribing repeats in Setaria viridis

Open the record for dataset details and reuse information.

publicJan 2025View details →
zenodo36/100

Duhumbi Stories - Transcribed, parsed, glossed, translated text files

<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publications:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <p>These Duhumbi stories&nbsp;can also be found in the &#39;Duhumbi Storybook&#39;&nbsp;(Monpasang Publications, ISBN 978-90-818610-1-4) published in autumn 2018. A separate Zenodo DOI also contains all the sound files (http://doi.org/10.5281/zenodo.1400495).</p> <p>The following list contains the sound file names, the shortcut code for the sentence names and the title of the story in Duhumbi, English and Hindi.</p> <ul> <li>[CHUK260413A4A] / OMAK / Dangpu budunbakaq tsawa / The origin of mankind / मानव जाती की उत्पत्ति together with&nbsp;[CHUK260413A4A] / MOD / Bengkhannaq dontha / The meaning of dreams / सपनों का मतलब</li> <li>[CHUK290412A8A] / CHLN / Duhum chakpaqkho tam &ndash; 1 / Settlement history of Duhum &ndash; 1/ दुहुम के निवास का इतिहास &ndash; 1</li> <li>[CHUK230512E1B] / CHT / Duhum chakpaqkho tam &ndash; 2 / Settlement history of Duhum &ndash; 2 / दुहुम मे निवास का इतिहास &ndash; 2</li> <li>[CHUK240314A1] / SPZP / Shawa Pema Zomba / शावा पेमा जोम्बा</li> <li>[CHUK230512F1A; CHUK230512F2A; CHUK230512F3A; CHUK230512F4A; CHUK230512F5A; CHUK230512F6A; CHUK230512F7A] and [CHUK230512G1A; CHUK230512G2A; CHUK230512G3A] / KDZ1 &ndash; KDZ10 / Khandro Drowa Zangmu / खांडरों द्रोवा जांग्मो</li> <li>[CHUK110614A1; CHUK110614B1; CHUK110614C1; CHUK110614D1; CHUK110614E1; CHUK110614F1; CHUK110614G1; CHUK110614H1; CHUK110614I1] / LGG1 &ndash; LGG9 / Ling Gesar Gepuwaq namthar / The life of King Ling Gesar / लिंग गेसर गेपु की जीवनी</li> <li>[CHUK110413A1A] / TNZY / Tshongpon Norbu Zangpo dangngaq waq uda / Tshongpon Norbu Zangpo and his son / त्शोंग्पोन नोरबू जांगपो और उसका बेटा</li> <li>[CHUK070115A1] / BMDC / Shadong dangngaq gomchen phawang / The macaque and the bat / बंदर और चमगादर</li> <li>[CHUK070115B] / BUDC / Pempelingngaq tam / The butterfly effect / तितली का असर</li> <li>[CHUK211015E1] / RSTT / Samtu dangngaq grongthang / The squirrel and the rat / गिलहरी और चूहा</li> <li>[CHUK130115E] / FBDC / Shalaqbaknyi men shikhennaq ama / The mother that fed the children with &lsquo;men&rsquo; / माँ जिसने बच्चों को &lsquo;मेन&rsquo; खिलाया</li> <li>[CHUK300115C1] / MTDC / Mani tam / मनी ताम</li> <li> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file &ldquo;Settings&rdquo;, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this dataset, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p> </li> </ul>

opencc-by-4.0Aug 2018View details →
zenodo36/100

The Cronica Pisana: Transcribing Together with FromThePage

<p>Video presentation and transcript of&nbsp;<em>The </em>Cronica Pisana: <em>Transcribing Together with FromThePage, </em>from &quot;Techniques and Tools for Teaching, Learning, and Researching Online: Manuscripts, Mapping, and Modeling,&quot; Medieval Academy of America Webinar Series, <em>Online Teaching for Medieval Studies: Philosophies, Learning Plans and Promising Tools</em>, July 21, 2020.</p>

opencc-by-4.0Oct 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record