Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
29
datasets available to search
ShareScore release 0.7.1
Dataset results
29 results for “part of speech”
Corpus of Occitan Written Traditional Folktales Annotated with Part-Of-Speech (OWT-Tag)
<p>This resource contains 5 extracts of texts in Occitan which were manually annotated with lemmas and parts-of-speech, following the Grace standard. It was produced during the ExpressioNarration project, funded by a Marie Curie Individual Fellowship, in order to evaluate the performance of an Occitan Part-Of-Speech tagger, Talismane, to the specifities of the corpus of the project called Oral Occitan (OcOr), also available on https://zenodo.org/record/1451753#.W78FJWOYSpo.<br> Each extract contains around 1500 words. They are extracted from 'Contes et proverbes populaires recueillis en armagnac et Contes populaires recueillis en agenais' de J.-F. Bladé, 'Coundes biarnés, couéilhuts aüs parsàas miéytadès dou péys dé Biarn' de J.-V. Lalanne, 'Contes populaires du Languedoc' de L. Lambert and 'Contes populaires recueillis dans la Grande-Lande' de F. Arnaudin.<br> The annotation process is described in the following article available on https://www.openscience.fr/IMG/pdf/iste_modocv1n1_2.pdf.</p>
A challenging data set for evaluating part-of-speech taggers
<p>This data set contains 2,227 sentences, with a part-of-speech (POS) tag specified for a single word in the sentence. The data file is a tab-separated text file where each row (after the header row) is formatted as follows:</p><p><i>sentence <TAB> POS tag <TAB> (optional) motivation</i></p><p>Note that, in the sentence (= a string of space-separated characters), the POS-tagged word is indicated by the POS tag in brackets, placed just after the word to which it refers. Example:</p><p><i>The road bends [VERB] to the right . VERB</i></p><p>In this example, the optional motivation is not included, as the tagged word can easily be identified as being a verb.</p>
A part-of-speech (POS) tagged corpus of Classical Tibetan
<p>This part-of-speech (POS) tagged corpus of Classical Tibetan was prepared in the course of the research project 'Tibetan in Digital Communication' (2012-2015) hosted at SOAS, University of London and funded by the UK's Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. This corpus includes the <em>Mdzaṅs blun</em> (9th century, canonical), the <em>Bu ston chos ḥbyuṅ</em> (13th century, ecclesiastical history), the <em>Mi la ras paḥi rnam thar</em> and <em>Mar paḥi rnam thar</em> (15th century, biography).</p>
A rule based Tibetan part-of-speech (POS) tagger for the creation of gold standard training data
<p>This rule based Tibetan part-of-speech (POS) tagger was prepared in the course of the research project 'Tibetan in Digital Communication' (2012-2015) hosted at SOAS, University of London and funded by the UK's Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. For a description of the tagger itself see Garrett et al. 2014. Note that the tagger must be used together with a lexicon (for example Hill & Garrett 2017a). One must use one's own script to tag all words with all tags in the lexicon and then apply the tagger to remove incorrect tags.</p> <p>On the associated corpus of 318,230 words (Hill & Garrett 2017b) the lexical tagger (i.e. simply applying all available tags to all words) tags 141,911 words with the correct unique tag, achieves as accuracy of 1.000 (by definition getting the right tag among others for each word) with an ambiguity of 2.73111. In contrast, the Rule Tagger tags 241,256 words with the correct unique tag, achieves an accuracy of 0.99893 and an ambiguity of 1.38577.</p> <p>Because this tagger does not achieve ambiguity 1.000 it is not suitable for tagging large scale corpora, but instead is useful for the creation of gold standard training data.</p> <p>N.B. In some rare cases the tagger removes all POS-tags.</p>
Hachidaishu Part-of-Speech Dataset
<p><strong>Full Changelog</strong>: https://github.com/yamagen/hachidaishu-pos/commits/1.0.1</p>
Cross-Register Projection for Headline Part of Speech Tagging
<p>POSH: The POS-tagged HeadlIne corpus was created for the paper “Cross-Register Projection for Headline Part of Speech Tagging” published in EMNLP 2021. This dataset contains headlines with gold annotated POS tags.</p> <p>The <em>GSCh</em> evaluation set is here: GSCh/gsc-headline-gold.test.conllu</p> <p>The smaller evaluation set of GSC headlines sampled uniformly at random (described in section 2.3) is here: gold_unconstrained_headlines/unifrand_gsc.test.conllu</p> <p>The POS-tagged NYT headlines described in section 2.3 are not shared directly as this text was drawn from the New York Times Annotated Corpus (LDC2008T19), and subject to license constraints. However, if you have access to and have untarred LDC2008T19, you can recover this evaluation set with:</p> <p> TAG_PATH="./unifrand_onlynyt.tags.json" # mapping from NYT headline span to gold POS tag</p> <p> python build_gold_nyt_headlines.py --nyt_dir /PATH/TO/ANNOTATED/NYT/CORPUS/ --tag_path ${TAG_PATH} --num_proc 4</p> <p>Increase the argument to --num_procs to process more shards from the NYT corpus in parallel and reduce build time.</p> <p>Under GSCproj we also share the <em>GSCproj</em> folds which we used to train and validate our models. These are not gold POS tags, and are shared purely for reproducibility sake.</p>
Main and extended tables for the 207-word Swadesh list of Early Sranan and Modern Sranan with parts of speech, semantic categories, source languages and semantic and lexical changes
<p>The dataset was made for the purposes of the author's master thesis, titled <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&lang=slv">"Socio-Cultural Motivations for the Acquisition of Lexical Items in Sranan Tongo’s Core Vocabulary"</a>. The dataset includes two worksheets. The first is titled "Main table", and it includes all the data, where each Swadesh gloss (1 to 207) is assigned one ID (No., first column), even if there are multiple Modern Sranan (MSr) equivalents. The second worksheet, titled "Extended table", includes additional IDs (No., first column) by hyphenating, so that each MSr equivalent has its separate ID number (e. g. gloss numbered 2 has 3 MSr equivalents, so these are now numbered 2-1, 2-2, and 2-3, respectively). <br> This allowed the author to also make a clearer distinction according to source languages, as the MSr equivalents for the same gloss sometimes come from different source languages. More about the methodology of the tables and their importance for the research is available in the thesis, available <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&lang=slv">at this link</a>. </p>
Creating speech zones with self-distributing acoustic swarms (Augmented Dataset Part 1 of 2)
<p>Datasets used in the paper: "Creating speech zones with self-distributing acoustic swarms"</p> <p>This deposit contains the <strong>first</strong> part of the augmented dataset containing simulated and real world collected data. The datasets contains 18000 training mixtures of 3-5 speakers, of which 6000 are simulated using PyRoomAcoustics, 6000 are created from synchronized real world recordings in an anechoic chamber, and 6000 are created from synchronized recordings in ordinary reverberant rooms.</p> <p>It also includes a validation set of 500 mixtures from reverberant rooms, and a testing set of 1000 mixtures from reverberant rooms.</p> <p>The source sounds are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker. The recordings are captured from an array of 7 microphones, as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams.</p> <p>You must download <strong>both</strong> the first and the second part of this dataset in order to use it properly.</p> <p>To uncompress the two datasets, download both and execute:</p> <p>```cat *.tar.gz.* | tar xvfz -```</p> <p>Please see the Readme for more information. Please see related identifiers for other datasets.</p>
Creating speech zones with self-distributing acoustic swarms (Augmented Dataset Part 2 of 2)
<p>Datasets used in the paper: "Creating speech zones with self-distributing acoustic swarms"</p> <p>This deposit contains the <strong>second</strong> part of the augmented dataset containing simulated and real world collected data. The datasets contains 18000 training mixtures of 3-5 speakers, of which 6000 are simulated using PyRoomAcoustics, 6000 are created from synchronized real world recordings in an anechoic chamber, and 6000 are created from synchronized recordings in ordinary reverberant rooms.</p> <p>It also includes a validation set of 500 mixtures from reverberant rooms, and a testing set of 1000 mixtures from reverberant rooms.</p> <p>The source sounds are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker. The recordings are captured from an array of 7 microphones, as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams.</p> <p>You must download <strong>both</strong> the first and the second part of this dataset in order to use it properly.</p> <p>To uncompress the two datasets, download both and execute:</p> <p>```cat *.tar.gz.* | tar xvfz -```</p> <p>Please see the Readme for more information. Please see related identifiers for other datasets.</p>
A part-of-speech (POS) lexicon of Classical Tibetan for NLP
<p>This part-of-speech (POS) lexicon of Classical Tibetan was prepared in the course of the research project 'Tibetan in Digital Communication' (2012-2015) hosted at SOAS, University of London and funded by the UK's Arts and Humanities Research Council (grant code: AH/J00152X/1). The data for verbs comes from a digitized version of <em>A Lexicon of Tibetan Verb Stems as Reported by the Grammatical Tradition</em> (Munich: Bayerische Akademie der Wissenschaften, 2010) by Nathan W. Hill. Otherwise data comes from the manually part-of-speech tagged training data produced by the corpus and a few lexical items specifically added by hand to improve rule based tagging.</p>
A Comprehensive Central Kurdish Sound Dataset for Robust Automatic Speech Recognition (Part 1).
<p>Exploring the intricacies of Speech Recognition Technology (SRT), our dataset encompasses a wide range of age demographics, spanning from adolescents to individuals in their fifties. This diverse dataset comprises a substantial collection of raw data, amounting to 1,739,089 entries. Within this dataset, a meticulous curation process has yielded a total of 1,683 hours of data, providing a thorough examination of language acquisition patterns across different age cohorts within the Central Kurdish linguistic domain.</p>
A New Annotation Scheme for the Sejong Part-of-speech Tagged Corpus
<p>We produce Sejong-style morphological analysis and part-of-speech tagging results which have been the de facto standard for Korean language processing by using UDPipe (http://ufal.mff.cuni.cz/udpipe) </p> <p> </p> <p>udpipe --tokenize --tag sjmorph.model input > output</p> <p>see https://github.com/jungyeul/sjmorph</p>
CUB-200 Speech captions (Part-I)
<p>Speech captions of CUB-200 database. These speech captions are synthesized by Tacotron-v2 according to the original textual descriptions.</p>
Developing an English course for Beginners with the topic Part of Speech using Google Classroom
<p>Self-paced learning means that you can study on your own time and schedule. You don't have to complete the same assignment or study at the same time as others. You can proceed from one topic or segment to the next at your own pace. Why we chose Google Classroom Media is because it's so easy to use. For people who are using Google classrooms for the first time, they will definitely have no trouble operating them.</p>
Target Speech Extraction Dataset for Knowledge Boosting (Part 1)
<p><strong>Part 1 of the Target Speech Extraction Dataset</strong> as described in <em>Knowledge boosting during low-latency inference</em> (Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. </p>
Target Speech Extraction Dataset for Knowledge Boosting (Part 2)
<p><strong>Part 2 of the Target Speech Extraction Dataset</strong> as described in <em>Knowledge boosting during low-latency inference</em> (Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. </p>
DDS (Device-Degraded Speech) Dataset - VCTK portion - Part 2
<p>DDS (Device-Degraded Speech) dataset provides aligned parallel recordings of high-quality speech (recorded in professional studios) and a large number of versions of low-quality speech, producing approximately 2,000 hours speech data. </p> <p>DDS is built on top of two datasets: DAPS and VCTK. We play clean speech recordings (4 hours from DAPS and 8 hours from VCTK) and re-record waveforms in nine environments (two offices, two conference rooms, three studios, one living room, one waiting room) on three different devices (one MEMS and two condenser microphones), producing 27 different recording conditions. Moreover, each version of condition consists of multiple recordings recorded at 6 different microphone positions to simulate various signal-to-noise ratio (SNR) and reverberation levels. </p> <p><strong>Arxiv:</strong> https://arxiv.org/abs/2109.07931</p> <p> </p> <p><strong>The whole dataset is split into 3 repositories (one part for DAPS portion, two parts for VCTK portion). This repository contains VCTK portion (part 2).</strong></p> <p><strong>For all repository links of DDS v0.8:</strong></p> <ul> <li><strong>DAPS portion:</strong> https://zenodo.org/record/5464104</li> <li><strong>VCTK portion part1:</strong> https://zenodo.org/record/5499506</li> <li><strong>VCTK portion part2:</strong> https://zenodo.org/record/5501697</li> </ul>
DDS (Device-Degraded Speech) Dataset - VCTK portion - Part 1
<p>DDS (Device-Degraded Speech) dataset provides aligned parallel recordings of high-quality speech (recorded in professional studios) and a large number of versions of low-quality speech, producing approximately 2,000 hours speech data. </p> <p>DDS is built on top of two datasets: DAPS and VCTK. We play clean speech recordings (4 hours from DAPS and 8 hours from VCTK) and re-record waveforms in nine environments (two offices, two conference rooms, three studios, one living room, one waiting room) on three different devices (one MEMS and two condenser microphones), producing 27 different recording conditions. Moreover, each version of condition consists of multiple recordings recorded at 6 different microphone positions to simulate various signal-to-noise ratio (SNR) and reverberation levels. </p> <p><strong>Arxiv:</strong> https://arxiv.org/abs/2109.07931</p> <p> </p> <p><strong>The whole dataset is split into 3 repositories (one part for DAPS portion, two parts for VCTK portion). This repository contains VCTK portion (part 1).</strong></p> <p><strong>For all repository links of DDS v0.8:</strong></p> <ul> <li><strong>DAPS portion:</strong> https://zenodo.org/record/5464104</li> <li><strong>VCTK portion part1:</strong> https://zenodo.org/record/5499506</li> <li><strong>VCTK portion part2:</strong> https://zenodo.org/record/5501697</li> </ul>
Differential Diagnosis Between Parkinson's Disease and Multiple System Atrophy Using Digital Speech Analysis - Part 2
ClinicalTrials.gov study NCT05807373. IPD Sharing: NO. Countries: 1. Publications: 0.
Developing an English course for beginners with the topic Part of Speech using Google Classroom
<p>Self-Paced Learning</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.