Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

29

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

29 results for “part of speech”

Learn how ShareScore rates datasets ↗
zenodo48/100

Corpus of Occitan Written Traditional Folktales Annotated with Part-Of-Speech (OWT-Tag)

<p>This resource contains 5 extracts of texts in Occitan which were manually annotated with lemmas and parts-of-speech, following the Grace standard. It was produced during the ExpressioNarration project, funded by a Marie Curie Individual Fellowship, in order to evaluate the performance of an Occitan Part-Of-Speech tagger, Talismane, to the specifities of the corpus of the project called Oral Occitan (OcOr), also available on https://zenodo.org/record/1451753#.W78FJWOYSpo.<br> Each extract contains around 1500 words. They are extracted from &#39;Contes et proverbes populaires recueillis en armagnac et Contes populaires recueillis en agenais&#39; de J.-F. Blad&eacute;, &#39;Coundes biarn&eacute;s, cou&eacute;ilhuts a&uuml;s pars&agrave;as mi&eacute;ytad&egrave;s dou p&eacute;ys d&eacute; Biarn&#39; de J.-V. Lalanne, &#39;Contes populaires du Languedoc&#39; de L. Lambert and &#39;Contes populaires recueillis dans la Grande-Lande&#39; de F. Arnaudin.<br> The annotation process is described in the following article available on https://www.openscience.fr/IMG/pdf/iste_modocv1n1_2.pdf.</p>

opencc-by-sa-4.0Oct 2018View details →
zenodo44/100

A challenging data set for evaluating part-of-speech taggers

<p>This data set contains 2,227 sentences, with a part-of-speech (POS) tag specified for a single word in the sentence. The data file is a tab-separated text file where each row &nbsp;(after the header row) is formatted as follows:</p><p><i>sentence &lt;TAB&gt; POS tag &lt;TAB&gt; (optional) motivation</i></p><p>Note that, in the sentence (= a string of space-separated characters), the POS-tagged word is indicated by the POS tag in brackets, placed just after the word to which it refers. Example:</p><p><i>The road bends [VERB] to the right . VERB</i></p><p>In this example, the optional motivation is not included, as the tagged word can easily be identified as being a verb.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

A part-of-speech (POS) tagged corpus of Classical Tibetan

<p>This part-of-speech (POS) tagged corpus of Classical Tibetan was prepared in the course of the research project &#39;Tibetan in Digital Communication&#39; (2012-2015) hosted at SOAS, University of London and funded by the UK&#39;s Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. This corpus includes the <em>Mdzaṅs blun</em> (9th century, canonical), the <em>Bu ston chos ḥbyuṅ</em> (13th century, ecclesiastical history), the <em>Mi la ras paḥi rnam thar</em> and <em>Mar paḥi rnam thar</em> (15th century, biography).</p>

opencc-by-4.0May 2017View details →
zenodo40/100

A rule based Tibetan part-of-speech (POS) tagger for the creation of gold standard training data

<p>This rule based Tibetan part-of-speech (POS) tagger was prepared in the course of the research project &#39;Tibetan in Digital Communication&#39; (2012-2015) hosted at SOAS, University of London and funded by the UK&#39;s Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. For a description of the tagger itself see Garrett et al. 2014. Note that the tagger must be used together with a lexicon (for example Hill &amp; Garrett 2017a). One must use one&#39;s own script to tag all words with all tags in the lexicon and then apply the tagger to remove incorrect tags.</p> <p>On the associated corpus of 318,230 words (Hill &amp; Garrett 2017b) the lexical tagger (i.e. simply applying all available tags to all words) tags 141,911 words with the correct unique tag, achieves as accuracy of 1.000 (by definition getting the right tag among others for each word) with an ambiguity of 2.73111. In contrast, the Rule Tagger tags 241,256 words with the correct unique tag, achieves an accuracy of 0.99893 and an ambiguity of 1.38577.</p> <p>Because this tagger does not achieve ambiguity 1.000 it is not suitable for tagging large scale corpora, but instead is useful for the creation of gold standard training data.</p> <p>N.B. In some rare cases the tagger removes all POS-tags.</p>

opencc-by-4.0May 2017View details →
zenodo40/100

Hachidaishu Part-of-Speech Dataset

<p><strong>Full Changelog</strong>: https://github.com/yamagen/hachidaishu-pos/commits/1.0.1</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Cross-Register Projection for Headline Part of Speech Tagging

<p>POSH: The POS-tagged HeadlIne corpus was created for the paper &ldquo;Cross-Register Projection for Headline Part of Speech Tagging&rdquo; published in EMNLP 2021.&nbsp; This dataset contains headlines with gold annotated POS tags.</p> <p>The <em>GSCh</em> evaluation set is here: GSCh/gsc-headline-gold.test.conllu</p> <p>The smaller evaluation set of GSC headlines sampled uniformly at random (described in section 2.3) is here: gold_unconstrained_headlines/unifrand_gsc.test.conllu</p> <p>The POS-tagged NYT headlines described in section 2.3 are not shared directly as this text was drawn from the New York Times Annotated Corpus (LDC2008T19), and subject to license constraints.&nbsp; However, if you have access to and have untarred LDC2008T19, you can recover this evaluation set with:</p> <p>&nbsp;&nbsp;&nbsp; TAG_PATH=&quot;./unifrand_onlynyt.tags.json&quot;&nbsp; # mapping from NYT headline span to gold POS tag</p> <p>&nbsp;&nbsp;&nbsp; python build_gold_nyt_headlines.py --nyt_dir /PATH/TO/ANNOTATED/NYT/CORPUS/ --tag_path ${TAG_PATH} --num_proc 4</p> <p>Increase the argument to --num_procs to process more shards from the NYT corpus in parallel and reduce build time.</p> <p>Under GSCproj we also share the <em>GSCproj</em> folds which we used to train and validate our models.&nbsp; These are not gold POS tags, and are shared purely for reproducibility sake.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Main and extended tables for the 207-word Swadesh list of Early Sranan and Modern Sranan with parts of speech, semantic categories, source languages and semantic and lexical changes

<p>The dataset was made for the purposes of the author&#39;s master thesis, titled <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&amp;lang=slv">&quot;Socio-Cultural Motivations for the Acquisition of Lexical Items in Sranan Tongo&rsquo;s Core Vocabulary&quot;</a>. The dataset includes two worksheets. The first is titled &quot;Main table&quot;, and it includes all the data, where each Swadesh gloss (1 to 207) is assigned one ID (No., first column), even if there are multiple Modern Sranan (MSr) equivalents. The second worksheet, titled &quot;Extended table&quot;, includes additional IDs (No., first column) by hyphenating, so that each MSr equivalent has its separate ID number (e. g. gloss numbered 2 has 3 MSr equivalents, so these are now numbered 2-1, 2-2, and 2-3, respectively).&nbsp;<br> This allowed the author to also make a clearer distinction according to source languages, as the MSr equivalents for the same gloss sometimes come from different source languages. More about the methodology of the tables and their importance for the research is available in the thesis, available <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&amp;lang=slv">at&nbsp;this link</a>.&nbsp;&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Creating speech zones with self-distributing acoustic swarms (Augmented Dataset Part 1 of 2)

<p>Datasets used in the paper:&nbsp;&quot;Creating speech zones with self-distributing acoustic swarms&quot;</p> <p>This deposit contains the <strong>first</strong> part of the augmented dataset containing simulated and real world collected data. The datasets contains 18000 training mixtures of 3-5 speakers, of which 6000 are simulated using PyRoomAcoustics, 6000 are created from&nbsp;synchronized real world recordings in an anechoic chamber, and 6000 are created from synchronized recordings in ordinary reverberant rooms.</p> <p>It also includes a validation set of 500 mixtures from reverberant rooms, and a&nbsp;testing set of 1000 mixtures from reverberant rooms.</p> <p>The source sounds&nbsp;are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker.&nbsp;The recordings are captured from an array of 7 microphones,&nbsp;as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams.</p> <p>You must download <strong>both</strong>&nbsp;the first and the second part of this dataset in order to use it properly.</p> <p>To uncompress the two datasets, download both and execute:</p> <p>```cat *.tar.gz.* | tar xvfz -```</p> <p>Please see the Readme for more information. Please see related identifiers for other datasets.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Creating speech zones with self-distributing acoustic swarms (Augmented Dataset Part 2 of 2)

<p>Datasets used in the paper:&nbsp;&quot;Creating speech zones with self-distributing acoustic swarms&quot;</p> <p>This deposit contains the <strong>second</strong> part of the augmented dataset containing simulated and real world collected data. The datasets contains 18000 training mixtures of 3-5 speakers, of which 6000 are simulated using PyRoomAcoustics, 6000 are created from synchronized real world recordings in an anechoic chamber, and 6000 are created from synchronized recordings in ordinary reverberant rooms.</p> <p>It also includes a validation set of 500 mixtures from reverberant rooms, and a&nbsp;testing set of 1000 mixtures from reverberant rooms.</p> <p>The source sounds&nbsp;are various utterances from the VCTK dataset. For real world data, the utterances are played over a Rokono Bass+ Mini Speaker.&nbsp;The recordings are captured from an array of 7 microphones,&nbsp;as they are recorded by our robotic swarm as it is distributed across the table. The recorded audio in the real world has been subjected to audio compression and decompression using the Opus Codec to enable multiple simultaneous streams.</p> <p>You must download <strong>both</strong>&nbsp;the first and the second part of this dataset in order to use it properly.</p> <p>To uncompress the two datasets, download both and execute:</p> <p>```cat *.tar.gz.* | tar xvfz -```</p> <p>Please see the Readme for more information. Please see related identifiers for other datasets.</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

A part-of-speech (POS) lexicon of Classical Tibetan for NLP

<p>This part-of-speech (POS) lexicon of Classical Tibetan was prepared in the course of the research project &#39;Tibetan in Digital Communication&#39; (2012-2015) hosted at SOAS, University of London and funded by the UK&#39;s Arts and Humanities Research Council (grant code: AH/J00152X/1). The data for verbs comes from a digitized version of <em>A Lexicon of Tibetan Verb Stems as Reported by the Grammatical Tradition</em> (Munich: Bayerische Akademie der Wissenschaften, 2010) by Nathan W. Hill. Otherwise data comes from the manually part-of-speech tagged training data produced by the corpus and a few lexical items specifically added by hand to improve rule based tagging.</p>

opencc-by-4.0May 2017View details →
zenodo36/100

A Comprehensive Central Kurdish Sound Dataset for Robust Automatic Speech Recognition (Part 1).

<p>Exploring the intricacies of Speech Recognition Technology (SRT), our dataset encompasses a wide range of age demographics, spanning from adolescents to individuals in their fifties. This diverse dataset comprises a substantial collection of raw data, amounting to 1,739,089 entries. Within this dataset, a meticulous curation process has yielded a total of 1,683 hours of data, providing a thorough examination of language acquisition patterns across different age cohorts within the Central Kurdish linguistic domain.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

A New Annotation Scheme for the Sejong Part-of-speech Tagged Corpus

<p>We produce Sejong-style morphological analysis and part-of-speech tagging results which have been the de facto&nbsp;standard for Korean language processing by using UDPipe (http://ufal.mff.cuni.cz/udpipe)&nbsp;</p> <p>&nbsp;</p> <p>udpipe --tokenize --tag sjmorph.model input &gt; output</p> <p>see&nbsp;https://github.com/jungyeul/sjmorph</p>

opencc-by-4.0May 2019View details →
zenodo32/100

CUB-200 Speech captions (Part-I)

<p>Speech captions of CUB-200 database. These speech captions are synthesized by Tacotron-v2 according to the original textual descriptions.</p>

opencc-by-4.0Aug 2020View details →
zenodo32/100

Developing an English course for Beginners with the topic Part of Speech using Google Classroom

<p>Self-paced learning means that you can study on your own time and schedule. You don&#39;t have to complete the same assignment or study at the same time as others. You can proceed from one topic or segment to the next at your own pace. Why we chose Google Classroom Media is because it&#39;s so easy to use. For people who are using Google classrooms for the first time, they will definitely have no trouble operating them.</p>

opencc-by-4.0Jan 2021View details →
zenodo32/100

Target Speech Extraction Dataset for Knowledge Boosting (Part 1)

<p><strong>Part 1 of the Target Speech Extraction Dataset</strong> as described in&nbsp;<em>Knowledge boosting during low-latency inference</em>&nbsp;(Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel &nbsp;technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a &nbsp;streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications.&nbsp;</p>

openJun 2024View details →
zenodo32/100

Target Speech Extraction Dataset for Knowledge Boosting (Part 2)

<p><strong>Part 2 of the Target Speech Extraction Dataset</strong> as described in&nbsp;<em>Knowledge boosting during low-latency inference</em>&nbsp;(Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel &nbsp;technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a &nbsp;streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications.&nbsp;</p>

openJul 2024View details →
zenodo32/100

DDS (Device-Degraded Speech) Dataset - VCTK portion - Part 2

<p>DDS (Device-Degraded Speech) dataset provides aligned parallel recordings of high-quality speech (recorded in professional studios) and a large number of versions of low-quality speech, producing approximately 2,000 hours speech data.&nbsp;</p> <p>DDS is built on top of two datasets: DAPS and VCTK. We play clean speech recordings (4 hours from DAPS and 8 hours from VCTK) and re-record waveforms in nine environments (two offices, two conference rooms, three studios, one living room, one waiting room) on three different devices (one MEMS and two condenser microphones), producing 27 different recording conditions. Moreover, each version of condition consists of multiple recordings recorded at 6 different microphone positions to simulate various signal-to-noise ratio (SNR) and reverberation levels.&nbsp;</p> <p><strong>Arxiv:</strong>&nbsp;https://arxiv.org/abs/2109.07931</p> <p>&nbsp;</p> <p><strong>The whole dataset is split into 3 repositories (one part for DAPS portion, two parts for VCTK portion). This repository contains VCTK portion (part 2).</strong></p> <p><strong>For all repository links of DDS v0.8:</strong></p> <ul> <li><strong>DAPS portion:</strong>&nbsp;https://zenodo.org/record/5464104</li> <li><strong>VCTK portion part1:</strong>&nbsp;https://zenodo.org/record/5499506</li> <li><strong>VCTK portion part2:</strong>&nbsp;https://zenodo.org/record/5501697</li> </ul>

openodc-bySep 2021View details →
zenodo32/100

DDS (Device-Degraded Speech) Dataset - VCTK portion - Part 1

<p>DDS (Device-Degraded Speech) dataset provides aligned parallel recordings of high-quality speech (recorded in professional studios) and a large number of versions of low-quality speech, producing approximately 2,000 hours speech data.&nbsp;</p> <p>DDS is built on top of two datasets: DAPS and VCTK. We play clean speech recordings (4 hours from DAPS and 8 hours from VCTK) and re-record waveforms in nine environments (two offices, two conference rooms, three studios, one living room, one waiting room) on three different devices (one MEMS and two condenser microphones), producing 27 different recording conditions. Moreover, each version of condition consists of multiple recordings recorded at 6 different microphone positions to simulate various signal-to-noise ratio (SNR) and reverberation levels.&nbsp;</p> <p><strong>Arxiv:</strong>&nbsp;https://arxiv.org/abs/2109.07931</p> <p>&nbsp;</p> <p><strong>The whole dataset is split into 3 repositories (one part for DAPS portion, two parts for VCTK portion). This repository contains VCTK portion (part 1).</strong></p> <p><strong>For all repository links of DDS v0.8:</strong></p> <ul> <li><strong>DAPS portion:</strong>&nbsp;https://zenodo.org/record/5464104</li> <li><strong>VCTK portion part1:</strong>&nbsp;https://zenodo.org/record/5499506</li> <li><strong>VCTK portion part2:</strong>&nbsp;https://zenodo.org/record/5501697</li> </ul>

openodc-bySep 2021View details →
ClinicalTrials.gov32/100

Differential Diagnosis Between Parkinson's Disease and Multiple System Atrophy Using Digital Speech Analysis - Part 2

ClinicalTrials.gov study NCT05807373. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
zenodo28/100

Developing an English course for beginners with the topic Part of Speech using Google Classroom

<p>Self-Paced Learning</p>

opencc-by-4.0Jan 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record