Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

87

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

87 results for “LANGUAGE PROCESSING”

Learn how ShareScore rates datasets ↗
ClinicalTrials.gov32/100

Hereditary Deficits in Auditory Processing Leading to Language Impairment

ClinicalTrials.gov study NCT00004570. IPD Sharing: Not stated. Countries: 1. Publications: 3.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Natural Language Processing and Quality Assessment in Primary Care

ClinicalTrials.gov study NCT01023243. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

Figure 1d from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 1d A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Fossilised animal skin (Natural History Museum 2009)

opencc-by-4.0Jul 2020View details →
zenodo28/100

Figure 1b from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 1b A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Pinned insect specimen (Natural History Museum 2018)

opencc-by-4.0Jul 2020View details →
zenodo28/100

Figure 1c from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 1c A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Microscope slide (Natural History Museum 2017)

opencc-by-4.0Jul 2020View details →
zenodo28/100

Figure 1a from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 1a A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Herbarium specimen (Natural History Museum 2007a)

opencc-by-4.0Jul 2020View details →
zenodo28/100

Figure 11 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 11 The distribution of languages across the specimen and herbaria. EN=English, FR=French, LA=Latin, ET=Estonian, DE=German, NL=Dutch, PT=Portuguese, ES=Spanish, SV=Swedish, RU=Russian, FI=Finnish, IT=Italian, ZZ=Unknown. The codes for the contributing herbaria are listed in Table 11 (from Dillen et al. 2019).

opencc-by-4.0Jul 2020View details →
zenodo28/100

Supplementary material 1 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Appendices

opencc-zeroJul 2020View details →
zenodo28/100

Figure 1e from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 1e A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Liquid preserved specimen (Natural History Museum 2010)

opencc-by-4.0Jul 2020View details →
zenodo28/100

Figure 2 from: Owen D, Groom Q, Hardisty A, Leegwater T, Livermore L, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e58030. https://doi.org/10.3897/rio.6.e58030

Figure 2 A possible semi-automatic digitisation workflow to extract data from the labels of collection specimens.

opencc-by-4.0Sep 2020View details →
zenodo28/100

On the Understandability of Temporal Properties Formalized in Linear Temporal Logic, Property Specification Patterns and Event Processing Language

<p>Experimental material &amp; data</p>

opencc-by-4.0Sep 2017View details →
zenodo28/100

MIGRATION PROCESSES TO THE LANGUAGE IN THE INFLUENCE TO HIMSELF CHARACTERISTIC

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

LEXICON OF THE UZBEK LANGUAGE AND THE PROCESS OF WORD ACQUISITION IN IT

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

Supplemental package of a study on Automated User Story Validation Using Natural Language Processing

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo28/100

Datasets for "Leveraging Machine Learning and Natural Language Processing Techniques for Agriculture Experiment Station Project Classification"

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo28/100

USE OF INNOVATIVE AND INTERACTIVE METHODS IN THE PROCESS OF TEACHING MOTHER LANGUAGE LITERATURE TO STUDENTS

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo28/100

Japhug for Natural Language Processing: a single-speaker audio corpus with transcriptions

<p><em>(fran&ccedil;ais ci-dessous)</em></p> <p>This archive contains a dataset (audio files and transcriptions) of a minority language, Japhug (Glottocode: japh1234; closest iso 639-3 code: jya). The archive contains a subset of the Japhug corpus of the Pangloss Collection: it is a single-speaker corpus, consisting of all the audio resources transcribed, for the main speaker of this corpus (Ms. Tshendzin).<br> The corpus is versioned, so that the experiments carried out on these resources (for linguistic research or for Natural Language Processing) are fully reproducible. All relevant information is contained in YAML files (.yml extension; one in French, one in English).<br> The data sub-folder contains the converted and demultiplexed audio files, as well as the annotations associated with each channel of the audio files.<br> The summary files contain, among other things, the list of graphemes used in the language (complex graphemes are particularly important), as well as information on the various resources (audio and annotations), such as their identifiers (DOIs) and links to the original files.<br> From a computational point of view, the list of DOIs of the audios and annotations described in this YAML file is sufficient to generate this corpus at a given time. A corpus like the present one can be viewed as the version, at a given time, of a set of documents in the Pangloss collection: a corpus as it stands at a precise version.</p> <p>Further information is available from&nbsp;<a href="https://gitlab.com/lacito/outilspangloss">https://gitlab.com/lacito/outilspangloss</a></p> <p>---------------</p> <p>Cette archive contient un jeu de donn&eacute;es (audios et transcriptions) d&rsquo;une langue &agrave; tradition orale, le japhug (Glottocode: japh1234; code iso 639-3 le plus proche : jya). L&rsquo;archive contient un sous-ensemble du corpus japhug de la collection Pangloss : c&rsquo;est un corpus monolocuteur, constitu&eacute; de l&rsquo;int&eacute;gralit&eacute; des ressources audio transcrites pour la locutrice principale de ce corpus (Mme Tshendzin).<br> Le corpus est versionn&eacute;, de sorte que les exp&eacute;riences men&eacute;es sur ces ressources (pour la linguistique ou pour le Traitement automatique des langues) soient reproductibles de fa&ccedil;on exacte (en pensant bien &agrave; joindre l&rsquo;algorithme : param&egrave;tres, r&eacute;partitions des fichiers dans les diff&eacute;rents ensembles, etc.). Toutes les informations pertinentes se trouvent dans les fichiers YAML (extension .yml ; un en fran&ccedil;ais, un autre en anglais).<br> Le sous-dossier des donn&eacute;es contient d&rsquo;une part les audios convertis et d&eacute;multiplex&eacute;s et d&rsquo;autre part les annotations associ&eacute;es &agrave; chaque canal desdits audios.<br> Les fichiers r&eacute;capitulatifs contiennent notamment la liste des graph&egrave;mes utilis&eacute;s dans cette langue (les graph&egrave;mes complexes sont particuli&egrave;rement importants), ainsi que des informations sur les diff&eacute;rentes ressources (audios et annotations), comme les identifiants (DOI), les liens vers les fichiers originaux, etc.<br> Au plan informatique, la liste des identifiants DOI des audios et annotations d&eacute;crits dans ce fichier YAML suffit pour g&eacute;n&eacute;rer ce corpus &agrave; un instant t. Un corpus comme celui-ci peut &ecirc;tre vu comme la version &agrave; l&rsquo;instant t d&rsquo;un ensemble de documents de la collection Pangloss : un corpus arr&ecirc;t&eacute; &agrave; une version pr&eacute;cise.<br> Pour plus de pr&eacute;cisions :&nbsp;<a href="https://gitlab.com/lacito/outilspangloss">https://gitlab.com/lacito/outilspangloss</a></p>

openSep 2021View details →
zenodo28/100

Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Rating Confusion Matrix

<p>Resulting rating confusion matrix&nbsp;for the experiment &quot;Evaluation of a simple score-based Natural Language Processing (NLP) algorithm&quot;.</p>

opencc-byMay 2023View details →
zenodo28/100

A Free Verbalization Method of Evaluating Sound Design: The Effectiveness of Artificially Intelligent Natural Language Processing Methods and Tools

<p>&quot;Robot&quot; voice sound files. Seventeen sound files were recorded in four formats; raw human voiceover (VO), and three types of robot voice: vocoded voice 1 (&ldquo;robo&rdquo;), vocoded voice 2 with music (&ldquo;kbd&rdquo;), and a &ldquo;beep&rdquo; voice. Each was recorded as 44.1kHz, 24-bit wav files in a professional recording studio. VO was recorded by professional voice actor DB Cooper, who has been the robot voice for several video games, as well as the voice of the DEE BMW internal car AI voice system. Cooper recorded three versions of the emotes on a Sennheiser MKH-416. Professional sound designer pdx Drescher, an expert in robot<br> and interface sound design, created three sets of robot voices from<br> the original voice files. With guidance from one of the authors,<br> pdx was tasked with trying different approaches to turning the VO<br> samples into three different types of robot voice while attempting<br> to maintain the meaning of the original sounds as described in the<br> list above through preserving the prosody/melodic contour of the<br> original. The first set, robo, used some clips from one of pdx&rsquo;s prior<br> robot voice projects and integrated them to approximate the emo-<br> tional intention of the VO. Clips were re-pitched, manipulated, and<br> modulated using ProTools plugins. For the kbd takes, VO sounds<br> were played into a Shure SM58 microphone. Vocoder patches mod-<br> ified the signal by voice formants, and the pitch was determined<br> by MIDI notes and pitch-bend controllers. Output of the synthe-<br> sizer was then edited with additional synth patches and effects (EQ,<br> modulation, etc.). We made particular use of a plugin called Envy<br> by Cargo Cult, which takes the volume, pitch, and EQ envelopes of<br> one sound (the original VO) and apply them to another sound. This<br> helped make the synth resemble the prosody of the original sound<br> to some degree. The beep sounds underwent a similar development<br> process as the robo takes, but with interface &ldquo;bleeps and bloops&rdquo;<br> derived from various sound effects libraries, including the Star Trek<br> LCARS soundset.</p> <p>The following sounds were recorded: 1. Warning calm (&ldquo;Uh-oh&rdquo;) 2. Warning alarm (&ldquo;ah!&rdquo;) 3. Wrong/<br> error (&ldquo;rrrrr&rdquo;) 4. Correct/good (&ldquo;yay&rdquo;) 5. Surprise (neutral) (&ldquo;Oh!&rdquo;) 6.<br> Surprise (good) (&ldquo;Oh!&rdquo;) 7. Surprise (bad) &ldquo;(ohhh&rdquo;) 8. Love/adoration<br> (&ldquo;awww&rdquo;) 9. Disgust (&ldquo;ew&rdquo;) 10. Contempt (&ldquo;ech&rdquo;) 11. Guilt (&ldquo;hmmm&rdquo;)<br> 12. Confused (&ldquo;huh?&rdquo;) 13. Laugh (&ldquo;ha ha&rdquo;) 14. Calculating (&ldquo;hmmm&rdquo;)<br> 15. Sigh 16. Giggle 17. Pain (&ldquo;ow &quot;)</p>

opencc-by-4.0Jul 2023View details →
zenodo28/100

Data for article entitled 'Prediction in SVO and SOV languages: Processing and typological considerations', published in Linguistics

<p>see the publication for the description of the data and analysis</p>

opencc-by-4.0Oct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record