Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
116
datasets available to search
ShareScore release 0.9.0
Dataset results
116 results for “Natural Language”
Supplemental package of a study on Automated User Story Validation Using Natural Language Processing
Open the record for dataset details and reuse information.
Datasets for "Leveraging Machine Learning and Natural Language Processing Techniques for Agriculture Experiment Station Project Classification"
Open the record for dataset details and reuse information.
Japhug for Natural Language Processing: a single-speaker audio corpus with transcriptions
<p><em>(français ci-dessous)</em></p> <p>This archive contains a dataset (audio files and transcriptions) of a minority language, Japhug (Glottocode: japh1234; closest iso 639-3 code: jya). The archive contains a subset of the Japhug corpus of the Pangloss Collection: it is a single-speaker corpus, consisting of all the audio resources transcribed, for the main speaker of this corpus (Ms. Tshendzin).<br> The corpus is versioned, so that the experiments carried out on these resources (for linguistic research or for Natural Language Processing) are fully reproducible. All relevant information is contained in YAML files (.yml extension; one in French, one in English).<br> The data sub-folder contains the converted and demultiplexed audio files, as well as the annotations associated with each channel of the audio files.<br> The summary files contain, among other things, the list of graphemes used in the language (complex graphemes are particularly important), as well as information on the various resources (audio and annotations), such as their identifiers (DOIs) and links to the original files.<br> From a computational point of view, the list of DOIs of the audios and annotations described in this YAML file is sufficient to generate this corpus at a given time. A corpus like the present one can be viewed as the version, at a given time, of a set of documents in the Pangloss collection: a corpus as it stands at a precise version.</p> <p>Further information is available from <a href="https://gitlab.com/lacito/outilspangloss">https://gitlab.com/lacito/outilspangloss</a></p> <p>---------------</p> <p>Cette archive contient un jeu de données (audios et transcriptions) d’une langue à tradition orale, le japhug (Glottocode: japh1234; code iso 639-3 le plus proche : jya). L’archive contient un sous-ensemble du corpus japhug de la collection Pangloss : c’est un corpus monolocuteur, constitué de l’intégralité des ressources audio transcrites pour la locutrice principale de ce corpus (Mme Tshendzin).<br> Le corpus est versionné, de sorte que les expériences menées sur ces ressources (pour la linguistique ou pour le Traitement automatique des langues) soient reproductibles de façon exacte (en pensant bien à joindre l’algorithme : paramètres, répartitions des fichiers dans les différents ensembles, etc.). Toutes les informations pertinentes se trouvent dans les fichiers YAML (extension .yml ; un en français, un autre en anglais).<br> Le sous-dossier des données contient d’une part les audios convertis et démultiplexés et d’autre part les annotations associées à chaque canal desdits audios.<br> Les fichiers récapitulatifs contiennent notamment la liste des graphèmes utilisés dans cette langue (les graphèmes complexes sont particulièrement importants), ainsi que des informations sur les différentes ressources (audios et annotations), comme les identifiants (DOI), les liens vers les fichiers originaux, etc.<br> Au plan informatique, la liste des identifiants DOI des audios et annotations décrits dans ce fichier YAML suffit pour générer ce corpus à un instant t. Un corpus comme celui-ci peut être vu comme la version à l’instant t d’un ensemble de documents de la collection Pangloss : un corpus arrêté à une version précise.<br> Pour plus de précisions : <a href="https://gitlab.com/lacito/outilspangloss">https://gitlab.com/lacito/outilspangloss</a></p>
Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Rating Confusion Matrix
<p>Resulting rating confusion matrix for the experiment "Evaluation of a simple score-based Natural Language Processing (NLP) algorithm".</p>
Natural Language Inference Dataset for Software Engineering
<p>Active research in requirements engineering and software engineering necessitates the application of Natural Language Processing (NLP) techniques to address unique challenges and enhance software quality. However, there is a dearth of effective Natural Language Inference (NLI) datasets for training neural network models to generate distributed sentence representations and tackle diverse NLP tasks. In this paper, we present a NLI dataset, tailored specifically to software engineering, empowers neural network models to effectively handle NLP tasks in this domain. The creation of this dataset involved meticulous annotation and careful consideration of diverse sources, including software documentation, user guides, App reviews and different articles related to software systems. Our dataset maintains compatibility with existing NLI datasets like Stanford Natural Language Inference, facilitating seamless adaptation of models without additional preprocessing.</p>
A Free Verbalization Method of Evaluating Sound Design: The Effectiveness of Artificially Intelligent Natural Language Processing Methods and Tools
<p>"Robot" voice sound files. Seventeen sound files were recorded in four formats; raw human voiceover (VO), and three types of robot voice: vocoded voice 1 (“robo”), vocoded voice 2 with music (“kbd”), and a “beep” voice. Each was recorded as 44.1kHz, 24-bit wav files in a professional recording studio. VO was recorded by professional voice actor DB Cooper, who has been the robot voice for several video games, as well as the voice of the DEE BMW internal car AI voice system. Cooper recorded three versions of the emotes on a Sennheiser MKH-416. Professional sound designer pdx Drescher, an expert in robot<br> and interface sound design, created three sets of robot voices from<br> the original voice files. With guidance from one of the authors,<br> pdx was tasked with trying different approaches to turning the VO<br> samples into three different types of robot voice while attempting<br> to maintain the meaning of the original sounds as described in the<br> list above through preserving the prosody/melodic contour of the<br> original. The first set, robo, used some clips from one of pdx’s prior<br> robot voice projects and integrated them to approximate the emo-<br> tional intention of the VO. Clips were re-pitched, manipulated, and<br> modulated using ProTools plugins. For the kbd takes, VO sounds<br> were played into a Shure SM58 microphone. Vocoder patches mod-<br> ified the signal by voice formants, and the pitch was determined<br> by MIDI notes and pitch-bend controllers. Output of the synthe-<br> sizer was then edited with additional synth patches and effects (EQ,<br> modulation, etc.). We made particular use of a plugin called Envy<br> by Cargo Cult, which takes the volume, pitch, and EQ envelopes of<br> one sound (the original VO) and apply them to another sound. This<br> helped make the synth resemble the prosody of the original sound<br> to some degree. The beep sounds underwent a similar development<br> process as the robo takes, but with interface “bleeps and bloops”<br> derived from various sound effects libraries, including the Star Trek<br> LCARS soundset.</p> <p>The following sounds were recorded: 1. Warning calm (“Uh-oh”) 2. Warning alarm (“ah!”) 3. Wrong/<br> error (“rrrrr”) 4. Correct/good (“yay”) 5. Surprise (neutral) (“Oh!”) 6.<br> Surprise (good) (“Oh!”) 7. Surprise (bad) “(ohhh”) 8. Love/adoration<br> (“awww”) 9. Disgust (“ew”) 10. Contempt (“ech”) 11. Guilt (“hmmm”)<br> 12. Confused (“huh?”) 13. Laugh (“ha ha”) 14. Calculating (“hmmm”)<br> 15. Sigh 16. Giggle 17. Pain (“ow ")</p>
Data from: What is in a general plan? using natural language processing to read 461 California city general plans
Open the record for dataset details and reuse information.
Data from: Natural language processing and recurrent network models for identifying genomic mutation-associated cancer treatment change from patient progress notes
Open the record for dataset details and reuse information.
Data from: On the links between nature's values and language
Open the record for dataset details and reuse information.
Leukocyte gene regulation and patterns of natural language use.
GEO Series GSE87656. Homo sapiens. 143 samples. Type: Expression profiling by array.
Leveraging Natural Language for Program Search and Abstraction Learning Graphics
<p>Program synthesis dataset containing graphics programs tasks and language annotations (synthetic and human annotated) for the Leveraging Natural Language for Program Search and Abstraction Learning (currently under review at NeurIPS 2020). Will be deanonymized upon review.</p>
Figure 13 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 13 Gold standard versus NER output.
Figure 12 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 12 An example of a specimen label.
Figure 10 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 10 Results per field from Google Cloud Vision.
A Comparison of Natural Language Understanding Platforms for Chatbots in Software Engineering
<p>The file contains the training dataset used in our study. And, it includes the results of all NLUs for both tasks.</p>
Harnessing Natural Language Processing to Deciphers Ancient Temple Texts
<p>This paper explores the use of Natural Language Processing (NLP) and Optical Character Recognition (OCR) to digitize and translate ancient Indian temple inscriptions. By applying machine learning to ancient languages like Sanskrit and Tamil, the research offers methods to decode and reconstruct eroded texts. Semantic analysis tools are also used to uncover the cultural and historical significance of the inscriptions. This interdisciplinary approach combines data science with cultural preservation, providing new ways to understand ancient texts.</p>
Data set to train a natural language classifier able to differentiate between 15 topics relevant to biodiversity informatics
<p><strong>Scope and size</strong><br> This data set is used to train a natural language processing classifier. The classifier shall be able to differentiate between 15 topics relevant to biodiversity informatics. The list of relevant topics was adapted from Searls (2012).</p> <p>The data set was split into training data, testing data (for tweaking and unit-testing the classifier) and validation data. Each data set is stored as PDF files in a separate directory.</p> <ul> <li>Training data (5494 pages)</li> <li>Test data (977 pages)</li> <li>Validation data (215 pages)</li> </ul> <p><strong>Data sources and licenses</strong><br> Details about the licenses for each data set can be found in the corresponding directories.</p> <ul> <li>Training data was compiled from MIT OpenCourseWare resources provided by MIT under a Creative Commons BY-NC-SA License.</li> <li>Testing data was compiled from MIT OpenCourseWare exams, provided by MIT under a Creative Commons License BY-NC-SA.</li> <li>Validation data was compiled from Wikipedia, provided under a Creative Commons License by Wikipedia editors and contributors.</li> </ul> <p><strong>Topic references</strong><br> Each topic references one or more MIT OpenCourseWare courses:</p> <ul> <li>Algorithms (Demaine, and Devadas, 2011)</li> <li>Artificial Intelligence (Winston, 2010)</li> <li>Building Dynamic Websites (Abelson, and Greenspun, 2003)</li> <li>Computational Biology (Kellis, 2015)</li> <li>Computer Graphics (Matusik, and Durand, 2012)</li> <li>Computer Science and Programming (Bell, Grimson, and Guttag, 2016)</li> <li>Databases (Madden, Morris, Stonebraker, and Curino, 2010)</li> <li>Data Structures (Demaine, 2012)</li> <li>Digital Image Processing (Clifford, Fisher, Greenberg, and Wells, 2007; Golland, 2005)</li> <li>Machine Learning (Singh, Jaakkola, and Mohammad, 2006)</li> <li>Machine Structures (Morris, and Madden, 2009)</li> <li>Natural Language Processing (Berwick, 2003; Collins, and Barzilay, 2005)</li> <li>Parallel Computing (Edelman, 2011)</li> <li>Software Engineering (Jackson, and Devadas, 2005)</li> <li>Structure and Interpretation of Computer Programs (Miller, and Goldman, 2016).</li> </ul> <p><strong>References</strong></p> <p>Harold Abelson, and Philip Greenspun. 6.171 Software Engineering for Web Applications. Fall 2003. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Ana Bell, Eric Grimson, and John Guttag. 6.0001 Introduction to Computer Science and Programming in Python. Fall 2016. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Robert Berwick. 6.863J Natural Language and the Computer Representation of Knowledge. Spring 2003. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Gari Clifford, John Fisher, Julie Greenberg, and William Wells. HST.582J Biomedical Signal and Image Processing. Spring 2007. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Michael Collins, and Regina Barzilay. 6.864 Advanced Natural Language Processing. Fall 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Erik Demaine. 6.851 Advanced Data Structures. Spring 2012. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Erik Demaine, and Srini Devadas. 6.006 Introduction to Algorithms. Fall 2011. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Alan Edelman. 18.337J Parallel Computing. Fall 2011. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Polina Golland. 6.881 Representation and Modeling for Image Analysis. Spring 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Daniel Jackson, and Srini Devadas. 6.170 Laboratory in Software Engineering. Fall 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Manolis Kellis. 6.047 Computational Biology. Fall 2015. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Robert Miller, and Max Goldman. 6.005 Software Construction. Spring 2016. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Samuel Madden, Robert Morris, Michael Stonebraker, and Carlo Curino. 6.830 Database Systems. Fall 2010. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Wojciech Matusik, and Frédo Durand. 6.837 Computer Graphics. Fall 2012. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Morris, Robert, and Madden, Samuel. 6.033 Computer System Engineering. Spring 2009. Massachusetts Institute of Technology: MIT OpenCourseWare, http://hdl.handle.net/1721.1/118791. License: Creative Commons BY-NC-SA.</p> <p>David B. Searls. An online bioinformatics curriculum. 2012. PLoS computational biology, 8(9), p.e1002632.</p> <p>Rohit Singh, Tommi Jaakkola, and Ali Mohammad. 6.867 Machine Learning. Fall 2006. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Patrick Winston. 6.034 Artificial Intelligence. Fall 2010. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p> </p>
data for: Inconsistencies Detection in Natural Language Requirements using ChatGPT: a Preliminary Evaluation
<p>Supplementary material for the paper: Inconsistencies Detection in Natural Language Requirements using ChatGPT: a Preliminary Evaluation</p> <p>In the "gold" sheets there are the requirements: originals requirements are marked with "1" in the third column, mutants are marked with "0". </p>
An Open Internet-based Survey and Natural Language Processing Project Analysing Written Monologues by Headache Patients
ClinicalTrials.gov study NCT05153876. IPD Sharing: NO. Countries: 1. Publications: 0.
Detecting Delayed Discharge in Acute Geriatric Unit Using Natural Language Processing
ClinicalTrials.gov study NCT04965480. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.