Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

103

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

103 results for “English language”

Learn how ShareScore rates datasets ↗
zenodo44/100

The Dublin Language Garden Perceptual Dialectology of Irish English Collection

<blockquote> <p><strong>Recommended citation for this dataset:</strong><br> Garnett, Vicky, &amp; Lucek, Stephen. (2020). The Dublin Language Garden Perceptual Dialectology of Irish English Collection (Version 1.0.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.4247829</p> </blockquote> <p>&nbsp;</p> <p><strong>About this Dataset</strong><br> The field of Perceptual Dialectology is&nbsp;an area of sociolinguistic study that investigates how non-linguists view different varieties of language.&nbsp; It often includes hand-drawn map exercises in which participants indicate where they believe various varieties are spoken, and their attitudes towards them.&nbsp;</p> <p>In 2015, as part of a public linguistics outreach event (the Dublin Language Garden) held at Trinity College Dublin, the authors created an activity for members of the public and collected hand-drawn maps from them that gave responses to the following tasks:</p> <p>a. Indicate where you come from on the map (using a red dot sticker)<br> b. Draw where you think the Dublin dialect occurs<br> c. Draw the boundaries of any other dialects you believe occur in Ireland<br> d. Tell us what you think are the features of those dialects<br> e. Tell us what you think are the characteristics of the people who speak those dialects.</p> <p>Participants of all ages were encouraged to take part, but only data from those over 18 were retained after the event and used in this data collection. &nbsp;Participants were all given information on how the data was to be anonymised, processed and published on a clearly displayed poster to read before they were given a map to complete the 5 tasks (listed above). &nbsp;No additional information about the participants, aside from that acquired through Task a, was collected.</p> <p>&nbsp;</p> <p><strong>File List:</strong></p> <ul> <li>_READ_ME - Dublin Language Garden Perceptual Dialectology of Irish English data.txt<br> Contains a detailed description of this dataset.<br> &nbsp;</li> <li>DLG_PDIE_KML_data_by_location.zip<br> This zipped folder contains the .kml data of multiple hand-drawn maps organised into folders by their location<br> &nbsp;</li> <li>DLG_PDIE_KML_data_by_part.zip<br> This zipped folder contains the .kml data of multiple hand-drawn maps organised into folders according to the participants.</li> </ul> <p>These folders have been organised in this way in order to make discoverability easier between the data. &nbsp;Users may wish to analyse the data only by the locations of the varieties identified by the participants. &nbsp;Other users may only be interested in the data given by specific participants, and therefore the folder that organises the data in this way may be of better use to them. &nbsp;Both folders, however, contain the same data, it is simply how they are organised.</p> <ul> <li>Garnett and Lucek DLG_PD_IE Qualitative Data (Nov 2020).xlsx<br> Spreadsheet featuring tabulated qualitative data taken from all maps<br> &nbsp;</li> <li>Sample Hand-drawn Maps.zip<br> Folder containing 2 sample hand-drawn maps from the participants to help contextualise the data presented here.</li> </ul> <p>&nbsp;</p> <p><strong>Any questions?</strong><br> Any enquiries regarding this dataset should be directed to either Vicky Garnett (<a href="mailto:garnetv@tcd.ie">garnetv@tcd.ie</a>) or Stephen Lucek (<a href="mailto:stephen.lucek@ucd.ie">stephen.lucek@ucd.ie</a>).</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Polifonia Corpus - Books Module Metadata - English Language (Full)

<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

The Threatening English Language (TEL) Corpus

<p>TEL is the Threatening English Language corpus. It is a collection of 309 written texts compiled from the publicly-available portion of CTARC (the Communicated Threat Assessment Research Corpus, compiled by Tammy Gales), MFT (the Malicious Forensic Texts corpus, compiled by Andrea Nini), and the written portion of CoJO (the Corpus of Judicial Opinions, compiled by Julia Muschalik). Additional texts are from ForensicLing.com (the forensic linguistic data site hosted by Tammy Gales and Dakota Wing). Basic metadata is supplied for each text where known from the original case research. We wish to thank our graduate student fellows who helped compile the texts and metadata: Nicole Harris, Annina van Riper, Zara Rabinko, and Zachary Boudreaux.</p> <p>Total texts: 309<br> Total estimated authors: 203<br> Total word count: 54,167</p> <p>METADATA KEY</p> <p>TG = Tammy Gales (public portion of CTARC)<br> AN = Andrea Nini (MFT)<br> JM = Julia Muschalik (written portion of CoJo)<br> FL = ForensicLing.com (Tammy Gales and Dakota Wing)</p> <p>Name###_## = file name, case number, text number within case<br> File name might be threat recipient or author; remaining info is about the author, where known</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Files and code for English dictionaries, gold and silver standard corpora for biomedical natural language processing related to SARS-CoV-2 and COVID-19

<p><span lang="EN-GB">Automated information extraction with natural language processing (NLP) tools is required to gain systematic insights from the large number of COVID-19 publications, reports and social media posts, which far exceed human processing capabilities.&nbsp;</span></p> <p><span lang="EN-GB">Here we present an NLP toolbox comprising COVID-19-related dictionaries and annotated corpora in English as well as useful code and workflows for their update and use. The dictionaries contain terms referring to the COVID-19 disease, the SARS-CoV-2 virus, its variants and common mutations, respectively. They were used together with the EasyNER NLP tool to extract and annotate all 764&nbsp;398 abstracts in the CORD-19 dataset, creating a very large silver standard corpus (named Lund-Annotated-CORD-19 corpus). This was complemented with a small gold standard corpus consisting of PubMed abstracts manually annotated for key entity classes such as disease, virus, symptom, protein/gene, cell type, chemical and species terms. </span></p> <p><span lang="EN-GB">The toolbox can support various text analysis tasks related to COVID-19 such as named entity recognition and co-mention analysis. A preliminary version of the toolbox, which was released early in the pandemic, was</span><span lang="EN-GB"> for example already used to create a COVID-19 knowledge graph and study the evolution and variation of COVID-19-related terminology. In addition, the toolbox can be applied in the development of other NLP tools, for example to train and evaluate large language models.</span></p> <p><span lang="EN-GB">When using the toolbox, please cite this record and the associated article.</span></p> <p>&nbsp;</p> <p>&nbsp;</p>

openJun 2022View details →
zenodo40/100

TweetC19SR-Eng - Manually annotated dataset of English language COVID-19 tweets containing self-reports of symptoms

<p><strong>In this work, we release two expert curated, manually annotated datasets of COVID-19 self-reported symptoms. The first dataset contains tweets in English and the second contains tweets in Spanish, both containing around 36,500 tweets in total. These datasets were used for the Sixth and Seventh Workshop on Social Media Mining For Health (2021 and 2022)</strong></p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Polifonia Corpus - Encyclopedic Module Metadata - English Language

<p>We make available the Metadata related to the Wikipedia pages that constitute the Encyclopedic Module of the Polifonia Textual Corpus. Metadata for this module includes, per each Wikipedia page, its Wikipedia ID, BabelNet ID, gloss, resource type (that can be named entity or concept), Lemmata, Sensekey, WikiData ID.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Figure 1. The design of the study-The Effect of English Learning Anxiety on Iranian High-School Students' English Language Achievement

<p>The dependent variable in this study was English<br> achievement, and the independent variable was language anxiety. The intervening variable is<br> English proficiency. The control variables are: age of subjects (second year high school students<br> with an average age of 17) and years of experience in English learning (a minimum of 5<br> consecutive years). The schematic design of the study is presented below (Figura 1).</p>

opencc-by-4.0Jun 2011View details →
zenodo40/100

Lost in Translation? Not for Large Language Models: Automated Divergent Thinking Scoring Performance Translates to Non-English Contexts (Datasets)

<p>Datasets for: Zielińska, A., Organisciak, P., Dumas, D., &amp; Karwowski, M. (2023). Lost in translation? Not for large language models: Automated divergent thinking scoring performance translates to non-English contexts. <em>Thinking Skills and Creativity, 50</em>, 101414. https://doi.org/10.1016/j.tsc.2023.101414</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

<p>This benchmark dataset is published with the article:&nbsp;</p> <p><em>Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Martin Katz, and Nikolaos Aletras. 2021.&nbsp;LexGLUE: A Benchmark Dataset for Legal Language Understanding in English. ArXiv.</em></p> <p><strong>Short Description</strong></p> <p>Inspired by the recent widespread use of the GLUE multi-task benchmark NLP dataset (Wang et al., 2018), the subsequent more difficult SuperGLUE (Wang et al., 2019), other previous multi-task NLP benchmarks (Conneau and Kiela,2018; McCann et al., 2018), and similar initiatives in other domains (Peng et al., &nbsp;2019), we introduce LexGLUE, a benchmark dataset to evaluate the performance of NLP methods in legal tasks. LexGLUE is based on seven existing legal NLP datasets:</p> <ul> <li>ECtHR Task A &nbsp;(Chalkidis et al., 2019)</li> <li>ECtHR Task B &nbsp;(Chalkidis et al., 2021a)</li> <li>SCOTUS (Spaeth et al., 2020)</li> <li>EUR-LEX (Chalkidis et al., 2021b)</li> <li>LEDGAR (Tuggener et al. (2020)</li> <li>UNFAIR-ToS (Lippi et al., 2019)</li> <li>CaseHOLD (Zheng et al., 2021)</li> </ul>

opencc-by-4.0Sep 2021View details →
zenodo40/100

An explorative study into the attitudes, skills, and knowledge of primary and post-Primary teachers in relation to the provision for students with English as an Additional Language (EAL).

<p>The data contained here was gathered for a quantitative and qualitative-based study&nbsp;to investigate the attitudes, skills, and knowledge of primary and post-Primary teachers in Ireland in relation to the provision for students with English as an Additional Language (EAL).&nbsp;The data collection period will begin 20th January and the 20th February 2021 and was approved by&nbsp;UCC Social Research Ethics Committee (2021-232). A data codebook is also provided to support the interpretation of the data.&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

BSL-Hansard: A parallel, multimodal corpus of English and interpreted British Sign Language data from parliamentary proceedings

<p>BSL-Hansard is a novel open source and multimodal resource composed by combining Sign Language video data in BSL and English text from the official transcription of British parliamentary sessions. This paper describes the method followed to compile BSL-Hansard including time alignment of text using the MAUS (Schiel, 2015) segmentation system, gives some statistics about this dataset, and suggests experiments. These primarily include end-to-end Sign Language-to-text translation, but is also relevant for broader machine translation, and speech and language processing tasks.</p> <p>This dataset will be useful for translation between BSL and English, or for studies in BSL or English down to the phonetic level.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Annotated corpus samples of resultative nomination constructions in 4 languages (Dutch, English, French, Spanish)

<p>This dataset includes an annotated corpus sample retrieved from Sketch Engine of 3200 (i.e. 800x4) resultative constructions in 4 languages (Dutch, English, French, Spanish) with 16 (i.e. 4x4) different nomination verbs (Dutch: kronen &#39;crown&#39;, verkiezen &#39;elect&#39;, promoveren &#39;promote&#39;, uitroepen &#39;proclaim; English: crown, elect, promote, proclaim; French: couronner &#39;crown&#39;, &eacute;lire &#39;elect&#39;, promouvoir &#39;promote&#39;, proclamer &#39;proclaim&#39;; Spanish: coronar &#39;crown&#39;, elegir &#39;elect&#39;, ascender &#39;promote&#39;, proclamar &#39;proclaim&#39;).</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

A dataset of late 1990s and early 2000s web banner ads on Chinese- and English-language web pages

<p>This dataset contains information about 22,915 unique banner ad images appearing on Chinese- and English-language web pages in the late 1990s and early 2000s. The dataset is mined from 1,384,355 archived web page snapshots downloaded from the Wayback Machine, representing 77,747 unique HTTP URLs. The URLs are collected from six printed Internet directory books published in mainland China and the United States between 1999 and 2001, as part of a larger research project on Chinese-language web archiving.</p> <p><br> For each banner ad image, the dataset provides standard image metadata such as file format and dimension. The dataset also provides the original URLs of the web pages where the banner ad image was found, timestamps of the archived web page snapshots containing the image, archived URLs of the image file, and, if available, archived URLs of web pages to which the ad image is linked. Additionally, the dataset provides text data obtained from the banner ad images using optical character recognition (OCR). We expect the dataset to be useful for researchers across a variety of disciplines and fields such as visual culture, history, media studies, and business and marketing.</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

SEEFLEX - The Corpus of Secondary English as a Foreign Language (EFL) Exams

<blockquote> <p>This version of the corpus may have been superseded by the&nbsp;<a href="https://github.com/tobib92/SEEFLEX/" target="_blank" rel="noopener">GitHub repository</a></p> </blockquote> <p>&nbsp;</p> <p><strong>Abstract:</strong></p> <p>This report presents the&nbsp;<em>Corpus of Secondary School English as a Foreign Language (EFL) Exams (SEEFLEX)</em>. In Germany, upper secondary school EFL exams feature recurring tasks targeting diverse text types. The&nbsp;<em>SEEFLEX</em> was developed to investigate how students complete these tasks linguistically and whether they meet the curricular requirements. The corpus contains data from 575 transcribed authentic curriculum-based examinations (n<sub>texts</sub> = 1979, total = ~625.000 words). The metadata include standardized receptive vocabulary assessments, a cognition scale, the participants&rsquo; reading habits, social background, and their language experience and proficiency. Extensive xml mark-up was added to investigate the influence of inter alia source material, structural text features, and selected language mistakes. An online repository provides full-text access as well as ample additional resources, including an interactive Shiny application to investigate register variation in the corpus.</p>

opencc-by-nc-sa-4.0Oct 2024View details →
zenodo36/100

English-language articles related to petrophysics retrieved from the SSCI sub-database of Web of Science core database (Time: 2000.01.01-2022.12.31)

<p>According to the research content of petrophysics, petrophysics can be divided into eight branches: rock electricity, rock acoustics, rock nuclear physics, rock mechanics, rock thermophysics, rock nuclear magnetic resonance spectroscopy(NMR), rock gravimetry (density), and rock magnetism. This dataset presents the scientific literature retrieved by selecting the "Science Citation Index Expanded(SCI-EXPANDED)-- 1982-present" sub-library from the core collection database of Web of Science.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Polifonia Corpus - Encyclopedic Module Data - English Language

<p>We release the data of the Encyclopedic Module of the Polifonia Textual Corpus (Wikipedia pages), collected by selecting from <a href="http://lcl.uniroma1.it/babeldomains/">BabelNet domains</a> all the <a href="https://www.wikipedia.org">Wikipedia</a> musical pages.</p> <p>Full description at <a href="/github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Wikidata Dump English-language Songs

<p>RDF dump of wikidata produced with <a href="//wdumps.toolforge.org/">wdumper</a>.</p><p><br><a href="//wdumps.toolforge.org/dump/2782">View on wdumper</a></p><p><b>entity count<b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0</b></b></p>

opencc-zeroOct 2022View details →
zenodo36/100

Sustainable Development Goals (SDGs) in English as a Foreign Language (EFL)

<p>The qualitative mixed-method intervention study, grounded in content analysis and comparative methodology, reveals that incorporating SDGs into teacher training programmes is key since there is a significant direct impact on society. The objective was firstly to create an SDG-based didactic proposal including inquiry-based learning as its pedagogical approach for developing critical thinking. Secondly, to study its effect on participants regarding raising awareness of SDGs and their projection to society as future teachers.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

English Language Learners

<p>English Language Learners reports the number and percentage of English Language Learner (ELL) students, per grade.</p>

opencc-by-4.0May 2023View details →
ClinicalTrials.gov36/100

English as a Second Language Health Literacy Program

ClinicalTrials.gov study NCT04125680. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record