Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
103
datasets available to search
ShareScore release 0.9.0
Dataset results
103 results for “English language”
Knowledge from non-English-language studies broadens contributions to conservation policy and helps to tackle bias in biodiversity data
Open the record for dataset details and reuse information.
Supplementary material 1 from: Zermoglio PF, Plos A, Acosta N, Amaya L, Escobar DA, Grattarola F, Mancina CA, Nuñez F, Plata CA, Quintero E, Vargas M (2020) Latin American Plea for Incorporation of Other, Non-English Languages in TDWG Standards Documentation. Biodiversity Information Science and Standards 4: e58973. https://doi.org/10.3897/biss.4.58973
Signatories to the petition for incorporation of other languages to the Biodiversity Information Standards (TDWG) standards and documentation
Recurrent Neural Network Language Models Always Learn English-Like Relative Clause Attachment
<p>This repository contains the raw results (by word information-theoretic measures for the experimental stimuli) and the LSTM models analyzed in <a href="https://www.aclweb.org/anthology/2020.acl-main.179/">Recurrent Neural Network Language Models Always Learn English-Like Relative Clause Attachment</a>. The models from the synthetic experiments are given in the synthetic archive, as well as the training data generation script. There is a README included that gives more details for recreating/evaluating results from those experiments.</p> <p>The naming convention for each model in the models directory is:<br> [Language]_hidden[Hidden Units]_batch[Batch Size]_dropout[Dropout Rate]_lr[Learning Rate]_[Model Number].pt</p> <p>Language: en for English and es for Spanish<br> Hidden Units: All models had two layers with 650 hidden units per layer<br> Batch Size: The size of the batch (128 for English, 64 for Spanish)<br> Dropout Rate: All models used a dropout rate of 0.2<br> Learning Rate: All models has a learning rate of 20<br> Model Number: Identifier of the model (English model 0 is the best model from <a href="https://github.com/facebookresearch/colorlessgreenRNNs">Gulordava et al. (2018)</a>) </p> <p> </p> <p> </p>
Interpreting from a Language of Wider Communication to a Language of Narrower Communication: The Case of English to Moghamo
<p>This study set out to examine interpreting from English – language of wider communication – to Moghamo –language of narrower communication. To achieve this end, three objectives were set: 1) state the difficulties they face in performing their duty as interpreters; 2) assess the impact on effective communication in the Moghamo context; and 3) show the impact of their interpretations on Moghamo language and receptors of the interpreting. The data were collected from several sources: documentary from public and private libraries, through interviews and questionnaires and observant participation in churches. The collected data were analysed based on some major linguistic branches: phonology, semantics, morphology and lexicology. The goal here is to prove how Moghamo has been affected by code-switching by natural interpreters and Moghamo speakers. The research led to the following main findings: 1) Natural interpreters have little or no knowledge of the code of ethics governing the profession, and have not even undergone any formal training in the art of interpreting. The consequence is that the audience is generally misinformed and even exploited; 2) The code-switching and other factors contribute significantly to the plethora of English loanwords used regularly in Moghamo; and 3) The string of loanwords regularly used in Moghamo is detrimental because this threatens the very existence and survival of the latter as an independent language. The offshoot of this practice is the existence of what the researcher terms a hybrid Moghamo language or the birth of a completely new language in due course. Being already an endangered language, There is therefore an urgent need for something to be done to stop Moghamo from getting extinct.</p>
Survey among English Language Teachers of Kazakhstan
<p>Survey among English Language Teachers of Kazakhstan on ELT.</p>
English to Yoruba Language Translation
<p>Dataset on the English to Yoruba short message service speech and text translator for android phones.</p> <p> </p>
Topic model of English-language fiction, 1880-1999, with 200 topics.
<p>A topic model of 29,341 volumes of fiction, written in English and published between 1880 and 1999. The underlying corpus was organized by Ted Underwood for an experiment on period and cohort effects in cultural change. Metadata is in finalcorpus.tsv (which also has rows for 10 volumes not actually included in the model). </p> <p>To identify volumes as fiction, we relied on the NovelTM Dataset of English-Language Fiction (https://culturalanalytics.org/article/13147-noveltm-datasets-for-english-language-fiction-1700-2009). To confirm birth years of authors and publication dates of books, we compared NovelTM metadata both to the Chicago Novel Corpus and to a copy of the US Copyright Registry, digitized by the New York Public Library (https://github.com/NYPL/catalog_of_copyright_entries_project).</p> <p>The corpus itself is in cohort4.txt.gz; each line represents a roughly 10,000-word "chunk" of a document. The first 15% and last 5% of pages in each volume were discarded; the remaining pages were divided into chunks of roughly equal size. Chunk id is the first token on each line; it is formed by taking a HathiTrust volume id and adding an underscore + sequential integer (chunk number). Removing the underscore and integer produces a "document id" that can be paired to the metadata. The words in the line are not presented in original order; they are taken from HathiTrust Extracted Features, which records only page-level word counts.</p> <p>The topic model was produced using MALLET (http://mallet.cs.umass.edu/index.php), and has 200 topics.</p> <p>The top words in each topic are listed in the "keys" file; document-topic proportions are listed in "doctopics."</p> <p>For more information on the construction of the corpus and the experiment it is designed to support, see https://github.com/tedunderwood/period-cohort and/or a permanent Zenodo object created from that repository.</p>
Annotated corpus sample of resultative constructions in cooking recipes in 4 languages (Dutch, English, French, Spanish)
<p>This dataset contains an annotated corpus sample of 4000 (i.e. 1000 per language) occurrences of resultative constructions (e.g. cut the onion thin, whisk the egg whites to a foam, roll the dough into a ball) in 4 languages (Dutch, English, French, Spanish) retrieved from a comparable multilingual corpus of cooking recipes.</p>
The results of model learning on base an annotated text, compiled on the basis of the English-language news feed of the Yuri Gagarin State Technical University of Saratov
<p>The results of model learning on base an annotated text, compiled from the English-language news feed of the Yuri Gagarin State Technical University of Saratov.</p> <p>This file can be used in conjunction with Data for Model Learning on base OPENNLP DOI 10.5281/zenodo.3550016</p> <p> </p> <p> </p>
1000 Disease Ontology terms and their Wikidata mappings to 17 mostly Indian languages and English
<p>This dataset contains the result of a SPARQL query run on the <a href="https://query.wikidata.org/">Wikidata Query Service</a> on 13 February 2020 around 22:25 UTC. They are archived here as a means to determine progress with the coverage of disease-related terms in languages other than English, particularly in languages of India.</p> <p>The <a href="https://query.wikidata.org/#%23%20Wikidata%20items%20for%20concepts%20that%20have%20a%20Disease%20Ontology%20ID%20%28P699%29%0A%23%20sorted%20by%20number%20of%20sitelinks%0A%23%20optionally%20with%20their%20Wikidata%20label%20in%20English%0A%23%20optionally%20with%20their%20Wikidata%20label%20in%20Hindi%2C%20Bangla%20and%20Swahili%0A%23%20optionally%20with%20their%20Wikidata%20label%20in%20Marathi%2C%20Telugu%2C%20Eastern%20Punjabi%2C%20Western%20Punjabi%2C%20Gujarathi%2C%20Maithili%2C%20Kannada%2C%20Odia%2C%20Bhojpuri%2C%20Tamil%2C%20Nepali%2C%20Urdu%2C%20Malayalam%2C%20Esperanto%0A%0ASELECT%20DISTINCT%20%3Fitem%20%0A%23%20English%20Wikidata%20label%0A%20%3FLabelEN%0A%0A%23%20%20%20%20English%20Wikipedia%20article%20title%0A%3FPageTitleEN%20%0A%0A%23%20Hindi%2C%20%20%20Bangla%20and%20Swahili%20Wikidata%20labels%0A%3FLabelHI%20%3FLabelBN%20%20%20%20%3FLabelSW%20%0A%0A%23%20Marathi%2C%20%20Telugu%2C%20Eastern%20Punjabi%2C%20Western%20Punjabi%2C%20Gujarathi%2C%20Maithili%2C%20Kannada%2C%20%20%20%20Odia%2C%20Bhojpuri%2C%20%20%20Tamil%2C%20%20Nepali%2C%20%20%20%20Urdu%2C%20Malayalam%2C%20Esperanto%0A%20%3FLabelMR%20%3FLabelTE%20%20%20%20%20%20%20%20%20%3FLabelPA%20%20%20%20%20%20%20%20%3FLabelPNB%20%20%20%3FLabelGU%20%3FLabelMAI%20%3FLabelKN%20%3FLabelOR%20%20%3FLabelBH%20%3FLabelTA%20%3FLabelNE%20%3FLabelUR%20%20%20%3FLabelML%20%20%20%3FLabelEO%0A%0A%3Fsitelinks%20%0A%0AWHERE%20%7B%0A%20%20%3Fitem%20wdt%3AP699%20%5B%5D%20.%20%20%0A%20%20%3Fitem%20wikibase%3Asitelinks%20%3Fsitelinks%20.%0A%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3F%3FLabelEN%20filter%20%28lang%28%3FLabelEN%29%20%3D%20%22en%22%29%20.%20%7D%20%0A%20%20OPTIONAL%20%7B%20%3Farticle%20schema%3Aabout%20%3Fitem%20%3B%20schema%3AisPartOf%20%3Chttps%3A%2F%2Fen.wikipedia.org%2F%3E%20%3B%20%20schema%3Aname%20%3FPageTitleEN%20.%20%7D%0A%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelHI%20filter%20%28lang%28%3FLabelHI%29%20%3D%20%22hi%22%29%20.%20%7D%20%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelBN%20filter%20%28lang%28%3FLabelBN%29%20%3D%20%22bn%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelSW%20filter%20%28lang%28%3FLabelSW%29%20%3D%20%22sw%22%29%20.%20%7D%0A%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelMR%20filter%20%28lang%28%3FLabelMR%29%20%3D%20%22mr%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelTE%20filter%20%28lang%28%3FLabelTE%29%20%3D%20%22te%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelPA%20filter%20%28lang%28%3FLabelPA%29%20%3D%20%22pa%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelPNB%20filter%20%28lang%28%3FLabelPNB%29%20%3D%20%22pnb%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelGU%20filter%20%28lang%28%3FLabelGU%29%20%3D%20%22gu%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelMAI%20filter%20%28lang%28%3FLabelMAI%29%20%3D%20%22mai%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelKN%20filter%20%28lang%28%3FLabelKN%29%20%3D%20%22kn%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelOR%20filter%20%28lang%28%3FLabelOR%29%20%3D%20%22or%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelBH%20filter%20%28lang%28%3FLabelBH%29%20%3D%20%22bh%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelTA%20filter%20%28lang%28%3FLabelTA%29%20%3D%20%22ta%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelNE%20filter%20%28lang%28%3FLabelNE%29%20%3D%20%22ne%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelUR%20filter%20%28lang%28%3FLabelUR%29%20%3D%20%22ur%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelML%20filter%20%28lang%28%3FLabelML%29%20%3D%20%22ml%22%29%20.%20%7D%0A%20%20OPTIONAL%20%7B%20%3Fitem%20rdfs%3Alabel%20%3FLabelEO%20filter%20%28lang%28%3FLabelEO%29%20%3D%20%22eo%22%29%20.%20%7D%0A%7D%0AORDER%20BY%20DESC%28%3Fsitelinks%29%0ALIMIT%201000">SPARQL query</a></p> <ul> <li>was for <ul> <li> <p>Wikidata items for concepts that have a <a href="https://www.wikidata.org/wiki/Property:P699">Disease Ontology ID (P699)</a></p> <ul> <li> <p>sorted by number of sitelinks</p> </li> <li> <p>optionally with their Wikidata label in English</p> </li> <li> <p>optionally with their Wikipedia article title in English</p> </li> <li> <p>optionally with their Wikidata label in Hindi, Bangla and Swahili</p> </li> <li> <p>optionally with their Wikidata label in Marathi, Telugu, Eastern Punjabi, Western Punjabi, Gujarathi, Maithili, Kannada, Odia, Bhojpuri, Tamil, Nepali, Urdu, Malayalam, Esperanto</p> </li> </ul> </li> </ul> </li> <li>is contained in the file SPARQL.txt,</li> </ul> <p>whereas the results are available in several formats, as provided by the Wikidata Query Service:</p> <ul> <li>query.csv</li> <li>query.tsv<br> query.html</li> <li>query.json.txt (Zenodo produced an error upon trying to upload the file as query.json, so I renamed it, which worked fine).</li> </ul> <p>A simplified version of the SPARQL query can also be fed into the TABernacle tool that <a href="https://tools.wmflabs.org/tabernacle/#/tab/sparql/SELECT%20%09%0A%09%3Fitem%0AWHERE%20%09%0A%7B%0A%20%20%3Fitem%20wdt%3AP699%20%5B%5D%20.%20%20%0A%20%20%3Fitem%20wikibase%3Asitelinks%20%3Fsitelinks%20.%0A%0A%7D%0AORDER%20BY%20DESC(%3Fsitelinks)/Len%2Chi%2Cbn%2Csw%2Cmr%2Cte%2Cpa%2Cpnb%2Cgu%2Cmai%2Ckn%2Cor%2Cbh%2Cta%2Cne%2Cur%2Cml%2Ceo">represents</a> the live data in a way that facilitates editing the missing pieces.</p>
Exploring the Dominance of the English Language on the Websites of EU Countries
<p>This Dataset, in 29 files of xlsx format, contains the data of all metrics and accumulated information as they are described in the methodology, results and discussion section of the research article "Exploring the Dominance of the English Language on the Websites of EU Countries".</p>
THE PHENOMENON OF POLYSEMY IN ENGLISH AND UZBEK LANGUAGES
Open the record for dataset details and reuse information.
METHODOLOGY FOR IMPROVING THE EFFECTIVENESS OF STUDENTS LEARNING THE ENGLISH LANGUAGE.
Open the record for dataset details and reuse information.
ISOMORPHIC AND ALLOMORPHIC FEATURES OF COMPLEX SENTENCES WITH ADVERBIAL CLAUSES OF PLACE IN ENGLISH AND UZBEK LANGUAGES.
Open the record for dataset details and reuse information.
"THE IMPORTANCE OF TEACHING YOUNG AGE LEARNERS ENGLISH LANGUAGE"
Open the record for dataset details and reuse information.
THE IMPORTANCE OF MODERN TECHNOLOGIEIS IN ENGLISH LANGUAGE TEACHING.
Open the record for dataset details and reuse information.
THE ROLE OF HOME READING IN FORMING FOREIGN LANGUAGE SPEECH SKILLS IN ENGLISH LESSONS
Open the record for dataset details and reuse information.
THE LINGUACULTURAL CONCEPT OF DOUBT: A COMPARATIVE ANALYSIS OF ENGLISH AND UZBEK LANGUAGES
Open the record for dataset details and reuse information.
COMPARATIVE ANALYSIS OF THE CONCEPT "LOVE" IN THE PHRASEOLOGY OF ENGLISH AND UZBEKI LANGUAGES
Open the record for dataset details and reuse information.
FACTORS FOR ELIMINATING PROBLEMS IN THE TRANSLATION OF PREDICATIVE DEVICES IN ENGLISH INTO UZBEKI LANGUAGE
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.