Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13
datasets available to search
ShareScore release 0.7.1
Dataset results
13 results for “word frequency”
Multi-LEX: a database of multi-word frequencies (French files)
<p>Written word frequency is a key variable used in many psycholinguistic studies and is central in explaining visual word recognition. Indeed, methodological advances on single word frequency estimates have helped to uncover novel language-related cognitive processes, fostering new ideas and studies. In an attempt to support and promote research on a related emerging topic, visual multi-word recognition, we extracted from the exhaustive Google Ngram datasets a selection of millions of multi-word sequences and computed their associated frequency estimate. Such sequences are presented with Part-of-Speech information for each individual word. An online behavioral investigation making use of the French 4-gram lexicon in a grammatical decision task was carried out. The results show an item-level frequency effect of word sequences. Moreover, the proposed datasets were found useful during the stimulus selection phase, allowing more precise control of the multi-word characteristics.</p>
Multi-LEX: a database of multi-word frequencies (English files)
<p>Written word frequency is a key variable used in many psycholinguistic studies and is central in explaining visual word recognition. Indeed, methodological advances on single word frequency estimates have helped to uncover novel language-related cognitive processes, fostering new ideas and studies. In an attempt to support and promote research on a related emerging topic, visual multi-word recognition, we extracted from the exhaustive Google Ngram datasets a selection of millions of multi-word sequences and computed their associated frequency estimate. Such sequences are presented with Part-of-Speech information for each individual word. An online behavioral investigation making use of the French 4-gram lexicon in a grammatical decision task was carried out. The results show an item-level frequency effect of word sequences. Moreover, the proposed datasets were found useful during the stimulus selection phase, allowing more precise control of the multi-word characteristics.</p>
Classification of word levels with usage frequency, expert opinions and machine learning
<p>This dataset includes classification of English words according to CEFR language levels. It can be used in various educational applications including determining levels of text that is appropriate for students learning English. </p> <p>For each word, part-of-speech, the word lemma and usage frequency is provided. For words that have no survey results, a machine learning based methodology is used to predict levels. These predictions are also included as a separate file. This data is released as part of the submission process to British Journal of Educational Technology Special Issue on Open Data.</p> <p>The included readme.pdf file contains a detailed description of data. </p>
Low-frequency Cortical Activity Reflects Context-dependent Parsing of Word Sequences
<p>The dataset and code used for the publication of "<a href="https://www.biorxiv.org/content/10.1101/2024.11.12.623335v1" target="_blank" rel="noopener"><strong>Low-frequency Cortical Activity Reflects Context-dependent Parsing of Word Sequences</strong></a>".</p> <div> <div><strong>Data path</strong></div> <div> <ol> <li>derivatives (./data/derivatives/): preprocessed data. <ul> <li>meg_quat-tsss-withEOG_100hz_0d3-40: meg data preprocessed at 100 Hz filtered between 0.3 and 40 Hz.</li> <li>FreeSurfer: anatomy data preprocessed</li> </ul> </li> <li>behavior_data (./behavior_data/)</li> <li>stimuli (./exp_procedure/material/): exp stimuli and procedure</li> <li>features (./data/features/): calculated phonological and linguistic features.</li> </ol> </div> <strong>Code path</strong></div> <div><br> <div>Preprocessing<strong>:</strong></div> <div> <ol> <li>[preprocessing.py]: preprocessing of behavior and MEG data.</li> <li>[recon.py]: reconstruction of T1 .niift data into surface data in FreeSurfer.</li> <li>[analysis_stimuli.ipynb]: preprocessing and feature extraction of stimuli.</li> </ol> </div> Analysis</div> <div> <ul> <li>[analysis_meg.ipynb]: spectral, time course, and RSA analysis.</li> </ul> </div>
Support for a novel, simple method for calculating word frequency of output on language production tasks
<p>Pre-recorded presentation for the International Workshop on Language Production 2021</p>
Frequency dataset for "Profile-based measures of lexical variation. Four case studies on variation in word choice between Belgian and Netherlandic Dutch."
<p>The dataset is structured according to:</p> <ul> <li>the lexical field (CLOTHING, TRAFFIC, IT, and EMOTION);</li> <li>the part of speech (noun or adjective);</li> <li>the corpus;</li> <li>the concept;</li> <li>the term.</li> </ul> <p>It first gives the absolute frequency as found in the corpus and also after it was disambiguated. The concept frequency and relative frequency is calculated based on the "disambiguated" absolute frequency.</p> <p>More information can be found in this dissertation:</p> <p>Daems, Jocelyne. 2022. <em>Profile-based measures of lexical variation. Four case studies on variation in word choice between Belgian and Netherlandic Dutch. </em>KU Leuven.</p>
Selected word frequencies for 20 short stories by three authors: Arthur Conan Doyle, Ernest Hemingway, Jack London
<p>Selected word frequencies for 20 short stories by three authors: Arthur Conan Doyle, Ernest Hemingway, Jack London. </p>
Word frequencies and text data for the study of Statistical Laws in Complex Systems
<p>The files contain the data reported in the manuscript "Statistical Laws in Complex Systems" and in the repository <a href="https://github.com/edugalt/StatisticalLaws">https://github.com/edugalt/StatisticalLaws</a></p> <p>The files contain list of word frequencies and filtered word texts, computed from the data provided in <a href="https://gutenberg.org/">Project Gutenberg</a>, <a href="https://en.wikipedia.org/wiki/Wikipedia:Database_download">Wikipedia</a>, and the <a href="https://books.google.com/ngrams/">Google N-gram database.</a></p>
Frequencies per million words for 5 epidemiologically relevant search terms in a dozen British 19th century newspapers
<p>COVID-19 is the first known coronavirus pandemic. Nevertheless, the seasonal circulation of the four milder coronaviruses of humans – OC43, NL63, 229E and HKU1 – raises the possibility that these viruses are the descendants of more ancient coronavirus pandemics. This proposal arises by analogy to the observed descent of seasonal influenza subtypes H2N2 (now extinct), H3N2 and H1H1 from the pandemic strains of 1957, 1968 and 2009, respectively. Recent historical revisionist speculation has focussed on the influenza pandemic of 1889-1892, based on molecular phylogenetic reconstructions that show the emergence of human coronavirus OC43 around that time, probably by zoonosis from cattle. If the "Russian influenza", as The Times named it in early 1890, was not influenza but caused by a coronavirus, the origins of the other three milder human coronaviruses may also have left a residue of clinical evidence in the 19th century medical literature and popular press. In this paper, we search digitised 19th century British newspapers for evidence of previously unsuspected coronavirus pandemics. We conclude that there is little or no corpus linguistic signal in the UK national press for large-scale outbreaks of unidentified respiratory disease for the period 1785 to 1890.</p>
Supporting data (pre-processed) for: "Frequency-tagged visual evoked responses track syllable effects in visual word recognition"
<p>Pre-processed data. </p> <p>Data were re-referenced off-line to the average of left and right mastoid electrodes, bandpass filtered from 5 to 100 Hz (4th order Butterworth filter) and then segmented to include 200 ms before and 2000 ms after stimulus onset. Epoched data were normalized based on a prestimulus period of 200 ms, and then evaluated according to a sample-by-sample procedure to remove noisy sensors that were replaced using spherical splines. Additionally, EEG epochs that contained data samples exceeding threshold (100 uV) were excluded on a sensor-by-sensor basis, including horizontal and vertical eye channels</p> <p>The original data:<br> Montani, Veronica. (2019). Supporting data for: "Frequency-tagged visual evoked responses track syllable effects in visual word recognition" [Data set]. Zenodo. http://doi.org/10.5281/zenodo.3260451</p>
Frequencies per million words for 5 epidemiologically relevant search terms in a dozen British 19th century newspapers
Open the record for dataset details and reuse information.
Frequency of positive words in grant applications
<p>The data was gathered to reproduce the methodology and findings presented in <a href="http://dx.doi.org/10.1136/bmj.l6573">Lerchenmueller et al. (2019)</a> in proposal texts that were submitted to different funding schemes offered by the Swiss National Science Foundation.</p> <h2>Data description</h2> <p>The data files (as .xlsx) contains three sheets:</p> <ol> <li>Career funding schemes (excluding fellowships): 1802 proposals included</li> <li>Spark funding scheme: 612 proposals included</li> <li>Project funding scheme (Projects): 5736 proposals included</li> </ol> <p>Each Sheet includes a data matrix where each row is a specific grant proposal. The unit of analysis are grant proposals and the texts used are the title and abstracts. The data is used in a project available from github (https://github.com/snsf-data/positive_language). [Paper soon to be submitted]</p> <p>The first 25 columns give us the 25 positive words and their respective counts in each of the analysed texts of the grant proposals. The <strong>positiv words </strong>(here the column names are) used are the following: amazing, assuring, astonishing, bright, creative, encouraging, enormous, excellent, favorable/ favourable, groundbreaking, hopeful, innovative, inspiring, inventive, novel, phenomenal, prominent, promising, reassuring, remarkable, robust, spectacular, supportive, unique, and unprecedented. Those were first proposed by <a href="https://www.bmj.com/content/351/bmj.h6467">Vinkers et al. (2015)</a>. </p> <p>Additionnally, the following columns are present in the data:</p> <ul> <li>sum_pos: The sum of the number of positive words in the texts. </li> <li>text_length and text_length100: the text length (count of words), and text length divided by 100.</li> <li>ResponsibleApplicantGender: the gender of the corresponding applicant (m or f)</li> <li>ResponsibleApplicantAge: the age of the corresponding applicant at submission (continuous)</li> <li>NationalityIsoCode: the nationality of the corresponding applicant (CH or not CH)</li> <li>IsApproved and IsFundable: binary variable indicating funding success, or whether the project would have been fundable given the grade with unlimited funding. </li> <li>Decision and CallYear: year of the call deadline, and year the funding decision was taken. </li> <li>ResearchInstitutionType: the type of institution the corresponding applicant is affiliated to (Cantonal University, ETH Domain, Other)</li> <li>which_lang: the language the proposal was written in (all english)</li> </ul> <h2>Text processing</h2> <p>The text corpus used in this analysis was also used as the basis for additional analyses. It therefore underwent a thorough cleaning with the help of the R-packages {tm} and {stringr}. After the pre-processing steps, the number of times each of the postitive words (see above) occured in the title and abstracts of the respective proposal is computed using a simple keyword search. </p> <h3>Pre-processing steps: </h3> <ul> <li>Punctuation was removed.</li> <li>All non-standard alphanumeric characters were removed. </li> <li>All characters were converted to lowercase. </li> <li>Extra white spaces were removed. </li> <li>Internet formatting was removed: URLs, email addresses, twitter formatting (words starting with # and @). </li> <li>Common English contractions were converted to their non-contracted form ("it's" --> "it is"). </li> <li>English language stopwords were removed. </li> </ul> <p> </p>
Supporting data for: "Frequency-tagged visual evoked responses track syllable effects in visual word recognition"
<p>The dataset consists of the original 17 .bdf files.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.