Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

475

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

475 results for “word”

Learn how ShareScore rates datasets ↗
OpenNeuro52/100

Component processes of word reading in adults and children

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
zenodo52/100

PsPM-EWO: Eye tracker (including pupillometry) measurements from emotional-words tasks

<p>This dataset includes eye tracker (including pupillometry) measurements for 37 healthy unmedicated participants (25 females and 12 males, age range: 18 - 34 years, mean age: 26.2 +/- 4.7 years) participating in an emotional-words task. In each session, participants were presented with 50 neutral and 50 negative five-letter nouns from the Berlin Affective Word List Reloaded (V&otilde; et al., 2009). Also included are task information, keypress responses, keypress response times.</p>

opencc-by-4.0Oct 2020View details →
zenodo52/100

CLDF dataset of the Enggano word list from 1895 in Stokhof and Almanar's (1987) Holle List

<p>The repository for the digitised Enggano word list from 1895 (see Stokhof and Almanar 1987 for the original source) that has been matched with the <a href="https://engganolang.github.io/digitised-holle-list/">digitised Holle List</a> (Rajeg 2023a; cf. Stokhof 1980), providing the English and Indonesian glosses for the Enggano forms. The data set <a href="https://github.com/engganolang/holle-list-enggano-1895/actions/workflows/cldf-validation.yml">conforms</a> to the Wordlist module of the Cross-Linguistic Data Format (<a href="https://cldf.clld.org/">CLDF</a>) (Forkel et al. 2018).</p> <p><em>The work in this repository is part of the <a href="https://gtr.ukri.org/project/8AB0C3DC-F1C9-4CFA-BB4D-5BE748213372">AHRC-funded research</a> on <strong>Lexical resources for Enggano, a threatened language of Indonesia</strong> (visit the <a href="https://enggano.ling-phil.ox.ac.uk/">central webpage of the Enggano research</a> and the specific repository of the <a href="https://portal.sds.ox.ac.uk/Lexical_resources_for_Enggano">Lexical Resources for Enggano</a> project as well the <a href="https://portal.sds.ox.ac.uk/Enggano/groups">main Enggano repository</a> on the University of Oxford's Sustainable Digital Scholarship (SDS))</em></p> <h1>Updates in version 2.0.0</h1> <p>The following items summarise the major updates in version 2.0.0:</p> <ul> <li> <p><strong>Adding <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/main/cldf/media.csv">MediaTable</a></strong> to accommodate <a href="https://github.com/engganolang/holle-list-enggano-1895/tree/main/img">images</a> in/for note ID &lt;26&gt; (commits <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/dab95401f128bd4203a81294f2e9f4620d45b145">dab9540</a> &amp; <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/a0040038e2577cb37ff8ad3ac68e8bbdedc26291">a004003</a> <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/a0040038e2577cb37ff8ad3ac68e8bbdedc26291/code/Enggano-Holle-List-with-NBL.R#L229-L232">at this line</a> and <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/2aab3ab385fc82c0613de2aacb5947ca4981b417/code/Enggano-Holle-List-with-NBL.R#L370-L397">these lines</a>)</p> </li> <li> <p><strong>Splitting multiple forms in a cell</strong> into their own rows, both for the original list and the forms in the Notes (commit <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/39cdc663843b265aa8f3c5bbdcb11628fcc17b5e">39cdc66</a> at <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/39cdc663843b265aa8f3c5bbdcb11628fcc17b5e#diff-ac46f8a3edb85868970d77f55bb86c5f0449feb25c37295895c4a9e560564301R83">this line</a> and <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/39cdc663843b265aa8f3c5bbdcb11628fcc17b5e#diff-ac46f8a3edb85868970d77f55bb86c5f0449feb25c37295895c4a9e560564301R83">this line</a>, and commit <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/a0040038e2577cb37ff8ad3ac68e8bbdedc26291">a004003</a> at <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/a0040038e2577cb37ff8ad3ac68e8bbdedc26291#diff-5b59e4a74b953f80c8867a60c422bd8405c9d6236bd61239a0d3f20b0d582b78R69">this line</a>)</p> </li> <li> <p><strong>Orthography transliteration</strong> into Enggano's common orthography and IPA (across several commits and [closed] issues [#1 #3 #4 #5 #7], but see <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/2aab3ab385fc82c0613de2aacb5947ca4981b417/code/Enggano-Holle-List-with-NBL.R#L32-L93">these lines</a> for retrieving the existing orthography profile and doing the editing, and <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/2aab3ab385fc82c0613de2aacb5947ca4981b417/code/Enggano-Holle-List-with-NBL.R#L220-L306">these lines</a> for running the transliteration using the <a href="https://cran.r-project.org/web/packages/qlcData/index.html">qlcData</a> R package [Moran &amp; Cysouw 2018; Cysouw 2024])</p> <ul> <li>In the <a href="https://github.com/engganolang/holle-list-enggano-1895/blob/main/cldf/forms.csv">FormTable</a>, the <code>Form</code> column contains the Enggano forms in their common orthography; the <code>Value</code> column contains their original transcription/orthography, with their tokenised/segmented formats available under the <code>Graphemes</code> column; the <code>Segments</code> column, finally, contains the segmented IPA transliteration of the Enggano forms (cf. #6 ). The <code>Comment</code> column is derived from the contents of the Notes. It includes, if any, Enggano forms in their original transcription followed by their segmented/tokenised forms in IPA in square brackets, their glosses in English (<strong>EN</strong>) and/or Indonesian (<strong>ID</strong>) inside the bracket, and finally the ID of the Notes in the original document inside angular brackets. The <code>English</code> and <code>Indonesian</code> columns respectively are glosses of the given language from the master/main Holle List (Stokhof 1980) that has been digitised (Rajeg 2023a).</li> <li>The output files of the orthography profiling and transliteration (commit <a href="https://github.com/engganolang/holle-list-enggano-1895/commit/2aab3ab385fc82c0613de2aacb5947ca4981b417">2aab3ab</a>) are available in <a href="https://github.com/engganolang/holle-list-enggano-1895/tree/main/data-raw">data-raw</a> with the file names prefixed with <code>ortho-...</code>.</li> </ul> </li> </ul> <h2>References</h2> <p>Cysouw, Michael. 2024. qlcData: Processing Data for Quantitative Language Comparison. https://cran.r-project.org/web/packages/qlcData/index.html. (25 December, 2024). Version 0.3</p> <p>Forkel, Robert, Johann-Mattis List, Simon J. Greenhill, Christoph Rzymski, Sebastian Bank, Michael Cysouw, Harald Hammarstr&ouml;m, Martin Haspelmath, Gereon A. Kaiping &amp; Russell D. Gray. 2018. Cross-Linguistic Data Formats, advancing data sharing and re-use in comparative linguistics. Scientific Data. Nature Publishing Group 5(1). 180205. https://doi.org/10.1038/sdata.2018.205.</p> <p>Moran, Steven &amp; Michael Cysouw. 2018. The Unicode cookbook for linguists: Managing writing systems using orthography profiles (Translation and Multilingual Natural Language Processing 10). Berlin: Language Science Press. https://doi.org/10.5281/zenodo.1296780.</p> <p>Rajeg, Gede Primahadi Wijaya. 2023a. Digitised, Searchable Holle List in Stokhof (1980) [Data set]. (1.3.0). Zenodo. https://doi.org/10.5281/ZENODO.7972273. https://engganolang.github.io/digitised-holle-list/. https://ora.ox.ac.uk/objects/uuid:a511951b-86fb-4019-94d4-280efa83de02</p> <p>Rajeg, Gede Primahadi Wijaya. 2023b. CLDF dataset of the Enggano word list from 1895 in Stokhof and Almanar's (1987) Holle List [Data set]. https://github.com/engganolang/holle-list-enggano-1895 https://doi.org/10.25446/oxford.23515788</p> <p>Stokhof, W. A. L., ed. 1980. Holle Lists, Vocabularies in Languages of Indonesia, Vol. 1: Introductory Volume. Vol. Materials in Languages of Indonesia. Canberra, A.C.T., Australia: Dept. of Linguistics, Research School of Pacific Studies, The Australian National University. https://core.ac.uk/reader/159464813.</p> <p>Stokhof, W. A. L., and Alma E. Almanar. 1987. Holle Lists, Vocabularies in Languages of Indonesia, Vol. 10/3: Islands Off the West Coast of Sumatra. Vol. Materials in Languages of Indonesia. Pacific Linguistics (Series d) 76. Canberra, A.C.T., Australia: Dept. of Linguistics, Research School of Pacific Studies, The Australian National University. http://hdl.handle.net/1885/144589.</p>

opencc-by-sa-4.0Dec 2022View details →
zenodo52/100

Word Embedding of Amazon Product Review Corpus

<p>A word embedding of the <a href="https://www.cs.uic.edu/~liub/FBS/sentiment-analysis.html#datasets">Amazon Product Review Corpus</a> (<a href="https://www.doi.org/10.1145/1341531.1341560">Jindal and Liu, 2008</a>).</p> <p>Created using <a href="https://code.google.com/archive/p/word2vec/">Word2Vec</a> in CBOW mode, 500 dimensions and window size 5.</p> <p>Words have been lemmatised and particle verbs have been merged into a single token (e.g. <code>calm_down</code>).</p> <ul> </ul> <p>&nbsp;</p> <p><strong>Attribution</strong></p> <p>This dataset was created as part of the following publication:</p> <p>Marc Schulder,&nbsp;Michael Wiegand,&nbsp;Josef Ruppenhofer&nbsp;and&nbsp;Benjamin Roth&nbsp;(2017).&nbsp;<strong>&quot;Towards Bootstrapping a Polarity Shifter Lexicon using Linguistic Features&quot;</strong>. Proceedings of the 8th International Joint Conference on Natural Language Processing (IJCNLP). Taipei, Taiwan, November 27 - December 3, 2017.&nbsp;<a href="https://doi.org/10.5281/zenodo.3365609">DOI: 10.5281/zenodo.3365609</a>.</p> <p>If you use the data in your research or work, please cite the publication.</p>

opencc-by-4.0Nov 2017View details →
OpenNeuro48/100

Transposition confusability during visual word recognition

Open the record for dataset details and reuse information.

openCC0Jan 2019View details →
zenodo48/100

Hungarian word association network

<p>Hungarian word association network. The network was collected from the word association database ConnectYourMind, which collected associations online between 2008 and 2014, primarily in Hungarian. The network has 24580 nodes and 72709 links, where nodes represent words. A directed link from node A to B indicates, that word B was given as a response to word A in a free association task. The links are weighted according to the number of instances where the specific response was given. More details about the construction of the database can be found in English in (Kovacs et al., 2021) and exhaustively in Hungarian in (Kovacs, 2013). The network is shared in edgelist format. Each row has three values A;B;C. Each row indicates a directed link from node/word A to node/word B with a weight of C. The file uses utf encoding to represent Hungarian characters. Data available according to Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license</p> <p><br> Kovacs L. Fogalmi rendszerek es lexikai halozatok a ment&aacute;lis lexikonban. 2., atdolgozott, bovitett kiadas. (In Hungarian) Budapest: &nbsp;Tinta. 2013.</p> <p>Kovacs L, Bota A, Hajdu L, Kresz M. Networks in the mind - what communities reveal about the structure of the lexicon. Open Linguistics. 2021 Jan 1;7(1):181-99.</p> <p>Kov&aacute;cs L, B&oacute;ta A, Hajdu L, Kr&eacute;sz M. Brands, networks, communities: How brand names are wired in the mind. PLoS ONE 2022 17(8): e0273192. https://doi.org/10.1371/journal.pone.0273192</p> <p>When using the data, please give a reference to the data itself and to at least one of above mentioned publications.&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo48/100

Diachronic word embeddings from 19th-century newspapers digitised by the British Library (1800-1919)

<p>Word vectors related to the paper&nbsp;<em>Machines in the media: semantic change in the lexicon&nbsp;of mechanization in 19th-century British newspapers&nbsp;</em>by Nilo Pedrazzini and Barbara McGillivray (2022).</p> <p>The embeddings were trained on a 4.2-billion-word corpus of 19th-century British newspapers using Word2Vec and the following parameters:</p> <pre><code>sg = True min_count = 1 window = 3 vector_size = 200 epochs = 5</code></pre> <p>The embeddings&nbsp;are divided into periods of ten years each, with the vectors from each decade aligned to the ones from the most recent decade (1910s) using Orthogonal Procrustes.</p> <p>See related GitHub repository for the full documentation:&nbsp;<a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a></p> <p>Project webpage (Living with Machines):&nbsp;<a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p>

opencc-by-4.0Oct 2022View details →
zenodo48/100

Multi-LEX: a database of multi-word frequencies (French files)

<p>Written word frequency is a key variable used in many psycholinguistic studies and is central in explaining visual word recognition. Indeed, methodological advances on single word frequency estimates have helped to uncover novel language-related cognitive processes, fostering new ideas and studies. In an attempt to support and promote research on a related emerging topic, visual multi-word recognition, we extracted from the exhaustive Google Ngram datasets a selection of millions of multi-word sequences and computed their associated frequency estimate. Such sequences are presented with Part-of-Speech information for each individual word. An online behavioral investigation making use of the French 4-gram lexicon in a grammatical decision task was carried out. The results show an item-level frequency effect of word sequences. Moreover, the proposed datasets were found useful during the stimulus selection phase, allowing more precise control of the multi-word characteristics.</p>

opencc-by-4.0Oct 2022View details →
zenodo48/100

Multi-LEX: a database of multi-word frequencies (English files)

<p>Written word frequency is a key variable used in many psycholinguistic studies and is central in explaining visual word recognition. Indeed, methodological advances on single word frequency estimates have helped to uncover novel language-related cognitive processes, fostering new ideas and studies. In an attempt to support and promote research on a related emerging topic, visual multi-word recognition, we extracted from the exhaustive Google Ngram datasets a selection of millions of multi-word sequences and computed their associated frequency estimate. Such sequences are presented with Part-of-Speech information for each individual word. An online behavioral investigation making use of the French 4-gram lexicon in a grammatical decision task was carried out. The results show an item-level frequency effect of word sequences. Moreover, the proposed datasets were found useful during the stimulus selection phase, allowing more precise control of the multi-word characteristics.</p>

opencc-by-4.0Oct 2022View details →
zenodo48/100

Behavioral and fMRI Data: Nurturing the reading brain: Home literacy practices are associated with children's neural response to printed words through vocabulary skills

<p>This is the behavioral and fMRI dataset described in &quot;Nurturing the reading brain: &nbsp;Home literacy practices are associated with children&rsquo;s neural response to printed words through vocabulary skills&quot;.&nbsp;</p> <p>Because of anonymization concerns within&nbsp;the framework of EU privacy regulations (<a href="https://gdpr-info.eu">GDPR</a>), we cannot provide raw MRI data. Therefore, the fMRI data consists of individual&nbsp;pre-processed volumes, normalized into the MNI&nbsp;template (see paper for details about the preprocessing pipeline). Anonymized behavioral data and first level analyses are also provided for each participant (SPM.mat file as well as beta, con, spmT, RPV and ResMS&nbsp;files). Note that the dataset&nbsp;also include runs and GLM results for a third task (Dots) that was not analyzed in the paper. Finally, the <a href="https://www.psychopy.org">PsychoPy</a> implementation of the tasks is also provided. If you have any questions, please send an email to jerome.prado [at] univ-lyon1.fr.&nbsp;</p> <p><strong>IMPORTANT:</strong></p> <p>In accordance with EU privacy regulations, we ask that you sign and return a Data Use Agreement (DUA) before downloading the data. You can download the DUA&nbsp;<a href="https://zenodo.org/record/4965716/files/DUA.pdf?download=1">here</a>. Please, sign it and send it to jerome.prado [at] univ-lyon1.fr.</p>

opencc-by-4.0Jul 2021View details →
zenodo48/100

MEWL: Few-shot multimodal word learning with referential uncertainty

<p><strong>Dataset Release for <a href="https://arxiv.org/abs/2306.00503">MEWL: Few-shot multimodal word learning with referential uncertainty&nbsp;(ICML 2023)&nbsp;</a></strong></p> <p><strong>GitHub:</strong> <a href="https://github.com/jianggy/MEWL">https://github.com/jianggy/MEWL</a></p> <p><strong>Abstract: </strong>Without explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This word learning capability is believed to be the most fundamental building block of multimodal understanding and reasoning. Despite recent advancements in multimodal learning, a systematic and rigorous evaluation is still missing for human-like word learning in machines. To fill in this gap, we introduce the MachinE Word Learning (MEWL) benchmark to assess how machines learn word meaning in grounded visual scenes. MEWL covers human&#39;s core cognitive toolkits in word learning: cross-situational reasoning, bootstrapping, and pragmatic learning. Specifically, MEWL is a few-shot benchmark suite consisting of nine tasks for probing various word learning capabilities. These tasks are carefully designed to be aligned with the children&#39;s core abilities in word learning and echo the theories in the developmental literature. By evaluating multimodal and unimodal agents&#39; performance with a comparative analysis of human performance, we notice a sharp divergence in human and machine word learning. We further discuss these differences between humans and machines and call for human-like few-shot word learning in machines.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Investigating the Effects of Embodiment on Emotional Categorization of Faces and Words in Children and Adults

<p>The three data files uploaded here contain the data used for the analyses in experiments 1a, 1b, and 2 as described in the article carrying the same title as this dataset, published in the journal Frontiers in Psychology. All analyses were carried out in SPSS version 22 as described in the published article.</p> <p>Article Abstract:</p> <p>The facial feedback hypothesis (FFH) indicates that besides being involved in the production of facial expressions, the musculature of the face also influences one&rsquo;s perception of emotional stimuli. Recently, this effect has been the focus of increased scrutiny as efforts to replicate a key study with adult participants supporting this hypothesis, using the so-called &ldquo;pen-in-the-mouth&rdquo; task, have not been successful at several labs. Our series of experiments attempted to investigate whether the assumed embodiment effect can be reproduced in a simplified emotional categorization task for emotional faces and words. We also wanted to test whether the embodiment effect can be detected in children because it is assumed that their bodily processes are especially closely linked with their sensory and cognitive processes. Our experiments involved child and adult participants categorizing faces and words as positive or negative as quickly as possible, while inducing a positive or negative facial or bodily state (holding a straw in the mouth such that a smile or a frown was generated, or creating a positive or negative body posture). The positive or negative facial and bodily states could therefore be either congruent or incongruent with the valence of the target face and word stimuli. Our results did not show any significant differences between the congruent and incongruent conditions in either children or adults. This suggests that embodiment effects either do not significantly impact valence-based categorization or are not strong enough to be detected by our approach considering the sample size in the present study.</p>

opencc-by-4.0Jan 2020View details →
zenodo44/100

Ogham Words

<p><strong>Ogham Words</strong></p> <p><em><code>words.csv</code></em>&nbsp;describes the words which are the basis for the extraction of the inscriptions.</p> <p>more at:&nbsp;<a href="https://github.com/ogi-ogham/oghamextractor/tree/master/words">https://github.com/ogi-ogham/oghamextractor/tree/master/words</a></p>

opencc-by-4.0Jan 2020View details →
zenodo44/100

ANPERC-source/SEG_Annual: The database of words and affiliations of the SEG Annual Conferences (1982 - 2019)

<p>The database of words and affiliations of the SEG Annual Conferences (1982 - 2019) v1.0 This repository includes data for the words and phrases frequency of occurrence analysis &quot;SEGgrams.sqlite&quot; and the database &quot;SEG_affiliations_data.sqlite&quot; consisting of the industry companies and academia of different countries that presented their research during the Society of Explorational Geophysicists Annual Conferences (1982 - 2019) with the corresponding number of affiliations for the whole period of study.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Word-in-Context Target Sense Verification

<pre>Formally, WiC is framed as a&nbsp;<strong>binary classification</strong>&nbsp;task. Each instance in WiC-TSV consists of a target word&nbsp;<em>w</em>&nbsp;with a corresponding target sense&nbsp;<em>s</em>&nbsp;represented by either its definition (subtask 1) or its hypernym/s (subtask 2), and a context&nbsp;<em>c</em>&nbsp;containing the target word&nbsp;<em>w</em>. The task aims to determine whether the meaning of the word&nbsp;<em>w</em>&nbsp;used in the context&nbsp;<em>c</em>&nbsp;matches the target sense&nbsp;<em>s</em>. In the following table there are some examples from the dataset. </pre> <p>&nbsp;</p> <p>Subtasks</p> <p>&nbsp;WiC-TSV has&nbsp;<strong>three subtasks</strong>&nbsp;- participants can submit results in any of the subtasks:</p> <p>Subtask 1: Definitions</p> <p>In Subtask 1 systems make use of&nbsp;<strong>definitions</strong>&nbsp;for deciding whether the target word in context corresponds to the given definition or not.</p> <p>Subtask 2: Hypernyms</p> <p>In Subtask 2 systems make use of&nbsp;<strong>hypernymy</strong>&nbsp;information for deciding whether the target word in context is a hyponym of the given hypernym or not.</p> <p>Subtask 3: Definitions + Hypernyms</p> <p>In subtask 3 systems can make use of&nbsp;<strong>both</strong>&nbsp;sources of information, i.e., definitions and hypernyms.</p>

opencc-by-4.0Mar 2020View details →
zenodo44/100

Supplementary material accompanying "Factoring lexical and phonetic phylogenetic characters from word lists"

<p>This repository contains the scripts and the data that were used to run the analyses for the paper &quot;Factoring lexical and phonetic phylogenetic characters from word lists&quot;. For details, please refer to the README.md file provided along with the dataset. If you run into problems replicating the analysis, please do not hesitate to contact the authors.</p>

opencc-by-4.0Nov 2015View details →
zenodo44/100

Word embeddings learnt on MEDLINE abstracts

<p>Accompanying a preprint manuscript and code repository, this folder contains both raw text data and learnt word embeddings. The data source is the set of MEDLINE articles published on or after 2000. Preprocessing consists of extraction of each article's title and abstract and some minor text processing. The result is a corpus of 10.5 million documents in a single 14 GB file. </p> <p>word2vec and fastText are used to learn word embeddings on this corpus and three sets of word embeddings are shared here: 1) word2vec skip-gram, 2) word2vec CBOW, and 3) fastText skip-gram. All three sets use the default parameters of the software (e.g. context=5) with the exception of hierarchical softmax optimization and dimension=200.</p> <p>Preprint manuscript: https://arxiv.org/abs/1705.06262<br> GitHub repository: https://github.com/vincentmajor/ctsa_prediction</p>

opencc-by-sa-4.0Jun 2017View details →
zenodo44/100

Phlorest phylogeny derived from Dunn et al. 2011 'Evolved structure of language shows lineage-specific trends in word-order universals'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Dunn M, Greenhill SJ, Levinson SC &amp; Gray RD. 2011. Evolved structure of language shows lineage-specific trends in word-order universals. Nature, 473(7345), 79-82.</p> </blockquote>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Original Recording of Freiburg Words for Testing Hearing with Speech

<p>The files contain recordings of the monosyllabic words, polysyllabic numbers, and CCITT noise (ITU, 1993) of the Freiburg Speech Test (Hahlbrock, 1953). The words are given in DIN 45621-1:1995. The 400 monosyllables are organized in 20 test lists à 20 nouns. They are stored in "Einsilbige Wörter.zip". The name of each wav-file includes the number of the test list (L01, L02, …) and the position of the word within the test list (W01, W02, …). The 100 polysyllables are organized in 10 test lists à 10 numbers. They are stored in "Mehrsilbige Wörter (Zahlen).zip" with similar nomenclature. DIN 45626-1:1995 describes the recordings. Speech signals were recorded in 1969 in the studios of Norddeutscher Rundfunk, Hamburg, Germany, with the speaker Claus Wunderlich (Brinkmann, 1974). The recordings were processed by Physikalisch-Technische Bundesanstalt, Braunschweig, Germany, and Polygram International, Hannover, Germany. Later, the recordings were digitalized by Siemens AG and distributed on compact disc by "Siemens Audiologische Technik GmbH" (legal successor WS Audiology A/S) with the title "Wörter für Gehörprüfung mit Sprache" ("Words for hearing tests with speech") under item no. 7970155. The words were cut out as accurately as possible, i.e., with as little background before and after each word as possible (see Winkler and Holube, 2016a). The CCITT noise was originally included for calibration purposes, but is often used as noise masker when the monosyllables are presented in background noise.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Sets of period sets for words of length n.

<p>We consider finite words of length n. Each word has a set of periods, but many words can have the same set of periods. For a definition of a period of a word, see [1]. A set of periods is a subset of the set {0, 1, ..., n-1}, but not all subsets of {0, 1, ..., n-1} are period sets. The set denoted Gamma(n) contains all possible period sets corresponding to at least one word of length n. For a definition of Gamma(n), see references [1] or [2]. For more details see reference [4].</p> <p>The series of text files provide the list of period set, one per line, for Gamma(n), for n = 1, 2, ..., 100.<br>Each line contains a list of integers sorted by increasing value: this list constitutes one period set. The &nbsp;separator symbol is a space. Hence, the number of non-empty lines in a file gives the cardinality of Gamma(n).<br>The period set are sorted by their basic period. &nbsp;For a definition of the notion of basic period, see [1] or [2].</p> <p>The files for n=61, ..., 100, where computed with the incremental algorithm described in [5].</p> <p>The sequence of the cardinalities of the set Gamma(n), is also called, the Number of distinct autocorrelations of binary words of length n, and corresponds to the sequence A005434 in the Encyclopedia of Integer Sequences (EOIS) link [3].</p> <p>The files have generic name formatted as follows: gamma.n.bps<br>where n is the word length, for n = 1, 2, ..., 100.</p> <p><strong>References</strong>:</p> <p>1. Eric Rivals, Sven Rahmann.<br>&nbsp; &nbsp;Combinatorics of Periods in Strings.<br>&nbsp; &nbsp;Proc. 28th International Colloquium on Automata, Languages, and Programming, Lecture Notes in Computer Science vol. 2076, p. 615-26., P. Orejas, P. G. Spirakis, J. van Leuween editors, Springer Verlag, Berlin, 2001.<br>&nbsp; &nbsp;doi: <a title="Publication 1" href="https://doi.org/10.1007/3-540-48224-5_51" target="_blank" rel="noopener">https://doi.org/10.1007/3-540-48224-5_51</a><br>2. Eric Rivals, Sven Rahmann.<br>&nbsp; &nbsp;Combinatorics of Periods in Strings.<br>&nbsp; &nbsp;Journal of Combinatorial Theory - Series A, 104(1), p. 95-113, October 2003<br>&nbsp; &nbsp;doi: <a title="Publication 2" href="https://doi.org/10.1016/s0097-3165(03)00123-7" target="_blank" rel="noopener">https://doi.org/10.1016/s0097-3165(03)00123-7</a><br>3. Entry A005434 from The On-Line Encyclopedia of Integer Sequences.<br>&nbsp; &nbsp;URL: <a title="Sequence entry A005434" href="https://oeis.org/A005434" target="_blank" rel="noopener">https://oeis.org/A005434</a><br>4. Autocorrelation of Strings. A comment on entries A005434 and A045690 of the Encyclopedia of Integer Sequences.<br>&nbsp; &nbsp;URL: <a title="Introduction to period sets (webpage)" href="https://www.lirmm.fr/~rivals/RESEARCH/PERIOD/" target="_blank" rel="noopener">https://www.lirmm.fr/~rivals/RESEARCH/PERIOD/</a><br>&nbsp;5. Eric Rivals. Incremental computation of the set of period sets. arXiv:2410.12077,&nbsp; 2024. <a title="Publication 5" href="https://doi.org/10.48550/arXiv.2410.12077" target="_blank" rel="noopener">https://doi.org/10.48550/arXiv.2410.12077</a><br><br></p>

opencc-by-4.0Sep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record