Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
475
datasets available to search
ShareScore release 0.9.0
Dataset results
475 results for “word”
Buddhist Chinese Word Embeddings
<p>Buddhist Chinese word embeddings trained with FastText on the Buddhist texts present in the Kanseki repository.</p> <p>There are four models present here (and the full binary output for one), each differing in how the Kanseki repository was segmented into tokens. The "chinese_model" was segmented into 1-grams (individual characters). The "chinese_model_word" was segmented into words using a dictionary of buddhist terms and phrases not found were segmented with the classical Chinese word segmenter distributed with the stanza python library (https://stanfordnlp.github.io/stanza/available_models.html). The "chinese_model_hybrid_char_term" was segmented with a glossary of Buddhist terms and phrases not found were divided into 1-grams. "chinese_model_hybrid_char_term_2" uses a more extensive glossary of Buddhist terms and missing sections are divided into 1-grams.</p> <p>All models are default 100 dimensional FastText models created for this pilot study: Felbur, Rafal, Marieke Meelen & Paul Vierthaler (2022), 'Crosslinguistic Semantic Textual Similarity of Buddhist Chinese and Classical Tibetan' in <em>Journal of Open Humanities Data</em>.</p> <p>This research was done with generous funding from the Open Philology project. This project (running 2018–2022) is funded by the European Research Council (ERC) under the Horizon 2020 program (Advanced Grant agreement No 741884). It is based at the Leiden University Institute for Area Studies.</p>
Social Sciences Word Embeddings in FastText
<p>These social science word embeddings in FastText have been created from 37,604 open access social science research papers from the social science access repository (https://www.gesis.org/ssoar/home). They are available in German and English.</p> <p>(skipgram model, n-grams with n≥3 and n≤6, different dimensions (100, 150, 200, 300, 500), five epochs, learning rate 0.05, five negative examples)</p> <p>Please cite:</p> <p>Schiffers, Ricardo, Dagmar Kern, and Daniel Hienert. 2022. "Evaluation of Word Embeddings for the Social Sciences." In <em>Proceedings of the 6th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature</em>, edited by Stefania Degaetano, Anna Kazantseva, Nils Reiter, and Stan Szpakowicz, 1-6. Gyeongju: Association for Computational Linguistics. <a href="https://aclanthology.org/2022.latechclfl-1.1">https://aclanthology.org/2022.latechclfl-1.1</a>.</p>
Tigrinya Analogy Test for evaluating Word Embeddings
<p><strong>Tigrinya Analogy Test for evaluating Word Embeddings</strong></p> <p>This is a Tigrinya version of the Google Analogy Test set, which is used to evaluate English word-embedding models. The analogy test is a well-established strategy to empirically evaluate the quality of word-embedding models. More information about the English task can be found at the <a href="https://aclweb.org/aclwiki/Google_analogy_test_set_(State_of_the_art)">ACL Wiki</a>.</p> <p>This data is was first machine translated then manually verified by a native speaker to reduce errors.</p> <p>Some aspects of the original analogy test is focused on English and may not transfer well to other languages, such as those related to grammar or morphology. Therefore, we have discarded examples that became irrelevant in Tigrinya when adapting the task. Finally, there are a total of <strong>18465</strong> entries in the Tigrinya Analogy Test set, while the source English data has <strong>19544</strong> entries.</p> <p>An entry is dropped if the translations led to one of the following conditions:</p> <ol> <li>If the source word pair map to one Tigrinya word, for example, lucky & luckiest both correspond to ዕድለኛ.</li> <li>If the source word results in a multi-word expression. For example, grandson (ወዲ ጓል / ወዲ ወዲ), granddaughter (ጓል ጓል / ጓል ወዲ). This because the typical word-embedding approaches such as <em>word2vec</em> are not designed to predict multi-word phrases.</li> </ol> <p> </p> <p><strong>Test Sections</strong></p> <p>The test includes a series of semantic and syntactic analogies divided up into subsections including world capitals, currencies, family, tense, and plurality. The test contains the following sections:</p> <ol> <li>capital-world</li> <li>currency</li> <li>city-in-state</li> <li>family</li> <li>gram1-adjective-to-adverb</li> <li>gram2-opposite</li> <li>gram3-comparative</li> <li>gram4-superlative</li> <li>gram5-present-participle</li> <li>gram6-nationality-adjective</li> <li>gram7-past-tense</li> <li>gram8-plural</li> <li>gram9-plural-verbs</li> </ol> <p> </p> <p><strong>Examples:</strong></p> <ul> <li>Semantic section of World Capitals: “ኣስመራ: ኤርትራ as ፓሪስ: ?” and if the model responds correctly it will return: “ፈረንሳ”.</li> <li>Semantic section of Family section: “ሰብኣይ: ሰበይቲ as ወዲ: ጓል”.</li> <li>Syntax section with tense, a sample analogy might be “Walk: Walked as Run: Ran”.<br> </li> </ul> <p><strong>Evaluation</strong></p> <p>The final accuracy of a model is the proportion of the questions that the model answers correctly.<br> Generally, a better-quality model would answer more questions correctly than a model of lower quality.<br> However, note that a model with low performance on this analogy test, might still contain useful information, but may not be robust or good enough for more complex tasks.</p> <p> </p> <p><strong>Limitations</strong></p> <ul> <li>The analogy test could be a good indicator of the quality of word-embeddings, but it should be used with caution when comparing models trained on varying domains of data. It shall not be expected to generalize equally to all domains.</li> <li>The final score can be affected by the size, vocabulary, and domain of the text with which the models are trained on. For example, this may not be a good benchmark to compare models trained on news text <em>vs</em> posts on social media.</li> <li>Even though a manual sanity check was performed, we note that the semi-automatic construction of the Tigrinya test set might contains errors. If you discover any, you are welcome to contribute back by either opening an <em>Issue</em> at the GitHub repo, <a href="https://github.com/fgaim/tigrinya-analogy-test">https://github.com/fgaim/tigrinya-analogy-test</a>.</li> </ul> <p> </p> <p><strong>Citation</strong></p> <p>If you use this resource in your research, please cite it accordingly.</p>
Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words
<p>Hi,KIA dataset is a shared short Wakeup Word database focusing on perceived emotion in speech The dataset contains <strong>488 </strong>Wakeup Word speech. </p> <p>For more detailed information about the dataset, please refer to our paper: Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words</p> <p><strong>File Description</strong></p> <ul> <li><em><strong>wav/</strong></em>: wav files. <ul> <li>Filename f`{gender}_{pid}_{scene}_{trial}_{emotion}.wav` The first letter was used to express emotion.<br> </li> </ul> </li> <li><em><strong>annotation/</strong></em>: Information related to annotation and human validation of the entire speech</li> <li> <p><em><strong>split</strong></em>: 8fold data split with {train, valid, test}.csv </p> </li> <li> <p><em><strong>handcraft:</strong></em> Features used for data EDA and baseline performance</p> </li> <li> <p><em><strong>best_weights:</strong></em> wav2vec2.0 context network finetuning weights for re-implementation. Due to file size, we attach only fold M1, F5</p> </li> </ul> <p> </p> <p><strong>Reference</strong></p> <ul> </ul> <p>Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words [[ArXiv](https://arxiv.org/abs/2211.03371)]</p> <p>```<br> @inproceedings{kim2022hi,<br> title={Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words},<br> author={Taesu Kim, SeungHeon Doh, Gyunpyo Lee, Hyung seok Jun, Juhan Nam, Hyeon-Jeong Suk},<br> booktitle={Proceedings of the 14th Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)},<br> year={2022}<br> }<br> ```</p>
Data for "Researchers and their data. A study based on the use of the word data in scholarly articles"
<p><em>Data</em> is one of the most used terms in scientific vocabulary. This article focusses on the relationship between data and research by analyzing the contexts of occurrence of the word <em>data</em> in a corpus of 72,471 research articles (1980-2012) from two distinct fields (Social sciences, Physical sciences). The aim is to shed light on the issues raised by research on data, namely the difficulty of defining what is considered as data, the transformations that data undergo during the research process and how they gain value for researchers who hold them. Relying on the distribution of occurrences throughout the texts and over time, it demonstrates that the word <em>data </em>mostly occurs at the beginning and at the end of research articles. Adjectives and verbs accompanying the noun <em>data</em> turn out to be even more important than <em>data</em> itself in specifying data. The increase in the use of possessive pronouns at the end of the articles reveals that authors tend to claim ownership of their data at the very end of the research process. Our research demonstrates that even if data handling operations are increasingly frequent, they are still described with imprecise verbs that do not reflect the complexity of these transformations.</p>
Lubrang Brokpa Lexicon - basic word list
<p>These sound files constitute the elicitation of the lexical entries of the Basic Word List in the Lubrang variety of Brokpa. Lubrang village is a recent (early 20th century) settlement of Brokpa speaker originating in Sakteng village of Bhutan. They left Bhutan due to its heavy taxation of the semi-nomadic Brokpa households and its anti-Gelukpa policies and settled in the then Tibetan-administered area on land belonging to the Khispi people of Lish village, to which they continue to pay an annual tax. Lubrang Brokpa should hence be close to Merak and Sakteng (Bhutan) Brokpa, and not so close to Nyukmadung and Senge Brokpa spoken closer by. Because of the speaker’s paternal background there may be some admixture with Dirang Tshangla.</p>
Red Gelao (Shajing) audio word list
<p>On July 27, 2012 in Fengyan, Shajing Township, Qianxi County, Guizhou, China, I recorded the last remaining Red Gelao speaker of Shajing Township, Li Tingju 李庭举 (age 80). She was accompanied by her middle-aged daughter, Li Zhongying 李忠英. Both were illiterate and could not even read or write their own names. Li Tingju was not easy to work with, as she declined to repeat each word three times into my microphone and only wanted to say random words that would randomly come to her memory. She was not fluent in Gelao, and only knew possibly a few dozen to a few hundred vocabulary words, as well as a few songs. Altogether, I only obtained 3 songs and less than 20-30 vocabulary words from her.</p> <p> </p>
Figure 3. Using the ADX agent for obtaining definitions, synonyms and antonyms, for a given word, during MS Word editing-ADX – Agent for Morphologic Analysis of Lexical Entries in a Dictionary
<p>In figure 3 we present a capture screen of using the ADX1 agent in editing the text in<br> Microsoft Word. It displays the definition of the current word, but it also generates the synonyms,<br> used in the application of some web searching rules, used by another intelligent agent, called ASR.<br> The ASR agent will automatically compose some search strings to use for a web search engine, like<br> Google, Yahoo, or other. For example, if the user will search the word zăpadă (snow) on Yahoo<br> search engine, then one can also generate a search for the word nea (synonym of zăpadă) by using<br> the following ASR rule:<br> # Yahoo search<br> IF<br> http://search.yahoo.com/search?p=^X^&fr=yfp-t-309&toggle=1&cop=mss&ei=UTF-8<br> THEN<br> http://search.yahoo.com/search?p=^Clasa(X)^&fr=yfp-t-309&toggle=1&cop=mss&ei=UTF-8</p>
Figure 1. Associations between objects and words-Intelligent Agent for Acquisition of the Mother Tongue Vocabulary
<p>So, the learning process is very complex, and it is a probabilistic one. Thus, a second<br> example is required, and even more than that. Normally, the mother speaks naturally to her child,<br> she speaks with love and affection, she doesn’t “judge” or “program” what to say to her child.<br> Therefore, another day she will tell her child, for example, “My darling son, let’s drink the milk<br> from the cup.”</p>
Table 1: The Number of Words and Lessons in Book Three
<p>Three hundred and thirty three lexical items included in the English textbook, namely,<br> Birjandy, P. & Norouzi, M. & Mahmoody, G. (2004). English Book Three, which is routinely<br> assigned by the Ministry of Education for EFL teaching in Iranian public high schools in Grade<br> three were taught to the participants in the present study. The book typically includes sections of<br> reading comprehension, grammar, pronunciation practice, dialogues, and a list of new words (with<br> no explanation of meaning or synonyms) in the ending part of each lesson. These wordlists (See<br> Appendix 1) were used for vocabulary instruction and subsequent vocabulary testing in the study.<br> To save space, the number of the new words in each lesson is presented in Table 1</p>
Attempted Speech on Single Words [intracortical array][T16][BG2]
<h2>Summary of the Data and Experiment</h2> <h3>General Overview</h3> <p>The data provided come from an attempted speech experiment involving a participant (T16) in the BrainGate2 clinical trial. The dataset includes neural recordings from four microelectrode arrays implanted in the left precentral gyrus. The primary focus is on the array located in the ventral precentral gyrus (6v), which is associated with speech-related activity.</p> <h3>Data Description</h3> <p>The neural data consists of two main streams:</p> <ol> <li>Log Local Field Potential (LFP) Power: <ul> <li>Description: Extracellular neural data from the lower frequency bands.</li> <li>Sampling Rate: Preprocessed to 2 kHz.</li> <li>Filtering: Bandpass filtered with a 150-450 Hz band.</li> <li>Scale: Log-scale</li> <li>Bin size: 20ms Summed</li> <li>Shape: 128457x64 (time by channels).</li> </ul> </li> <li>Binned Threshold Crossings: <ul> <li>Description: Extracellularly collected neural spiking activity related to attempted speech.</li> <li>Threshold: -4.5 RMS</li> <li>Bin size: 20ms Summed</li> <li>Shape: 128457x64 (time by channels)</li> </ul> </li> </ol> <p>The included data are within a single mat file with the keys `trial_info_df`, `lfp_matrix`, `spikes_matrix`, and `timevector`.</p> <h3>Trial Information</h3> <p>The `trial_info_df` key contains information about each trial in the form of a pandas DataFrame with the following columns:</p> <ul> <li>block_num</li> <li>cue</li> <li>start_time</li> <li>go_cue_time</li> <li>end_time</li> </ul> <h3>Participant and Experiment Information</h3> <ul> <li>Participant T16: Right-handed, 52-year-old woman with tetraplegia and dysarthria due to a pontine stroke.</li> <li>Implantation: Four 64-channel intracortical microelectrode arrays in the left precentral gyrus.</li> <li>Recording Day: Data collected 69 days post-implantation.</li> <li>Recording Platform: Backend for Realtime Asynchronous Neural Decoding (BRAND) platform.</li> <li>Spiking Data Extraction: Linear regression referencing, bandpass filtering (250-5000 Hz), and threshold crossings identification (-4.5 RMS threshold).</li> <li>LFP power Extraction: Band pass filtering (150 - 450 Hz), and notch filtering at harmonics of 60 Hz.</li> </ul> <h3>Experimental Task</h3> <ul> <li>Task: Cued speech task where Participant T16 vocalized words presented on a screen.</li> <li>Word Bank: Compiled from a 50-word vocabulary by Moses et al.</li> <li>Trial Structure:<br> <ul> <li>A red square appeared below a word.</li> <li>After 1500 ms, the square turned green, cueing the participant to vocalize the word.</li> <li>The trial ended after the participant finished speaking, with a 1000 ms interval before the next trial.</li> </ul> </li> </ul> <h2>Loading and Handling the Data</h2> <p>The data can be loaded into Python using `scipy.io.loadmat`:</p> <h3>Loading the Data</h3> <blockquote> <p>import scipy.io<br>import pandas as pd</p> <p># Load the .mat file<br>data = scipy.io.loadmat('path_to_mat_file.mat')</p> <p># Extracting the trial information<br>trial_info_df = pd.DataFrame(data['trial_info_df'])<br>trial_info_df.columns = ['block_num', 'cue', 'start_time', 'go_cue_time', 'end_time']</p> <p># Extracting neural data matrices<br>lfp_matrix = data['lfp_matrix']<br>spikes_matrix = data['spikes_matrix']</p> <p># Extracting time vector<br>timevector = data['timevector']</p> </blockquote> <h3>Wrapping Trial Information in a DataFrame</h3> <blockquote> <p>trial_info_df = pd.DataFrame(data['trial_info_df'], columns=['block_num', 'cue', 'start_time', 'go_cue_time', 'end_time'])</p> </blockquote> <h3>Exploring the Data</h3> <ul> <li>Log Local Field Potential (LFP) Power: `lfp_matrix`</li> <li>Threshold Crossings (Spikes): `spikes_matrix`</li> <li>Time Vector: `timevector`</li> </ul> <h2>Support</h2> <p><span>This work was supported by NIH-NINDS/OD DP2NS127291 (CP), NIH-NICHD F32HD112173 (SRN), NIH K08NS060223 (MS), R01NS112942 (MS), and RF1NS125026 (MS). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health, or the Department of Veterans Affairs, or the United States Government.</span></p> <p><span>*Caution: Investigational Device. Limited by federal law to investigational use. </span></p> <p> </p>
Synthesized audio of 300 core words of 42 Indo-European languages
<p>The speech sounds of 300 core words in this repository are synthesized using the text-to-speech engine in Microsoft <a href="https://speech.microsoft.com/" rel="nofollow">Azure AI Speech Studio</a>, which encompasses 42 Indo-European languages. Each word is synthesized in both male and female voices, resulting in a total of 25,200 audio clips (300 words × 42 languages × 2 genders). All audio clips are in 16-bit, 16 kHz, mono WAV format, with leading and trailing silences trimmed.</p>
Word Embeddings for the Software Engineering Domain
<p>A .bin file for a word2vec model pre-trained on 15GB of Stack Overflow posts. </p> <p>For more details refer to the following paper:</p> <p>Efstathiou, V., Chatzilenas, C., Spinellis, D., 2018. "Word Embeddings for the Software Engineering Domain". In <em>Proceedings of the 15th International Conference on Mining Software Repositories.</em> ACM</p>
Duhumbi Phonology - Syllabic and Word Stress
<p>In Duhumbi, stress in general is non-distinctive, prosodic, and relatively unpronounced. In both disyllabic and polysyllabic words, stress falls on the first syllable. This also holds for polymorphemic lexemes, such as inflected words with suffixes. In glossary items in the Duhumbi lexicon, stress is indicated by a stress mark [ˈ] before the stressed syllable, whenever it is not predictable. Phrase stress is generally initial and falling towards the end of the phrase, combined with dependant-head word order in which subject/object precede the verb and we find postpositions rather than prepositions, although nouns always precede adjectives, but adverbs generally precede verbs. Word order is relatively flexible. </p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: bodttim (at) gmail (dot) com</p>
Spanish 3B words Word2Vec Embeddings
<p>Ready to use gensim Word2Vec embedding models for the Spanish language. Models are created using a window of +/- 5 words, discarding those words with less than 5 instances and creating a vector of 400 dimensions for each word. The text used to create the embeddings has been recovered from news, Wikipedia, the Spanish BOE, web crawling and open literary sources. The used text has a total of 3.257.329.900 words and 18.852.481.207 characters.</p> <p>We support two types of models: Gensim full models (complete_model.zip) and KeyedVectors (keyed_vectors.zip). You can check the differences between them in the following URL: <a href="https://radimrehurek.com/gensim/models/keyedvectors.html">https://radimrehurek.com/gensim/models/keyedvectors.html</a></p> <p>To load the full model use: model = Word2Vec.load("complete.model")<br> To load the KeyedVectors use: word_vectors = KeyedVectors.load('complete.kv', mmap='r')</p> <p>More info about the models can be found in: <a href="https://github.com/aitoralmeida/spanish_word2vec">https://github.com/aitoralmeida/spanish_word2vec</a></p>
Molecular Biology Open Access Pubmed Word and Sentence Representations
<p><strong>Natural Language Embeddings about Molecular Biology</strong></p> <p>This dataset is concerned with developing a tailored training data set for word and sentence embedding based on biomedical text that has some component associated with molecular work (as opposed to the other range of work indexed in PubMed like non molecular clinical work, studies of human behavior, etc). </p> <p><strong>Raw Data</strong></p> <p>In order to develop natural language embeddings (for words and sentences), we queried PMC and MEDLINE for molecular papers only by using high-level MeSH terms to restrict interest to papers with a molecular focus. We used the following MeSH terms:</p> <ul> <li>Cells [A11]</li> <li>Multiprotein Complexes [D05.500]</li> <li>Protein Aggregates [D05.875]</li> <li>Hormones [D06]</li> <li>Enzymes and Coenzymes [D08]</li> <li>Carbohydrates [D08]</li> <li>Lipids [D10]</li> <li>Amino Acids, Peptides and Proteins [D12]</li> <li>Nucleic Acids, Nucleotides and Nucleosides [D13]</li> <li>Biological Factors [D23]</li> <li>Pharmaceutical Preparations [D26]</li> <li>Metabolism [G03]</li> <li>Genetic Phenomena [G06]</li> </ul> <p>Queries for these terms use the following string:</p> <blockquote> <p>"cells"[MeSH Terms] OR "Multiprotein Complexes"[mh] OR "Protein Aggregates"[mh] OR "Hormones, Hormone Substitutes, and Hormone Antagonists"[mh] OR "Enzymes and Coenzymes"[mh] OR "Carbohydrates"[mh] OR "Lipids"[mh] OR "Amino Acids, Peptides, and Proteins"[mh] OR "Nucleic Acids, Nucleotides, and Nucleosides"[mh] OR "Biological Factors"[mh] OR "Pharmaceutical Preparations"[mh] OR "Metabolism"[mh] OR "Cell Physiological Phenomena"[mh] OR "Genetic Phenomena"[mh]</p> </blockquote> <p>PubMed returns 11,447,521 abstracts. PMC returns, 1,720,266 documents, 509,722 of these are open access. We downloaded, parsed and concatenated 403,825 PMC open access documents into a single file `molecular_oa_pmc.tsv`. This is a 33GB TSV file with the following columns:</p> <ul> <li>File:Paragraph - a unique identifier for each paragraph</li> <li>SentenceId - the local number of the sentence in the document</li> <li>Sentence Text - tokenized text of the sentence (based on <a href="https://github.com/ClearTK/cleartk/blob/master/cleartk-token/src/main/java/org/cleartk/token/tokenizer/TokenAnnotator.java">ClearTk's TokenAnnotator.java</a>)</li> <li>Codes - <code>exLink</code> for the presence of a citation, <code>inLink</code> for the presence of link to a Figure</li> <li>Figures - Figure codes</li> <li>Headings - High level section of the paper</li> <li>Offset_Begin - offset of the start of the sentence within the paper</li> <li>Offset_End - offset of the start of the sentence within the paper</li> </ul> <p>We repeated the same process for PubMed abstracts to generate a 3.6G file (`molecular_oa_medline.tsv`) with three columns:</p> <ol> <li>Pubmed ID</li> <li>A Boolean value indicating whether the article is a review</li> <li>Text</li> </ol> <p>We concatenated the text columns of these two files into a single 30GB file (`molecular_oa.txt`) where each line is a single sentence and the text is fully tokenized. </p> <p>These three files are archived in `molecular_oa_raw_text.tar.gz`.</p> <p><strong>Fasttext Embedding</strong></p> <p>We trained a fasttext model on the raw training data (https://fasttext.cc/) using the standard `skipgram` parameter. A gzipped copy of the word embeddings is included in `fasttext.model.vec.gz` </p> <p><strong>GloVe Embedding</strong></p> <p>We trained GloVe models on the raw training data (https://nlp.stanford.edu/projects/glove/). A gzipped copy of the best performing word embeddings is included in `bio_GloVe_300.tar.gz` </p>
Corpus of Spanish Word-in-Noise Confusions
<p>The dataset represents a large-scale corpus of noise-induced robust misperceptions in Spanish. The corpus contains 3235 consistent misperceptions, selected for the corpus if at least 6 listeners reported the same response from a group of 15 listeners. The dataset consists of a metadata table, separate audio waveforms for the speech and noise signals that led to each confusion, and masker waveforms.</p> <p>The corpus was described in the following journal article: http://dx.doi.org/10.1121/1.4905877 </p>
Monthly word embeddings for Twitter random sample (English, 2012-2018)
<p>This dataset contains monthly word embeddings created from the tweets available via the statuses/sample endpoint of the Twitter Streaming API from 2012 to 2018. Full details of the creation of the dataset are given in <a href="https://www.aclweb.org/anthology/D19-1007/">Room to Glo: A Systematic Comparison of Semantic Change Detection Approaches with Word Embeddings</a>. </p> <p>The md5sum of the gzipped tarball file is a76888ffec8cc7aebba09d365ca55ace .</p>
Data for: "Considering unique, shared, and dominant brain activation in the VWFA and LOC: A comparison of separate and combined word and picture naming"
<p>FMRI data underlying analyses for Experiments 1 and 2 in "Considering unique, shared, and dominant brain activation in the VWFA and LOC: A comparison of separate and combined word and picture naming".</p>
Recordings of spoken Dagbani for the letters zero to ten and the words yes and no
<p>The dataset contains recordings of the words zero to ten, and yes and no, spoken in Dagbani. The recordings have been made by the Tingoli and Nyankpala communities of the northern region of Ghana.</p> <p>The dataset is unfiltered and may contain noise or incorrect recordings.</p> <p> </p> <p>The number of available recordings are as follows.</p> <table> <tbody> <tr> <td>Number of recordings</td> <td>English translation</td> <td>Spoken Dagbani*</td> </tr> <tr> <td>720</td> <td>yes</td> <td>ii, iin</td> </tr> <tr> <td>710</td> <td>no</td> <td>aayi, ai, iiyi</td> </tr> <tr> <td>218</td> <td>zero</td> <td>**</td> </tr> <tr> <td>218</td> <td>one</td> <td>yini, zagyini, zaɣ’ yini</td> </tr> <tr> <td>223</td> <td>two</td> <td>ayi</td> </tr> <tr> <td>220</td> <td>three</td> <td>ata</td> </tr> <tr> <td>223</td> <td>four</td> <td>anahi</td> </tr> <tr> <td>220</td> <td>five</td> <td>anu</td> </tr> <tr> <td>219</td> <td>six</td> <td>ayobu, ayɔbu</td> </tr> <tr> <td>224</td> <td>seven</td> <td>ayopoin, apoin, ayɔpɔin, apɔin</td> </tr> <tr> <td>222</td> <td>eight</td> <td>anii</td> </tr> <tr> <td>219</td> <td>nine</td> <td>awei, awɛi</td> </tr> <tr> <td>162</td> <td>ten</td> <td>pia</td> </tr> </tbody> </table> <p>* the list may not be complete, other forms may exist and be spoken</p> <p>** no clear single definition</p> <p> </p> <p>A second separate set of `old' recordings is included with the following recordings.</p> <table> <tbody> <tr> <td>Number of recordings</td> <td>English translation</td> </tr> <tr> <td>146</td> <td>yes</td> </tr> <tr> <td>142</td> <td>no</td> </tr> </tbody> </table> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.