Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

475

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

475 results for “word”

Learn how ShareScore rates datasets ↗
zenodo44/100

Database of Cross-Linguistic Norms, Ratings, and Relations for Words and Concepts as CLDF dataset

<p>Cite the source of the dataset as:</p> <blockquote> <p>Tjuka, Annika, Robert Forkel, and Johann-Mattis List. 2022. Linking Norms, Ratings, and Relations of Words and Concepts Across Multiple Language Varieties. Behavior Research Methods 54. 864–884. DOI: 10.3758/s13428-021-01650-1</p> </blockquote>

opencc-by-4.0Nov 2022View details →
zenodo44/100

A list of Swedish words that have experienced historical semantic changes

<p>This list contains a set of Swedish words that have experienced semantic change during the past centuries. The list has been collected during the VR funded project <a href="https://languagechange.org/">Towards Computational Lexical Semantic Change Detection</a>, ( 2018-01184) and is a work in progress. Because the work is currently on pause, we have chosen to release the list as is in the hope of facilitating collaboration and use in other research.</p> <p>The list has four columns in the following format:</p> <pre><code>Ord*, Vilken betydelseförändring har skett*, När skedde förändringen, Källor (exv SAOL) word, what change occurred, when the change occur, reference</code></pre> <p><br> Not all fields are filled for every word. Where there are multiple change periods, there are multiple lines with empty values for word and what change occurred, see the example with <em>egendomlig </em>below.</p> <pre><code>egendomlig,som utgör (ngns) egendom &gt; karaktäristisk (positiv) &gt; speciell (negativt!!),"A. Sen 1600tal, ",SAOB ,,B. Sen 1850, ,,"C. Sen ??, efter 190",</code></pre> <p>The .xlsx file contains links to the references.</p> <p>The resources are freely available for education, research and other non-commercial purposes.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

From the Horse's Mouth: The Words We Use to Teach Diverse Student Groups Across Three Continents

<p>Word frequency pairs for courses A, B, C from:&nbsp;</p> <p>Brett A. Becker, Daniel Gallagher, Paul Denny, James Prather, Colleen Gostomski, Kelli Norris, and Garrett Powell. 2022. From the Horse&rsquo;s Mouth: The&nbsp;Words We Use to Teach Diverse Student Groups Across Three Continents.&nbsp;In Proceedings of the 53rd ACM Technical Symposium on Computer Science&nbsp;Education V. 1 (SIGCSE 2022), March 3&ndash;5, 2022, Providence, RI, USA. ACM,&nbsp;New York, NY, USA, 7 pages. https://doi.org/10.1145/3478431.3499392</p> <p><strong>When referring to this dataset, please cite the above article. That contains the DOI of this dataset. Please do not cite this dataset directly without citing the article.</strong></p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

CWID-hi: A Dataset for Complex Word Identification in Hindi Text

<p>This dataset was created by conducting a human intelligence test, wherein native and non-native Hindi speakers annotated words they could not understand in Hindi text. They were then asked to rank the complexity of these words along with their synonyms. A word that received an average rank of &lt;=3 (out of 5) is labeled 1 and the word that received an average rank of &gt;3 is labeled 0. 1 indicates complex and 0 indicates simple.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Words are monuments data-only archive

<p>These are the data that go with the paper <a href="https://besjournals.onlinelibrary.wiley.com/doi/full/10.1002/pan3.10302">&quot;Words are monuments: Patterns in US national park place names perpetuate settler colonial mythologies including white supremacy&quot;</a>&nbsp;(available 6 April 2022) by the above authors.&nbsp;Our code (and these data) are available at&nbsp;<a href="https://doi.org/10.5281/zenodo.5712009">https://doi.org/10.5281/zenodo.5712009</a>. This archive is here for folks who just want the data. This is the&nbsp;<strong>full dataset with name explanations</strong> and references for the explanations, traditional Indigenous place names for settler colonial place names (where&nbsp;available),&nbsp;additional categories for sorting, and more.&nbsp;</p> <p>Note: .csv files can be opened in Excel and Google Sheets.</p> <p>column info:</p> <p>A: Unique ID per row</p> <p>B: National Park place name is in</p> <p>C: Place name according to NPS visitor map</p> <p>D: Feature type (mountain, island, picnic area, etc.)</p> <p>E: Name type (person, non-human animal, plant, etc.)</p> <p>F: Natural or human constructed (human constructed includes so-called &quot;ruins&quot;, as well as visitor centers, etc.)</p> <p>G: Is the word from an Indigenous or western language?</p> <p>H: Is it a traditional Indigenous place name?</p> <p>I: If it is Indigenous is it the name of an Indigenous person or people?</p> <p>J: Is it a translation of a traditional Indigenous PN?</p> <p>K: Word meaning class (similar to column E)</p> <p>L: Erasure -- see paper for definitions of each class below and decision trees</p> <ul> <li>Yes</li> <li>potentially</li> <li>no information</li> <li>not erasure - evidence it is a traditional Indigenous PN or settler built with western PN</li> <li>translation of traditional IPN)</li> </ul> <p>M: Dimensions of racism and colonialism&nbsp;-- see paper for definitions of each class below and decision trees</p> <ul> <li>No (traditional Indigenous PN, western built with western PN, or erasure as only problem)</li> <li>Name itself promotes racist ideas and/or violence against a group</li> <li>Named after person who supported racist ideas (but not physically violent)</li> <li>Named for a person who directly or use their power to indirectly perpetrate&nbsp;violence against a racial group</li> <li>Western use of Indigenous name (Appropriation)</li> <li>Other - truly does not fit any other classes</li> <li>No info - cannot find explanation</li> <li>Colonialism - memorializes colonialism</li> <li>Relevant western use of Indigenous name (i.e., appropriation of traditional name)</li> </ul> <p>N: Derogatory</p> <ul> <li>Yes</li> <li>Potentially</li> <li>No info</li> </ul> <p>O: explanation of name found in research</p> <p>P: Link to resources explaining name (for full citation for books cited, e.g., &quot;name year&quot; entries, see Table S1 in the paper).</p> <p>Q: Link to resources explaining name (if second link or source available)</p> <p>R: Indigenous name found in research</p> <p>v1.0.0 did not include columns O-R by mistake; corrected in this version update.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Classical Tibetan Word Embeddings

<p>Classical Tibetan word embeddings trained with FastText based on the 2018 version of the BDRC corpus, a segmented version of which is available on Zenodo:</p> <p>Meelen, Marieke, &amp; Roux, &Eacute;lie. (2020). The Annotated Corpus of Classical Tibetan (ACTib) - Version 2.0 (Segmented &amp; POS-tagged) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.3951503</p> <p>This is the first version trained with default FastText settings (100D) for a pilot study on Chinese-Tibetan crosslinguistic Semantic Textual Similarity:</p> <p>Felbur, Rafal, Marieke Meelen &amp; Paul Vierthaler (2022), &#39;Crosslinguistic Semantic Textual Similarity of Buddhist Chinese and Classical Tibetan&#39; in <em>Journal of Open Humanities Data</em>.</p> <p>This research was done with generous funding from the Open Philology project. This project (running 2018&ndash;2022) is funded by the European Research Council (ERC) under the Horizon 2020 program (Advanced Grant agreement No 741884). It is based at the Leiden University Institute for Area Studies.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Deep Reference Mining from Scholarly Literature in the Arts and Humanities - Pre-trained word embeddings

<p>Pre-trained word vectors of dimensionality 100 and 300 for the publication:&nbsp;Deep Reference Mining from Scholarly Literature in the Arts and Humanities, submitted to Frontiers in Digital Humanities.</p> <p>The corpus of scholarly publications from which these vectors were trained is under copyright, therefore we publish these vectors for reproducibility. Please refer to the publication&#39;s repository for further details:&nbsp;<a href="https://github.com/dhlab-epfl/LinkedBooksDeepReferenceParsing">https://github.com/dhlab-epfl/LinkedBooksDeepReferenceParsing</a>.</p> <p>These vectors were trained using Gensim 3.1.0. The corpus was preprocessed as follows:</p> <ol> <li>word tokenization with NLTK word_punct tokenizer.</li> <li>digits were converted into the $NUM$ token</li> <li>words less frequent than 5 times, for every document,&nbsp;were converted to the $UNK$ token</li> <li>vectors were trained using the function:&nbsp;Word2Vec(window=5, min_count=5, sg=1)</li> </ol>

opencc-by-4.0Feb 2018View details →
zenodo44/100

RUSSE'2018: Human-Annotated Sense-Disambiguated Word Contexts for Russian

<p>This dataset contains human-annotated sense identifiers for 2562 contexts of 20 words used in the <a href="https://russe.nlpub.org/2018/wsi/">RUSSE&#39;2018</a> shared task on Word Sense Induction and Disambiguation for the Russian language; part of the&nbsp;<em>bts-rnc</em>&nbsp;evaluation dataset. These sense identifiers are disambiguated as according to the sense inventory of the <a href="http://gramota.ru/slovari/info/bts/">Large Explanatory Dictionary of Russian</a>.</p> <p>The annotation is done on December 1, 2017, on the&nbsp;<a href="https://tolokanyandex.com/">Yandex.Toloka</a>&nbsp;crowdsourcing platform. In particular, 80 pre-annotated contexts are used for&nbsp;training the human annotators, 2562 contexts are annotated by humans such that each&nbsp;context was annotated by 9 different annotators. The annotation reliability&nbsp;is indicated by a high value of Krippendorff&#39;s&nbsp;&alpha; = 0.83. After the annotation, every context was additionally inspected (&ldquo;curated&rdquo;) by the organizers of the shared task.</p> <p>The following words are represented:&nbsp;<em>акция</em> (action / stock), <em>байка</em> (yarn / tale), <em>гвоздика</em> (carnation / nail), <em>гипербола</em> (hyperbole), <em>град</em> (avalanche), <em>гусеница</em> (grub), <em>домино</em> (domino), <em>кабачок</em> (marrow / pub), <em>капот</em> (hood), <em>карьер</em> (mine / career), <em>кок</em> (cook), <em>крона</em> (top / crown), <em>круп</em> (croup), <em>мандарин</em> (mandarine), <em>рок</em> (fate / rock), <em>слог</em> (syllable), <em>стопка</em> (glass, stack), <em>таз</em> (bowl), <em>такса</em> (rate / badger-dog), <em>шах</em> (shah / check).</p> <p>The following files are included in this dataset:</p> <ul> <li>Toloka assignments (training:&nbsp;<em>tasks-train.tsv</em>, annotation: <em>tasks-test.tsv</em>)</li> <li>Toloka output (non-aggregated: <em>assignments_01-12-2017.tsv.xz</em>, aggregated:&nbsp;<em>aggregated_results_pool_1036853__2017_12_01.tsv</em>)</li> <li>annotator agreement report (<em>agreement.txt</em>)</li> <li>curated report (<em>report-curated.tsv.xz</em> and a supplementary file&nbsp;<em>tasks-eval.tsv.xz</em>)</li> <li>the final aggregated dataset (<em>bts-rnc-crowd.tsv</em>)</li> </ul> <p>The <em>bts-rnc-crowd.tsv</em>&nbsp;file has the following format: <em>id</em>, <em>lemma</em>, <em>sense_id</em>, <em>left</em> hand side context, <em>word</em> form, <em>right</em> hand side context, list of&nbsp;<em>senses</em>. The encoding is UTF-8 and the line breaks are LF (UNIX).</p>

opencc-by-sa-4.0Jan 2018View details →
zenodo44/100

Counting Words That Count: NLP for exploring Romanian Parliament Transcripts

<p>The data is obtained by scraping the cdep.ro website and contains 500k+ instances of speech from the parliament podium from 1996 to 2019. (Up to 2001 only the Chamber of Deputies published transcripts, after jan. 2001&nbsp;Senate data is also included.)&nbsp;<br> <br> Columns:&nbsp;</p> <p>&#39;index&#39; - incremented integer as row number in order of scraping</p> <p>&#39;title&#39;, - title of the scraped page, usually contains the name of the chamber and the exact data</p> <p>&#39;name&#39;, - the name of the speaker, preappended with Mr. or Mrs.&nbsp;</p> <p>&#39;speech&#39;, - the content of the speech,&nbsp;&nbsp;</p> <p>&#39;gender&#39;, - the gender of the speaker</p> <p>&#39;url&#39; - the url to the profile of the speaker (useful for extending the data)</p> <p>&nbsp;</p> <p>CDEPs2.csv - Contains all transcripts, prone to parsing errors. 100% of data.</p> <p>validated-1.csv - Consists of 99% of original data. Less than 1% dropped for convenience. Ready to use.</p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

Kusunda 250 Word List Audio Files

<p>This data set contains 2 zip files and 3 pdf files containing the cut sound files of the elicited 250-concept word list (plus numerous additional concepts and lexical items) recorded&nbsp;with the last two Kusunda speakers from Nepal, Gyani Maiya Sen Kusunda and Kamala Khatri (Sen Kusunda), on in late July and early August&nbsp;2019, in Kathmandu, Nepal.</p> <p>As per our knowledge, these are&nbsp;the first recordings of Kusunda that are available in the public domain.</p> <p>The zip files, when unpacked, will reveal several hundreds of cut sound files. At least the concepts from the list are triple repeated. There are also numerous additional concepts on recording. The pdf files contain&nbsp;explanations regarding the elicited concepts as well as the way in which the forms of the two speakers were used to come to an underlying form, a &#39;reconstruction&#39; of some sorts.&nbsp;</p> <p>The file numbers of the individual sound files refer to both the speakers (GM, K), to the date of recording and the exact origin file number. More information regarding a certain concept or pronunciation can be found in these respective origin sound files.</p> <p>Please note that the actual transcriptions in IPA may differ between the file names of the individual sound recordings, the individual speaker&#39;s pdf file, and the aggregated pdf file. We request users of these data to take the transcriptions in the aggregated pdf file as the most recent ones, or alternatively contact us for the most recent, updated transcriptions.</p> <p>This research was funded by a 2,000 USD grant from&nbsp;the Endangered Language Fund (<a href="http://www.endangeredlanguagefund.org/">http://www.endangeredlanguagefund.org/</a>), a 700 euro contribution by the European Research Council Starting Grant 715618 &ldquo;Computer-Assisted Language Comparison&rdquo; (<a href="http://calc.digling.org/">http://calc.digling.org</a>), and a total of 2,320 euros raised through crowdfunding at GoFundMe (<a href="https://www.gofundme.com/f/saving-the-kusunda-language-in-nepal">https://www.gofundme.com/f/saving-the-kusunda-language-in-nepal</a>). Many thanks to all generous contributors.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes <strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), (paid) public screening other than for educational purposes, storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Create Commons Attribution 4.0 International / Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>We&nbsp;greatly value feedback, suggestions, advice, analysis etc. based on this material which will help in the description of the Kusunda language, especially any comments and suggestions that will enable the revitalisation of the language, including a standardisation of the phonology and a phonologically consistent but also practical orthography in both देवनागरी Devanāgarī&nbsp;and Roman script.</p> <p>Uday Raj Aaley: aarambhkhabar (at) gmail (dot) com</p> <p>Tim Bodt: bodttim&nbsp;(at) gmail (dot) com</p>

opencc-by-4.0Aug 2019View details →
zenodo44/100

Words or terms? Saṃjñā annotated dataset

<p>These data were used for the study published in:</p> <p>Lugli, Ligeia. 2019. Words or terms? Models of terminology and the translation of Buddhist Sanskrit vocabulary. In Alice Collett (ed.) Buddhism and Translation: Historical and Contextual Perspectives, New York: SUNY.</p> <p>data include:<br> 1. concordance lines for saṃj&ntilde;ā used for the study mentioned above. The concordance lines have been exported from the Sketch Engine and come from an automatically segmented corpus (segmenter = Lugli&#39;s version 1). They have not been proofread and contain segmentation errors.&nbsp;<br> 2. csv file with Lugli&#39;s semantic annotation of the concordance lines for saṃj&ntilde;ā. The data was annotated by Ligeia Lugli in 2017; part of the data constitutes a much revised version of a dataset originally prepared by Roberto Garcia for the Buddhist Translators Workbench in 2016.<br> 3. a pre-publication version of the study</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>The creation of these data was funded by the British Academy through a Newton International Fellowship; the research was conducted at King&#39;s College London.<br> &nbsp;</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

Swahili word analogy dataset

<p>Swahili Analogy dataset contains pairs of words that are organized in 4&#39;s to facilitate word analogy test. Word analogy test is used to evaluate the quality of word representation vectors from a language model. The dataset contains 12,864 questions that have been organized in 12 categories.</p>

opencc-by-4.0Nov 2019View details →
zenodo44/100

French Word Sense Disambiguation with Princeton WordNet Identifiers

<p>This is a dataset for the Word Sense Disambiguation of French using Princeton WordNet identifiers. It contains two training corpora : the SemCor and the WordNet Gloss Corpus, both automatically translated from their original English version, and with sense tags automatically aligned. It contains also a test corpus : the task 12 of SemEval 2013, originally sense annotated with BabelNet identifiers, converted into Princeton WordNet 3.0.</p>

opencc-by-4.0Nov 2019View details →
zenodo44/100

CLDF dataset derived from Liú et al.'s "Collection of Basic Words in Chinese Dialects" from 2007

<p>Cite the source of the dataset as:</p> <blockquote> <p>Líu, L.; Wáng, H.; Bǎi, Y. (2007): Xiàndài Hànyǔ fāngyán héxīncí, tèzhēng cíjí 现代汉语方言核心词·特征词集 [Collection of basic vocabulary words and characteristic dialect words in modern Chinese dialects]. Nánjīng: Fènghuáng.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

CLDF dataset derived from Ivani's "Basic Words in Suansu" from 2019

<p>Cite the source of the dataset as:</p> <blockquote> <p>Ivani, J. K. (2019): A first overview on Suansu, a Tibeto-Burman language from Northeastern India. Talk, held at the 29th conference of the Southeast Asian Linguistic Society (27-29 May, Tokyo). https://zenodo.org/record/3383006</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

SWL-LSE: SignaMed Word-Level LSE, a Dataset of Spanish Sign Language Health Signs

<h2>SWL-LSE Dataset</h2> <p>The SWL-LSE dataset is coined from SignaMed Word-Level LSE (Lengua de Signos Espa&ntilde;ola -Spanish Sign Language).</p> <h2>Overview</h2> <p>The dataset consists of 8,000 sign sequences from 300 different sign classes related to the health domain. Each class is represented by an RGB video that serves as the dictionary sign. These dictionary signs were reproduced by 124 signers, including deaf individuals, interpreters, and L2 Spanish Sign Language (LSE) students, using their webcams or mobile phones via the SignaMed platform (<a href="https://signamed.web.app" target="_new" rel="noopener">https://signamed.web.app</a>). For privacy reasons, only the skeleton data is shared.</p> <p>The process of collecting the dataset is described in:</p> <p>V&aacute;zquez-Enr&iacute;quez, M.; Alba-Castro, J.L.; P&eacute;rez-P&eacute;rez, A.; Cabeza-Pereiro, C.; Doc&iacute;o-Fern&aacute;ndez, L. SignaMed: a Cooperative<br>Bilingual LSE-Spanish Dictionary in the Healthcare Domain. In Proceedings of the Proceedings of the LREC-COLING 2024<br>11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources; Efthimiou, E.;&nbsp;Fotinea, S.E.; Hanke, T.; Hochgesang, J.A.; Mesch, J.; Schulder, M., Eds., Torino, Italia, 2024; pp. 386&ndash;394.&nbsp;</p> <p>The dataset itself and the pipeline for training and executing a baseline model based on skeletons is described in this github (https://github.com/mvazquezgts/SWL-LSE), and this paper:</p> <p>V&aacute;zquez-Enr&iacute;quez, M.; Alba-Castro, J.L.; Doc&iacute;o-Fern&aacute;ndez, L.; Rodr&iacute;guez-Banga, E. SWL-LSE: A Dataset of Spanish Sign Language Health Signs with an ISLR Baseline Method. Technologies 2024, 12(10), 205, D.O.I:10.3390/technologies12100205</p> <h2>Files</h2> <h3>1. VIDEOS_REF.zip</h3> <ul> <li><strong>Description</strong>: RGB videos recorded in lab conditions that represent each sign-class</li> <li><strong>Total files</strong>: 300</li> </ul> <h3>2. videos_ref_annotations.csv</h3> <ul> <li><strong>Description</strong>: CSV file with the correspondence between the name of the video, its class ID and gloss in spanish: FILENAME,CLASS_ID,LABEL.</li> <li><strong>Total files</strong>: 1</li> </ul> <h3>3. ANNOTATIONS.zip</h3> <ul> <li><strong>Description</strong>: 3 CSV files with train, validation and test file-class correspondences: FILENAME,CLASS_ID</li> <li><strong>Total files</strong>: 3</li> </ul> <h3>4. MEDIAPIPE.zip</h3> <ul> <li><strong>Description</strong>: Pickle files containing the full output of Mediapipe using their Heavy model. Each .pkl file contains the outputs of Mediapipe Holistic legacy, Mediapipe Pose and Mediapipe Hands. Each file is package as a dictionary: dict_keys(['pose', 'hands', 'holistic_legacy'])</li> <li><strong>Total files</strong>: 8000</li> </ul> <h2>Usage</h2> <p>Researchers and practitioners in pattern recognition, machine learning, and sign language linguistics may find this dataset valuable for:</p> <ul> <li>Training/testing machine learning models for isolated sign language recognition or gesture recognition.</li> <li>Analyzing patterns on signs realization</li> </ul> <h2>Acknowledgments</h2> <p>This dataset is a collaborative effort of the next research goups and entities:</p> <ul> <li><a href="http://gtm.uvigo.es/en/">Group of Multimedia Technologies (GTM)</a> from the <a href="https://atlanttic.uvigo.es/en">atlanTTic Research Center</a> of <a href="http://www.uvigo.es/">University of Vigo</a> (Spain)</li> <li><a href="http://grades.uvigo.gal/">Group of Discourse and Society (GRADES)</a> from the <a href="https://fft.uvigo.es/en/">School of Philology and Translation</a> of <a href="http://www.uvigo.es/">University of Vigo</a> (Spain)</li> <li><a href="http://www.faxpg.es/">Federation of Deaf People Galician Associations (FAXPG)</a></li> <li><a href="https://fundacioncnse-dilse.org">Fundaci&oacute;n CNSE-DILSE</a></li> </ul> <p>Gratitude is extended to them for their contributions and support.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Kam-Niger-Congo comparative word list

<p>This is a comparative word list containing data collected with the Leipzig-Jakarta word list, intended to compare basic vocabulary between Kam and other Niger-Congo languages. It contains reconstructions for a variety of proto-languages already available in the literature (e.g. Jukunoid, Mumuyic, Proto-Bantu, Proto-Gbe, Proto-Potou-Akanic, and Proto-Fula-Sereer), as well as the author's own quasi-reconstructions for Niger-Congo, Benue-Congo, and Delta-Cross and cognate judgements.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

CLDF dataset derived from Wang's "Basic Words in Chinese Dialects" from 2004

<p>Cite the source of the dataset as:</p> <blockquote> <p>Wang, F. 2004. BCD: basic words of Chinese dialects. Unpublished dataset. [Digital version in: List, J.-M. (2015): Network perspectives on Chinese dialect history. Bulletin of Chinese Linguistics 8. 42-67.]</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Data for: Unmasking the Effects of Orthography, Semantics, and Phonology on 2AFC Visual Word Perceptual Identification

<p>This data was used in analyses for &quot;Unmasking the Effects of Orthography, Semantics, and Phonology on 2AFC Visual Word Perceptual Identification&quot;.</p>

opencc-by-4.0Aug 2020View details →
zenodo44/100

Middle Dutch syllabified words

<p><strong>Specifics of the data:</strong></p> <ul> <li>Text file (<em>syllabified_crm.txt</em>)&nbsp;containing&nbsp;43,710&nbsp;syllabified Middle Dutch words, taken&nbsp;from the&nbsp;<em>Corpus Van Reenen-Mulder</em>. This&nbsp;corpus, created by Pieter van Reenen en Maaike Mulder&nbsp;at the Free University Amsterdam, contains&nbsp;about 2,500 Middle Dutch&nbsp;charters. It has&nbsp;about 750,000 tokens.&nbsp;The charters were written in the Netherlands&nbsp;and&nbsp;Flanders between 1300&nbsp;and&nbsp;1400.</li> <li>The 43,710&nbsp;syllabified words in this list is the total amount of unique words from the <em>Corpus Van Reenen-Mulder</em>. Some tokens from this corpus were, however, excluded when assembling the data set&nbsp;due to the fact that they contained diacritic symbols to indicate abbreviations, clitics,&nbsp;or unclear parts in the original charter.</li> <li>A dash-symbol (-)&nbsp;is used as separator.</li> <li>Apart from the entire data set, this DOI also includes: <ul> <li>A pdf-file visualizing the data set</li> <li>The splits used for the automatic syllabification experiment by Haverals, Kestemont &amp; Karsdorp (2018).</li> <li>A gold standard&nbsp;out-of-corpus sample of 1,748 Middle Dutch words, taken at random from the <em>Cd-rom Middelnederlands</em>, also used in the above-mentioned syllabification experiment</li> </ul> </li> </ul>

opencc-by-sa-4.0Dec 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record