Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

475

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

475 results for “word”

Learn how ShareScore rates datasets ↗
zenodo44/100

The ICDAR 2003 Informal Competition for the Recognition of On-line Words: The Unipen-ICROW-03 benchmark set - Version 0.0

<p>Proposal for an informal benchmark on word recognition. See for the related ImUnipen collection<br> of word images from on-line vectorial handwriting data:&nbsp;https://zenodo.org/record/1195059</p> <p>At the time (ICDAR 2003) there was not a lot of interest so the project was not pursued.</p> <p>Lambert Schomaker - February 2023</p> <p>_______________________________________________________________________________</p> <p>The ICDAR 2003 Informal Competition for the Recognition of On-line Words:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;The Unipen-ICROW-03 benchmark set&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Version 0.0</p> <p>Lambert Schomaker / International Unipen Foundation</p> <p>The ICROW suite of test files for the recognition of isolated on-line<br> free-style (handprint, mixed and cursive) words has been<br> composed. Different tablets, nationalities and languages<br> are involved. Only the ASCII set is used within word labels.</p> <p>The set contains:</p> <p>&nbsp; &nbsp;13119 written words<br> &nbsp; &nbsp; &nbsp;884 unique lexical word entries<br> &nbsp; &nbsp; &nbsp; 72 writers&nbsp;</p> <p>Language: Dutch, English, Italian.<br> Nationalities: Dutch, Irish, Italian, + mixed</p> <p>The benchmark test is a good estimator for&nbsp;<br> &quot;walk-up&quot; recognition performance.</p> <p>[Note: some of the writers (NIC-Pc95*.dat set) are present in the<br> UNIPEN R01/V07 distribution, but the actual words are unseen&nbsp;<br> outside of the Int. Unipen Foundation.]</p> <p>Please note the Copyright notice in the&nbsp;<br> accompanying file &#39;Copyright&#39;</p> <p>Wed Jul 16 21:20:10 CEST 2003</p> <p>Lambert Schomaker</p> <p>---------------------------------------------------------------------------</p> <p>Instructions for the ICDAR 2003 informal competition for<br> the recognition of on-line words.</p> <p>1 - unpack the .tgz file<br> 2 - use the UNIPEN files as input for your recognizer.<br> 3 - report, for each writer, a file &lt;writer-id&gt;.res</p> <p>&nbsp; Example: do-my-recognizer &lt; NIC-Hi93b-marc.dat &gt; NIC-Hi93b-marc.res</p> <p>Format of the .res file.</p> <p>No XML for this moment: simplicity does it.</p> <p>We assume that the recognizer is able to produce a top-10 list<br> of likely words, sorted from most likely to least likely.<br> The output for each word is on a single line. The correct<br> target word is in the first column.</p> <p>&lt;targetword 1&gt; &lt;best word hyp.&gt; &lt;2nd-best word hyp.&gt; ... &lt;10th-best word hyp&gt;<br> &lt;targetword 2&gt; &lt;best word hyp.&gt; &lt;2nd-best word hyp.&gt; ... &lt;10th-best word hyp&gt;</p> <p>Example with two words:</p> <p>summertime &nbsp; slumbertime slipknot summertime somatome spumante simulative semitone schoolmate sermonette semimature<br> Aberdeen &nbsp; &nbsp; Adamson Aberdeen Addison Armageddon Abyssinian Araban Albanian Alabamian Abraham Adelaide</p> <p><br> 4 - pack the &nbsp;*.res files in a .tgz or .zip file and send them<br> &nbsp; &nbsp; to schomaker@ai.rug.nl<br> &nbsp; &nbsp; All *.dat files need to be processed.</p> <p>LS.<br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2003View details →
zenodo44/100

Zero marking and word order of core arguments

<p>This material contains the dataset from the&nbsp;<a href="https://version.helsinki.fi/hals/sinnemaki/sinnemaki2010">gitlab repository</a>&nbsp;of the following article, with corrections and some additions. Please cite the article when using the data.</p> <p>Sinnem&auml;ki, Kaius 2010. Word order in zero-marking languages. <em>Studies in Language</em> 34(4): 869&ndash;912.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

cultural-ai/wordsmatter: Words Matter: a knowledge graph of contentious terms

<p>The choice of words describing cultural heritage can cause debates. It is especially sensitive when artefacts relate to different cultures and peoples who have been historically marginalised. Words chosen by archivists or curators may transmit stereotypes. The cultural heritage community has produced knowledge on potentially stereotyping and offensive terminology in heritage collections. At the same time, their knowledge is difficult to incorporate into existing online collections unless this knowledge is structured and machine-readable.</p> <p>The Words Matter Knowledge Graph represents domain expert knowledge on discussions about contentious terminology in the cultural sector. In the knowledge graph, 75 English and 83 Dutch contentious terms are linked to explanations of their usage and suggested alternatives from domain experts. There are also related matches between contentious terms and sources from external datasets: Wikidata, Princeton WordNet, Open Dutch WordNet, and Getty Art &amp; Architecture Thesaurus.</p> <p>This Zenodo publication includes the CULCO scheme used to model contentious terms in the knowledge graph. The scheme documentation is <a href="https://cultural-ai.github.io/wordsmatter/" target="_blank" rel="noopener">available on a separate page</a>.</p> <p>This knowledge graph is <a href="https://amsterdam.wereldmuseum.nl/en/about-wereldmuseum-amsterdam/research/words-matter-publication" target="_blank" rel="noopener">based</a> on the publication &ldquo;Words Matter: An Unfinished Guide to Word Choices in the Cultural Sector&rdquo; by the National Museum of World Cultures (NMVW).&nbsp;</p> <p><a href="https://doi.org/10.1007/978-3-031-33455-9_30" target="_blank" rel="noopener">Read more</a> about this work in the paper "A Knowledge Graph of Contentious Terminology for Inclusive Representation of Cultural Heritage" (2023) by Andrei Nesterov,&nbsp;Laura Hollink,&nbsp;Marieke van Erp &amp;&nbsp;Jacco van Ossenbruggen.</p> <p>In this version:</p> <ul> <li>the CULCO scheme documentation is updated</li> <li>versioning is fixed</li> <li>typos are corrected</li> </ul>

opencc-by-sa-4.0Mar 2023View details →
zenodo44/100

Raw and post-processed data for the study of prosodic cues to word boundaries in a segmentation task using reverse correlation

<p>The current dataset provides all the stimuli (folder <strong>../01-Stimuli/</strong>), raw data (folder <strong>../02-Raw-data/</strong>) and post-processed data (<strong>../03-Post-proc-data/</strong>) used in a prosody reverse correlation study with the title &quot;prosodic cues to word boundaries in a segmentation task using reverse correlation&quot; by the same authors. The listening experiment was implemented using one-interval trials with target words of the structure l&#39;aX (option 1) and la&#39;X (option 2). The experiment was designed and implemented using the <a href="https://github.com/aosses-tue/fastACI">fastACI toolbox</a> under the name &#39;segmentation&#39;. A between-subject design was used with a total of 47 participants, who evaluated one of five conditions, LAMI (N=16), LAPEL (N=18), LACROCH (N=5), LALARM (N=5), and LAMI_SHIFTED (N=3). More details are given in the related publication (to be submitted to JASA-EL in May 2023).</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Diachronic and diatopic word embeddings from newspapers digitised by the British Library (1830-1889): North and South England

<p>Diachronic&nbsp;word embeddings (decade-level) trained with Word2Vec (via Gensim) on different geographic subcorpora of the Heritage Made Digital British and the Living with Machines&nbsp;historical newspaper collections:</p> <p>- North England (north.zip)</p> <p>- South England (south.zip)</p> <p>At the moment, for each subcorpus, Word2Vec models are available for each decade in the period 1830-1889. More models are on the way for the following:</p> <p>- each decade in the periods 1780-1829 and 1890-1920 for both North and South England.</p> <p>- diachronic models for the following regions: Scotland, Wales, and Midlands.</p> <p>The models were trained using the following parameters:</p> <pre><code>sg = True min_count = 1 window = 5 vector_size = 200 epochs = 5</code></pre> <p>Like the embeddings in&nbsp;<a href="https://doi.org/10.5281/zenodo.7181682">this repository</a>, the model for each decade was aligned to the most recent one with Orthogonal Procrustes.</p> <p>See related GitHub repository for the full documentation:&nbsp;<a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a>.</p> <p>Project website (Living with Machines):&nbsp;<a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p> <p>Data related to:&nbsp;Nilo Pedrazzini &amp; Barbara McGillivray,&nbsp;<em>Diachronic and diatopic word embeddings from British historical newspapers</em>, presented at&nbsp;AIUCD (Convegno dell&rsquo;Associazione per l&rsquo;Informatica Umanistica e la Cultura Digitale) in Siena (Italy), June 2023.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

EyeLink 1000 raw eye tracking data - Reading numbers is harder than reading words: An eye-tracking study

<p>Reading Arabic numerals is a fundamentally different activity compared to word reading. This study aimed to investigate the eye movements of normal-reading adults when reading aloud short and long Arabic numerals (with or without a thousand separator) compared to matched-in-length words and pseudowords.</p> <p>This dataset contains the raw data of the article &quot;Reading numbers is harder than reading words: An eye-tracking study&quot; published in the journal <em>Acta Psychologica</em> (https://doi.org/10.1016/j.actpsy.2023.103942)</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Transliteration Model for Egyptian Words

<p>Transliteration (word) models for hieroglyphic texts created from The Ramses Transliteration Corpus, V. 2019-09-01 based on the Ramses Online data published by Serge Rosmorduc (in https://gitlab.cnam.fr/gitlab/rosmorse/ramses-trl and https://doi.org/10.5281/zenodo.4954597) and from the AES - Ancient Egyptian Sentences; Corpus of Ancient Egyptian sentences for corpus-linguistic research based on the Thesaurus Linguae Aegyptiae data published by Simon Schweitzer (in https://github.com/simondschweitzer/aes).</p> <p>The creation of the models is explained in: Jauhiainen, H. &amp; T. Jauhiainen 2023. Transliteration Model for Egyptian Words.&nbsp;<em>Digital Humanities in the Nordic and Baltic Countries Publications</em>, <em>5</em>(1), 149-164. https://doi.org/10.5617/dhnbpub.10659.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

How many words is a picture worth? Attention allocation on thumbnails versus title text regions: Dataset

<p>Dataset for the following publication:&nbsp;https://jainlab.cise.ufl.edu/eyetrack-onlineux.html</p> <p>How many words is a picture worth? Attention allocation on thumbnails versus title text regions, Yandandul, Chaitra and Paryani, Sachin and Le, Madison and Jain, Eakta, ACM Symposium on Eye Tracking Research &amp; Applications. (ETRA)</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo40/100

Nyokon word list audio recordings

<p>This deposit contains the original recorded lists of words of the Nyokon language (ISO 639-3: nvo) gathered in the process of writing the following publication:&nbsp;</p> <p>Lovestrand, Joseph. 2011. Notes on Nyokon phonology (Bantu A.45, Cameroon). SIL Cameroon. http://silcam.org/download.php?stid=&amp;folder=documents&amp;file=Nyokon_notes_phonology-Lovestrand_2011.pdf.</p> <p>A PDF copy of the publication is included in the deposit [Notes on Nyokon phonology - Lovestrand (October 2011).pdf]</p> <p>Rough phonetic transcriptions and annotations of the word lists are in SML format in the file: [Nyokon-dekereke.txt]</p> <p>This file was designed in the freeware Dekereke program, available online at: https://casali.canil.ca/</p> <p>The audio recordings can be divided into four categories.&nbsp;</p> <p>1. 1700-item word list</p> <p>Files named with a range of four-digit numbers, e.g. 0001-0005.wav, are the original audio files. The numbers refer to the &quot;SIL comparatige African wordlist&quot;. Not all 1700 items were elicted and recorded.&nbsp;</p> <p>Roberts, James &amp; Keith Snider. 2006. SIL comparative African wordlist (SILCAWL). SIL&nbsp;<br> Electronic Working Papers 2006-005. 49. http://www.sil.org/silewp/abstract.asp?ref=2006-005</p> <p>2. Individual words for Dekereke&nbsp;</p> <p>Files name with a single four-digit number followed by an English gloss, e.g. 0001 body, are audio clips taken from the origial recordings so that the clips can be listened to in the Dekereke program.&nbsp;<br> Not all words were clipped for this purpose. The recordings of many words are only availabel in the original recordigns.</p> <p>3. Contrastive Pairs&nbsp;</p> <p>Files named ContrastivePairs followed by a number are files that were recorded later in the analysis to check on assumed and suspected phonemic contrast. Most of the words repeat those found elsewhere.&nbsp;</p> <p>4. Consent&nbsp;</p> <p>There are a number of files recording the oral consent of the Nyokon speakers who participated in the word list elicitation sessions.&nbsp;<br> As acknolwedged in Lovestrand (2011) the Nyokon speaker who participated are:</p> <p>NGOUNG Isaac (chief of Ambann)<br> NGAGNI Amos Jules (chief of Ahoung)<br> AMBANG Maurice<br> EMBOM Pierre<br> ENAM Samuel<br> FOUTH Brice Rodrigue<br> HEU Emmanuel<br> HEU Paul<br> IMBO Hermine Doroth&eacute;e<br> KAMANDA Jean Achille<br> KIARI Andr&eacute; Jules<br> KOUMA Fanny<br> MBIRNANG Thomas Blaise<br> MOUOL Catherine<br> NGAGNI Emmanuel<br> YAMBASA Andr&eacute;</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Luxembourgish word embedding (User comments from RTL.lu)

<p>This dataset is a word embedding model trained on Luxembourgish user comments from the media platform RTL.lu. It contains data from roughly 544k Luxembourgish texts published&nbsp;between December 2008 and December 2018. See the documentation file for detailed info.</p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

Children's learning new word-picture pairs during physical exercise in the classroom

<p>This dataset includes memory performance and questionnaire data. In three experiments school aged children performed physical exercises in the classroom while learning new word-picture associations during motion or sitting conditions (repeated measures). The words were in an artificial language. <em>Free recall </em>memory performance was collected from a group of children (aged 8-9 years) in <em>running</em> and sitting conditions (experiment 1). <em>Recognition memory</em> performance was collected from a different group of children (aged 8-9 years) in <em>running</em> and sitting conditions (experiment 2). <em>Recognition memory</em> performance was collected from a different group of children (aged 8-9 years) in <em>stepping</em> and sitting conditions (experiment 3). Experiment 2 showed higher recognition memory performance during the sitting compared to running condition. Gender, English as an additional language (EAL), and whether the children received extra support for language work was recorded before the experiment. A survey after the experiment&nbsp;asked children which condition they preferred. All data were collected by Naomi Kefford in one school in Surrey (UK) during the MSc thesis at the School of Psychology at the University of Surrey. Naomi Kefford was supervised by Peter Klaver.&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Bilingual English-German word embedding models for scientific text

<p>This data set contains three word embedding models, constructed from the same training corpus of English and German parallel scientific texts (abstracts and research project descriptions). All text was pre-processed by language-specific stemming with the Porter stemming algorithm, removing numbers, and lower-casing.</p> <p>The first model is a 1000-dimensional Latent Semantic Analysis model, constructed from concatenating the English and German texts. The input data was a&nbsp;m&times;n (297,852&times;923,864) document-term matrix of tf-idf weights. This was processed with truncated SVD. There are two files, the word vectors in file lsa_1000_Vmat.csv (the V* term by latent factors matrix of right singular values) and the dimension weights in lsa_1000_d_weights.csv (the 1000 values of the diagonal of the <span class="math-tex">\(\Sigma\)</span> matrix.</p> <p>lsa_1000_Vmat.csv has two fields, the term and its vector representation in LSA space, separated by a &quot;|&quot; character. The structure looks like this:</p> <p>tarifplural|{5.00599733151825e-08,-1.43071379136936e-08,8.32862290483082e-08,-6.08010721687266e-08,1.15831140150142e-07,-2.46470313387358e-08,3.43215595753282e-07,6.24301666802575e-07,-2.62907158945831e-07,-1.04120313981517e-07,4.5864574355164e-07,-2.31799632277312e-07,8.37354377858843e-07,8.22507467711628e-07,4.07585381069368e-07,-4.26358988941922e-08,-8.38652991154651e-07,1.98091851171759e-07,-3.94768548759816e-08,-4.28802181962385e-07, ...}</p> <p>The other two models are a basic Random Indexing and a Reflective Random Indexing model, contained in same file, RI_training.csv. Both models have 1000 dimensions. The data structure is as follows.</p> <ul> <li>language: either &quot;en&quot; (English) or &quot;de&quot; (German), the language of the term</li> <li>term: the term as a character string</li> <li>term_collection_count: integer, number of times the term occurred in the training data</li> <li>c_vector: vector of 1000 reals, RI context vector of the term. formatted like this: &quot;{0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0.12309149,0,0,-0.12309149,0,0,0,0,0,0,0,0,0,0,0,0,0,0, ...}&quot;</li> <li>n_docs: integer, number of different documents which contained the term</li> <li>c_vector_o2: vector of 1000 reals, RRI context vector of the term, formatted like c_vector above</li> </ul> <p>1,034,860 rows.</p> <p>All files are aggressively compressed with GNU gzip and will require much more disk space when uncompressed.&nbsp; Note the special formatting of the vector numeric variables, which are different for the two models.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

The database of words and affiliations of the SEG Annual Conferences (1982 - 2019)

<p>The database of the SEG Annual meetings v2.0<br> This repository includes data for the words and phrases frequency of occurrence analysis &quot;SEGgrams.sqlite&quot; and the database &quot;SEG_affiliations_data.sqlite&quot; consisting of the industry companies and different countries academia that presented their research during the 38 Society of Explorational Geophysicists Annual Conferences (1982 - 2019) with the corresponding number of affiliations for the whole period of study.</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Word Sense Change Testset

<p><strong>Overview</strong></p> <p>This testset consists of 23 terms which have experienced word sense change during the past centuries. The main changes for each term were found using Wikipedia, dictionary.com and the Oxford English Dictionary. We consider major changes in usage as well as changes to sense. In cases where multiple (fine-grained) senses were available, we opted to accept the widest sense. E.g. for the term <em>rock</em> we consider a music sense without any distinction between different types of <em>rock music</em>, because our dataset is unlikely to have fine-grained sense differentiations. If a clear time point cannot be pinpointed, we choose the earliest possible. For comparison purposes we also chose a set of 11 terms that have experienced minimal change during the investigated period, i.e., stable terms.</p> <p> </p> <p><strong>Supplementary material</strong></p> <p><em>1. testset.txt </em></p> <p>Contains a list of all terms and the different change types for each term with a short description of the sense and change.</p> <p> </p> <p><em>2. Files of the kind "TERM.txt"</em></p> <p>The header tells us the term, which clustering coefficient was used, which similarity threshold and which similarity measure.</p> <p>A path starts with "Path:". </p> <p>A unit starts with "UNIT:"</p> <p>and the numbers following indicate 1. the number of years that the unit spans, and then a list of all years that the internal clusters stem from.</p> <p>E.g., UNIT: 83 1785, 1787, 1790, 1793, 1798, 1801, 1823, 1867, spanns 83 years and consists of clusters from year  1785, 1787, 1790 etc.</p> <p>Indentation shows the tree structure, more indentation means lower level branch in the tree.</p> <p>As an example, in AEROPLANE.txt unit UNIT: 23 1908, 1909, 1910, 1911, 1914, 1918, 1930, 1908 is the root node and the unit is related to UNIT: 27 1916, 1919, 1924, 1932, 1942, 1916.</p> <p> </p> <p><strong>Interesting findings</strong></p> <p>The longest units and paths are found for stable terms, e.g., <em>newspaper</em>. These are statistically significantly longer than the average units and paths for terms that later evolve.</p> <p><em>Newspaper</em> has a unit that spans 145 years and the first path spans from 1852 - 2007.</p> <p> </p> <p><em>FLIGHT.txt</em></p> <p>For the term <em>flight</em> we find that the first unit captures a name, <em>Flight &amp; Robson</em> who were organ builders.</p> <p>The second unit (it its own path) represents the <em>flight</em> over a hurdle: UNIT: 28 1868, 1869, 1870, 1877, 1885, 1889, 1890, 1892, 1893, 1894, 1895</p> <p>There is a unit (it its own path) that represents the <em>flight</em> of a cricket ball: UNIT: 29 1938, 1957, 1966</p> <p>Finally, the last path represents <em>flight</em> as in a means of transportation, in particular for holidays, starting with  UNIT: 19 1962, 1970, 1973, 1980</p> <p> </p> <p><em>TAPE.txt</em></p> <p>The first path for tape is a path related to <em>sowing tape</em>.</p> <p>Then there is a second path starting with  UNIT: 38 1970, 1974, 2007 that takes up the <em>musical tape</em>.</p> <p>The last path end in the same units that the second path ends in, also related to the <em>musical tape</em>.</p> <p>The <em>music tape</em> and the <em>sowing tape</em> should be related because of their shape, but we cannot find any relation as there are few or no overlapping terms.</p> <p> </p>

opencc-by-4.0Jun 2017View details →
zenodo40/100

The Effect of Typing Efficiency and Suggestion Accuracy on Usage of Word Suggestions and Entry Speed

<p>Data collected during our experiments investigating the effect of suggestion accuracy and typing efficiency on usage of word suggestions, and entry speed</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Thai Word Embeddings (word2vec) Trained on Oscar Corpus

<p>A large Thai word2vec model trained on Oscar corpus and tokenized and normalized with&nbsp;PyThaiNLP. The model can be loaded using gensim, it is saved in binary format.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Early Irish Analogy Dataset for Word Embedding Evaluation

<p>An embedding evaluation dataset for Early Irish described in the paper "<a href="https://aclanthology.org/2023.insights-1.10.pdf">Do not Trust the Experts: How the Lack of Standard Complicates <span>NLP</span> for Historical <span>I</span>rish</a>".</p> <p>Traditionally, analogy datasets are based on pairwise semantic proportion, and therefore every question has a single correct answer. Given the high level of variation in historical languages, such a strict definition of a correct answer seems unjustified. Therefore, Early Irish Analogy Dataset follows the <a href="https://vecto.space/projects/BATS/">Bigger Analogy Test Set (BATS)</a> and provides several correct answers to each analogy question.&nbsp;</p> <p>Morphological and spelling variation data are extracted from the <a href="https://dil.ie/">eDIL</a>, a historical dictionary of medieval Irish. Unlike BATS, no distinction is made between inflection types due to eDIL's structure. The raw data amounted to 2,370 spelling variation and 9,690 morphological variation questions, from which 150 examples were randomly selected for each of the subsets to be comparable in size with the synonym and antonym subsets. The synonym and antonym subsets are translations of the correspondent BATS parts obtained by reverse-searching the eDIL and proofread by four expert evaluators. The dataset includes 98 entries in the synonym subset and 109 entries in the antonym subset, upon which three or more experts agreed.</p>

opencc-by-4.0Feb 2024View details →
zenodo40/100

BanglaWriting Words Dataset: A Collection of Isolated Word Images from the BanglaWriting multi-purpose Bangla offline-handwriting dataset (WoBW)

<p>The WoBW (Words from BanglaWriting) dataset is a curated collection of isolated word images, adapted from the original BanglaWriting dataset (url:&nbsp;https://data.mendeley.com/datasets/r43wkvdk4w/1).</p> <p>WoBW focuses on individual words extracted from handwritten Bangla text samples in the BanglaWriting corpus, making it a valuable resource for research in word-level Bangla handwriting recognition and related natural language processing tasks.</p> <p>Mridha, Dr. M. F.; Quwsar Ohi, Abu; Ali, M. Ameer; Emon, Mazedul Islam; Kabir, Md Mohsin (2020), &ldquo;BanglaWriting: A multi-purpose offline Bangla handwriting dataset&rdquo;, Mendeley Data, V1, doi: 10.17632/r43wkvdk4w.1</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Twitter Dataset - Over 200,000 Tweets containing the word "Vaccine" for research porpuses

<p>This dataset contains 220,085 tweets containing the word vaccine between December 9th and December 18th 2021 at different times during each day, extracted using the Twitter API v2. Each tweet was extracted at least 3 days after its initial posting time in order to register 3 days of engagements, and it doesn&#39;t include retweets.</p> <p>Includes:</p> <ul> <li>Tweet ID</li> <li>Text</li> <li>Author ID</li> <li>Date</li> <li>Like count</li> <li>Retweet count</li> <li>Quote count</li> <li>Reply count</li> <li>User data (Followers, Following, Tweet count, Account creation date, Verified status)</li> </ul> <p>Usernames are hidden for privacy reasons</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Pennsylvania German word list (lemmatized and POS-annotated)

<p>The file presents the words used in the Pennsylvania German part of the ENDE corpus (www.deitsch.eu). The list contains every lemma with its associated word forms documented in the corpus, comprised of&nbsp;1761 lemmata and 2704 word forms.</p> <p>The ENDE corpus (&ldquo;English-Deitsch&nbsp;translation corpus&rdquo;) is the first POS-annotated and searchable text corpus in Pennsylvania German (= Deitsch;&nbsp;ISO language code: pdc), aligned to the English source texts. Despite many digital texts in Deitsch are available on the internet, there are, so far, no digital corpora for this language. This is due mainly to the lack of a generally recognized standard variety which could serve as a reference point for the linguistic analysis needed for lemmatization and annotation.</p> <p>Lemmatization was done with the help of different lexicographic resources (https://www.deitsch.eu/news/view/9) most of which follow other spelling conventions. A fair number of word forms,&nbsp;especially English loanwords of some sort, cannot be found in the dictionaries. Moreover, the&nbsp;variety used here&nbsp;is characterized by a high variability regarding not only the spelling but also other aspects of the&nbsp;language.</p> <p>Part-of-speech tags were assigned manually (see tagsets A and B below). These tagsets for part-of-speech annotation of Deitsch texts are based on the 2017 version of the STTS system created and widely used for German (https://ids-pub.bsz-bw.de/frontdoor/deliver/index/docId/6063/file/Westpfahl_Schmidt_Jonietz_Borlinghaus_STTS_2_0_2017.pdf), which has been slightly modified and adapted to the corpus texts written in the Plain Deitsch variety. Tagset A gives a broader view and refers to the lemma level, tagset B is more fine-grained and suitable for&nbsp;&nbsp;the single word forms documented in the corpus. Only those tags are listed which are actually employed for the annotation of the corpus texts. Foreign items not integrated in the Deitsch text flow (e.g. English quotations) have been omitted.</p> <p>For more details about the corpus and the project please refer to the above mentioned website.</p>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record