Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.9.0
Dataset results
6 results for “Syllabics”
Handwritten Cree Syllabics
<p>An open dataset of labelled Cree syllabic handwriting samples.</p> <p>The dataset consists of handwriting sample images with corresponding Tesseract box files containing the label and area of each syllabic.</p>
Swahili syllabic Alphabet
<p>The syllabic alphabet outlines all the possible combination of consonants and vowels that serve as a basis for all the Swahili words. Therefore, the Swahili syllabic alphabet enumerates all the possible syllables that are used to construct Swahili words. The syllables are considered the smallest unit in Swahili and could consists of a vowel preceded by one to three consonants though there are special syllables made of single consonants or vowels. To derive the Swahili syllabic alphabet, we used the syllabification rules by and the digraphs and trigraphs proposed by Masengo. The syllables include prefixes that serve as Swahili noun class markers or subject prefixes (a, wa, vi, ki, m, mi), tense prefixes (na, li, ta, nge), relative prefixes (o, mo, ko, po, cho, vyo, lo, ye) and object markers (ki, vi, m, wa, mwu, ya, ji, mu, kwa, zi).</p>
Georgian: Syllabic Structure
<p>Lecture on syllabic structure, complex onsets and the sonority hierarchy of Georgian, containing exercises.</p> <p>This lecture is part of the lecture series:</p> <p><em>Glottothèque: Languages of the Anatolia, Caucasus, Iran, Mesopotamia; grammatical snippets online </em>(electronic resource). Bamberg, Cambridge, Göttingen, Moskow, Nicosia, Paris: LACIM network, at https://spw.uni-goettingen.de/projects/lacim/, edited by Christiane Bulut, Anaïd Donabédian-Demopoulos, Geoffrey Haig, Geoffrey Khan, Pollet Samvelian, Stavros Skopeteas, Nina Sumbatova.</p>
Relationship Between Poetic Meter and Meaning in Accentual-Syllabic Verse (data and replication code)
<ul> <li> <p><strong>main.py</strong>: script to train both lda and word2vec models</p> </li> <li> <p><strong>main.ipynb</strong>: Jupyter Notebook containing all the analyses reported in the paper</p> </li> <li> <p><strong>pos.ipynb</strong>: clustering based on frequencies of parts-of-speech</p> </li> <li><strong>corpora</strong>: contains original data for Czech, English, and Dutch poetry in JSON (proprietary German and Russian not included)</li> </ul> <pre><code>{ <= Each item in the following lists corresponds to particular poem and holds: 'words': [] <= list of lemmata found in the poem 'pos_tags': [] <= their POS-tags (Positional Morphological Tags for Czech, MyStem for Russian, TreeTagger tagsets for other corpora) 'meters': [[]] <= list of meters found in poem 'years': [] <= year when poem published (year when author born in case of English) 'n_words': [] <= number of words 'n_lines': [] <= number of lines 'authors': [] <= author of the poem 'titles': [] <= title of the poem 'schemes': [] <= line-ending schemes } </code></pre> <ul> <li><strong>dicts</strong>: contains Gensim dictionary files for all 5 corpora</li> </ul> <ul> <li><strong>fig</strong>: contains all resulting figures</li> </ul> <ul> <li><strong>json</strong> <strong>> metadata:</strong> contains all metadata on poems in particular corpora</li> </ul> <pre><code>{ <= Each item in the following lists corresponds to particular poem and holds: 'meters': [[]] <= list of meters found in poem 'years': [] <= year when poem published (year when author born in case of English) 'n_words': [] <= number of words 'n_lines': [] <= number of lines 'authors': [] <= author of the poem 'titles': [] <= title of the poem } </code></pre> <ul> <li><strong>json > topics:</strong> contains topic probabilities in particular poems</li> </ul> <pre><code>[ <= each item corresponds to particular poem and comprise 100-dimensional dict { 'topic title': its probability in poem } ] </code></pre> <ul> <li><strong>json > pos:</strong> contains POS relative frequencies in particular poems</li> </ul> <pre><code>[ <= each item corresponds to particular poem { 'POS': its frequency } ] </code></pre> <ul> <li><strong>json > w2v:</strong> contains mapping of lemmata and their neighbours in word2vec models</li> </ul> <ul> <li><strong>models</strong>: contains pretrained lda and word2vec models (Gensim)</li> <li><strong>regression</strong>: contains data and code to produce S5_table and S6_fig</li> </ul>
Duhumbi Phonology - Syllabic and Word Stress
<p>In Duhumbi, stress in general is non-distinctive, prosodic, and relatively unpronounced. In both disyllabic and polysyllabic words, stress falls on the first syllable. This also holds for polymorphemic lexemes, such as inflected words with suffixes. In glossary items in the Duhumbi lexicon, stress is indicated by a stress mark [ˈ] before the stressed syllable, whenever it is not predictable. Phrase stress is generally initial and falling towards the end of the phrase, combined with dependant-head word order in which subject/object precede the verb and we find postpositions rather than prepositions, although nouns always precede adjectives, but adverbs generally precede verbs. Word order is relatively flexible. </p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: bodttim (at) gmail (dot) com</p>
Data from: Neural tracking of syllabic and phonemic time scale
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.