Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.7.1
Dataset results
11 results for “swahili”
Swahili : News Classification Dataset
<p>Swahili is spoken by 100-150 million people across East Africa. In Tanzania, it is one of two national languages (the other is English) and it is the official language of instruction in all schools. News in Swahili is an important part of the media sphere in Tanzania.</p> <p>News contributes to education, technology, and the economic growth of a country, and news in local languages plays an important cultural role in many Africa countries. In the modern age, African languages in news and other spheres are at risk of being lost as English becomes the dominant language in online spaces.<br> <br> The Swahili news dataset was created to reduce the gap of using the Swahili language to create NLP technologies and help AI practitioners in Tanzania and across Africa continent to practice their NLP skills to solve different problems in organizations or societies related to Swahili language. Swahili News were collected from different websites that provide news in the Swahili language. I was able to find some websites that provide news in Swahili only and others in different languages including Swahili.<br> <br> The dataset was created for a specific task of text classification, this means each news content can be categorized into six different topics (Local news, International news , Finance news, Health news, Sports news, and Entertainment news). The dataset comes with a specified train/test split. The train set contains 75% of the dataset and test set contains 25% of the dataset.</p> <p><strong>Acknowledgment</strong>: This project was supported by the <a href="https://www.k4all.org/project/language-dataset-fellowship/">AI4D language dataset fellowship</a> through K4All and <a href="https://zindi.africa/">Zindi Africa</a>.</p>
Swahili syllabic Alphabet
<p>The syllabic alphabet outlines all the possible combination of consonants and vowels that serve as a basis for all the Swahili words. Therefore, the Swahili syllabic alphabet enumerates all the possible syllables that are used to construct Swahili words. The syllables are considered the smallest unit in Swahili and could consists of a vowel preceded by one to three consonants though there are special syllables made of single consonants or vowels. To derive the Swahili syllabic alphabet, we used the syllabification rules by and the digraphs and trigraphs proposed by Masengo. The syllables include prefixes that serve as Swahili noun class markers or subject prefixes (a, wa, vi, ki, m, mi), tense prefixes (na, li, ta, nge), relative prefixes (o, mo, ko, po, cho, vyo, lo, ye) and object markers (ki, vi, m, wa, mwu, ya, ji, mu, kwa, zi).</p>
Swahili word analogy dataset
<p>Swahili Analogy dataset contains pairs of words that are organized in 4's to facilitate word analogy test. Word analogy test is used to evaluate the quality of word representation vectors from a language model. The dataset contains 12,864 questions that have been organized in 12 categories.</p>
Swahili Image Captioning Dataset
<p>The SwaFlickr8k dataset is an extension of the well-known Flickr8k dataset, specifically designed for image captioning tasks. It includes a collection of images and corresponding captions written in Swahili. With 8,091 unique images and 40,455 captions, this dataset provides a valuable resource for research and development in the field of image understanding and language processing, particularly in the context of Swahili language.</p>
Unannoted Swahili data
<p>The Unannoted Swahili dataset contains 28,000 unique words with 6.84M, 970k, and 2M words for the train, development and test partitions respectively which represent the ratio 80:10:10. The dataset is meant for language modeling. The entire dataset is lowercased, has no punctuation marks and, the start and end of sentence markers have been incorporated to facilitate easy tokenization during language modeling. The train partition is the largest in order to support unsupervised learning of word representations while the hyper-parameters are adjusted based on the performance on the development partition before evaluating the language model on the test partition. </p>
Language modeling data for Swahili
<p>The Swahili dataset developed specifically for language modeling task. The dataset contains 28,000 unique words with 6.84M, 970k, and 2M words for the train, valid and test partitions respectively which represent the ratio 80:10:10. The entire dataset is lowercased, has no punctuation marks and, the start and end of sentence markers have been incorporated to facilitate easy tokenization during language modeling. The train partition is the largest in order to support unsupervised learning of word representations while the hyper-parameters are adjusted based on the performance on the valid partition before evaluating the language model on the test partition.</p>
Swahili Staff
Top section of a 19th or 20th-century Swahili staff, now in the collection of the Minneapolis Institute of Art. More information about the object here: https://collections.artsmia.org/art/12117/staff-swahili A point cloud version of the object is here: https://skfb.ly/6qE9t Source: Objaverse 1.0 / Sketchfab
Swahili Image Captioning Dataset
<p>The SwaFlickr8k dataset is an extension of the well-known Flickr8k dataset, specifically designed for image captioning tasks. It includes a collection of images and corresponding captions written in Swahili. With 8,091 unique images and 40,455 captions, this dataset provides a valuable resource for research and development in the field of image understanding and language processing, particularly in the context of the Swahili language.</p>
Clinic Waiting Room-based Study of Swahili Language Artificial Intelligence-driven Symptom Assessments in Tanzanian Primary Health Care Facilities
ClinicalTrials.gov study NCT04958577. IPD Sharing: NO. Countries: 1. Publications: 2.
Swahili Staff - point cloud
Top section of a 19th or 20th-century Swahili staff, now in the collection of the Minneapolis Institute of Art. This point cloud version was inspired by Sketchfab's #plantpointschallenge contest, but it's not a contest entry. More information about the object here: https://collections.artsmia.org/art/12117/staff-swahili The mesh version of this object is here on Sketchfab: https://skfb.ly/6qEvt Source: Objaverse 1.0 / Sketchfab
SWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASET
<p>This dataset consist of swahili and code switched English-Swahili tweets labeled for hate speech type,the target and language.</p> <p>The dataset can be accessed upon request from the authors via email addresses:</p> <p> -nodhianbo@gmail.com</p> <p> -endeshanelly@gmail.com</p> <p> -mscci00083@student.maseno.ac.ke</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.