Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8
datasets available to search
ShareScore release 0.7.1
Dataset results
8 results for “low resource languages”
Machine Assisted Translation of Wikipedia Articles into Low Resource Languages
<p><strong>Wikipedia is the largest encyclopedia ever assembled with the vision of enabling every human being to freely share in the sum of all knowledge. Wikipedia currently has a total of more than six million articles and over 17 billion words in its English edition. Unfortunately, millions of people cannot access this resource because it’s not available in their language. For instance, at the moment there are only 218 Tigrinya Wikipedia and 15,018 Amharic Wikipedia articles.</strong></p> <p><strong>In this project, we investigate the problem of translating Wikipedia articles from a high resource language into low resource languages using human-in-the-loop MT systems. In particular, we investigate different approaches to translate a sample of English Wikipedia articles into Tigrinya and Amharic. Currently, this repository contains 100k English Wikipeida abstracts translated using Lesan (https://lesan.ai) into Amharic and Tigrinya.</strong><br> <br> </p> <p><strong>Structure of data directory:</strong></p> <p><strong>data<br> ├── human<br> └── mt<br> ├── google<br> ├── lesan<br> │ ├── am.txt<br> │ ├── en.txt<br> │ └── ti.txt<br> └── microsoft</strong><br> </p>
Human review for post-training improvement of low-resource language performance in large language models
<p>Large language models (LLMs) have significantly improved natural language processing, holding the potential to support health workers and their clients directly. Unfortunately, there is a substantial and variable drop in performance for low-resource languages. Here we present results from an exploratory case study in Malawi, aiming to enhance the performance of LLMs in Chichewa through innovative prompt engineering techniques. By focusing on practical evaluations over traditional metrics, we assess the subjective utility of LLM outputs, prioritizing end-user satisfaction. Our findings suggest that tailored prompt engineering may improve LLM utility in underserved linguistic contexts, offering a promising avenue to bridge the language inclusivity gap in digital health interventions.</p>
Lesan: Machine Translation for Low Resource Languages
<p>Human evaluation dataset to evaluate machine translation systems to and from Amharic, English and Tigrinya.</p>
Human review for post-training improvement of low-resource language performance in large language models
Open the record for dataset details and reuse information.
[Replication Package] Enhancing Code Generation for Low-Resource Languages: No Silver Bullet
<p>This repository contains scripts and results related to the work <em>"Enhancing Code Generation for Low-Resource Languages: No Silver Bullet".<br><br></em><strong>Link to the GitHub repository: <a href="https://github.com/Devy99/low-resource-study">https://github.com/Devy99/low-resource-study</a></strong><em><br></em></p>
Chat-GPT MT: Competitive for High- (but not Low-) Resource Languages
<p>System outputs for the paper Chat-GPT MT: Competitive for High- (but not Low-) Resource Languages. Contains outputs for ChatGPT zero-shot and five-shot, GPT-4 five-shot, and Google Translate API</p>
An explorative Investigation into Neural Machine Translation: the Case of Low-Resource Language Pairs in Burkina Faso
<p>We explore the caveats of a promising research field, namely neural machine translation, in the context of preserving indigenous languages in Africa. We face the challenge of dealing with low-resource language pairs. Methodically, we employ some literature approaches and frameworks to learn translation models from Bible data in Moore (a major language in Burkina Faso) and French. Our experiments indeed confirm previous findings in the literature that vanilla neural machine translation models are ineffective for low resource language pairs. Surprisingly, however, we also found that we are not able to even remotely match performance recorded by the state of the art adapted methods for low-resource language pair in the literature. Nevertheless, although<br> word alignment, Byte Pair Encoding and Adam optimization did not successfully bring reasonable performance, we note that there are many remaining insightful approaches for low-resource language pairs.</p>
Afro-MNIST: Synthetic generation of MNIST-style datasets for low-resource languages
<p>We present Afro-MNIST, a set of synthetic MNIST-style datasets for four orthographies used in Afro-Asiatic and Niger-Congo languages: Ge`ez (Ethiopic), Vai, Osmanya, and N'Ko.<br> These datasets serve as ``drop-in'' replacements for MNIST. We hope that MNIST-style datasets will be developed for other numeral systems, and that these datasets vitalize machine learning education in underrepresented nations in the research community.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.