Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
59
datasets available to search
ShareScore release 0.9.0
Dataset results
59 results for “sign languages”
Mocap video examples for the analysis of Sign Language motion
<p>These mocap videos support my PhD thesis "Extracting human characteristics from motion: the case of identity in Sign Language" carried out from October 2018 to October 2021. The original mocap data is taken from the <a href="https://www.ortolang.fr/market/corpora/mocap1/">MOCAP1</a> corpus of French Sign Language. The videos have been generated using Python code available as part of the <a href="https://github.com/felixbgd/PLmocap">PLmocap</a> library.</p>
Mocap video examples for the analysis of Sign Language motion
<p>These mocap videos support my PhD thesis "Extracting human characteristics from motion: the case of identity in Sign Language" carried out from October 2018 to October 2021. The original mocap data is taken from the <a href="https://www.ortolang.fr/market/corpora/mocap1/">MOCAP1</a> corpus of French Sign Language. The videos have been generated using Python code available as part of the <a href="https://github.com/felixbgd/PLmocap">PLmocap</a> library.</p>
Mocap video examples for the analysis of Sign Language motion
<p>These mocap videos support my PhD thesis "Extracting human characteristics from motion: the case of identity in Sign Language" carried out from October 2018 to October 2021. The original mocap data is taken from the <a href="https://www.ortolang.fr/market/corpora/mocap1/">MOCAP1</a> corpus of French Sign Language. The videos have been generated using Python code available as part of the <a href="https://github.com/felixbgd/PLmocap">PLmocap</a> library.</p> <p> </p>
Greek Elementary Sign Language Dataset scripts for loading dataset
<p>Entails py.scpripts for downloading the course material, loading video and text datasets, for each course of “Greek Elementary Sign Language Dataset”.</p>
GSLW - Greek Sign Language in the Wild Dataset
<p>GSLW: The dataset has been recorded under various background and lighting variations and with different smartphones. The camera position and orientation are gently varied among subsequent recordings to increase video diversity. <strong>15 cases</strong> of hearing-impaired people dealing with public services have been recorded. The resulting test dataset has <strong>1,736 videos</strong> that were annotated both at individual gloss and sentence level.<br> <br> Instructions:<br> <br> <strong><em> GSLW.csv</em></strong> file has the paths, gloss and sentence annotations of the videos.</p> <p><strong><em> load_gslw.py</em></strong> has a python function to load the GSLW.csv file for training or inference.<br> <br> Citation: <br> If you use our dataset please cite our work :</p> <pre>@article{gsl, author={N. M. {Adaloglou} and T. {Chatzis} and I. {Papastratis} and A. {Stergioulas} and G. T. {Papadopoulos} and V. {Zacharopoulou} and G. {Xydopoulos} and K. {Antzakas} and D. {Papazachariou} and P. n. {Daras}}, journal={IEEE Transactions on Multimedia}, title={A Comprehensive Study on Deep Learning-based Methods for Sign Language Recognition}, year={2021}, volume={}, number={}, pages={1-1}, doi={10.1109/TMM.2021.3070438}}</pre>
LOOKing for Multi-word Expressions in American Sign Language
Open the record for dataset details and reuse information.
BdSL47: A complete dataset of sign alphabet and digits of Bangla Sign Language (BdSL) using depth information via MediaPipe
Open the record for dataset details and reuse information.
Computational phylogenetics reveal the history of sign languages
Open the record for dataset details and reuse information.
Videos to accompany publication "The Synthesis of Complex Shape Deployments in Sign Language" (Filhol & McDonald, 2020)
<p>Videos to accompany publication "The Synthesis of Complex Shape Deployments in Sign Language" (Filhol & McDonald, 2020)</p>
Mexican Sign Language Alphabet (static signs only)
<h2>Dataset Description: Mexican Sign Language Alphabet</h2><p> </p><h3>Overview:</h3><p>The Mexican Sign Language Alphabet dataset is a comprehensive collection of static signs representing the Mexican Sign Language (LSM) alphabet. This dataset is designed to facilitate research and development in sign language recognition and understanding. It includes 21 distinct signs, covering all letters except J, K, Ñ, Q, X, and Z, which are dynamic signs.</p><p> </p><h3>Data Representation:</h3><ul><li>The dataset comprises images of hand signs, captured against a uniform green background.</li><li>The signs are represented as static images, allowing for a clear and standardized view of each sign.</li><li>A total of 20 participants were involved in capturing the dataset, ensuring a diverse range of sign representations.</li></ul><p> </p><h3>Data Collection:</h3><ul><li>The data collection process employed a green screen and constant illumination to maintain consistent visual quality across all images.</li><li>To capture a wide range of variations, the dataset is divided into three distinct groups: A, B, and C.<ul><li>Group A contains images with low variation, primarily focusing on light rotations among all degrees of freedom.</li><li>Group B features images with rotations along three axes (yaw, roll, pitch), introducing variations in sign orientation.</li><li>Group C comprises images with high variation, emphasizing pronounced movements along all degrees of freedom.</li></ul></li></ul><p> </p><h3>Data Partitioning:</h3><ul><li>The dataset is structured into train and test partitions, ensuring a reliable evaluation of model performance.</li><li>The train partition contains data from 18 randomly selected participants, while the test partition includes data from the remaining 2 participants.</li></ul><p> </p><h3>Directory Structure:</h3><ul><li>The root folder contains subdirectories for the three distinct variation groups: lss-abc-A, lsm-abc-B, lsm-C.</li><li>Each of these variation groups is further divided into "train" and "test" partitions.</li><li>Within the "train" and "test" partitions, you will find subdirectories labeled with class names, representing individual signs (e.g., A, B, C, D, ..., Y).</li><li>Inside each class directory, you will find images in jpg format, depicting the corresponding sign. These images are organized for training and evaluation purposes.</li></ul><p> </p>
AzSLD - Azerbaijani Sign Language Dataset
<p>The Azerbaijani Sign Language Dataset (AzSLD) is a comprehensive, large dataset designed to facilitate the development and evaluation of machine learning models for the recognition and translation of Azerbaijani Sign Language (AzSL). </p> <p>AzSLD is the first publicly available dataset focused on Azerbaijani Sign Language. It contributes to the global effort to improve accessibility for the deaf and hard-of-hearing community in Azerbaijan. The dataset aims to bridge the gap between technology and accessibility by providing high-quality data for researchers, developers, and practitioners working on sign language recognition or translation systems.</p> <p>The data collection costs are covered by the "Strengthening Data Analytics Research and Training Capacity through Establishment of dual Master of Science in Computer Science and Master of Science in Data Analytics (MSCS/DA) degree program at ADA University" project, funded by BP and the Ministry of Education of the Republic of Azerbaijan.</p> <h3><strong>Dataset Composition</strong></h3> <p>AzSLD is organized into three primary components:</p> <h4>1. AzSLD_Sentences</h4> <p>This component contains video sequences of complete sentences in AzSL. It is designed to capture the fluidity and contextual nature of sign language, providing data for more complex language modeling tasks. It includes over 60 hours of high-definition video recordings, annotated with timestamped glosses for 500 distinct classes, enabling precise analysis and robust model training. Ground truth annotations of sentences for each class were added in a separate file. The videos were performed by 18 to 25 different signers, with a slight imbalance among them. </p> <p>2. AzSLD_Words<br>This component comprises a collection of short video samples representing frequently used words in Azerbaijani Sign Language. It is divided into two subsets:</p> <ul> <li>AzSLD_Words_100: Contains 100 commonly used words in AzSL.</li> <li>AzSLD_Words_200: Extends the first subset, including all 100 words from AzSLD_Words_100 along with an additional 100 words, for a total of 200 words.</li> </ul> <p>Folder names indicate the ground truth labels for the ease of word-level model evaluation.</p> <h4>3. AzSLD_Fingerspelling</h4> <p>This component includes over 14,000 video and image samples of letters of the Azerbaijani alphabet. Each sign is captured from multiple angles to ensure comprehensive coverage of dactylology in AzSL. This component is ideal for tasks involving letter recognition and the integration of fingerspelling into broader sign language recognition systems.</p> <h3><strong>Key Features</strong></h3> <h4>Double-View Recordings</h4> <p>The dataset includes 10,104 synchronized video recordings from two camera angles to capture both frontal and side views of hand and body movements, ensuring that the subtle nuances of sign language are well-represented.</p> <h4>Diverse Signers</h4> <p>The dataset features recordings from a diverse group of native AzSL signers, encompassing variations in age, gender, and signing style. This diversity is crucial for training models that are robust to variations in signing.</p> <h4>Detailed Annotations</h4> <p>Each video is annotated with comprehensive metadata, including the sign’s label (dactyl, word, or sentence), signer ID, and timestamped glosses for sentence-level signs. </p> <h4>High-Quality Data Format</h4> <p>The dataset comprises RGB videos in high-definition (HD) resolution at 35 frames per second, accompanied by JSON files containing annotations and metadata. The data is systematically organized into folders by category for ease of navigation.</p> <p><strong>Ethical Transparency</strong></p> <p>All participants provided informed consent for collecting, publishing, and using the data, ensuring compliance with ethical research standards.</p> <p><strong>Accessibility</strong></p> <p>The AzSLD is available under Creative Commons Attribution 4.0 International with free access for academic research through Zenodo.</p> <p><strong>Citation</strong>: When using AzSLD in your research, please cite the following paper:</p> <p>Alishzade, N., Hasanov, J. (2025). AzSLD: Azerbaijani sign language dataset for fingerspelling, word, and sentence translation with baseline software, Data in Brief, Volume 58, 2025, 111230, ISSN 2352-3409, <a title="https://url.au.m.mimecastprotect.com/s/szU6C2xMQziEvMn1kFBi9S5WqA6?domain=doi.org" href="https://url.au.m.mimecastprotect.com/s/szU6C2xMQziEvMn1kFBi9S5WqA6?domain=doi.org" target="_blank" rel="noopener noreferrer">https://doi.org/10.1016/j.dib.2024.111230</a>.</p> <p>The preprint is available at: <a href="https://arxiv.org/abs/2411.12865" target="_blank" rel="noopener">https://arxiv.org/abs/2411.12865</a> </p> <p><strong>Contact</strong>:<br>For questions, feedback, or contributions, please contact the project team at: <a rel="noopener">slr.project.ada@gmail.com</a></p>
Greek Elementary Sign Language Dataset
<p>Entails the course material of the first years of elementary school in Greece. It includes 29.698 signed phrases that are present in the 33 issues of 13 distinct textbooks of the A, B and C years of Primary school.</p> <p>The Elementary Dataset consists of the following courses:</p> <ul> <li>9507 videos of Greek Language (1st, 2nd and 3rd year)</li> <li>6599 videos of Mathematics (1st, 2nd and 3rd year)</li> <li>4163 videos of Anthology of Greek Literacy (1st, 2nd, 3rd and 4th year)</li> <li>5528 videos of Environmental Studies (1st, 2nd and 3rd year)</li> <li>2069 videos of History (3rd year)</li> <li>1832 Videos of Religious Study (3rd year)</li> </ul> <p>Version 2 - Major Features and Improvements</p> <ul> <li>Removed Duplicated Videos</li> <li>Improved Transcriptions</li> <li>Audio extraction (for Greek STT tasks)</li> </ul>
Greek Text to 3D Trajectories Sign Language Dataset
<p>Entails the 3D human pose trajectories of the Greek Elementary Sign Language Dataset.</p>
American Sign Languages 50x50 images
<p>The Dataset is collected with the Webcam Considering 1100 image files for each folder containing A to Z letters. A total of 28,600 Files are present in the Folders.</p>
Arabic words sign language video dataset
<p>The dataset consists of 3000 unaltered videos captured by five volunteers, with each volunteer performing 20 repetitions of 30 signs, resulting in approximately 600 videos per volunteer. These videos were recorded without stabilization tools, reflecting the prevalent use of smartphones with built-in cameras. They encompass diverse resolutions, locations, places, and backgrounds, providing a comprehensive representation of real-life scenarios.</p>
Assessing the Impact of Vidéo Remote Sign Language Interpreting in Healthcare
ClinicalTrials.gov study NCT05966623. IPD Sharing: NO. Countries: 1. Publications: 2.
Health Protection and Promotion of Sign Language Interpreters Through Implementation of Total Worker Health®
ClinicalTrials.gov study NCT06058949. IPD Sharing: Not stated. Countries: 1. Publications: 2.
Effect of Infant Sign Training on Speech-language Development
ClinicalTrials.gov study NCT06143254. IPD Sharing: NO. Countries: 1. Publications: 7.
Virtual Reality Glasses Integrated With Sign Language on Dental Anxiety Among Children With Hearing Impairment During Pulpotomy Procedure
ClinicalTrials.gov study NCT06153823. IPD Sharing: Not stated. Countries: 1. Publications: 1.
INCLUDE: A Large Scale Dataset for Indian Sign Language Recognition
<p><strong>Dataset Details</strong>: The INCLUDE dataset has 4292 videos (the paper mentions 4287 videos but 5 videos were added later). The videos used for training are mentioned in train.csv (3475), while that used for testing is mentioned in test.csv (817 files). Each video is a recording of 1 ISL sign, signed by deaf students from St. Louis School for the Deaf, Adyar, Chennai.</p> <p>INCLUDE50 has 766 train videos and 192 test videos.</p> <p><strong>Train-Test Split:</strong> Please download the train-test split for INCLUDE and INCLUDE50 from here: <a href="https://drive.google.com/file/d/1tGjOR9xbk279ZUYnmk3BE-ktcWBdS24w/view?usp=sharing">Train-Test Split</a></p> <p><strong>Publication Link</strong><em>: <a href="https://dl.acm.org/doi/10.1145/3394171.3413528">https://dl.acm.org/doi/10.1145/3394171.3413528</a></em></p> <p><strong>AI4Bharat website</strong><em>: <a href="https://ai4bharat.org/include-dataset">https://sign-language.ai4bharat.org/</a></em></p> <p><strong>Download Instructions</strong></p> <p>For ease of access, we have prepared a Shell Script to download all the parts of the dataset and extract them to form the complete INCLUDE dataset.</p> <p><em>You can find the script here: <a href="http://bit.ly/include_dl">http://bit.ly/include_dl</a></em></p> <p><strong>Paper Abstract:</strong> <em>Indian Sign Language (ISL) is a complete language with its own grammar, syntax, vocabulary and several unique linguistic attributes. It is used by over 5 million deaf people in India. Currently, there is no publicly available dataset on ISL to evaluate Sign Language Recognition (SLR) approaches. In this work, we present the Indian Lexicon Sign Language Dataset - INCLUDE - an ISL dataset that contains 0.27 million frames across 4,287 videos over 263 word signs from 15 different word categories. INCLUDE is recorded with the help of experienced signers to provide close resemblance to natural conditions. A subset of 50 word signs is chosen across word categories to define INCLUDE-50 for rapid evaluation of SLR methods with hyperparameter tuning. The best performing model achieves an accuracy of 94.5% on the INCLUDE-50 dataset and 85.6% on the INCLUDE dataset</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.