Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21
datasets available to search
ShareScore release 0.9.0
Dataset results
21 results for “Sign Language dataset”
Greek Text to Trajectories Sign Language Dataset
<p>Entails the 2D human pose trajectories of Greek Elementary Sign Language Dataset and Greek News Sign Language Dataset (31681 examples).</p>
SWL-LSE: SignaMed Word-Level LSE, a Dataset of Spanish Sign Language Health Signs
<h2>SWL-LSE Dataset</h2> <p>The SWL-LSE dataset is coined from SignaMed Word-Level LSE (Lengua de Signos Española -Spanish Sign Language).</p> <h2>Overview</h2> <p>The dataset consists of 8,000 sign sequences from 300 different sign classes related to the health domain. Each class is represented by an RGB video that serves as the dictionary sign. These dictionary signs were reproduced by 124 signers, including deaf individuals, interpreters, and L2 Spanish Sign Language (LSE) students, using their webcams or mobile phones via the SignaMed platform (<a href="https://signamed.web.app" target="_new" rel="noopener">https://signamed.web.app</a>). For privacy reasons, only the skeleton data is shared.</p> <p>The process of collecting the dataset is described in:</p> <p>Vázquez-Enríquez, M.; Alba-Castro, J.L.; Pérez-Pérez, A.; Cabeza-Pereiro, C.; Docío-Fernández, L. SignaMed: a Cooperative<br>Bilingual LSE-Spanish Dictionary in the Healthcare Domain. In Proceedings of the Proceedings of the LREC-COLING 2024<br>11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources; Efthimiou, E.; Fotinea, S.E.; Hanke, T.; Hochgesang, J.A.; Mesch, J.; Schulder, M., Eds., Torino, Italia, 2024; pp. 386–394. </p> <p>The dataset itself and the pipeline for training and executing a baseline model based on skeletons is described in this github (https://github.com/mvazquezgts/SWL-LSE), and this paper:</p> <p>Vázquez-Enríquez, M.; Alba-Castro, J.L.; Docío-Fernández, L.; Rodríguez-Banga, E. SWL-LSE: A Dataset of Spanish Sign Language Health Signs with an ISLR Baseline Method. Technologies 2024, 12(10), 205, D.O.I:10.3390/technologies12100205</p> <h2>Files</h2> <h3>1. VIDEOS_REF.zip</h3> <ul> <li><strong>Description</strong>: RGB videos recorded in lab conditions that represent each sign-class</li> <li><strong>Total files</strong>: 300</li> </ul> <h3>2. videos_ref_annotations.csv</h3> <ul> <li><strong>Description</strong>: CSV file with the correspondence between the name of the video, its class ID and gloss in spanish: FILENAME,CLASS_ID,LABEL.</li> <li><strong>Total files</strong>: 1</li> </ul> <h3>3. ANNOTATIONS.zip</h3> <ul> <li><strong>Description</strong>: 3 CSV files with train, validation and test file-class correspondences: FILENAME,CLASS_ID</li> <li><strong>Total files</strong>: 3</li> </ul> <h3>4. MEDIAPIPE.zip</h3> <ul> <li><strong>Description</strong>: Pickle files containing the full output of Mediapipe using their Heavy model. Each .pkl file contains the outputs of Mediapipe Holistic legacy, Mediapipe Pose and Mediapipe Hands. Each file is package as a dictionary: dict_keys(['pose', 'hands', 'holistic_legacy'])</li> <li><strong>Total files</strong>: 8000</li> </ul> <h2>Usage</h2> <p>Researchers and practitioners in pattern recognition, machine learning, and sign language linguistics may find this dataset valuable for:</p> <ul> <li>Training/testing machine learning models for isolated sign language recognition or gesture recognition.</li> <li>Analyzing patterns on signs realization</li> </ul> <h2>Acknowledgments</h2> <p>This dataset is a collaborative effort of the next research goups and entities:</p> <ul> <li><a href="http://gtm.uvigo.es/en/">Group of Multimedia Technologies (GTM)</a> from the <a href="https://atlanttic.uvigo.es/en">atlanTTic Research Center</a> of <a href="http://www.uvigo.es/">University of Vigo</a> (Spain)</li> <li><a href="http://grades.uvigo.gal/">Group of Discourse and Society (GRADES)</a> from the <a href="https://fft.uvigo.es/en/">School of Philology and Translation</a> of <a href="http://www.uvigo.es/">University of Vigo</a> (Spain)</li> <li><a href="http://www.faxpg.es/">Federation of Deaf People Galician Associations (FAXPG)</a></li> <li><a href="https://fundacioncnse-dilse.org">Fundación CNSE-DILSE</a></li> </ul> <p>Gratitude is extended to them for their contributions and support.</p>
Greek News Sign Language Dataset - Part A
<p>Part A entails 1.000 signed phrases of crime-related news stories broadcasted in Greece.</p>
Greek News Sign Language Dataset - Part B
<p>Part B entails 989 signed phrases of crime-related news stories broadcasted in Greece.</p>
Costarican Sign Language (LESCO) emergency-based signs dataset
<p>This dataset was part of Juan Zamora-Mora's doctoral dissertation on the recognition of Costarican Sign Language (LESCO) in emergency situations from Aspen University. The dataset is composed of 39 signs. There are three videos for each sign on each folder. Videos have been cropped and are on average 1 second long. This dataset contains a total of mp4 117 videos. </p>
ArabicSL-Net: A Benchmark Video Dataset for Arabic Words Sign Language
<p>The data was captured by mobile camera in four main organization namely Bank , Cafe , Hospital , and Train station. The ArabicSL-Net initially consists of 307 words recorded in approximately 30,000 videos. For each organization, we capture the most representative words that are used in those places. For Bank data, we have a total of 76 of words, while Cafe data contains 54 words. For Hospital, we collects videos for 102 words, and collects videos for 71 words in Train station.</p>
Greek Elementary Sign Language Dataset
<p>Entails the course material of the first years of elementary school in Greece. It includes 29.653 signed phrases that are present in the 33 issues of 13 distinct textbooks of the A, B and C years of Primary school.</p> <p>The Elementary Dataset consists of the following courses:</p> <p>9499 videos of Greek Language (1st, 2nd and 3rd year)<br> 6581 videos of Mathematics (1st, 2nd and 3rd year)<br> 4158 videos of Anthology of Greek Literacy (1st, 2nd, 3rd and 4th year)<br> 5521 videos of Environmental Studies (1st, 2nd and 3rd year)<br> 2067 videos of History (3rd year)<br> 1825 Videos of Religious Study (3rd year)</p> <p><br> Version 3 - Major Features and Improvements</p> <p>Converted Video format and Video FPS<br> Removed English words and characters<br> Improved Transcriptions</p>
SignBD-Word: Video-Based Bangla Word-Level Sign Language Dataset
<p>Bangla sign language (BdSL) is a complete and independent natural sign language with its own linguistic characteristics. While there exists video datasets for well-known sign languages, there is currently no available dataset for word-level BdSL. In this study, we present a video-based word-level dataset for Bangla sign language, called SignBD-Word, consisting of 6000 sign videos representing 200 unique words. The dataset includes full and upper-body views of the signers, along with 2D body pose information. This dataset can also be used as a benchmark for testing sign video classification algorithms.<br><br>Official Train Test Spllit (for both RGB and bodypose) can be found from the following link: <br>https://sites.google.com/view/signbd-word/dataset<br><br>This dataset is part of the following paper:<br>A. Sams, A. H. Akash and S. M. M. Rahman, "SignBD-Word: Video-Based Bangla Word-Level Sign Language and Pose Translation," 2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT), Delhi, India, 2023, pp. 1-7, doi: 10.1109/ICCCNT56998.2023.10306914.<br><br>Download the corresponding paper from this link:<br>https://asnsams.github.io/Publications.html</p>
BdSL47: A complete dataset of sign alphabet and digits of Bangla Sign Language (BdSL) using depth information via MediaPipe
<p><strong>BdSL47</strong> is the first open-access complete dataset in Bangla Sign Language that contains hand signs from both 10 sign digits (from sign ০ to sign ৯) and 37 sign alphabet (from sign অ to sign ँ).</p> <p>Dataset summary :</p> <ul> <li>100 RGB images per sign (total 47 signs) from each of 10 users</li> <li>Total input images : 100×47×10 = 47000</li> <li>Input images are processed via MediaPipe, which provided <ul> <li>an output image with hand key-points being detected</li> <li>3D coordinate values of 21 predefined key-points</li> <li>Total 63 coordinate values for each sample</li> </ul> </li> <li>The values are stored in csv files</li> <li>1 CSV file contains values from 100 samples of 1 sign from 1 user</li> <li>Total CSV files : 47×10 = 470</li> </ul> <p>The dataset has been made public for further research purposes. It is also available upon request <a href="https://drive.google.com/drive/u/8/folders/1wmJUlgWUrWNnOvzuL8Ci82Hm3zUx4wS-" rel="noopener">here</a>.</p>
Greek Elementary Sign Language Dataset scripts for loading dataset
<p>Entails py.scpripts for downloading the course material, loading video and text datasets, for each course of “Greek Elementary Sign Language Dataset”.</p>
GSLW - Greek Sign Language in the Wild Dataset
<p>GSLW: The dataset has been recorded under various background and lighting variations and with different smartphones. The camera position and orientation are gently varied among subsequent recordings to increase video diversity. <strong>15 cases</strong> of hearing-impaired people dealing with public services have been recorded. The resulting test dataset has <strong>1,736 videos</strong> that were annotated both at individual gloss and sentence level.<br> <br> Instructions:<br> <br> <strong><em> GSLW.csv</em></strong> file has the paths, gloss and sentence annotations of the videos.</p> <p><strong><em> load_gslw.py</em></strong> has a python function to load the GSLW.csv file for training or inference.<br> <br> Citation: <br> If you use our dataset please cite our work :</p> <pre>@article{gsl, author={N. M. {Adaloglou} and T. {Chatzis} and I. {Papastratis} and A. {Stergioulas} and G. T. {Papadopoulos} and V. {Zacharopoulou} and G. {Xydopoulos} and K. {Antzakas} and D. {Papazachariou} and P. n. {Daras}}, journal={IEEE Transactions on Multimedia}, title={A Comprehensive Study on Deep Learning-based Methods for Sign Language Recognition}, year={2021}, volume={}, number={}, pages={1-1}, doi={10.1109/TMM.2021.3070438}}</pre>
BdSL47: A complete dataset of sign alphabet and digits of Bangla Sign Language (BdSL) using depth information via MediaPipe
Open the record for dataset details and reuse information.
AzSLD - Azerbaijani Sign Language Dataset
<p>The Azerbaijani Sign Language Dataset (AzSLD) is a comprehensive, large dataset designed to facilitate the development and evaluation of machine learning models for the recognition and translation of Azerbaijani Sign Language (AzSL). </p> <p>AzSLD is the first publicly available dataset focused on Azerbaijani Sign Language. It contributes to the global effort to improve accessibility for the deaf and hard-of-hearing community in Azerbaijan. The dataset aims to bridge the gap between technology and accessibility by providing high-quality data for researchers, developers, and practitioners working on sign language recognition or translation systems.</p> <p>The data collection costs are covered by the "Strengthening Data Analytics Research and Training Capacity through Establishment of dual Master of Science in Computer Science and Master of Science in Data Analytics (MSCS/DA) degree program at ADA University" project, funded by BP and the Ministry of Education of the Republic of Azerbaijan.</p> <h3><strong>Dataset Composition</strong></h3> <p>AzSLD is organized into three primary components:</p> <h4>1. AzSLD_Sentences</h4> <p>This component contains video sequences of complete sentences in AzSL. It is designed to capture the fluidity and contextual nature of sign language, providing data for more complex language modeling tasks. It includes over 60 hours of high-definition video recordings, annotated with timestamped glosses for 500 distinct classes, enabling precise analysis and robust model training. Ground truth annotations of sentences for each class were added in a separate file. The videos were performed by 18 to 25 different signers, with a slight imbalance among them. </p> <p>2. AzSLD_Words<br>This component comprises a collection of short video samples representing frequently used words in Azerbaijani Sign Language. It is divided into two subsets:</p> <ul> <li>AzSLD_Words_100: Contains 100 commonly used words in AzSL.</li> <li>AzSLD_Words_200: Extends the first subset, including all 100 words from AzSLD_Words_100 along with an additional 100 words, for a total of 200 words.</li> </ul> <p>Folder names indicate the ground truth labels for the ease of word-level model evaluation.</p> <h4>3. AzSLD_Fingerspelling</h4> <p>This component includes over 14,000 video and image samples of letters of the Azerbaijani alphabet. Each sign is captured from multiple angles to ensure comprehensive coverage of dactylology in AzSL. This component is ideal for tasks involving letter recognition and the integration of fingerspelling into broader sign language recognition systems.</p> <h3><strong>Key Features</strong></h3> <h4>Double-View Recordings</h4> <p>The dataset includes 10,104 synchronized video recordings from two camera angles to capture both frontal and side views of hand and body movements, ensuring that the subtle nuances of sign language are well-represented.</p> <h4>Diverse Signers</h4> <p>The dataset features recordings from a diverse group of native AzSL signers, encompassing variations in age, gender, and signing style. This diversity is crucial for training models that are robust to variations in signing.</p> <h4>Detailed Annotations</h4> <p>Each video is annotated with comprehensive metadata, including the sign’s label (dactyl, word, or sentence), signer ID, and timestamped glosses for sentence-level signs. </p> <h4>High-Quality Data Format</h4> <p>The dataset comprises RGB videos in high-definition (HD) resolution at 35 frames per second, accompanied by JSON files containing annotations and metadata. The data is systematically organized into folders by category for ease of navigation.</p> <p><strong>Ethical Transparency</strong></p> <p>All participants provided informed consent for collecting, publishing, and using the data, ensuring compliance with ethical research standards.</p> <p><strong>Accessibility</strong></p> <p>The AzSLD is available under Creative Commons Attribution 4.0 International with free access for academic research through Zenodo.</p> <p><strong>Citation</strong>: When using AzSLD in your research, please cite the following paper:</p> <p>Alishzade, N., Hasanov, J. (2025). AzSLD: Azerbaijani sign language dataset for fingerspelling, word, and sentence translation with baseline software, Data in Brief, Volume 58, 2025, 111230, ISSN 2352-3409, <a title="https://url.au.m.mimecastprotect.com/s/szU6C2xMQziEvMn1kFBi9S5WqA6?domain=doi.org" href="https://url.au.m.mimecastprotect.com/s/szU6C2xMQziEvMn1kFBi9S5WqA6?domain=doi.org" target="_blank" rel="noopener noreferrer">https://doi.org/10.1016/j.dib.2024.111230</a>.</p> <p>The preprint is available at: <a href="https://arxiv.org/abs/2411.12865" target="_blank" rel="noopener">https://arxiv.org/abs/2411.12865</a> </p> <p><strong>Contact</strong>:<br>For questions, feedback, or contributions, please contact the project team at: <a rel="noopener">slr.project.ada@gmail.com</a></p>
Greek Elementary Sign Language Dataset
<p>Entails the course material of the first years of elementary school in Greece. It includes 29.698 signed phrases that are present in the 33 issues of 13 distinct textbooks of the A, B and C years of Primary school.</p> <p>The Elementary Dataset consists of the following courses:</p> <ul> <li>9507 videos of Greek Language (1st, 2nd and 3rd year)</li> <li>6599 videos of Mathematics (1st, 2nd and 3rd year)</li> <li>4163 videos of Anthology of Greek Literacy (1st, 2nd, 3rd and 4th year)</li> <li>5528 videos of Environmental Studies (1st, 2nd and 3rd year)</li> <li>2069 videos of History (3rd year)</li> <li>1832 Videos of Religious Study (3rd year)</li> </ul> <p>Version 2 - Major Features and Improvements</p> <ul> <li>Removed Duplicated Videos</li> <li>Improved Transcriptions</li> <li>Audio extraction (for Greek STT tasks)</li> </ul>
Greek Text to 3D Trajectories Sign Language Dataset
<p>Entails the 3D human pose trajectories of the Greek Elementary Sign Language Dataset.</p>
Arabic words sign language video dataset
<p>The dataset consists of 3000 unaltered videos captured by five volunteers, with each volunteer performing 20 repetitions of 30 signs, resulting in approximately 600 videos per volunteer. These videos were recorded without stabilization tools, reflecting the prevalent use of smartphones with built-in cameras. They encompass diverse resolutions, locations, places, and backgrounds, providing a comprehensive representation of real-life scenarios.</p>
INCLUDE: A Large Scale Dataset for Indian Sign Language Recognition
<p><strong>Dataset Details</strong>: The INCLUDE dataset has 4292 videos (the paper mentions 4287 videos but 5 videos were added later). The videos used for training are mentioned in train.csv (3475), while that used for testing is mentioned in test.csv (817 files). Each video is a recording of 1 ISL sign, signed by deaf students from St. Louis School for the Deaf, Adyar, Chennai.</p> <p>INCLUDE50 has 766 train videos and 192 test videos.</p> <p><strong>Train-Test Split:</strong> Please download the train-test split for INCLUDE and INCLUDE50 from here: <a href="https://drive.google.com/file/d/1tGjOR9xbk279ZUYnmk3BE-ktcWBdS24w/view?usp=sharing">Train-Test Split</a></p> <p><strong>Publication Link</strong><em>: <a href="https://dl.acm.org/doi/10.1145/3394171.3413528">https://dl.acm.org/doi/10.1145/3394171.3413528</a></em></p> <p><strong>AI4Bharat website</strong><em>: <a href="https://ai4bharat.org/include-dataset">https://sign-language.ai4bharat.org/</a></em></p> <p><strong>Download Instructions</strong></p> <p>For ease of access, we have prepared a Shell Script to download all the parts of the dataset and extract them to form the complete INCLUDE dataset.</p> <p><em>You can find the script here: <a href="http://bit.ly/include_dl">http://bit.ly/include_dl</a></em></p> <p><strong>Paper Abstract:</strong> <em>Indian Sign Language (ISL) is a complete language with its own grammar, syntax, vocabulary and several unique linguistic attributes. It is used by over 5 million deaf people in India. Currently, there is no publicly available dataset on ISL to evaluate Sign Language Recognition (SLR) approaches. In this work, we present the Indian Lexicon Sign Language Dataset - INCLUDE - an ISL dataset that contains 0.27 million frames across 4,287 videos over 263 word signs from 15 different word categories. INCLUDE is recorded with the help of experienced signers to provide close resemblance to natural conditions. A subset of 50 word signs is chosen across word categories to define INCLUDE-50 for rapid evaluation of SLR methods with hyperparameter tuning. The best performing model achieves an accuracy of 94.5% on the INCLUDE-50 dataset and 85.6% on the INCLUDE dataset</em></p>
FluentSigners-50: a signer independent benchmark dataset for Sign Language Processing
<p>A new large-scale Kazakh-Russian Sign Language dataset (FluentSigners-50) as a new Continuous Sign Language Recognition benchmark. FluentSigners-50 proposes to address three shortcomings of commonly used datasets: continuous signing, signer variety, and native signers. FluentSigners-50's main advantage is in its large signer variety: age (ranging from 8 to 57 years old), gender (18 male and 32 female), clothing, skin tone, body proportions, disability (deaf or hard of hearing), and fluency. Additionally, as the dataset was crowd-sourced: the participants were using a variety of their own recording devices (such as smartphones and web cameras), it resulted in a large variety of backgrounds, lighting conditions, camera quality, frame rates, camera aspect ratios, and angles. Finally, FluentSigners-50 contains recordings of 50 contributors that use sign language on a daily basis: either deaf, hard of hearing, hearing CODA (Child of Deaf Adults), and hearing SODA (Sibling of a Deaf Adult). As a result, the dataset contains a high degree of linguistic variability, including phonetic, phonological, lexical, and syntactic variations. It thus is a better training set for recognition of natural signing.</p> <p>The FluentSigners-50 dataset consists of everyday conversational phrases and sentences in KRSL, the sign language used in the Republic of Kazakhstan. KRSL is closely related to Russian Sign Language (RSL) and some other sign languages of the ex-Soviet Union. While no official research comparing KRSL with RSL exists, our observations based on our experience researching both languages are that they show a substantial lexical overlap and are entirely mutually intelligible. The sentences and phrases of FluentSigners-50 represent the following sentence types: statements, polar questions, wh-questions, and requests.</p> <p>All FluentSigners-50 contributors use sign language on a daily basis as they are either deaf (N=32), hard of hearing (N=6), hearing SODA (N=3), or hearing CODA (N=9). Native signers are signers who have been exposed to signed languages since birth because their parents are deaf. While the early acquisition may be necessary for the development of native language abilities, other factors, particularly the quality of language input, may play a role. According to this distinction, FluentSigners-50 has 30 CODA contributors (including nine hearing signers) and 20 who are not CODA (16 deaf, one hard of hearing, and three hearing SODA). Nevertheless, we decided to name our dataset FluentSigners-50 because all of our contributors use sign language daily, and it is their primary language of communication. They all came from various regions of Kazakhstan and are of different age and gender groups.</p>
Dataset for Usability Study of a Sign Language Learning Tool
Open the record for dataset details and reuse information.
SMILE Swiss German Sign Language Dataset
<p><strong>Description</strong></p> <p>The SMILE Swiss German Sign Language Dataset consists of videos, joint coordinates and annotations of 100 isolated signs of a Swiss German Sign Language (Deutschschweizerische Gebärdensprache, DSGS) vocabulary production test. All items were produced multiple times by 16 adult L1 signers and 22 adult L2 learners of DSGS. Associated linguistic transcriptions and annotations are available for second path data of 10 adult L1 signers and 18 adult L2 learners of DSGS.</p> <p>The dataset has been created in the context of developing an assessment system for lexical signs of DSGS in the SNSF project SMILE.</p> <p>More precisely, for each participant, the following files are available:</p> <ul> <li>Kinect color video (.mp4); 1920x1080 Pixels @ 30 FPS;</li> <li>Kinect Pose Information (.csv); 25 Joints; 3D Joint Coordinates and Angles;</li> <li>OpenPose output (.json); 2D Joint Coordinates and Confidences;</li> <li>iLex annotation files (.xml); linguistic annotations.</li> </ul> <p> </p> <p><strong>Reference</strong></p> <p>If you use this database, please cite the following publication:</p> <p><em>Sarah Ebling, Necati Cihan Camgöz, Penny Boyes Braem, Katja Tissi, Sandra Sidler-Miserez, Stephanie Stoll, Simon Hadfield, Tobias Haug, Richard Bowden, Sandrine Tornay, Marzieh Razavi, and Mathew Magimai-Doss. SMILE Swiss German Sign Language Dataset. In Proceedings of the 11th Language Resources and Evaluation Conference (LREC 2018), pages 4221–4229, 2018.</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.