Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
59
datasets available to search
ShareScore release 0.9.0
Dataset results
59 results for “sign languages”
Greek Text to Trajectories Sign Language Dataset
<p>Entails the 2D human pose trajectories of Greek Elementary Sign Language Dataset and Greek News Sign Language Dataset (31681 examples).</p>
SWL-LSE: SignaMed Word-Level LSE, a Dataset of Spanish Sign Language Health Signs
<h2>SWL-LSE Dataset</h2> <p>The SWL-LSE dataset is coined from SignaMed Word-Level LSE (Lengua de Signos Española -Spanish Sign Language).</p> <h2>Overview</h2> <p>The dataset consists of 8,000 sign sequences from 300 different sign classes related to the health domain. Each class is represented by an RGB video that serves as the dictionary sign. These dictionary signs were reproduced by 124 signers, including deaf individuals, interpreters, and L2 Spanish Sign Language (LSE) students, using their webcams or mobile phones via the SignaMed platform (<a href="https://signamed.web.app" target="_new" rel="noopener">https://signamed.web.app</a>). For privacy reasons, only the skeleton data is shared.</p> <p>The process of collecting the dataset is described in:</p> <p>Vázquez-Enríquez, M.; Alba-Castro, J.L.; Pérez-Pérez, A.; Cabeza-Pereiro, C.; Docío-Fernández, L. SignaMed: a Cooperative<br>Bilingual LSE-Spanish Dictionary in the Healthcare Domain. In Proceedings of the Proceedings of the LREC-COLING 2024<br>11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources; Efthimiou, E.; Fotinea, S.E.; Hanke, T.; Hochgesang, J.A.; Mesch, J.; Schulder, M., Eds., Torino, Italia, 2024; pp. 386–394. </p> <p>The dataset itself and the pipeline for training and executing a baseline model based on skeletons is described in this github (https://github.com/mvazquezgts/SWL-LSE), and this paper:</p> <p>Vázquez-Enríquez, M.; Alba-Castro, J.L.; Docío-Fernández, L.; Rodríguez-Banga, E. SWL-LSE: A Dataset of Spanish Sign Language Health Signs with an ISLR Baseline Method. Technologies 2024, 12(10), 205, D.O.I:10.3390/technologies12100205</p> <h2>Files</h2> <h3>1. VIDEOS_REF.zip</h3> <ul> <li><strong>Description</strong>: RGB videos recorded in lab conditions that represent each sign-class</li> <li><strong>Total files</strong>: 300</li> </ul> <h3>2. videos_ref_annotations.csv</h3> <ul> <li><strong>Description</strong>: CSV file with the correspondence between the name of the video, its class ID and gloss in spanish: FILENAME,CLASS_ID,LABEL.</li> <li><strong>Total files</strong>: 1</li> </ul> <h3>3. ANNOTATIONS.zip</h3> <ul> <li><strong>Description</strong>: 3 CSV files with train, validation and test file-class correspondences: FILENAME,CLASS_ID</li> <li><strong>Total files</strong>: 3</li> </ul> <h3>4. MEDIAPIPE.zip</h3> <ul> <li><strong>Description</strong>: Pickle files containing the full output of Mediapipe using their Heavy model. Each .pkl file contains the outputs of Mediapipe Holistic legacy, Mediapipe Pose and Mediapipe Hands. Each file is package as a dictionary: dict_keys(['pose', 'hands', 'holistic_legacy'])</li> <li><strong>Total files</strong>: 8000</li> </ul> <h2>Usage</h2> <p>Researchers and practitioners in pattern recognition, machine learning, and sign language linguistics may find this dataset valuable for:</p> <ul> <li>Training/testing machine learning models for isolated sign language recognition or gesture recognition.</li> <li>Analyzing patterns on signs realization</li> </ul> <h2>Acknowledgments</h2> <p>This dataset is a collaborative effort of the next research goups and entities:</p> <ul> <li><a href="http://gtm.uvigo.es/en/">Group of Multimedia Technologies (GTM)</a> from the <a href="https://atlanttic.uvigo.es/en">atlanTTic Research Center</a> of <a href="http://www.uvigo.es/">University of Vigo</a> (Spain)</li> <li><a href="http://grades.uvigo.gal/">Group of Discourse and Society (GRADES)</a> from the <a href="https://fft.uvigo.es/en/">School of Philology and Translation</a> of <a href="http://www.uvigo.es/">University of Vigo</a> (Spain)</li> <li><a href="http://www.faxpg.es/">Federation of Deaf People Galician Associations (FAXPG)</a></li> <li><a href="https://fundacioncnse-dilse.org">Fundación CNSE-DILSE</a></li> </ul> <p>Gratitude is extended to them for their contributions and support.</p>
Greek News Sign Language Dataset - Part A
<p>Part A entails 1.000 signed phrases of crime-related news stories broadcasted in Greece.</p>
Greek News Sign Language Dataset - Part B
<p>Part B entails 989 signed phrases of crime-related news stories broadcasted in Greece.</p>
Costarican Sign Language (LESCO) emergency-based signs dataset
<p>This dataset was part of Juan Zamora-Mora's doctoral dissertation on the recognition of Costarican Sign Language (LESCO) in emergency situations from Aspen University. The dataset is composed of 39 signs. There are three videos for each sign on each folder. Videos have been cropped and are on average 1 second long. This dataset contains a total of mp4 117 videos. </p>
Figure 1. The Structure of the Automatic Translate Voice to Sign Language Animation System-Development an Automatic Speech to Facial Animation Conversion for Improve Deaf Lives
<p>All technologies of voice recognition, speaker identification and verification, each has its<br> own advantages and disadvantages and may requires different treatments and techniques. The<br> choice of which technology to use is application-specific. At the highest level, all voice recognition<br> systems contain two main modules: feature extraction and feature matching. Feature extraction is<br> the process that extracts a small amount of data from the voice signal that can later be used to<br> represent each word. Feature matching involves the actual procedure to identify the unknown word<br> by comparing extracted features from his/her voice input with the ones from a set of known words.<br> A wide range of possibilities exist for parametrically representing the speech signal for the<br> voice recognition task, such as Linear Prediction Coding (LPC), RASTA-PLP and Mel-Frequency<br> Cepstrum Coefficients (MFCC).</p>
Systematic mapping data for translation-enabling technologies for sign languages
<p>These data correspond to the papers selected for a systematic mapping of that translation-enabling technologies for sign languages.</p>
lexibank/powerma: Evolutionary Dynamics in the Dispersal of Sign Languages
<p>CLDF dataset accompanying the study "Evolutionary Dynamics in the Dispersal of Sign Languages"</p>
Thesis summary in South African Sign Language
<p>The summary of the doctoral thesis: Community-Based Co-Design for Accessible Health Information for Deaf People in a Context with Societal Complexity is presented in South Africa Sign Language to provide information accessible to Deaf people. </p>
ArabicSL-Net: A Benchmark Video Dataset for Arabic Words Sign Language
<p>The data was captured by mobile camera in four main organization namely Bank , Cafe , Hospital , and Train station. The ArabicSL-Net initially consists of 307 words recorded in approximately 30,000 videos. For each organization, we capture the most representative words that are used in those places. For Bank data, we have a total of 76 of words, while Cafe data contains 54 words. For Hospital, we collects videos for 102 words, and collects videos for 71 words in Train station.</p>
Greek Elementary Sign Language Dataset
<p>Entails the course material of the first years of elementary school in Greece. It includes 29.653 signed phrases that are present in the 33 issues of 13 distinct textbooks of the A, B and C years of Primary school.</p> <p>The Elementary Dataset consists of the following courses:</p> <p>9499 videos of Greek Language (1st, 2nd and 3rd year)<br> 6581 videos of Mathematics (1st, 2nd and 3rd year)<br> 4158 videos of Anthology of Greek Literacy (1st, 2nd, 3rd and 4th year)<br> 5521 videos of Environmental Studies (1st, 2nd and 3rd year)<br> 2067 videos of History (3rd year)<br> 1825 Videos of Religious Study (3rd year)</p> <p><br> Version 3 - Major Features and Improvements</p> <p>Converted Video format and Video FPS<br> Removed English words and characters<br> Improved Transcriptions</p>
A collection of Sign Language Narations by Virtual Humans in the Greek Sign Language
<p>A collection of Sign Language Narations by Virtual Humans in the Greek Sign Language in the context of the Mastic Pilot of Mingei.</p>
BSL-Hansard: A parallel, multimodal corpus of English and interpreted British Sign Language data from parliamentary proceedings
<p>BSL-Hansard is a novel open source and multimodal resource composed by combining Sign Language video data in BSL and English text from the official transcription of British parliamentary sessions. This paper describes the method followed to compile BSL-Hansard including time alignment of text using the MAUS (Schiel, 2015) segmentation system, gives some statistics about this dataset, and suggests experiments. These primarily include end-to-end Sign Language-to-text translation, but is also relevant for broader machine translation, and speech and language processing tasks.</p> <p>This dataset will be useful for translation between BSL and English, or for studies in BSL or English down to the phonetic level.</p>
Sign Languages Phylogeny
<p>Code for the replication of the results from the article "Computational phylogenetics reveal the history of sign languages", and to apply the same analysis to different datasets. This is a replicate from the GitHub repository GClarte/SignLanguagesPhylogeny, additional details can be found in the README.</p>
Videos accompanying article "Facial Expressions for Sign Language Synthesis using FACSHuman and AZee"
<p>Accompanying videos for the article "Facial Expressions for Sign Language Synthesis using FACSHuman and AZee".</p> <ul> <li>all_action_units.mp4 - Shows mesh deformations for all the action units used.</li> <li>all_expressions.mp4 - Shows all the expressions synthesized using the study.</li> <li>big_threatening_hot.mp4 - Shows comparison of a sign with and without facial expressions.</li> </ul>
SignBD-Word: Video-Based Bangla Word-Level Sign Language Dataset
<p>Bangla sign language (BdSL) is a complete and independent natural sign language with its own linguistic characteristics. While there exists video datasets for well-known sign languages, there is currently no available dataset for word-level BdSL. In this study, we present a video-based word-level dataset for Bangla sign language, called SignBD-Word, consisting of 6000 sign videos representing 200 unique words. The dataset includes full and upper-body views of the signers, along with 2D body pose information. This dataset can also be used as a benchmark for testing sign video classification algorithms.<br><br>Official Train Test Spllit (for both RGB and bodypose) can be found from the following link: <br>https://sites.google.com/view/signbd-word/dataset<br><br>This dataset is part of the following paper:<br>A. Sams, A. H. Akash and S. M. M. Rahman, "SignBD-Word: Video-Based Bangla Word-Level Sign Language and Pose Translation," 2023 14th International Conference on Computing Communication and Networking Technologies (ICCCNT), Delhi, India, 2023, pp. 1-7, doi: 10.1109/ICCCNT56998.2023.10306914.<br><br>Download the corresponding paper from this link:<br>https://asnsams.github.io/Publications.html</p>
BdSL47: A complete dataset of sign alphabet and digits of Bangla Sign Language (BdSL) using depth information via MediaPipe
<p><strong>BdSL47</strong> is the first open-access complete dataset in Bangla Sign Language that contains hand signs from both 10 sign digits (from sign ০ to sign ৯) and 37 sign alphabet (from sign অ to sign ँ).</p> <p>Dataset summary :</p> <ul> <li>100 RGB images per sign (total 47 signs) from each of 10 users</li> <li>Total input images : 100×47×10 = 47000</li> <li>Input images are processed via MediaPipe, which provided <ul> <li>an output image with hand key-points being detected</li> <li>3D coordinate values of 21 predefined key-points</li> <li>Total 63 coordinate values for each sample</li> </ul> </li> <li>The values are stored in csv files</li> <li>1 CSV file contains values from 100 samples of 1 sign from 1 user</li> <li>Total CSV files : 47×10 = 470</li> </ul> <p>The dataset has been made public for further research purposes. It is also available upon request <a href="https://drive.google.com/drive/u/8/folders/1wmJUlgWUrWNnOvzuL8Ci82Hm3zUx4wS-" rel="noopener">here</a>.</p>
Supplementary materials for "Phonetic differences between affirmative and feedback head nods in German Sign Language (DGS): A pose estimation study"
<div> <pre>This is the supplementary data for the article "Phonetic differences between affirmative and feedback head nods in German Sign Language (DGS): A pose estimation study" by Anastasia Bauer, Anna Kuder, Marc Schulder and Job Schepens.<br><br>The supplementary data consists of three components, stored in separate directories:<br>- <code>annotations/</code>: The manual annotations of head nod categories, produced by Anna Kuder and Anastasia Bauer.<br>- <code>pose_analysis/</code>: Code and input/output files for the pose-based automatic analysis of phonetic attributes head nods, produced by Marc Schulder.<br>- <code>statistical_analysis/</code>: Code for the statistical analysis of the other two components and for the creation of related figures, produced by Job Schepens.<br><br>For further details, see the README files of the respective directories.</pre> </div>
lingpy/sign-language-evolution-paper: Evolutionary Dynamics in the Dispersal of Sign Languages
<p>Supplement for study on Manual Alphabet evolution.</p>
Mocap video examples for the analysis of Sign Language movements
<p>These mocap videos support my PhD thesis "Extracting human characteristics from motion: the case of identity in Sign Language" carried out from October 2018 to October 2021. The original mocap data is taken from the <a href="https://www.ortolang.fr/market/corpora/mocap1/">MOCAP1</a> corpus of French Sign Language. The videos have been generated using Python code available as part of the <a href="https://github.com/felixbgd/PLmocap">PLmocap</a> library.</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.