Skip to main content
zenodoopen

SWL-LSE: SignaMed Word-Level LSE, a Dataset of Spanish Sign Language Health Signs

<h2>SWL-LSE Dataset</h2> <p>The SWL-LSE dataset is coined from SignaMed Word-Level LSE (Lengua de Signos Espa&ntilde;ola -Spanish Sign Language).</p> <h2>Overview</h2> <p>The dataset consists of 8,000 sign sequences from 300 different sign classes related to the health domain. Each class is represented by an RGB video that serves as the dictionary sign. These dictionary signs were reproduced by 124 signers, including deaf individuals, interpreters, and L2 Spanish Sign Language (LSE) students, using their webcams or mobile phones via the SignaMed platform (<a href="https://signamed.web.app" target="_new" rel="noopener">https://signamed.web.app</a>). For privacy reasons, only the skeleton data is shared.</p> <p>The process of collecting the dataset is described in:</p> <p>V&aacute;zquez-Enr&iacute;quez, M.; Alba-Castro, J.L.; P&eacute;rez-P&eacute;rez, A.; Cabeza-Pereiro, C.; Doc&iacute;o-Fern&aacute;ndez, L. SignaMed: a Cooperative<br>Bilingual LSE-Spanish Dictionary in the Healthcare Domain. In Proceedings of the Proceedings of the LREC-COLING 2024<br>11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources; Efthimiou, E.;&nbsp;Fotinea, S.E.; Hanke, T.; Hochgesang, J.A.; Mesch, J.; Schulder, M., Eds., Torino, Italia, 2024; pp. 386&ndash;394.&nbsp;</p> <p>The dataset itself and the pipeline for training and executing a baseline model based on skeletons is described in this github (https://github.com/mvazquezgts/SWL-LSE), and this paper:</p> <p>V&aacute;zquez-Enr&iacute;quez, M.; Alba-Castro, J.L.; Doc&iacute;o-Fern&aacute;ndez, L.; Rodr&iacute;guez-Banga, E. SWL-LSE: A Dataset of Spanish Sign Language Health Signs with an ISLR Baseline Method. Technologies 2024, 12(10), 205, D.O.I:10.3390/technologies12100205</p> <h2>Files</h2> <h3>1. VIDEOS_REF.zip</h3> <ul> <li><strong>Description</strong>: RGB videos recorded in lab conditions that represent each sign-class</li> <li><strong>Total files</strong>: 300</li> </ul> <h3>2. videos_ref_annotations.csv</h3> <ul> <li><strong>Description</strong>: CSV file with the correspondence between the name of the video, its class ID and gloss in spanish: FILENAME,CLASS_ID,LABEL.</li> <li><strong>Total files</strong>: 1</li> </ul> <h3>3. ANNOTATIONS.zip</h3> <ul> <li><strong>Description</strong>: 3 CSV files with train, validation and test file-class correspondences: FILENAME,CLASS_ID</li> <li><strong>Total files</strong>: 3</li> </ul> <h3>4. MEDIAPIPE.zip</h3> <ul> <li><strong>Description</strong>: Pickle files containing the full output of Mediapipe using their Heavy model. Each .pkl file contains the outputs of Mediapipe Holistic legacy, Mediapipe Pose and Mediapipe Hands. Each file is package as a dictionary: dict_keys(['pose', 'hands', 'holistic_legacy'])</li> <li><strong>Total files</strong>: 8000</li> </ul> <h2>Usage</h2> <p>Researchers and practitioners in pattern recognition, machine learning, and sign language linguistics may find this dataset valuable for:</p> <ul> <li>Training/testing machine learning models for isolated sign language recognition or gesture recognition.</li> <li>Analyzing patterns on signs realization</li> </ul> <h2>Acknowledgments</h2> <p>This dataset is a collaborative effort of the next research goups and entities:</p> <ul> <li><a href="http://gtm.uvigo.es/en/">Group of Multimedia Technologies (GTM)</a> from the <a href="https://atlanttic.uvigo.es/en">atlanTTic Research Center</a> of <a href="http://www.uvigo.es/">University of Vigo</a> (Spain)</li> <li><a href="http://grades.uvigo.gal/">Group of Discourse and Society (GRADES)</a> from the <a href="https://fft.uvigo.es/en/">School of Philology and Translation</a> of <a href="http://www.uvigo.es/">University of Vigo</a> (Spain)</li> <li><a href="http://www.faxpg.es/">Federation of Deaf People Galician Associations (FAXPG)</a></li> <li><a href="https://fundacioncnse-dilse.org">Fundaci&oacute;n CNSE-DILSE</a></li> </ul> <p>Gratitude is extended to them for their contributions and support.</p>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
8
Access
20
Reuse readiness
8
Engagement
4