Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
45
datasets available to search
ShareScore release 0.9.0
Dataset results
45 results for “catalan”
Catalan United Nations v1.0 test set
<p>Catalan version [1] of the test set from the United Nations v1.0 [2]. The translation was performed in two steps: we did a first automatic translation from the Spanish test set version into Catalan and then a professional translator post-edited the output.</p> <p><br> [1] Marta R. Costa-Jussà, Noé Casas, Carlos Escolano, and José A. R. Fonollosa. 2019. Chinese-Catalan: A Neural Machine Translation Approach Based on Pivoting and Attention Mechanisms. <em>ACM Trans. Asian Low-Resour. Lang. Inf. Process.</em> 18, 4, Article 43 (August 2019), 8 pages. DOI:https://doi.org/10.1145/3312575</p> <p>[2] Michal Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016. The United Nations parallel corpus v1.0. In<br> Proceedings of the LREC, 2016</p>
Ontolex-lemon and TIAD versions of Apertium Aragonese-Catalan dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> <ul> <li>2015-2016, Francis M. Tyers</li> <li>2009-2016, Juan Pablo Martínez</li> <li>2009-2010, Jimmy O'Regan</li> <li>2005-, Universitat d'Alacant (Transducens group)</li> <li>contributors of apertium-es-ca and apertium-spa-arg</li> </ul>
Ontolex-lemon and TIAD versions of Apertium Catalan-Sardinian dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> Copyright (c) 2017 Gianfranco Fronteddu Copyright (c) 2010, 2017 Hèctor Alòs i Font Copyright (c) 2010 Francis Tyers Copyright (C) 2005--2010 Universitat d'Alacant (Dades de català)
Ontolex-lemon and TIAD versions of Apertium English-Catalan dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2017-2019, Marc Riera Irigoyen 2012-2017, Xavi Ivars 2007-2017, Francis M. Tyers 2016, Bror Hultberg 2007-2016, Mikel L. Forcada 2009-2015, Kevin Brubeck Unhammer 2014, Mariola Alaixa Gadea Ferrando 2007-2014, Sergio Ortiz Rojas 2007-2012, Jim O'Regan 2012, Anthony J. Bentley 2012, Hèctor Alòs i Font 2011, Trond Trosterud 2008-2011, Gema Ramírez Sánchez 2007-2011, Mireia Ginestí Rosell 2009, Oscar Senra Gómez 2008, Jacob Nordfalk 2007, Felipe Sánchez Martínez
Ontolex-lemon and TIAD versions of Apertium Catalan-Italian dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2011-2019, Hèctor Alòs i Font 2009-2018, Francis M. Tyers 2018, Marc Riera Irigoyen 2018, Sushain Cherivirala 2009-2014, Jim O'Regan 2013, Trond Trosterud 2012, Anthony J. Bentley 2010-2011, Carme Armentano-Oller 2010, Mireia Ginestí Rosell 2009-2010, Antonio Toral
Ontolex-lemon and TIAD versions of Apertium Portuguese-Catalan dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2019-2020, Hèctor Alòs i Font 2016, Bror Hultberg 2008-2016, Francis M. Tyers 2012, Anthony J. Bentley 2009, Pasquale Minervini 2009, Mireia Ginestí Rosell 2008, Jim O'Regan 2008, Carme Armentano-Oller 2008, Mikel L. Forcada
Ontolex-lemon and TIAD versions of Apertium Spanish-Catalan dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2017-2019, Jaume Ortolà i Font 2010-2019, Xavi Ivars 2018-2019, Marc Riera Irigoyen 2017-2019, Hèctor Alòs i Font 2007-2019, Gema Ramírez Sánchez 2013-2019, Kevin Brubeck Unhammer 2017-2019, Donís Seguí 2018, Alberto Navalon 2007-2018, Mikel L. Forcada 2007-2018, Francis M. Tyers 2018, Sushain Cherivirala 2007-2014, Sergio Ortiz Rojas 2008-2012, Jim O'Regan 2012, Anthony J. Bentley 2009-2012, Miquel Esplà 2009-2011, Mireia Farrus Cabeceran 2011, Oscar Senra Gómez 2007-2010, Mireia Ginestí Rosell 2010, Garbine 2008, Jacob Nordfalk 2007-2008, Carme Armentano-Oller 2007, Yisas 2007, Carme Pla
Ontolex-lemon and TIAD versions of Apertium Romanian-Catalan dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2018-2019, Marc Riera Irigoyen 2007-2018, Francis M. Tyers 2018, Xavi Ivars 2013-2015, Kevin Brubeck Unhammer 2012, Anthony J. Bentley 2008-2010, Jim O'Regan 2009, Pasquale Minervini 2009, Mireia Ginestí Rosell 2008, Jacob Nordfalk 2008, Dipesquirol 2007-2008, Enrique Benimeli Bofarull 2008, Gema Ramírez Sánchez
DATA BASE-OFFENSIVE DISCOURSE ON TWITTER: VOX IN THE CATALAN PARLIAMENTARY ELECTIONS
<p>DATA BASE-OFFENSIVE DISCOURSE ON TWITTER: VOX IN THE CATALAN PARLIAMENTARY ELECTIONS</p>
Expert annotations for the Catalan Common Voice (v13)
<h2>Dataset Description</h2> <p>- Homepage: <a href="https://projecteaina.cat/tech/">https://projecteaina.cat/tech/</a>]<br>- Point of Contact: langech@bsc.es</p> <h3>Dataset Summary</h3> <p>These are the annotations made by a team of experts on the speakers with more than 1200 seconds recorded in the Catalan set of the <a href="https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0/tree/main/transcript/ca">Common Voice dataset (v13)</a>.</p> <p>The annotators were initially tasked with evaluating all recordings associated with the same individual. Following that, they were instructed to annotate the speaker's accent, gender, and the overall quality of the recordings.</p> <p>The accents and genders taken into account are the ones used until version 8 of the Common Voice corpus.</p> <p>See annotations for more details.</p> <h3>Supported Tasks and Leaderboards</h3> <p>Gender classification, Accent classification.</p> <h3>Languages</h3> <p>The dataset is in Catalan (ca).</p> <h2>Dataset Structure</h2> <h3>Instances</h3> <p>Two xlsx documents are published, one for each round of annotations.</p> <p>The following information is available in each of the documents:</p> <p><br><code>{</code><br><code> 'speaker ID': '1b7fc0c4e437188bdf1b03ed21d45b780b525fd0dc3900b9759d0755e34bc25e31d64e69c5bd547ed0eda67d104fc0d658b8ec78277810830167c53ef8ced24b', </code><br><code> 'idx': '31', </code><br><code> 'same speaker': {'AN1': 'SI',</code><br><code> 'AN2': 'SI',</code><br><code> 'AN3': 'SI',</code><br><code> 'agreed': 'SI',</code><br><code> 'percentage': '100'}, </code><br><code> 'gender': {'AN1': 'H',</code><br><code> 'AN2': 'H',</code><br><code> 'AN3': 'H',</code><br><code> 'agreed': 'H',</code><br><code> 'percentage': '100'}, </code><br><code> 'accent': {'AN1': 'Central',</code><br><code> 'AN2': 'Central',</code><br><code> 'AN3': 'Central',</code><br><code> 'agreed': 'Central',</code><br><code> 'percentage': '100'}, </code><br><code> 'audio quality': {'AN1': '4.0',</code><br><code> 'AN2': '3.0',</code><br><code> 'AN3': '3.0',</code><br><code> 'agreed': '3.0',</code><br><code> 'percentage': '66',</code><br><code> 'mean quality': '3.33',</code><br><code> 'stdev quality': '0.58'}, </code><br><code> 'comments': {'AN1': '',</code><br><code> 'AN2': 'pujades i baixades de volum',</code><br><code> 'AN3': 'Deu ser d'alguna zona de transició amb el central, perquè no fa una reducció total vocàlica, però hi té molta tendència'}, </code><br><code>}</code></p> <p> </p> <p>We also publish the document Guia anotació parlants.pdf, with the guidelines the annotators recieved.</p> <h3>Data Fields</h3> <ul> <li>speaker ID (string): An id for which client (voice) made the recording in the Common Voice corpus</li> <li> idx (int): Id in this corpus</li> <li> AN1 (string): Annotations from Annotator 1</li> <li> AN2 (string): Annotations from Annotator 2</li> <li> AN3 (string): Annotations from Annotator 3</li> <li> agreed (string): Annotation from the majority of the annotators</li> <li>percentage (int): Percentage of annotators that agree with the agreed annotation</li> <li>mean quality (float): Mean of the quality annotation</li> <li>stdev quality (float): Standard deviation of the mean quality</li> </ul> <h3>Data Splits</h3> <p>The corpus remains undivided into splits, as its purpose does not involve training models.</p> <h2>Dataset Creation</h2> <h3>Curation Rationale</h3> <p>During 2022, a campaign was launched to promote the Common Voice corpus within the Catalan-speaking community, achieving remarkable success. However, not all participants provided their demographic details such as age, gender, and accent. Additionally, some individuals faced difficulty in self-defining their accent using the standard classifications established by specialists.</p> <p>In order to obtain a balanced corpus with reliable information, we have seen the the necessity of enlisting a group of experts from the University of Barcelona to provide accurate annotations.</p> <p>We release the complete annotations because transparency is fundamental to our project. Furthermore, we believe they hold philological value for studying dialectal and gender variants.</p> <h3>Source Data</h3> <p>The original data comes from the [Catalan sentences of the Common Voice corpus](https://commonvoice.mozilla.org/en/datasets).</p> <p><strong>Initial Data Collection and Normalization</strong></p> <p>We have selected speakers who have recorded more than 1200 seconds of speech in the Catalan set of the <a href="https://commonvoice.mozilla.org/en/datasets">version 13 of the Common Voice corpus</a>.</p> <p><strong>Who are the source language producers?</strong></p> <p>The original data comes from the <a href="https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main/transcript/ca">Catalan sentences of the Common Voice corpus</a>.</p> <h3>Annotations</h3> <p><strong>Annotation process</strong></p> <p>Starting with <a href="https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main/transcript/ca">version 13 of the Common Voice corpus</a> we identified the speakers (273) who have recorded more than 1200 seconds of speech. </p> <p>A team of three annotators was tasked with annotating:</p> <ul> <li>if all the recordings correspond to the same person</li> <li>the gender of the speaker</li> <li>the accent of the speaker</li> <li>the quality of the recording</li> </ul> <p>They conducted an initial round of annotation, discussed their varying opinions, and subsequently conducted a second round.</p> <p>We release the complete annotations because transparency is fundamental to our project. Furthermore, we believe they hold philological value for studying dialectal and gender variants.</p> <p><strong>Who are the annotators?</strong></p> <p>The annotation was entrusted to the [CLiC (Centre de Llenguatge i Computació)](https://clic.ub.edu/en/que-es-clic) team from the University of Barcelona. <br>They selected a group of three annotators (two men and one woman), who received a scholarship to do this work. </p> <p>The annotation team was composed of:</p> <ul> <li>Annotator 1: 1 female annotator, aged 18-25, L1 Catalan, student in the Modern Languages and Literatures degree, with a focus on Catalan.</li> <li>Annotators 2 & 3: 2 male annotators, aged 18-25, L1 Catalan, students in the Catalan Philology degree.</li> <li>1 female supervisor, aged 40-50, L1 Catalan, graduate in Physics and in Linguistics, Ph.D. in Signal Theory and Communications.</li> </ul> <p>To do the annotation they used a Google Drive spreadsheet</p> <h3>Personal and Sensitive Information</h3> <p>The Common Voice dataset consists of people who have donated their voice online. We don't share here their voices, but their gender and accent. <br>You agree to not attempt to determine the identity of speakers in the Common Voice dataset.</p> <h2>Considerations for Using the Data</h2> <h3>Social Impact of Dataset</h3> <p>The ID come from the Common Voice dataset, that consists of people who have donated their voice online.</p> <p><em>You agree to not attempt to determine the identity of speakers in the Common Voice dataset.</em></p> <p>The information from this corpus will allow us to train and evaluate well balanced Catalan ASR models. Furthermore, we believe they hold philological value for studying dialectal and gender variants.</p> <h3>Discussion of Biases</h3> <p>Most of the voices of the common voice in Catalan correspond to men with a central accent between 40 and 60 years old. The aim of this dataset is to provide information that allows to minimize the biases that this could cause.</p> <p>For the gender annotation, we have only considered "H" (male) and "D" (female).</p> <h3>Other Known Limitations</h3> <p>[N/A]</p> <h2>Additional Information</h2> <h3>Dataset Curators</h3> <p>Language Technologies Unit at the Barcelona Supercomputing Center (langtech@bsc.es)</p> <p>This work has been promoted and financed by the Generalitat de Catalunya through the <a href="https://projecteaina.cat/">Aina project</a>.</p> <h3>Licensing Information</h3> <p>This dataset is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a> license.</p> <p>It can be used for any purpose, whether academic or commercial, under the terms of the license. <br>Give appropriate credit, provide a link to the license, and indicate if changes were made.</p> <h3>Citation Information</h3> <p><a href="../badge/DOI/10.5281/zenodo.11104388.svg">DOI</a></p> <h3>Contributions</h3> <p>The annotation was entrusted to the <a href="https://stel2.ub.edu/el-servei">STeL</a> team from the University of Barcelona.</p>
Catalan Government Crawling
<p>The Catalan Government Crawling Corpus is a 39-million-token web corpus of Catalan built from the web. It has been obtained by crawling the .gencat domain and subdomains, belonging to the Catalan Government during September and October 2020. It consists of 39.117.909 tokens, 1.565.433 sentences and 71.043 documents. Documents are separated by single new lines. It is a subcorpus of the <a href="http://zenodo.org/record/4519349">Catalan Textual Corpus</a>.</p> <p>We license the actual packaging of this data under a <a href="http://creativecommons.org/publicdomain/zero/1.0/">CC0 1.0 Universal License</a>.</p> <p><strong>Notice and take down policy</strong></p> <p>Notice: Should you consider that our data contains material that is owned by you and should therefore not be reproduced here, please:</p> <ul> <li>Clearly identify yourself, with detailed contact data such as an address, telephone number or email address at which you can be contacted.</li> <li>Clearly identify the copyrighted work claimed to be infringed.</li> <li>Clearly identify the material that is claimed to be infringing and information reasonably sufficient to allow us to locate</li> </ul> <p> </p> <p>If you use this resource in your work, please cite our latest paper:</p> <p>@inproceedings{armengol-estape-etal-2021-multilingual,<br> title = "Are Multilingual Models the Best Choice for Moderately Under-resourced Languages? {A} Comprehensive Assessment for {C}atalan",<br> author = "Armengol-Estap{\'e}, Jordi and<br> Carrino, Casimiro Pio and<br> Rodriguez-Penagos, Carlos and<br> de Gibert Bonet, Ona and<br> Armentano-Oller, Carme and<br> Gonzalez-Agirre, Aitor and<br> Melero, Maite and<br> Villegas, Marta",<br> booktitle = "Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021",<br> month = aug,<br> year = "2021",<br> address = "Online",<br> publisher = "Association for Computational Linguistics",<br> url = "https://aclanthology.org/2021.findings-acl.437",<br> doi = "10.18653/v1/2021.findings-acl.437",<br> pages = "4933--4946",<br>}</p>
COPA-ca: Choice of plausible alternatives in Catalan
<p>The COPA-ca dataset (Choice of plausible alternatives in Catalan) is a professional translation of the English COPA dataset into Catalan, commissioned by BSC LangTech Unit.</p> <p>The dataset consists of 1000 premises, each given a question and two choices with a label encoding which of the choices is more plausible given the annotator.</p> <p>This work is licensed under a Attribution-ShareAlike 4.0 International License.</p> <p>This work was funded by the Departament de la Vicepresidència i de Polítiques Digitals i Territori de la Generalitat de Catalunya within the framework of Projecte AINA.</p>
ChatSubs: A dataset of movie dialogues in Spanish, Catalan, Basque and Galician
<p><strong>Description</strong>: The ChatSubs dataset contains dialogues in Spanish and three co-official languages of Spain (Catalan, Basque, and Galician). It was obtained from OpenSubtitles and processed to generate clearly segmented dialogues and turns. The dataset consists of 206,706 JSON files, with over 20 million dialogues and 96 million turns, making it one of the largest dialogue corpora available. It serves as an excellent resource for research teams interested in training dialogue models in Spanish, Catalan, Basque, and Galician.</p> <p><strong>License</strong>: <a href="https://creativecommons.org/licenses/by-nc/4.0/">CC BY-NC 4.0</a>.</p>
Mechanism of Action of Vichy Catalan Water
ClinicalTrials.gov study NCT01334840. IPD Sharing: Not stated. Countries: 1. Publications: 9.
Impact of the Social Determinants of Health in the Central Catalan Region
ClinicalTrials.gov study NCT04151056. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Catalan Referendum Twitter corpus
Open the record for dataset details and reuse information.
Figures 4-7 from: Palacios Vargas J, Catalan E (2013) A new genus and species of Tullbergiidae (Collembola) from the Pacific Mexican coast. ZooKeys 326: 91-97. https://doi.org/10.3897/zookeys.326.5451
Figures 4-7 - Mexicaphorura guerrensis sp. n. 4 dorsal chaetotaxy of body 5 tibiotarsus III 6 ventral abdominal chaetotaxy 7 male genital plate.
Figures 1-3 from: Palacios Vargas J, Catalan E (2013) A new genus and species of Tullbergiidae (Collembola) from the Pacific Mexican coast. ZooKeys 326: 91-97. https://doi.org/10.3897/zookeys.326.5451
Figures 1-3 - Mexicaphorura guerrensis sp. n. 1 dorsal antennal segments I to IV, with detail of the ventral sensillum of Ant. III 2 chaetotaxy of labrum 3 chaetotaxy of labium. a,b,d,d,e = sensilla on Ant. IV, so = subapical organite, ms =microsensillum, sc = thick sensory clubs on Ant. III, sr = sensory rods, if = integumentary fold, vsc = ventral sensory club on Ant. III.
Linguistic laws in speech: the case of Catalan and Spanish
Open the record for dataset details and reuse information.
Title: Leaf transcriptomics of Catalan A. thaliana demes under alkaline and carbonated stress at 3h and 48 hours.
GEO Series GSE164502. Arabidopsis thaliana. 72 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.