Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

45

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

45 results for “Catalan”

Learn how ShareScore rates datasets ↗
zenodo24/100

Catalan basketball results

<p>Dataset containing results for basketball competitions in Catalonia.<br> The data depicted belongs to the Catalan Federation of Basketball, with their site at www.basquetcatala.cat</p> <p>The dataset contains one record for each match, with the following fields:<br> - temporada: start/end year of the season<br> - ambit: scope of the competition, either for the whole territory or within a province<br> - categoria: category of the involved teams<br> - competicio: name of the competition or phase within it<br> - grup: some competitions have a high number of competing teams, they are split in several groups then<br> - jornada: round within the competition<br> - data: date and time of the match<br> - equip_local: name of the local team<br> - punts_local: points obtained by the local team<br> - equip_visitant: name of the visitor team<br> - punts_visitant: points obtained by the visitor team</p> <p>There are some missing values in the points fields, with a hyphen instead of a numeric value.<br> There&#39;s also a missing category name, containing the fragment of link where it&#39;s pointed to in the site. &#39;42018/950&#39; instead of a name.</p> <p>Applicable license is Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International</p>

opencc-by-4.0Apr 2020View details →
zenodo24/100

CaSERa-catalan-stance-emotions-raco

<h3>Dataset Summary</h3> <p>The CaSERa dataset is a Catalan corpus from the forum Rac&oacute; Catal&agrave; annotated with Emotions and Dynamic Stance. The dataset contains 15.782 unique sentences grouped in 10.745 pairs of sentences, paired as parent messages and replies to these messages.</p> <p>We provide the following files and folders:</p> <ul> <li>The dataset folder contains the final dataset (an aggregation of the dynamic stance and emotion identification tasks), as well as the dataloader and README file as provided in Hugging Face.</li> <li>The annotations folders with one file for each of the two annotation task: dynamic stance and emotion identification.</li> <li>The annotation guidelines folders with a file for each task.</li> </ul> <h3>Languages</h3> <p>The dataset is in Catalan (ca-ES).</p> <h3>Dataset Structure</h3> <p>Each instance in the dataset is a pair of parent-reply messages annotated with the relation between the two messages (the dynamic stance). For each message there is an individual id and the emotions identified in the message.</p> <p>Example:<br>&nbsp; &nbsp;&nbsp;<br>{<br>"id_conversation": "782135",&nbsp;<br>"id_reply": "782135_2_2",&nbsp;<br>&nbsp; "parent_text": "Alguns petits apunts que s'haurien de tenir en compte en la creaci&oacute; d'aquesta hipot&egrave;tica grada jove: -Ren&uacute;ncia total i expl&iacute;cita a la viol&egrave;ncia, aquest per mi &eacute;s el punt clau i b&agrave;sic que la directiva tindria en compte, sense ren&uacute;ncia no hi ha grup. -Edat de la gent: indispensable establir un m&iacute;nim i un m&agrave;xim d'edat per adquirir entrades de la grada jove, s'ha d'acabar amb els freaks de 40-50 anys amb esperit jove i ganes d\'animar (sona discriminatori, per&ograve; crec que &eacute;s &ograve;bvi que s'ha de limitar l'edat) -Seguretat b&agrave;sica: Agradi o no, s'hauria d'incrementar molt&iacute;ssim la seguretat privada (&eacute;s un dels punts que no agrada a la directiva). Els indesitjables de sempre no tolerarien cap grada jove sense ells. -Per tant doncs, absteniu-vos i oblideu-vos tots els hooligans potencials (n'hi ha a grapats per tot Catalunya) de crear una grada hooligan amb skins, punkis i resta d'est&egrave;tiques "caracter&iacute;stiques", tamb&eacute; seria un punt clau (confirmat per un directiu!) En definitiva, ara per ara veig for&ccedil;a dif&iacute;cil i inviable la creaci&oacute; d'una grada jove nombrosa i contundent, s'hauria d'optar per opcions m&eacute;s "descafe&iuml;nades" rotllo Sang Cul&eacute; o Dracs.",&nbsp;<br>&nbsp; "reply_text": "Doncs si es cre&eacute;s una grada jove ja et dic jo que sompliria d'skins, i els primers en entrari serien els Boixos que han expulsat del seu lloc. Tot i a&igrave;x&ograve; que hi hagin skins no vol dir que hi hagi viol&egrave;ncia, no crec que es peguessin amb els del mateix equip. Si es cre&eacute;s un grada jove l'ambient al camp nou seria brutal, si 200 Boixos feien bivrar el camp nou, imaginat a 4000 perones o la gent que hi capigu&eacute;s a la grada jove.",&nbsp;<br>&nbsp; &nbsp; "dynamic_stance": "Elaborate",&nbsp;<br>&nbsp; "parent_emotion": ["fear", "distrust", "anticipation"],&nbsp;<br>&nbsp; "reply_emotion": ["anger", "sadness", "fear", "distrust", "anticipation"]<br>}</p> <h3>Data Splits</h3> <p>The dataset does not contain splits.</p> <h3>Dataset Creation</h3> <p>We created this corpus to contribute to the development of language models in Catalan, a low-resource language.</p> <h3>Source Data</h3> <p>The data was collected using the messages of the forum Rac&oacute; Catal&agrave; by the Barcelona Supercomputing Center.</p> <h3>Initial Data Collection and Normalization</h3> <p>The data was collected selecting random messages from 13 of the thematic sections in Rac&oacute; Catal&agrave; that had at least 15 tokens and at most 300. Then, we kept the messages that had at least one replying message with the same length requirement. We got a maximum of 3 replying messages per parent message.<br>Who are the source language producers?</p> <p>The source language producers are users of Rac&oacute; Catal&agrave;.</p> <h3>Annotations</h3> <p>- &nbsp; Emotions are annotated in a multi-label fashion. The labels can be: Anger, Anticipation, Disgust, Fear, Joy, Sadness, Surprise, Distrust, and No emotion.</p> <p>- &nbsp; Dynamic stance is annotated per pair. The labels can be: Agree, Disagree, Elaborate, Query, Neutral, Unrelated, NA.</p> <h3>Annotation process</h3> <p>- &nbsp; For emotions there were 3 annotators. The gold labels are an aggregation of all the labels annotated by the 3. The IAA calculated with Fleiss' Kappa per label was, on average, 38.73.</p> <p>- &nbsp; For dynamic stance there were 4 annotators. If at least 3 of the annotators disagreed, a fifth annotator chose the gold label. The overall Fleiss' Kappa between the 4 annotators was 57.63, and the average Fleiss' Kappa of the annotators with the gold labels is 85.98.</p> <h3>Who are the annotators?</h3> <p>All the annotators are native speakers of Catalan.</p> <h3>Personal and Sensitive Information</h3> <p>The data was annonymised to remove user names and emails, which were changed to random Catalan names. The mentions to the chat itself have also been changed.</p> <h3>Social Impact of Dataset</h3> <p>We hope this corpus contributes to the development of language models in Catalan, a low-resource language.</p> <h3>Discussion of Biases</h3> <p>We are aware that, since the data comes from a public forum, this will contain biases, hate speech and toxic content. We have not applied any steps to reduce their impact.</p> <h3>Dataset Curators</h3> <p>Language Technologies Unit (LangTech) at the Barcelona Supercomputing Center.</p> <p>This work/research has been promoted and financed by the Government of Catalonia through the Aina project.</p> <h3>Licensing Information</h3> <p>Creative Commons Attribution 4.0 International.</p> <h3>Citation Information</h3> <p>@inproceedings{figueras-etal-2023-dynamic,<br>&nbsp; &nbsp; title = "Dynamic Stance: Modeling Discussions by Labeling the Interactions",<br>&nbsp; &nbsp; author = "Figueras, Blanca &nbsp;and<br>&nbsp; &nbsp; &nbsp; Baucells, Irene &nbsp;and<br>&nbsp; &nbsp; &nbsp; Caselli, Tommaso",<br>&nbsp; &nbsp; editor = "Bouamor, Houda &nbsp;and<br>&nbsp; &nbsp; &nbsp; Pino, Juan &nbsp;and<br>&nbsp; &nbsp; &nbsp; Bali, Kalika",<br>&nbsp; &nbsp; booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2023",<br>&nbsp; &nbsp; month = dec,<br>&nbsp; &nbsp; year = "2023",<br>&nbsp; &nbsp; address = "Singapore",<br>&nbsp; &nbsp; publisher = "Association for Computational Linguistics",<br>&nbsp; &nbsp; url = "https://aclanthology.org/2023.findings-emnlp.432",<br>&nbsp; &nbsp; doi = "10.18653/v1/2023.findings-emnlp.432",<br>&nbsp; &nbsp; pages = "6503--6515",<br>}</p> <p>@inproceedings{gonzalez-agirre-etal-2024-building-data,<br>&nbsp; &nbsp; title = "Building a Data Infrastructure for a Mid-Resource Language: The Case of {C}atalan",<br>&nbsp; &nbsp; author = "Gonzalez-Agirre, Aitor &nbsp;and<br>&nbsp; &nbsp; &nbsp; Marimon, Montserrat &nbsp;and<br>&nbsp; &nbsp; &nbsp; Rodriguez-Penagos, Carlos &nbsp;and<br>&nbsp; &nbsp; &nbsp; Aula-Blasco, Javier &nbsp;and<br>&nbsp; &nbsp; &nbsp; Baucells, Irene &nbsp;and<br>&nbsp; &nbsp; &nbsp; Armentano-Oller, Carme &nbsp;and<br>&nbsp; &nbsp; &nbsp; Palomar-Giner, Jorge &nbsp;and<br>&nbsp; &nbsp; &nbsp; Kulebi, Baybars &nbsp;and<br>&nbsp; &nbsp; &nbsp; Villegas, Marta",<br>&nbsp; &nbsp; editor = "Calzolari, Nicoletta &nbsp;and<br>&nbsp; &nbsp; &nbsp; Kan, Min-Yen &nbsp;and<br>&nbsp; &nbsp; &nbsp; Hoste, Veronique &nbsp;and<br>&nbsp; &nbsp; &nbsp; Lenci, Alessandro &nbsp;and<br>&nbsp; &nbsp; &nbsp; Sakti, Sakriani &nbsp;and<br>&nbsp; &nbsp; &nbsp; Xue, Nianwen",<br>&nbsp; &nbsp; booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",<br>&nbsp; &nbsp; month = may,<br>&nbsp; &nbsp; year = "2024",<br>&nbsp; &nbsp; address = "Torino, Italia",<br>&nbsp; &nbsp; publisher = "ELRA and ICCL",<br>&nbsp; &nbsp; url = "https://aclanthology.org/2024.lrec-main.231",<br>&nbsp; &nbsp; pages = "2556--2566",<br>}</p> <h3>Contact Information</h3> <p>For further information, please send an email to langtech@bsc.es.</p>

opencc-by-4.0Dec 2023View details →
zenodo24/100

CaSET-catalan-stance-emotions-twitter

<h3>Dataset Summary</h3> <p>The CaSET dataset is a Catalan corpus of Tweets annotated with Emotions, Static Stance, and Dynamic Stance. The dataset contains 11k unique sentences on five controversial topics, grouped in 6k pairs of sentences, paired as parent messages and replies to these messages.</p> <p>We provide the following files and folders:</p> <ul> <li>The dataset folder contains the final dataset (an aggregation of the dynamic stance, static stance, and emotion identification tasks), as well as the dataloader and README file as provided in Hugging Face.</li> <li>The annotations folders with one file for each of the three annotation tasks.</li> <li>The annotation guidelines folders with a file for each task.</li> </ul> <h3>Supported Tasks and Leaderboards</h3> <p>This dataset can be used to train models for emotion detection, static stance detection, and dynamic stance detection.</p> <h3>Languages</h3> <p>The dataset is in Catalan (ca-ES).</p> <h3>Dataset Structure</h3> <p>Each instance in the dataset is a pair of parent-reply messages, annotated with the relation between the two messages (the dynamic stance) and the topic of the messages. For each message there is the id to retrieve it with the Twitter API, the emotions identified in the message, and the relation between the message and the topic (static stance). The text fields have to be retrieved using the Twitter API.</p> <h3>Data Instances</h3> <p>{<br>"id_parent": "1413960970066710533",&nbsp;<br>"id_reply": "1413968453690658816",&nbsp;<br>&nbsp; "parent_text": "",&nbsp;<br>&nbsp; "reply_text": "",&nbsp;<br>&nbsp; &nbsp; "topic": "vaccines",&nbsp;<br>&nbsp; &nbsp; "dynamic_stance": "Disagree",&nbsp;<br>&nbsp; "parent_stance": "FAVOUR",&nbsp;<br>&nbsp; "reply_stance": "AGAINST",&nbsp;<br>&nbsp; "parent_emotion": ["distrust", "joy", "disgust"],&nbsp;<br>&nbsp; "reply_emotion": ["distrust"]<br>}</p> <h3>Data Splits</h3> <p>The dataset does not contain splits.</p> <h3>Dataset Creation</h3> <p>We created this corpus to contribute to the development of language models in Catalan, a low-resource language.<br>Source Data</p> <p>The data was collected using the Twitter API by the Barcelona Supercomputing Center.</p> <h3>Initial Data Collection and Normalization</h3> <p>The data was collected based on a list of keywords related to the five topics included in the dataset: vaccines, rent regulation, surrogate pregnancy, airport expansion, and a TV show rigging. Specific periods in which the topic was under discussion were also selected.</p> <h3>Who are the source language producers?</h3> <p>The source language producers are users of Twitter.</p> <h3>Annotations</h3> <ul> <li>Emotions are annotated in a multi-label fashion. The labels can be: Anger, Anticipation, Disgust, Fear, Joy, Sadness, Surprise, Distrust, and No emotion. CA</li> <li>Static stance is annotated per message. The labels can be: FAVOUR, AGAINST, NEUTRAL, NA.&nbsp; &nbsp;</li> <li>Dynamic stance is annotated per pair. The labels can be: Agree, Disagree, Elaborate, Query, Neutral, Unrelated, NA.</li> </ul> <h3>Annotation process</h3> <ul> <li>For emotions there were 3 annotators. The gold labels are an aggregation of all the labels annotated by the 3. The IAA calculated with Fleiss' Kappa per label was, on average, 45.38.</li> <li>For static stance there were 2 annotators, in the cases of disagreement a third annotated chose the gold label. The overall Fleiss' Kappa between the 2 annotators is 82.71.</li> <li>For dynamic stance there were 4 annotators. If at least 3 of the annotators disagreed, a fifth annotator chose the gold label. The overall Fleiss' Kappa between the 4 annotators was 56.51, and the average Fleiss' Kappa of the annotators with the gold labels is 85.17.</li> </ul> <h3>Who are the annotators?</h3> <p>All the annotators are native speakers of Catalan.<br><br>Social Impact of Dataset</p> <p>We hope this corpus contributes to the development of language models in Catalan, a low-resource language.</p> <h3>Discussion of Biases</h3> <p>We are aware that, since the data comes from social media, this will contain biases, hate speech and toxic content. We have not applied any steps to reduce their impact.</p> <h3>Other Known Limitations</h3> <p>The dataset has to be downloaded using the Twitter API, therefore some instances might be lost.</p> <h3>Dataset Curators</h3> <p>Language Technologies Unit (LangTech) at the Barcelona Supercomputing Center.</p> <p>This work has been promoted and financed by the Generalitat de Catalunya through the Aina project.</p> <h3>Licensing Information</h3> <p>Creative Commons Attribution 4.0.</p> <h3>Citation Information</h3> <p>@inproceedings{figueras-etal-2023-dynamic,<br>&nbsp; &nbsp; title = "Dynamic Stance: Modeling Discussions by Labeling the Interactions",<br>&nbsp; &nbsp; author = "Figueras, Blanca &nbsp;and<br>&nbsp; &nbsp; &nbsp; Baucells, Irene &nbsp;and<br>&nbsp; &nbsp; &nbsp; Caselli, Tommaso",<br>&nbsp; &nbsp; editor = "Bouamor, Houda &nbsp;and<br>&nbsp; &nbsp; &nbsp; Pino, Juan &nbsp;and<br>&nbsp; &nbsp; &nbsp; Bali, Kalika",<br>&nbsp; &nbsp; booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2023",<br>&nbsp; &nbsp; month = dec,<br>&nbsp; &nbsp; year = "2023",<br>&nbsp; &nbsp; address = "Singapore",<br>&nbsp; &nbsp; publisher = "Association for Computational Linguistics",<br>&nbsp; &nbsp; url = "https://aclanthology.org/2023.findings-emnlp.432",<br>&nbsp; &nbsp; doi = "10.18653/v1/2023.findings-emnlp.432",<br>&nbsp; &nbsp; pages = "6503--6515",<br>}</p> <p>@inproceedings{gonzalez-agirre-etal-2024-building-data,<br>&nbsp; &nbsp; title = "Building a Data Infrastructure for a Mid-Resource Language: The Case of {C}atalan",<br>&nbsp; &nbsp; author = "Gonzalez-Agirre, Aitor &nbsp;and<br>&nbsp; &nbsp; &nbsp; Marimon, Montserrat &nbsp;and<br>&nbsp; &nbsp; &nbsp; Rodriguez-Penagos, Carlos &nbsp;and<br>&nbsp; &nbsp; &nbsp; Aula-Blasco, Javier &nbsp;and<br>&nbsp; &nbsp; &nbsp; Baucells, Irene &nbsp;and<br>&nbsp; &nbsp; &nbsp; Armentano-Oller, Carme &nbsp;and<br>&nbsp; &nbsp; &nbsp; Palomar-Giner, Jorge &nbsp;and<br>&nbsp; &nbsp; &nbsp; Kulebi, Baybars &nbsp;and<br>&nbsp; &nbsp; &nbsp; Villegas, Marta",<br>&nbsp; &nbsp; editor = "Calzolari, Nicoletta &nbsp;and<br>&nbsp; &nbsp; &nbsp; Kan, Min-Yen &nbsp;and<br>&nbsp; &nbsp; &nbsp; Hoste, Veronique &nbsp;and<br>&nbsp; &nbsp; &nbsp; Lenci, Alessandro &nbsp;and<br>&nbsp; &nbsp; &nbsp; Sakti, Sakriani &nbsp;and<br>&nbsp; &nbsp; &nbsp; Xue, Nianwen",<br>&nbsp; &nbsp; booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",<br>&nbsp; &nbsp; month = may,<br>&nbsp; &nbsp; year = "2024",<br>&nbsp; &nbsp; address = "Torino, Italia",<br>&nbsp; &nbsp; publisher = "ELRA and ICCL",<br>&nbsp; &nbsp; url = "https://aclanthology.org/2024.lrec-main.231",<br>&nbsp; &nbsp; pages = "2556--2566",<br>}</p> <h3>Contact information</h3> <p>For further information, please send an email to langtech@bsc.es.</p>

opencc-by-4.0Dec 2023View details →
zenodo20/100

Newswire Catalan Corpus

<p>The Catalan Newswire Corpus is a 163-million-token corpus of Catalan newswire text built from three major Catalan news providers: <a href="https://www.acn.cat/">Ag&egrave;ncia Catalana de Not&iacute;cies</a>, <a href="https://www.naciodigital.cat/">Naci&oacute; Digital</a> and <a href="https://www.vilaweb.cat/">Vilaweb</a>.</p> <p>It consists of 163.248.451 tokens, 6.317.202 sentences and 410.218 documents. Documents are separated by single new lines.</p> <p>We license the actual packaging of these data under a <a href="https://creativecommons.org/licenses/by-nc-nd/4.0/deed.en">Attribution-NonCommercial-NoDerivatives 4.0 License</a>.</p> <p>Copyright (c) 2022 Text Mining Unit at BSC</p>

opencc-by-nc-nd-4.0Nov 2022View details →
zenodo16/100

Digital competence profiles of first year, pre-service teachers. Analysis in the Catalan university system. RAWDATA

<p>Dataset for the ARMIF research project 2018-19</p>

restrictedcc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record