Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

45

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

45 results for “catalan”

Learn how ShareScore rates datasets ↗
zenodo48/100

CORINE land cover - Catalan land cover taxonomy dictionary for metropolitan Barcelona

<p>This dataset includes Catalan Land Cover categories that are relevant for metropolitan Barcelona and their translation to CORINE land cover categories.</p> <p>This work has been developed for the ERC project <a href="https://urbag.eu/">URBAG</a></p> <p>Original Catalan land cover map: <a href="https://www.creaf.uab.es/mcsc/">https://www.creaf.uab.es/mcsc/ </a></p> <p>CORINE land cover: <a href="https://land.copernicus.eu/pan-european/corine-land-cover">https://land.copernicus.eu/pan-european/corine-land-cover</a></p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium Esperanto-Catalan dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2005--2007, Universitat d'Alacant (Transducens group) (c) 2007, Universitat Pompeu Fabra (c) 2009--, Hèctor Alòs i Font

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium French-Catalan dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2005, Universitat d'Alacant, (Transducens group) (c) 2007 Eleka Ingenieritza Linguistikoa S.L. Prompsit Language Engineering S.L. (c) 2009 Jimmy O'Regan 2017-2020, Hèctor Alòs i Font 2015-2018, Francis M. Tyers 2017-2018, Xavi Ivars 2017, Bernard Chardonneau 2017, Kevin Brubeck Unhammer

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium Occitan (post 1500)-Catalan dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2005 Universitat d'Alacant (Transducens group) (c) 2005--2008 Generalitat de Catalunya (c) 2005 Universitat Pompeu Fabra (IULA) (c) 2007--2008 Prompsit Language Engineering S.L.

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Wikipedia: wikipedia-ca (Catalan)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>

opencc-by-sa-4.0Aug 2024View details →
zenodo44/100

Actuacions castelleres als Països Catalans l'any 2023

<p>Actuacions castelleres als Pa&iuml;sos Catalans l'any 2023</p> <p>Cont&eacute;: Lloc, data, motiu, colles participants</p> <p>Font de les dades:&nbsp;<a href="https://castellscat.cat/ca/base-de-dades" rel="nofollow">Base de Dades Coordinadora &ndash; Jove Tarragona (BDCJ)</a></p> <p>Les dades obtingudes en la cerca es poden utilitzar sempre i quan s&rsquo;identifiqui la seva proced&egrave;ncia:&nbsp;<strong>Base de Dades Coordinadora &ndash; Jove Tarragona (BDCJ)</strong>.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Translated Wikipedia Biographies (English-Catalan)

<p>A professional translation of the Translated Wikipedia Biographies dataset, designed to analyze common gender errors in machine translation like incorrect gender choices in anaphora resolutions, possessives and gender agreement, commissioned by BSC LangTech Unit.</p> <p>License type: CC BY-SA 4.0</p> <p>This work was funded by the Departament de la Vicepresid&egrave;ncia i de Pol&iacute;tiques Digitals i Territori de la Generalitat de Catalunya within the framework of Projecte AINA.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

CatVolc: A new database of geochemical and geochronological data of volcanic-related materials from the Catalan Volcanic Zone (Spain)

<p>The Catalan Volcanic Zone (CVZ) (NE Spain) consists of an intraplate alkaline volcanic zone associated with the opening of the Western Mediterranean and the development of the European Rift System. Volcanic activity in the CVZ started in the L&rsquo;Empord&agrave; area (ca. &gt; 12 - 8 Ma), extended to La Selva (7.9 - 1.7 Ma), and finally migrated to the Garrotxa Volcanic Field (&lt; 0.7 - 0.01Ma). Despite the scientific interest in the CVZ since the early 19th century, certain aspects remain poorly constrained. These include a full understanding of the spatial and temporal evolution of the magma plumbing system and ascent mechanisms, as well as the chronology of volcanism across the CVZ. Addressing these unresolved questions requires geochemical, petrological, and geochronological data, which, in the case of the CVZ, are scattered and have never been integrated or analyzed within a unified framework. Here, we present the CatVolc (Catalan Volcanism) database, which compiles available geochemical and geochronological data of volcanic-related materials of the CVZ. &nbsp;For each sample, the CatVolc database lists general information about the sampling site, sample lithology, whole-rock analyses (including major and trace elements), isotopic ratios, mineral chemistry, and radiometric/thermoluminescence dating information, if available. A preliminary analysis of the information contained in the CatVolc database highlights the critical limitations of the current state of knowledge and allows suggesting potential future directions for volcanic-driven investigations in the CVZ. Additionally, the results obtained validate the CatVolc database as a key tool for comprehending the spatial and temporal evolution of the magmatic system(s) and volcanic activity in the CVZ, particularly in the Garrotxa Volcanic Field. This aspect is critical for advancing in the assessment of the volcanic hazards in the region and for gaining a comprehensive understanding of future volcanic activity.</p> <p>The current database version consists of three MS Excel files dedicated to primary magmatic rocks (<em>CatVolc_magmatic_rocks.xlsx</em>), xenoliths (<em>CatVolc_xenoliths.xlsx</em>) and radiometric/thermoluminescence dating information (<em>CatVolc_dating.xlsx</em>). Each MS Excel file is structured around a main table (<em>Samples_general_info</em>) containing general information (e.g., location, age, sampling site) of the listed samples, and a certain number of secondary tables.&nbsp; Secondary tables included in the MS Excel files for magmatic rocks and xenoliths report: (i) whole-rock geochemistry (<em>Major_elements</em> and <em>Trace_elements</em>); (ii) isotopic relations data (<em>Isotopic_relations</em>); and (iii) mineral and volcanic glass chemistry (<em>Amphibole</em><strong>, </strong><em>Feldspar</em><strong>, </strong><em>Felspathoid</em><strong>, </strong><em>Glass</em><strong>, </strong><em>Mica</em><strong>, </strong><em>Olivine</em><strong>, </strong><em>Pyroxene</em><strong>, </strong><em>Oxides</em><strong>, </strong><em>Serpentines</em><strong> </strong>and<strong> </strong><em>Sulphides</em><em>)</em>. In addition to the <em>Samples_general_info </em>table, the <em>CatVolc_dating.xlsx </em>also includes radiometric/thermoluminescence dating information (<em>Dating</em>). Finally, in all three Excel files, we have incorporated a table with consulted references (<em>References</em><em>) </em>and a glossary of the acronyms used for the parameters included in the main and secondary tables (<em>Codes</em>).&nbsp;</p> <p>&nbsp;</p>

opencc-by-nc-4.0Sep 2023View details →
zenodo40/100

VeLeCa: a Verbal Lexicon of Catalan

<p>VeLeCa is an inflected lexicon of Catalan verbal inflection containing the phonological form of 174,200 word forms from 3,484 lexemes and their respective lexical and morphosyntactic values and frequencies.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Annotated Corpora of Historical Catalan (HisCat) - Llibre dels Fets

<p>This repository is part of the Annotated Corpora of Historical Catalan (HisCat). It contains the first POS-tagged text that is partially manually corrected and used to train Old Catalan POS taggers, described in the following paper:</p> <p>Meelen, Marieke &amp; Pujol i Campeny, Afra, (2021) &#39;Old Catalan Morphosyntax: developing an annotated corpus&#39; in <em>Journal of Open Humanities Data</em>.</p> <p>This POS-tagged text is the 13th century <em>Llibre dels Fets</em>, a historical chronicle. The version of the text used for this project is</p> <p>Bruguera, J. (1991). <em>El Llibre dels Fets del Rei en Jaume</em>. Barcelona: Barcino.</p> <p>as prepared for the <em>Corpus Informatitzat del Catal&agrave; Antic</em></p> <p>Torruella, J., P&eacute;rez Saldanya, M., &amp; Martines, J. (2009). <em>Corpus Informatitzat del Catal&agrave; Antic</em>. URL: <a href="http://cica.cat/">http://cica.cat/</a>.</p> <p>The subcorpus counts with 164,096 POS-annotated tokens (165,538 tokens including punctuation and folio markers), of which 60,000 have been manually corrected. This subcorpus contains a total of and 4,506 main clauses. POS tagging of this text was done with the Memory-Based Tagger by TiMBL (<a href="https://languagemachines.github.io/mbt/">https://languagemachines.github.io/mbt/</a>). The code accompanying the paper can be found on GitHub: <a href="https://github.com/lothelanor/catalancorpora">https://github.com/lothelanor/catalancorpora</a>). In addition to memory-based tagging, have tried neural-based tagging with TARGER (<a href="https://github.com/achernodub/targer">https://github.com/achernodub/targer</a>) for which we created word embeddings that can be found on <a href="https://doi.org/10.5281/zenodo.5615556">Zenodo</a>. Results for memory-based tagging were better, however, which is why this version is uploaded here.</p> <pre>&nbsp;</pre>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Figure 3 in Presence of larvae of lampreys, Lampetra sp. (Cephalaspidomorphi, Petromyzontiformes), in a French Catalan basin

Figure 3. – Neighbourg-Joining barcoding tree with the COI marker (642 bp) of European Lampetra spp. and Petromyzon marinus (94 specimens) iden- tifying the three ammocoetes caught at Nefiach. Numbers at nodes correspond to bootstrap values. Grey box refers to the complex [Lampetra fluviatilis + Lampetra planeri].

opencc-by-4.0Dec 2018View details →
zenodo40/100

Figure 1 in Presence of larvae of lampreys, Lampetra sp. (Cephalaspidomorphi, Petromyzontiformes), in a French Catalan basin

Figure 1. – The Têt River in the Pyrénées-Orientales department and the lamprey location at Nefiach (square).

opencc-by-4.0Dec 2018View details →
zenodo40/100

Figure 2. – A in Presence of larvae of lampreys, Lampetra sp. (Cephalaspidomorphi, Petromyzontiformes), in a French Catalan basin

Figure 2. – A: Three of the six lamprey larvae, or ammocoetes, belonging to the genus Lampetra caught in the Têt River at Nefiach (MNHN 2016-0363). B, C: Enlargement of the head and tail showing the absence of pigmentation respectively on the snout and on the extremity of the caudal fin (arrows). Scale bars = 1 cm.

opencc-by-4.0Dec 2018View details →
zenodo40/100

Linked collectors and determiners for: Salpida species off the Catalan coast (NW Mediterranean) during summer 2003 and 2004. Neustonic and epipelagic planktonic sampling.

Natural history specimen data linked to collectors and determiners held within, "Salpida species off the Catalan coast (NW Mediterranean) during summer 2003 and 2004. Neustonic and epipelagic planktonic sampling". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/44aae8d4-cfcf-40fc-a3b0-274260069431">https://bionomia.net/dataset/44aae8d4-cfcf-40fc-a3b0-274260069431</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/44aae8d4-cfcf-40fc-a3b0-274260069431">https://gbif.org/dataset/44aae8d4-cfcf-40fc-a3b0-274260069431</a>. Formatted as a Frictionless Data package.

opencc-zeroJan 2024View details →
zenodo40/100

Catalan General Crawling

<p>The Catalan General Crawling Corpus is a 435-million-token web corpus of Catalan built from the web. It has been obtained by crawling the 500 most popular .cat and .ad domains during July 2020. It consists of 434.817.705 tokens, 19.451.691 sentences and 1.016.114 documents. Documents are separated by single new lines. It is a subcorpus of the&nbsp;<a href="../record/4519349">Catalan Textual Corpus</a>.</p> <p>We license the actual packaging of this data under a <a href="https://creativecommons.org/licenses/by/4.0/">Attribution 4.0 International License</a>.</p> <p><strong>Notice and take down policy</strong></p> <p>Notice: Should you consider that our data contains material that is owned by you and should therefore not be reproduced here, please:</p> <ul> <li>Clearly identify yourself, with detailed contact data such as an address, telephone number or email address at which you can be contacted.</li> <li>Clearly identify the copyrighted work claimed to be infringed.</li> <li>Clearly identify the material that is claimed to be infringing and information reasonably sufficient to allow us to locate</li> </ul> <p>Copyright (c) 2021 Text Mining Unit at BSC</p> <p>&nbsp;</p> <p>If you use this resource in your work, please cite our latest paper:</p> <p>@inproceedings{armengol-estape-etal-2021-multilingual,<br>&nbsp; &nbsp; title = "Are Multilingual Models the Best Choice for Moderately Under-resourced Languages? {A} Comprehensive Assessment for {C}atalan",<br>&nbsp; &nbsp; author = "Armengol-Estap{\'e}, Jordi &nbsp;and<br>&nbsp; &nbsp; &nbsp; Carrino, Casimiro Pio &nbsp;and<br>&nbsp; &nbsp; &nbsp; Rodriguez-Penagos, Carlos &nbsp;and<br>&nbsp; &nbsp; &nbsp; de Gibert Bonet, Ona &nbsp;and<br>&nbsp; &nbsp; &nbsp; Armentano-Oller, Carme &nbsp;and<br>&nbsp; &nbsp; &nbsp; Gonzalez-Agirre, Aitor &nbsp;and<br>&nbsp; &nbsp; &nbsp; Melero, Maite &nbsp;and<br>&nbsp; &nbsp; &nbsp; Villegas, Marta",<br>&nbsp; &nbsp; booktitle = "Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021",<br>&nbsp; &nbsp; month = aug,<br>&nbsp; &nbsp; year = "2021",<br>&nbsp; &nbsp; address = "Online",<br>&nbsp; &nbsp; publisher = "Association for Computational Linguistics",<br>&nbsp; &nbsp; url = "https://aclanthology.org/2021.findings-acl.437",<br>&nbsp; &nbsp; doi = "10.18653/v1/2021.findings-acl.437",<br>&nbsp; &nbsp; pages = "4933--4946",<br>}</p>

opencc-by-4.0Mar 2021View details →
zenodo40/100

ParlamentParla - Speech corpus of Catalan Parliamentary sessions

<p>This is the <a href="http://github.com/CollectivaT-dev/ParlamentParla">ParlamentParla</a> speech corpus for Catalan prepared by <a href="https://collectivat.cat/">Col&middot;lectivaT</a>. The audio segments were extracted from recordings the Catalan Parliament (<a href="https://www.parlament.cat/">Parlament de Catalunya</a>) plenary sessions, which took place between 2007/07/11 - 2018/07/17. We aligned the transcriptions with the recordings and extracted the corpus. The content belongs to the Catalan Parliament and the data is released conforming their <a href="https://www.parlament.cat/pcat/serveis-parlament/avis-legal/">terms of use</a>.</p> <p>Preparation of this corpus was partly supported by the <a href="http://cultura.gencat.cat/">Department of Culture</a> of the Catalan autonomous government, and the v2.0 was supported by the Barcelona Supercomputing Center, within the framework of the project <a href="http://aina.gencat.cat/">AINA</a> of the <a href="https://politiquesdigitals.gencat.cat/">Departament de Pol&iacute;tiques Digitals</a>.</p> <p>As of v2.0 the corpus is separated into 211 hours of clean and 400 hours of other quality segments. Furthermore, each speech segment is tagged with its speaker and each speaker with their gender. The statistics are detailed in the readme file.</p> <p>For more information, go to <a href="https://github.com/CollectivaT-dev/ParlamentParla">https://github.com/CollectivaT-dev/ParlamentParla</a> or mail info@collectivat.cat.</p> <p><strong>Revision log:</strong></p> <ul> <li> <p><em>2.0:</em> Major changes in the file structure; speaker ids with respective<br> genders added. The speakers of train, test and dev corpora do not overlap.<br> A major increase in size with a total time of 611 hours 43 minutes.</p> </li> <li> <p><em>1.0:</em> Much better quality due to improved segmentation, corpus separated<br> into clean and other.</p> </li> <li> <p><em>0.2:</em> First public release of approx. 320 hours.</p> </li> </ul>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Catalan CBOW Word Embeddings in Floret

<p><strong>Embeddings with the Catalan Textual Corpus</strong></p> <p>The embeddings have been trained with a Catalan textual corpus of&nbsp; more than 34GB of data using&nbsp;<a href="https://arxiv.org/pdf/1607.04606.pdf">floret</a>&nbsp;with the following&nbsp;&nbsp;hyperparameters:</p> <blockquote> <p>&nbsp; &nbsp; mode: str = &quot;floret&quot;,<br> &nbsp; &nbsp; model: str = &quot;cbow&quot;,<br> &nbsp; &nbsp; dim: int = 300,<br> &nbsp; &nbsp; mincount: int = 10,<br> &nbsp; &nbsp; minn: int = 5,<br> &nbsp; &nbsp; maxn: int = 6,<br> &nbsp; &nbsp; neg: int = 10,<br> &nbsp; &nbsp; hashcount: int = 2,<br> &nbsp; &nbsp; bucket: int = 50000,<br> &nbsp; &nbsp; thread: int = 128,</p> </blockquote> <p>The Catalan Textual Corpus used to train this embeddings, is the extended version of&nbsp;&nbsp;the initial available corpora described in&nbsp;<a href="https://arxiv.org/pdf/2107.07903.pdf">Armengol-Estap&eacute; et al. (2021)</a>. This new version includes:</p> <table> <caption>&nbsp;</caption> <thead> <tr> <th scope="col">Corpus</th> <th scope="col">Size in GB</th> </tr> </thead> <tbody> <tr> <td>CaCrawlat</td> <td>13.00</td> </tr> <tr> <td>Wikipedia</td> <td>1.10</td> </tr> <tr> <td>DOGC</td> <td>0.78</td> </tr> <tr> <td>Catalan Open Subtitles</td> <td>0.02</td> </tr> <tr> <td>Catalan Oscar</td> <td>4.00</td> </tr> <tr> <td>CaWaC</td> <td>3.60</td> </tr> <tr> <td>Cat. General Crawling</td> <td>2.50</td> </tr> <tr> <td>Cat. Goverment Crawling</td> <td>0.24</td> </tr> <tr> <td>ACN</td> <td>0.42</td> </tr> <tr> <td>Padicat</td> <td>0.63</td> </tr> <tr> <td>RacoCatal&agrave;</td> <td>8.10</td> </tr> <tr> <td>Naci&oacute;Digital</td> <td>0.42</td> </tr> <tr> <td>VilaWeb</td> <td>0.06</td> </tr> </tbody> </table> <p>From the new corpora, VilaWeb and Naci&oacute;Digital come from digital newspapers, Padicat is composed of crawlings of the Biblioteca de Catalunya, and CaCrawlat comes from the Biblioteca Nacional de Espa&ntilde;a (BNE).</p> <p>The processing took place&nbsp;on an HPC&nbsp;<a href="https://www.bsc.es/innovation-and-services/technical-information-cte-amd">node</a>&nbsp;equipped&nbsp;with an AMD EPYC 7742 (@ 2.250GHz) processor with 128 threads.</p> <p><strong>How to use</strong></p> <p>First initialize the spacy&nbsp;vectors from the floret table (.floret file):</p> <pre><code class="language-bash">spacy init vectors ca floret_embeddings_ca.floret floret_embeddings_ca --mode floret</code></pre> <pre><code class="language-python">import spacy # Load the floret vectors floret_embeddings = spacy.load("floret_embeddings_ca") # Get the embeddings of some words castanyes = floret_embeddings.vocab["castanyes"] flors = floret_embeddings.vocab["flors"] primavera = floret_embeddings.vocab["primavera"] tardor = floret_embeddings.vocab["tardor"] # Get some similarities print(flors.similarity(tardor)) print(flors.similarity(primavera)) # flors should be more similar to primavera than tardor. print(castanyes.similarity(primavera)) print(castanyes.similarity(tardor)) # castanyes should be more similar to tardor than primavera.</code></pre> <p><strong>Intended Uses and Limitations</strong></p> <p>At the time of submission, no measures have been taken to estimate the bias and toxicity embedded in the model. However, we are well aware that our models may be biased since the corpora have been collected using crawling techniques on multiple web sources. We intend to conduct research in these areas in the future, and if completed, this&nbsp;card will be updated.</p> <p><strong>Authors</strong></p> <p>The Text Mining Unit from Barcelona Supercomputing Center.</p> <p><strong>Contact Information</strong></p> <p>For further information, send an email to <a href="mailto:aina@bsc.es">aina@bsc.es</a>.</p> <p><strong>Funding</strong></p> <p>This work was funded by the <a href="https://politiquesdigitals.gencat.cat/ca/inici/index.html">Departament de la Vicepresid&egrave;ncia i de Pol&iacute;tiques Digitals i Territori de la Generalitat de Catalunya</a>&nbsp; within the framework of <a href="https://politiquesdigitals.gencat.cat/ca/economia/catalonia-ai/aina">Projecte AINA</a>.</p> <p><strong>Copyright</strong></p> <p>Copyright (c) 2022&nbsp;Text Mining Unit &nbsp;- Barcelona Supercomputing Center.</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Zipf's laws of meaning in Catalan Datasets

<p>This are the datasets referenced in the paper &quot;Zipf&rsquo;s laws of meaning in Catalan&quot;. The datasets are:</p> <p><strong>DIEC2_CTILC_senseCG</strong>. It contains the following information:</p> <ul> <li>lema: lemma in Catalan</li> <li>n_sentits: number of meanings for that lemma,</li> <li>freq: frequency for that lemma.</li> <li>rank of that lemma.</li> </ul> <p>Those lemmas are contained in both of the following corpuses:</p> <p>CTILC: https://ctilc.iec.cat.</p> <p>DIEC: Institut d&rsquo;Estudis Catalans. <em>Diccionari de la llengua catalana</em> [en l&iacute;nia]. 2a ed. Barcelona: Edicions 62: Enciclop&egrave;dia Catalana, 2007. [1a ed., 1995] &lt;<a href="https://dlc.iec.cat/">https://dlc.iec.cat/</a>&gt; [Consulta: October 2020]</p> <p><strong>DIEC2_GLISSANDO_senseCG. </strong></p> <p>It contains the following information:</p> <ul> <li>lema: lemma in Catalan</li> <li>n_sentits: number of meanings for that lemma,</li> <li>freq: frequency for that lemma.</li> <li>rank of that lemma.</li> </ul> <p>Those lemmas are contained in both of the following corpuses:</p> <p>Glissando: http://catalog.elra.info/en-us/repository/browse/ELRA-S0407/</p> <p>DIEC: Institut d&rsquo;Estudis Catalans. <em>Diccionari de la llengua catalana</em> [en l&iacute;nia]. 2a ed. Barcelona: Edicions 62: Enciclop&egrave;dia Catalana, 2007. [1a ed., 1995] &lt;<a href="https://dlc.iec.cat/">https://dlc.iec.cat/</a>&gt; [Consulta: October 2020]</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo36/100

FestCat speech synthesis dataset in Catalan, raw at 48kHz, 16bits/sample

<p>This dataset contains audio recordings from the 10 speakers of the FestCat project and an additional speaker &quot;uri&quot;.</p> <p>The recordings are available in the &quot;-raw&quot; compressed archives and are provided in the following format:</p> <pre><code class="language-bash">SAMPLERATE=48000 # Hz BITSPERSAMPLE=16 ENCODING="SIGNED_INTEGER" AUDIOHEADER="RAW" # (no header) # For the prompts and utts: TEXT_ENCODING="ISO-8859-15"</code></pre> <p>These sampling conditions were downsampled from the original recordings captured at 96kHz and using 24 bits/sample for all the FestCat speakers. The additional speaker &quot;uri&quot; had recordings 16kHz and were here upsampled to 48kHz.</p> <p>The recordings have been automatically segmented. The segmentation results are available at the &quot;-utts&quot; archives.</p> <p>The text prompts are available in the &quot;-prompts&quot; archive.</p> <p>The text information is encoded using ISO-8859-15.</p> <p>&nbsp;</p>

opencc-by-sa-3.0Oct 2014View details →
zenodo36/100

TeCla: Text Classification Catalan dataset

<p><em>Corpus de not&iacute;cies en catal&agrave; per a classificaci&oacute; textual, extret del web de l&#39;<a href="http://www.acn.cat">Ag&egrave;ncia Catalana de Not&iacute;cies</a> sota llic&egrave;ncia CC-BY-NC-ND</em></p> <p>TeCla (Text Classification) is a Catalan News corpus for thematic multi-class Text Classification tasks. The present version (2.0) contains 113.376 articles classified under a hierarchical class structure consisting of a coarse-grained and a fine-grained class. Each of the 4 coarse-grained classes accept a subset of fine-grained ones, 53 in total.</p> <p>The source data is crawled from the ACN (Catalan News Agency) site: <a href="http://www.acn.cat">http://www.acn.cat</a>, and used under CC-BY-NC-ND 4.0 licence. The dataset is released under the same licence, and is intended exclusively for training Machine Learning models.</p> <p>This dataset was developed by BSC TeMU as part of the AINA project, and intended as part of CLUB (Catalan Language Understanding Benchmark).</p>

opencc-by-nc-nd-4.0Mar 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record