Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

824

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

824 results for “spanish”

Learn how ShareScore rates datasets ↗
zenodo44/100

Photonics4All Bookmark Crime (Spanish)

<p>The purpose of the bookmarks for the project Photonics4All is to increase the public awareness of photonics and especially of the technological advances of photonics which have changed and improved everyday life (basic technology introduction).<br> <br> How can light help solve crimes?<br> <br> Light sources used to illuminate crime scenes help investigators solve crimes. Light is used by forensic detectives to help locate evidence such as latent fingerprints, bodily fluids, hair and fibres, bruises, wound patterns, shoe and foot imprints, gunshot residues or drug traces.  All of these clues fluoresce - or glow brightly - under selectively coloured light.  Photography is also vital to help collect and store this evidence. </p> <p>All thanks to Photonics! </p>

opencc-by-4.0Jul 2015View details →
zenodo44/100

Photonics4All Bookmark Needle (Spanish)

<p>The purpose of the bookmarks for the project Photonics4All is to increase the public awareness of photonics and especially of the technological advances of photonics which have changed and improved everyday life (basic technology introduction).<br> <br> How can light replace a needle?</p> <p><br> We no longer need to use a needle to monitor the level of oxygen in your blood!   We can use light emitting diodes (LEDs) attached to the top of your finger - and a light detector underneath to measure the amount of light passing through your finger.  As Hemoglobin - the proteins in red blood cells which carry oxygen - absorbs light we can determine whether you have enough oxygen in your blood. More advanced devices can also monitor your heart-rate and blood pressure.  We'll even be using light to measure your blood-sugar level in the future. All thanks to Photonics!</p> <p> </p>

opencc-by-4.0Jul 2015View details →
zenodo44/100

Photonics4All Bookmark Chip (Spanish)

<p>The purpose of the bookmarks for the project Photonics4All is to increase the public awareness of photonics and especially of the technological advances of photonics which have changed and improved everyday life (basic technology introduction).<br> <br> How light makes computers and phones smaller and faster?</p> <p>Did you know that we use light to fabricate the electronic chips in computers and mobile phones? Recent developments in photolithography where light is used to control where conductive metal is placed on the chips - have enabled us to put more transistors than there are people on earth!  Transistors are responsible for controlling the path of electricity/information through a chip.  These technological developments have led to improving the speed, size and energy consumption of our chips, making them smaller and more efficient.</p> <p> </p>

opencc-by-4.0Jul 2015View details →
zenodo44/100

IA Tweets Analysis Dataset (Spanish)

<h3><strong>Cite as</strong></h3> <p><em><strong>Guerrero-Contreras, G., Balderas-D&iacute;az, S., Serrano-Fern&aacute;ndez, A., &amp; Mu&ntilde;oz, A. (2024, June). Enhancing Sentiment Analysis on Social Media: Integrating Text and Metadata for Refined Insights. In 2024 International Conference on Intelligent Environments (IE) (pp. 62-69). IEEE.</strong></em></p> <h3>General Description</h3> <p>This dataset comprises 4,038 tweets in Spanish, related to discussions about artificial intelligence (AI), and was created and utilized in the publication "Enhancing Sentiment Analysis on Social Media: Integrating Text and Metadata for Refined Insights," (<a href="https://doi.org/10.1109/IE61493.2024.10599899" target="_blank" rel="noopener">10.1109/IE61493.2024.10599899</a>) presented at the 20th International Conference on Intelligent Environments. It is designed to support research on public perception, sentiment, and engagement with AI topics on social media from a Spanish-speaking perspective. Each entry includes detailed annotations covering sentiment analysis, user engagement metrics, and user profile characteristics, among others.</p> <h3>Data Collection Method</h3> <p>Tweets were gathered through the Twitter API v1.1 by targeting keywords and hashtags associated with artificial intelligence, focusing specifically on content in Spanish. The dataset captures a wide array of discussions, offering a holistic view of the Spanish-speaking public's sentiment towards AI.</p> <h3>Dataset Content</h3> <ul> <li><strong>ID</strong>: A unique identifier for each tweet.</li> <li><strong>text</strong>: The textual content of the tweet. It is a string with a maximum allowed length of 280 characters.</li> <li><strong>polarity</strong>: The tweet's sentiment polarity (e.g., Positive, Negative, Neutral).</li> <li><strong>favorite_count</strong>: Indicates how many times the tweet has been liked by Twitter users. It is a non-negative integer.</li> <li><strong>retweet_count</strong>: The number of times this tweet has been retweeted. It is a non-negative integer.</li> <li><strong>user_verified</strong>: When true, indicates that the user has a verified account, which helps the public recognize the authenticity of accounts of public interest. It is a boolean data type with two allowed values: True or False.</li> <li><strong>user_default_profile</strong>: When true, indicates that the user has not altered the theme or background of their user profile. It is a boolean data type with two allowed values: True or False.</li> <li><strong>user_has_extended_profile</strong>: When true, indicates that the user has an extended profile. An extended profile on Twitter allows users to provide more detailed information about themselves, such as an extended biography, a header image, details about their location, website, and other additional data. It is a boolean data type with two allowed values: True or False.</li> <li><strong>user_followers_count</strong>: The current number of followers the account has. It is a non-negative integer.</li> <li><strong>user_friends_count</strong>: The number of users that the account is following. It is a non-negative integer.</li> <li><strong>user_favourites_count</strong>: The number of tweets this user has liked since the account was created. It is a non-negative integer.</li> <li><strong>user_statuses_count</strong>: The number of tweets (including retweets) posted by the user. It is a non-negative integer.</li> <li><strong>user_protected</strong>: When true, indicates that this user has chosen to protect their tweets, meaning their tweets are not publicly visible without their permission. It is a boolean data type with two allowed values: True or False.</li> <li><strong>user_is_translator</strong>: When true, indicates that the user posting the tweet is a verified translator on Twitter. This means they have been recognized and validated by the platform as translators of content in different languages. It is a boolean data type with two allowed values: True or False.</li> </ul> <h3>Potential Use Cases</h3> <p>This dataset is aimed at academic researchers and practitioners with interests in:</p> <ul> <li>Sentiment analysis and natural language processing (NLP) with a focus on AI discussions in the Spanish language.</li> <li>Social media analysis on public engagement and perception of artificial intelligence among Spanish speakers.</li> <li>Exploring correlations between user engagement metrics and sentiment in discussions about AI.</li> </ul> <h3>Data Format and File Type</h3> <p>The dataset is provided in CSV format, ensuring compatibility with a wide range of data analysis tools and programming environments.</p> <h3>License</h3> <p>The dataset is available under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting sharing, copying, distribution, transmission, and adaptation of the work for any purpose, including commercial, provided proper attribution is given.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

(Rawdata) How do Spanish educational researchers use X's platform to promote the dissemination of scientific knowledge: a descriptive study: a descriptive study

<p>Rawdata used in the article 'How do Spanish educational researchers use X's platform to promote the dissemination of scientific knowledge: a descriptive study', from the project Comscienciaeduspain (FCT-20-15761), executed with the collaboration of the Spanish Foundation for Science and Technology &ndash; Ministry of Science and Innovation.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

MESINESP2 Corpora: Annotated data for medical semantic indexing in Spanish

<p>Gold Standard annotations of the MESINESP2 corpora (training, development and test sets).&nbsp;</p> <p><strong>Please cite this paper if you use this dataset:</strong></p> <pre><code class="language-bash">@inproceedings{gasco2021overview, title={Overview of BioASQ 2021-MESINESP track. Evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials}, author={Gasco, Luis and Nentidis, Anastasios and Krithara, Anastasia and Estrada-Zavala, Darryl and Murasaki, Renato Toshiyuki and Primo-Pe{\~n}a, Elena and Bojo Canales, Cristina and Paliouras, Georgios and Krallinger, Martin and others}, year={2021}, organization={CEUR Workshop Proceedings} }</code></pre> <p>&nbsp;</p> <p><strong>Introduction</strong></p> <p>The main aim of MESINESP2 is to promote the development of practically relevant semantic indexing tools for biomedical content in non-English language. We have generated a manually annotated corpus, where domain experts have labeled a set of scientific literature, clinical trials, and patent abstracts. All the&nbsp;documents were labeled with DeCS descriptors, which is a structured controlled vocabulary created by BIREME to index scientific publications on BvSalud,&nbsp;the largest database of scientific documents in Spanish, which hosts records from the databases LILACS, MEDLINE, IBECS, among others.&nbsp;</p> <p>MESINESP track at BioASQ9 explores the efficiency of systems for assigning DeCS to different types of biomedical documents. To that purpose, we have divided the task into three subtracks depending on the document type. Then,&nbsp;for each one we generated an annotated corpus which was provided to participating teams:</p> <ul> <li><strong>[Subtrack 1 corpus] MESINESP-L &ndash; Scientific Literature:&nbsp;</strong>It contains all Spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.</li> <li><strong>[Subtrack 2 corpus] <strong>MESINESP-T- Clinical Trials&nbsp;</strong></strong>contains records from&nbsp;<a href="https://reec.aemps.es/reec/public/web.html">Registro Espa&ntilde;ol de Estudios Cl&iacute;nicos (REEC)</a>. REEC doesn&#39;t&nbsp;provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC&nbsp;<a href="https://github.com/luisgasco/REECapi">API</a>.&nbsp;</li> <li><strong>[Subtrack 3 corpus] MESINESP-P &ndash; Patents:&nbsp;</strong>This corpus&nbsp;includes patents in Spanish extracted from Google Patents which have the IPC code &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;.</li> </ul> <p>In addition, we also provide a set of complementary data such as: the DeCS terminology file, a silver standard with the participants&#39; predictions to the task background set and the entities of medications, diseases, symptoms and medical procedures extracted from the BSC NERs documents.</p> <p>&nbsp;</p> <p><strong>Files structure:</strong></p> <p><strong>Silver_Standard_Mesinesp2.zip </strong>contains two separate sections. On the one hand, the union of the labels of the best model of each participating team as long as this model had obtained at least an F-score of 0.2 (folder <em>join</em>). On the other hand, the predictions of the best models of each participant have been included individually and anonymized&nbsp;(folder <em>separated</em>).&nbsp;This silver standard contains a set of <em>8642 scientific articles</em>, <em>1537 text sections from Clinical Practice Guidelines</em>, a set of <em>8458 text segments from Medication Data Sheets</em>, <em>461 clinical trials from REEC and 5170 patents</em>.&nbsp;</p> <p><strong>Subtrack1-Scientific_Literature.zip</strong> contains the corpora generated for subtrack 1. Content:</p> <ul> <li>Subtrack1: <ul> <li>Train:&nbsp; <ul> <li>training_set_track1_all.json: Full training set for subtrack 1.&nbsp;</li> <li>training_set_track1_only_articles.json:&nbsp;Articles training set for subtrack 1.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack1.json:&nbsp;</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack1.json: Test set for subtrack 1.&nbsp;</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack2-Clinical_Trials.zip</strong> contains the corpora generated for subtrack 2. Content:</p> <ul> </ul> <ul> <li>Subtrack2: <ul> <li>Train <ul> <li>training_set_subtrack2.json: Training set for subtrack 2.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack2.json:&nbsp;Manually annotated&nbsp;development set for subtrack 2.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack2.json: Test set for subtrack 2.</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack3-Patents.zip</strong> contains the corpora generated for subtrack 3. Content:</p> <ul> </ul> <ul> <li>Subtrack3: <ul> <li>Development <ul> <li>development_set_subtrack3.json:&nbsp;Manually annotated&nbsp;development set for subtrack 3.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack3.json: Test set for subtrack 3.</li> </ul> </li> </ul> </li> </ul> <p><strong>Additional data.zip&nbsp;</strong>contains the corpora with additional data for each subtrack of MESINESP2.</p> <p><strong>DeCS2020.tsv</strong> contains a DeCS table with the following structure:</p> <ul> <li>DeCS code</li> <li>Preferred descriptor (the preferred label in the Latin Spanish DeCS 2020&nbsp;set)</li> <li>List of synonyms (the descriptors and synonyms from&nbsp; Latin Spanish DeCS 2020&nbsp;set, separated by pipes.</li> </ul> <p><strong>DeCS2020.obo&nbsp;</strong>contains the *.obo file with the hierarchical relationships between DeCS descriptors.</p> <p>*Note: The <em>obo </em>and <em>tsv </em>files with DeCS2020 descriptors contain some additional COVID19 descriptors that will be included in future versions of DeCS. These items were provided by the Pan American Health Organization (PAHO), which has kindly shared this content to improve the results of the task by taking these descriptors into account.</p> <p>&nbsp;</p> <p><strong>Data format&nbsp;description</strong></p> <p>The&nbsp;<strong>input text files</strong>&nbsp;for the MESINESP track are JSON files with the following structure:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "title": "Metas de control de la presión arterial e impacto sobre desenlaces cardiovasculares en pacientes con diabetes mellitus tipo 2: un análisis crítico de la literatura", "abstractText": "La hipertensión arterial en individuos con diabetes mellitus tipo2 incrementa el riesgo de eventos cardiovasculares. Las guías internacionales de manejo recomiendan iniciar tratamiento farmacológico con valores de presión arterial &gt;140/90mmHg Sin embargo, no existe un punto de corte óptimo a partir del cual se logre reducir los eventos cardiovasculares sin originar eventos adversos; un rango de presión arterial &gt;130/80 y &lt;140/90mmHg parece ser el adecuado. Estos valores pueden alcanzarse mediante intervenciones no farmacológicas (dieta, ejercicio) y farmacológicas (por fármacos que hayan demostrado reducir eventos cardiovasculares). La elección de uno o varios fármacos debe ser individualizada, de acuerdo con factores como etnia, edad, comorbilidades asociadas, entre otros", "journal": "Clín. investig. arterioscler. (Ed. impr.)", "year": 2019, "db": "IBECS", "decsCodes": [ "D006973", "D000959", "D002318", "D003924", "D012307" ] } ] }</code></pre> <p>MESINESP&nbsp;<strong>entity mention files</strong>&nbsp;contain automatically generated mention annotations of medications, diseases, syntoms and medical procedures with the following JSON format:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "diseases": [ {"span": "hipertensión arterial", "start": "3", "end": "24"}, {"span": "diabetes mellitus tipo2", "start": "43", "end": "66"}, {"span": "eventos cardiovasculares", "start": "91", "end": "115"}], "medications": [], "procedures": [], "symptoms": []}] } ] }</code></pre> <p>&nbsp;</p> <p><strong>Dataset description:</strong><br> These corpora contain the data for each of the subtracks of MESINESP2 shared-task:</p> <ul> <li><strong>[Subtrack 1] MESINESP-L &ndash; Scientific Literature&nbsp;</strong>: &nbsp; <ul> <li><em><strong>Training set:&nbsp;</strong></em>It contains all spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.&nbsp;We have filtered out empty abstracts and non-Spanish abstracts.&nbsp;&nbsp;We have built the training dataset with the data crawled on 01/29/2021. This means that the data is a snapshot of that moment and that may change over time since LILACS and IBECS usually add or modify indexes after the first inclusion in the database.&nbsp;We distribute two different datasets: <ul> <li><strong>Articles training set:&nbsp;</strong>This corpus contains the set of 237574 Spanish scientific papers in VHL that have at least one DeCS code assigned to them.</li> <li><strong>Full training set</strong>: This corpus contains the whole set of 249474 Spanish documents from VHL that have at leas one DeCS code assigned to them.</li> </ul> </li> <li><strong>Development set:&nbsp;</strong>We provided a development set manually indexed by our expert annotators (not VHL ones). This dataset includes 1065 articles annotated with DeCS by three expert indexers in this controlled vocabulary. The articles were initially indexed by 7 annotators, after analyzing the Inter-Annotator Agreement among their annotations we decided to select the 3 best ones, considering their annotations the valid ones to build the test set. From those 1065 records: <ul> <li>213 articles were annotated by more than one annotator. We have selected de union between annotations.</li> <li>852 articles were annotated by only one of the three selected annotators with better performance.</li> </ul> </li> <li><strong>Test set:</strong> We provide a test set containing 491 abstracts&nbsp;from LILACS and IBECS. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 2] <strong>MESINESP-T- Clinical Trials</strong></strong>: &nbsp; <ul> <li><strong>Training set:&nbsp;</strong>The training dataset contains records from&nbsp;<a href="https://reec.aemps.es/reec/public/web.html">Registro Espa&ntilde;ol de Estudios Cl&iacute;nicos (REEC)</a>. REEC doesn&#39;t&nbsp;provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC&nbsp;<a href="https://github.com/luisgasco/REECapi">API</a>.&nbsp;Clinical trials are not indexed with DeCS terminology, we have used as training data a set of 3560 clinical trials that were automatically annotated in the first edition of MESINESP and that were published as a&nbsp;<a href="https://zenodo.org/record/3946558#.YFHyhZ1KiUk">Silver Standard outcome</a>. Because the performance of the models used by the participants was variable, we have only selected predictions from runs with a MiF higher than 0.41, which corresponds with the submission of the best team.&nbsp;</li> <li><strong>Development set: </strong>We provide a development set manually indexed by expert annotators. This dataset includes 147 clinical trials annotated with DeCS by seven expert indexers in this controlled vocabulary.</li> <li><strong>Test set:&nbsp;</strong>The test dataset contains a collection of 248 items. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 3] MESINESP-P &ndash; Patents:&nbsp;</strong> <ul> <li><strong>Development set: </strong>We provide a Development set manually indexed by expert annotators. This dataset includes 115 patents in Spanish extracted from Google Patents which have the IPC code &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;. We have selected these patents based on semantic similarity to the MESINESP-L training set to facilitate model generation and to try to improve model performance.</li> <li><strong>Test set:&nbsp;</strong>We provide a&nbsp;<strong>test set</strong>&nbsp;containing 119 records that correspond to a subset of patents published in Spanish with the IPC codes &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;.Similarly to the development set, we selected these records based on semantic similarity to the MESINESP-L training set.&nbsp;We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>Additional data:</strong> <ul> <li>&nbsp;We provide this information to the participants as additional data in the &ldquo;Additional Data&rdquo; folder. For each training, development, and test set there is an additional JSON file with the structure shown <a href="https://temu.bsc.es/mesinesp2/resources/">here</a>. Each file contains&nbsp;entities related to medications, diseases, symptoms, and medical procedures extrated with the BSC NERs.</li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Summary statistics:</strong></p> <table align="center"> <caption>MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP-L</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">Training</th> <td>237574</td> <td>1988684</td> <td>22434</td> <td>43106663</td> </tr> <tr> <th scope="row">Development</th> <td>1065</td> <td>11283</td> <td>3750</td> <td>211420</td> </tr> <tr> <th scope="row">Test</th> <td>491</td> <td>5398</td> <td>2124</td> <td>93645</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <th scope="row">Training</th> <td>3560</td> <td>52257</td> <td>3940</td> <td>4133166</td> </tr> <tr> <th scope="row">Development</th> <td>147</td> <td>2038</td> <td>771</td> <td>146791</td> </tr> <tr> <th scope="row">Test</th> <td>248</td> <td>3271</td> <td>905</td> <td>267031</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <th scope="row">Development</th> <td>109</td> <td>1092</td> <td>520</td> <td>38564</td> </tr> <tr> <th scope="row">Test</th> <td>119</td> <td>1176</td> <td>629</td> <td>9065</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> </tbody> </table> <p>&nbsp; </p><table align="center"> <caption>General MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">MESINESP-L</th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>243313</td> <td>2065199</td> <td>22641</td> <td>48006345</td> </tr> </tbody> </table> <p></p> <p><strong>Related resources:</strong></p> <ul> <li><a href="http://temu.bsc.es/mesinesp2/">MESINESP2&nbsp;Web</a></li> <li><a href="https://github.com/BioASQ/Evaluation-Measures">Evaluation library</a></li> <li><a href="http://metodologia.lilacs.bvsalud.org/download/E/LILACS-4-ManualIndexacao-es.pdf">Annotation guidelines</a></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhCNKd8zhgD0rLwbhxGqF_wX">Participating teams Youtube Videos</a></li> <li><a href="http://ceur-ws.org/Vol-2936/">Proceedings of BioASQ@CLEF2021</a></li> <li><a href="http://bioasq.org/">BioASQ Web</a></li> </ul> <p>&nbsp;</p> <p>For further information, please&nbsp;email us at luis.gasco@bsc.es</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

Spanish actresses photo gallery

<p>This is a dataset with the basic wikipedia information of Spanish actresses and the url of the photo.</p> <p>&nbsp;</p>

opencc-by-3.0Nov 2021View details →
zenodo44/100

Wind Speed vs Spanish Power Prices

<p>Average, min and max daily OMIE power prices (Spanish market) with corresponding wind average speed and maximum speed for each day. Units: &euro;/MWh (Power Price), km/h (wind speed).</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Dataset for sentiment analysis in Spanish

<p>This dataset is automatically generated by webscraping from sites such as Tripadvisor or Google Maps reviews. In these sites, the users post comments with ratings, allowing us to have tagged data. The code that generated this dataset can be found at the following URL:</p> <p><a href="https://github.com/fjramirezv/sentiment-webscraping">https://github.com/fjramirezv/sentiment-webscraping</a></p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Spanish Biomedical Crawled Corpus

<p>The largest Spanish biomedical and heath corpus to date gathered from a massive Spanish health domain crawler over more than 3,000 URLs were downloaded and preprocessed. All the collected data have been preprocessed to produce the CoWeSe (Corpus Web Salud Espa&ntilde;ol) resource, a large-scale and high-quality corpus intended for biomedical and health NLP in Spanish.</p> <p>Enlarged version with less restrictive document and sentence deduplication.</p> <p><strong>Citation</strong></p> <p>If you use this resource in your work, please cite our paper:</p> <pre>@misc{carrino2021spanish, title={Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models}, author={Casimiro Pio Carrino and Jordi Armengol-Estap&eacute; and Ona de Gibert Bonet and Asier Guti&eacute;rrez-Fandi&ntilde;o and Aitor Gonzalez-Agirre and Martin Krallinger and Marta Villegas}, year={2021}, eprint={2109.07765}, archivePrefix={arXiv}, primaryClass={cs.CL} } </pre> <p>Copyright (c) 2022 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Polifonia Corpus - Periodicals Module Metadata - Spanish Language

<p>We release the Metadata of the Periodicals module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at&nbsp;<a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

DATASET: Genotyping by sequencing of the common bean Spanish Diversity Panel

<p>Genotyping by sequencing of 308 common bean lines included in the Spanish Diversity Panel. The ApeKI restriction enzyme was used. The sequencing reads were aligned using the reference genome V2.1&nbsp;(https://phytozome.jgi.doe.gov/pz/portal.html#!info?alias=Org_Pvulgaris).&nbsp;A total of 11,763 SNP markers are included in this dataset after filtering for&nbsp;missing values (&lt; 10%) and minor allele frequency (MAF&gt; 0.05).&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Annotated Data in Spanish for Toxicity and Insults in Digital Social Networks

<p>This repository contains data sets and materials for a gold standard elaboration on toxicity and incivility in the digital sphere based on human coding to benchmark algorithmic classification tasks with transformers and LLMs. <strong>The labelling progress is 62%</strong>.</p> <p>We are labelling two samples of novel datasets of political digital interactions on Twitter (rebranded as X). The first set comprises almost 5 million data points from three Latin American protest events: (a) protests against the coronavirus and judicial reform measures in Argentina during August 2020; (b) protests against education budget cuts in Brazil in May 2019; and (c) the social outburst in Chile stemming from protests against the underground fare hike in October 2019. We are focusing on interactions in Spanish to elaborate a gold standard for digital interactions in this language, therefore, we prioritise Argentinian and Chilean data. The second set contains more than 31 million messages and more than 9 million interactions between 2010 and 2022, covering the election of members of the first Constitutional Convention in Chile, the drafting process and the referendum in which the proposal was rejected.</p> <p>This project is generously funded by the <strong>OpenAI Academic Programme</strong>, <strong>2024 FAE-UDP Research Grant</strong>, and partially by the <strong>St Hilda's College Muriel Wise Fund at the University of Oxford</strong>. The <a href="https://training-datalab.com/"><strong>Training Data Lab</strong></a> research group also logistically supports this project.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

ACTIV-ES: a comparable Spanish corpus comprised of film dialogue from Argentine, Mexican and Spanish productions

<p><strong>DESCRIPTION</strong>: ACTIV-ES is a comparable Spanish corpus comprised of film dialogue from Argentine, Mexican and Spanish productions. Titles for each of these three countries were seeded from the Internet Movie Database, subtitle data for the hearing impaired was provided by Opensubtitles.org and was post-processed to correct/remove subtitle, OCR and diacritic artifacts and annotated for part-of-speech.</p> <p>The data is available in two main formats: 1) running text for each document and 2) 1:5 gram aggregate files. Each format includes a plain text and part-of-speech annotated version. Document names reflect the language code, country, year, title, type, genre (first genre listed in the IMDb), and IMDb ID.</p> <p>For more information about the development and evaluation of these resources and to cite this work refer to:</p> <p>Francom, J., Hulden, M. and Ussishkin, A.. (2014) ACTIV-ES: a comparable, cross-dialect corpus of &#39;everyday&#39; Spanish from Argentina, Mexico, and Spain. In Proceedings of the Ninth Annual Language Resources and Evaluation Conference, Reykjavik, Iceland. European Language Resources Association (ELRA).</p> <p>In <strong>version .02</strong> of the tagged running format corpus in the /eagles directory has been added which includes the EAGLES tagset. This tagset is much more fleshed out than the simplified tagset in the /tagged directory. For information on the tagset refer here: <a href="http://nlp.lsi.upc.edu/freeling/doc/tagsets/tagset-es.html">http://nlp.lsi.upc.edu/freeling/doc/tagsets/tagset-es.html</a>.</p>

opengpl-2.0Nov 2018View details →
zenodo44/100

MeSDiCon - Medical Spanish Disease and symptom name Collection lexicon (unfiltered initial version)

<p>The MeSDiCon - (Medical Spanish Disease and symptom name Collection lexicon) consists of a list or gazetteer of candidate names of diseases and symptoms mentioned in Spanish clinical texts. Thus MeSDiCon serves as a lexical resource or dictionary for automatic detection of disease/symptom mentions, as well as indexing or classification of medical texts with such concept types.</p> <p>This collection was generated in a five step procedure:</p> <ol> <li>Automatic detection of mentions of disease/symptom terms in biomedical texts in English (including mapping/normalization to MeSH terms or OMIM identifiers).</li> <li>Generation of a unique name list from the detected concept mentions.</li> <li>Basic filtering of non- disease/symptom names or highly ambiguous mentions-abbreviations using basic characteristics like name morphology and length criteria.</li> <li>Automatic translation of name lists form English to Spanish using a medical machine translation system (see Soares, F. and Krallinger, M. BSC Participation in the WMT Translation of Biomedical Abstracts. In <em>Proceedings of the Fourth Conference on Machine Translation, Volume 3: Shared Task Papers, </em>pp. 175-178 2019; https://zenodo.org/record/3346802)</li> <li>Automatic mention lookup of translated names in a collection of 20 million Spanish clinical notes (primary care and pediatrics).</li> </ol> <p>Every term in MeSDiCon is&nbsp;identified by a text span (in Spanish), a target terminology namespace to which it was automatically mapped (MeSH or OMIM)&nbsp;and its corresponding concept identifier in that target terminology. Moreover, we provide for every text span the absolute term frequency, i.e. the number of matches in the corpus of 20 million clinical notes and the number of documents or notes in which it was automatically.</p> <p>Important note: no manual filtering of the MeSDiCon was carried out, implying that some entries might comprise errors, either due to the initial name recognition and concept mapping in English or due to wrong automatic translations into Spanish.</p> <p>The MeSDiCon resource is provided in two formats:</p> <ul> <li>TSV. Data is separated by tabs (\t). Every row of the file has&nbsp;the following fields:</li> </ul> <pre><code>terminology identifier translatedTerm termCount documentCount</code></pre> <ul> <li>JSON.&nbsp;Records are stored as a list of JSON objects. They have the following fields:</li> </ul> <pre><code>{ "terminology":"MESH", "identifier":"D025861", "translatedTerm":"Trastornos de la coagulación", "termFrequency":9, "documentFrequency":9 }</code></pre> <p>&nbsp;</p> <p>Copyright (c) 2019 Secretar&iacute;a de Estado para el Avance Digital</p>

opencc-by-4.0Nov 2019View details →
zenodo44/100

MeSCCon - Medical Spanish Chemical compound, drug and medication Name Lexicon (unfiltered version)

<p>The MeSCCon (Medical Spanish Chemical compound, drug and medication Name Lexicon) consists of a list or gazetteer of candidate names of chemicals, drugs, and medications mentioned in Spanish clinical texts. Thus MeSCCon serves as a lexical resource or dictionary for automatic detection of chemical/drug mentions, as well as indexing or classification of medical texts with such concept types.</p> <p>This collection was generated in a five step procedure:</p> <ol> <li>Automatic detection of mentions of chemicals and drugs in biomedical texts in English (including mapping/normalization to MeSH terms or ChEBI identifiers).</li> <li>Generation of a unique name list from the detected concept mentions.</li> <li>Basic filtering of non-chemical names or highly ambiguous mentions-abbreviations using basic characteristics like name morphology and length criteria.</li> <li>Automatic translation of name lists from English to Spanish using a medical machine translation system (see Soares, F. and Krallinger, M. BSC Participation in the WMT Translation of Biomedical Abstracts. In <em>Proceedings of the Fourth Conference on Machine Translation, Volume 3: Shared Task Papers, </em>pp. 175-178 2019; https://zenodo.org/record/3346802)</li> <li>Automatic mention lookup of translated names in a collection of 20 million Spanish clinical notes (primary care and pedriatrics).</li> </ol> <p>Every term in MeSCCon is&nbsp;identified by a text span (in Spanish), a target terminology namespace to which it was automatically mapped (MeSH or ChEBI)&nbsp;and the&nbsp;corresponding concept identifier in that terminology.</p> <p>Moreover, we provide for every text span the absolute term frequency, i.e. the number of matches in the corpus of 20 million clinical notes and the number of documents or notes in which it was found.</p> <p>Important note: no manual filtering of the MeSCCon was carried out, implying that some entries might comprise errors, either due to the initial name recognition and concept mapping in English or due to wrong automatic translations into Spanish.</p> <p>The MeSCCon resource is provided in two formats:</p> <ul> <li>TSV. Data is separated by tabs (\t). Every row of the file has&nbsp;the following fields:</li> </ul> <pre><code>terminology identifier translatedTerm termCount documentCount</code></pre> <ul> <li>JSON. Records are stored as a list of JSON objects. They have the following fields:</li> </ul> <pre><code class="language-javascript">{ "terminology":"MESH", "identifier":"D009020", "translatedTerm":"clorhidrato de morfina", "termFrequency":1, "documentFrequency":1 }</code></pre> <p>&nbsp;</p> <p>Copyright (c) 2019 Secretar&iacute;a de Estado para el Avance Digital</p>

opencc-by-4.0Nov 2019View details →
zenodo44/100

SWL-LSE: SignaMed Word-Level LSE, a Dataset of Spanish Sign Language Health Signs

<h2>SWL-LSE Dataset</h2> <p>The SWL-LSE dataset is coined from SignaMed Word-Level LSE (Lengua de Signos Espa&ntilde;ola -Spanish Sign Language).</p> <h2>Overview</h2> <p>The dataset consists of 8,000 sign sequences from 300 different sign classes related to the health domain. Each class is represented by an RGB video that serves as the dictionary sign. These dictionary signs were reproduced by 124 signers, including deaf individuals, interpreters, and L2 Spanish Sign Language (LSE) students, using their webcams or mobile phones via the SignaMed platform (<a href="https://signamed.web.app" target="_new" rel="noopener">https://signamed.web.app</a>). For privacy reasons, only the skeleton data is shared.</p> <p>The process of collecting the dataset is described in:</p> <p>V&aacute;zquez-Enr&iacute;quez, M.; Alba-Castro, J.L.; P&eacute;rez-P&eacute;rez, A.; Cabeza-Pereiro, C.; Doc&iacute;o-Fern&aacute;ndez, L. SignaMed: a Cooperative<br>Bilingual LSE-Spanish Dictionary in the Healthcare Domain. In Proceedings of the Proceedings of the LREC-COLING 2024<br>11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources; Efthimiou, E.;&nbsp;Fotinea, S.E.; Hanke, T.; Hochgesang, J.A.; Mesch, J.; Schulder, M., Eds., Torino, Italia, 2024; pp. 386&ndash;394.&nbsp;</p> <p>The dataset itself and the pipeline for training and executing a baseline model based on skeletons is described in this github (https://github.com/mvazquezgts/SWL-LSE), and this paper:</p> <p>V&aacute;zquez-Enr&iacute;quez, M.; Alba-Castro, J.L.; Doc&iacute;o-Fern&aacute;ndez, L.; Rodr&iacute;guez-Banga, E. SWL-LSE: A Dataset of Spanish Sign Language Health Signs with an ISLR Baseline Method. Technologies 2024, 12(10), 205, D.O.I:10.3390/technologies12100205</p> <h2>Files</h2> <h3>1. VIDEOS_REF.zip</h3> <ul> <li><strong>Description</strong>: RGB videos recorded in lab conditions that represent each sign-class</li> <li><strong>Total files</strong>: 300</li> </ul> <h3>2. videos_ref_annotations.csv</h3> <ul> <li><strong>Description</strong>: CSV file with the correspondence between the name of the video, its class ID and gloss in spanish: FILENAME,CLASS_ID,LABEL.</li> <li><strong>Total files</strong>: 1</li> </ul> <h3>3. ANNOTATIONS.zip</h3> <ul> <li><strong>Description</strong>: 3 CSV files with train, validation and test file-class correspondences: FILENAME,CLASS_ID</li> <li><strong>Total files</strong>: 3</li> </ul> <h3>4. MEDIAPIPE.zip</h3> <ul> <li><strong>Description</strong>: Pickle files containing the full output of Mediapipe using their Heavy model. Each .pkl file contains the outputs of Mediapipe Holistic legacy, Mediapipe Pose and Mediapipe Hands. Each file is package as a dictionary: dict_keys(['pose', 'hands', 'holistic_legacy'])</li> <li><strong>Total files</strong>: 8000</li> </ul> <h2>Usage</h2> <p>Researchers and practitioners in pattern recognition, machine learning, and sign language linguistics may find this dataset valuable for:</p> <ul> <li>Training/testing machine learning models for isolated sign language recognition or gesture recognition.</li> <li>Analyzing patterns on signs realization</li> </ul> <h2>Acknowledgments</h2> <p>This dataset is a collaborative effort of the next research goups and entities:</p> <ul> <li><a href="http://gtm.uvigo.es/en/">Group of Multimedia Technologies (GTM)</a> from the <a href="https://atlanttic.uvigo.es/en">atlanTTic Research Center</a> of <a href="http://www.uvigo.es/">University of Vigo</a> (Spain)</li> <li><a href="http://grades.uvigo.gal/">Group of Discourse and Society (GRADES)</a> from the <a href="https://fft.uvigo.es/en/">School of Philology and Translation</a> of <a href="http://www.uvigo.es/">University of Vigo</a> (Spain)</li> <li><a href="http://www.faxpg.es/">Federation of Deaf People Galician Associations (FAXPG)</a></li> <li><a href="https://fundacioncnse-dilse.org">Fundaci&oacute;n CNSE-DILSE</a></li> </ul> <p>Gratitude is extended to them for their contributions and support.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

CT-EBM-SP - Corpus of Clinical Trials for Evidence-Based-Medicine in Spanish (version 2)

<p>A collection of <strong>1200 texts</strong> (292173 tokens) about<strong> clinical trials studies</strong> and <strong>clinical trials announcements</strong> in <strong>Spanish</strong>:</p> <p>- 500 abstracts from journals published under a Creative Commons license, e.g. available in PubMed or the Scientific Electronic Library Online (SciELO).<br>- 700 clinical trials announcements published in the European Clinical Trials Register and Repositorio Espa&ntilde;ol de Estudios Cl&iacute;nicos.</p> <p>Texts were annotated with the following entities types:</p> <p>- <strong>Semantic groups from the Unified Medical Language System</strong>:&nbsp;<br>&nbsp; &bull; ANAT: anatomy<br>&nbsp; &bull; CHEM: pharmacological and chemical substances<br>&nbsp; &bull; DEVI: medical devices<br>&nbsp; &bull; DISO: pathologic conditions&nbsp;<br>&nbsp; &bull; LIVB: living beings, included the human being<br>&nbsp; &bull; PHYS: physiological processes<br>&nbsp; &bull; PROC: lab tests, diagnostic or therapeutic procedures<br>- <strong>Medical drug information</strong>:<br>&nbsp; &bull; Contraindicated: a contraindicated drug or treatment<br>&nbsp; &bull; Dose: dose or strength<br>&nbsp; &bull; Form: dosage form<br>&nbsp; &bull; Route: administration route or mode<br>- <strong>Temporal expressions</strong> &nbsp;<br>&nbsp; &bull; Age<br>&nbsp; &bull; Date<br>&nbsp; &bull; Duration<br>&nbsp; &bull; Frequency<br>&nbsp; &bull; Time<br>- <strong>Miscellaneous medical entities</strong>:&nbsp;<br>&nbsp; &bull; Concept: abstract concepts, statistical tests or measurement scales<br>&nbsp; &bull; Food: foods or drinks<br>&nbsp; &bull; Observation: medical observations or clinical findings<br>&nbsp; &bull; Quantifier_or_Qualifier: quantifier or qualifier adjective<br>&nbsp; &bull; Result_or_Value: result or value of a measurement, laboratory analysis or procedure<br>- <strong>Negation/Speculation</strong>: &nbsp;<br>&nbsp; &bull; Neg_cue: negation cue<br>&nbsp; &bull; Negated: negated event<br>&nbsp; &bull; Spec_cue: speculation cue<br>&nbsp; &bull; Speculated: speculated or uncertain event<br>- <strong>Attributes</strong>:&nbsp;<br>&nbsp; &bull; Temporality:<br>&nbsp; &nbsp; ◦ History_of: past event<br>&nbsp; &nbsp; ◦ Future: future event<br>&nbsp; &bull; Experiencer:<br>&nbsp; &nbsp; ◦ Patient: patient or participant on a clinical trial<br>&nbsp; &nbsp; ◦ Family_member<br>&nbsp; &nbsp; ◦ Other: other person different from the patient or the family member</p> <p>86 389 entities and 16 590 attributes were annotated. 10% of the corpus was doubly annotated, and high inter-annotator agreement (IAA) values were achieved: F1-score = 0.84% for entities; and F1-score = 0.88% for attributes (both in strict match).&nbsp;</p> <p>The dataset includes the <strong>texts and annotations used for the human evaluation</strong> of the medical named entity tool:</p> <p>- 100 clinical trial announcements from EudraCT not used for system development: we provide files of the version revised by medical professionals (Reference folder)<br>- 100 clinical cases with Creative Commons license: we provide files with the files revised by medical professionals (Reference folder). These data come from:</p> <p>&nbsp; &nbsp;&bull; Urgencias Bidasoa (https://urgenciasbidasoa.wordpress.com/casos-clinicos-3/)<br>&nbsp; &nbsp;&bull; Hipocampo.org (https://www.hipocampo.org/)<br>&nbsp; &nbsp;&bull; Cases published by Sociedad Andaluza de Medicina Familiar y Comunitaria (SAMFyC): we are greatly thankful for giving us permission to use these cases and we acknowledge that the copyright belongs to the authors' contents. Clinical cases were extracted from books published from 2016 to 2022 (https://www.samfyc.es/tipos-publicacion/publicaciones/).<br>&nbsp; &nbsp;<br>If you use these data, please, acknowledge the copyright and intellectual property rights to the authors' contents.</p> <p>The dataset is freely distributed for research and educational purposes under a Creative Commons Non-Commercial Attribution (CC-BY-NC-A) License.</p> <p>If you use the CT-EBM-SP vs. 2 dataset, please, cite as follows:</p> <p>Campillos-Llanos, L., A. Valverde-Mateos &amp; A. Capllonch-Carrion (2024) Hybrid natural language processing tool for semantic annotation of medical texts in Spanish. BMC Bioinformatics. BioMed Central.</p>

opencc-by-nc-4.0Sep 2024View details →
zenodo44/100

Wikipedia: wikipedia-es (Spanish)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>

opencc-by-sa-4.0Aug 2024View details →
zenodo44/100

Aggregate Dataset on Descriptive Representation in the Spanish Parliament (2016-2023)

<p>This is the dataset aggregated at the legislative/party level by the CSIC team from the individual-level data provided by the Sciences Po and CSIC teams on descriptive representation in the Spanish lower chamber of Parliament for WP4 of the ActEU project.&nbsp;</p>

opencc-by-4.0Sep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record