Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

824

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

824 results for “spanish”

Learn how ShareScore rates datasets ↗
zenodo44/100

TRANSIT long-distance multimodal trips model results - a Spanish case study

<p>The files downloaded present the results per scenario considered obtained with the agent-based model developed for the assessment of the&nbsp;Intermodal Timetable Synchronisation solution proposed in the scope of the TRANSIT project (<a href="https://www.transit-h2020.eu/">https://www.transit-h2020.eu/</a>).</p> <p>An agent-based modelling framework called <a href="https://github.com/StefanoPenazzi/jtap/tree/main">J-TAP</a>&nbsp;has been developed and put at work to implement a Spanish long-distance multimodal trips model. The enhanced version&nbsp;of J-TAP used in this work can be found on a&nbsp;<a href="https://github.com/NommonSolutionsAndTechnologies/jtap">github repository</a>.</p> <p>The case study is focused on modelling the long-distance travel patterns of the residents in the Valencia (Spain)&nbsp;area. The destinations considered include the whole of Spain. The period under study is a full year from March 2019 to February 2020 (both inclusive).The files are structured in three different scenarios:</p> <ol> <li><strong>CS01&nbsp;- Baseline</strong>. The current state of the network is considered and the actual long-distance travel patterns are obtained.</li> <li><strong>CS02 - HSR connection with Madrid-Barajas airport</strong>. The long-distance travel patterns are modelled with hard measures, the high-speed rail is connected to Madrid-Barajas airport.</li> <li><strong>CS03 - HSR connection with Madrid-Barajas airport and timetable synchronisation</strong>. The effects of the timetable synchronisation are modelled.</li> </ol> <p>Each scenario includes the following files:</p> <ul> <li>ctapModelParameters. A folder containing all the information extracted from the&nbsp;<a href="https://neo4j.com/product/graph-data-science/?utm_program=emea-prospecting&amp;utm_source=google&amp;utm_medium=cpc&amp;utm_campaign=emea-search-offers&amp;utm_adgroup=dynamic&amp;utm_content=dynamic&amp;utm_placement=&amp;utm_network=g&amp;gclid=Cj0KCQiAwJWdBhCYARIsAJc4idAo4CEi9lU8TXwmBym8MNHpEIZHPBs3x_4phxbu76y1XKbYlFoZCjIaAiGhEALw_wcB">neo4j</a>&nbsp;graph database created to model the multimodal network and the agents.&nbsp;J-TAP contains packages that simplify network creation in the graph database. This&nbsp;information is stored in .json files (e.g., &quot;Os2DsTravelCostParameter.json&quot; contains the generalised cost for each OD pair and transport mode, &quot;AttractivenessParameter.json&quot; contains the&nbsp;attractiveness by destination, activity,&nbsp;time of the year and agent, etc.). The solver included in the J-TAP framework uses this information to calculate the agents plans.</li> <li>population.json. The result&nbsp;of the J-TAP optimisation. It includes the fitness value for each agent plan evaluated during the J-TAP execution. The best plan for each agent is selected as the plan performed by the agent. An agent plan includes: <ul> <li>activities - Sequence of activities.</li> <li>locations - Sequence of locations</li> <li>ts - Initial time of the activity</li> <li>te - Final time of the activity</li> </ul> </li> <li>LinkTimeFlow.csv. It is obtained after processing the previous file. It contains the number of agents using each link in the network (i.e., road, rail, air and cross links) in each time interval. The first column represents&nbsp;the link id and the rest of columns indicates the number of agents in each interval.</li> </ul> <p>The J-TAP simulation framework is explained in detail in TRANSIT&#39;s deliverable&nbsp;<a href="http://www.nommon-files.es/transit/TRANSIT-D5.1_Modelling_Framework_v02.00.00.pdf">D5.1. TRANSIT Modelling and Simulation Framework</a>&nbsp;and the complete description of the case studies and scenarios tested is included in TRANSIT&#39;s deliverable&nbsp;<a href="http://www.nommon-files.es/transit/TRANSIT-D6.1_Assessment_of_Intermodal_Concepts_00.02.00.pdf">D6.1. Impact Assessment of New Intermodal Concepts and Passenger Information Services: Conclusions and Recommendations</a>.</p> <p>Thank you for downloading the dataset! It would be very helpful if you share your view on the data show with us.&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Spanish semantic fields

<p>A database with 73.000 frequent spanish words classified in semantic fields in a three level hierarchy</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Nos_Machine translation test suite Spanish-Galician

<p>334 spanish-galician parallel sentences for the evaluation of machine translation. Sentences are classified in categories based on challenging linguistic phenomena.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Nos_Machine translation gold standard Spanish-Galician 2

<p>1998 carefully curated Galician-Spanish parallel sentences for the evaluation of automatic machine translation. Sentence pairs are classified according to whether the translations are highly literal or not.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Nos_Machine translation gold standard Spanish-Galician 1

<p>1998 carefully curated Galician-Spanish parallel sentences for the evaluation of automatic machine translation. Sentence pairs are classified according to whether the translations are highly literal or not.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Spanish Workers' Statute Legal Relations and RDF

<p>These datasets result from&nbsp;extracting legal events and relationships from the Spanish Workers&#39; Statutes and structuring the extracted data in an RDF graph. After a set of experiments conducted with GPT-3.5 using&nbsp;scarce annotated data from previous works&nbsp;<a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6432">[1]</a>, a 5-shot learning approach was applied to the full text of the Spanish Workers&rsquo; Statute. Approximately 1500 relations&nbsp;were extracted in a JSON format, structured into a dataset, and represented in an RDF graph.</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

OneNet project - T9.3 - Spanish demo open data

<p>-Anonymized technical data from flexible resources participating in OneNet Spanish demonstration&nbsp;&nbsp;</p> <p>-Market results assessed by the local market platforms considering market bids and DSOs requirement for the OneNet Spanish demonstration</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Hydrodynamic and morphological information, and the absolute variations of the vulnerability indices for the period 2000-2015 of the Spanish Iberia Peninsula estuaries.

<p>The dataset included in this repository was obtained during the project entitled &#39;Sensibilidad f&iacute;sica y biotic de los estuarios peninsulares al cambio global (SENSES)&#39; funded by &#39;Fundaci&oacute;n Biodiversidad&#39;, PRCV00487. The data were used in&nbsp;the research article&nbsp;&#39;Sensitivity of Iberian estuaries to changes in sea water temperature, salinity, river-flow, mean sea level, and tidal amplitudes&#39; submitted to <em>Estuarine, Coastal and Shelf Science</em>.</p> <p>Brief description of dataset:</p> <p>For each estuary, the following parameters were calculated</p> <ul> <li>Fachade: the location of the estuary</li> <li>Area (km<sup>2</sup>)&nbsp;</li> <li>D (m): water depth at the mouth of the estuary in 2000 and 2015</li> <li>Tidal Prim (m<sup>3</sup>)</li> <li>Q<sub><em>f</em></sub>&nbsp;(m<sup>3</sup>/s): river flow in 2000 and 2015</li> <li><em>a&nbsp;</em>(m): tidal amplitude of the free surface elevation in 2000 and 2015</li> <li>∆<em>U</em>&nbsp;(m/s): absolute variation of the tidal current amplitude between 2000 and 2015</li> <li>∆<em>E</em>&nbsp;(W/m<sup>2</sup>): absolute variation of the tidal energy flux propagation index between 2000 and 2015</li> <li>∆<em>Ri&nbsp;</em>: absolute variation of the bulk Richardson number index between 2000 and 2015</li> <li>∆<em>SI</em>: absolute variation of the salinity intrusion index between 2000 and 2015</li> </ul> <p>A wide description of the parameters can be found in Serrano, M. A. et al (submitted to <em>Estuarine, Coastal and Shelf Science</em>)</p> <p>Contact person: mserranog@ugr.es</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

FastText Spanish Medical Embeddings

<p>[Plan TL/medicine/word embeddings] Word embeddings generated from Spanish corpora that include: (a) the full-text in Spanish available in SciELO.org (until December/2018), (b) all articles from the following Wikipedia categories: Pharmacology, Pharmacy, Medicine and Biology (during December/2018) and (c) the concatenation of the previous two corpora.</p> <p>We used fastText to train the word embeddings.</p> <p>For more information, we refer to the corresponding article:&nbsp;<a href="https://www.aclweb.org/anthology/W19-1916/">https://www.aclweb.org/anthology/W19-1916/</a></p> <p>Copyright (c) 2021 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0Apr 2020View details →
zenodo40/100

Figures 16-19 in Roncus elbulli (Arachnida, Pseudoscorpiones), a new species from Cap de Creus Nature Park (Catalonia, Spain), with a key to the Spanish species of the genus Roncus

Figures 16-19. Roncus elbulli sp. n., female paratype, Cala Canadell. SEM photographs: 16. left palp, dorsal view; 17. chelal microsetae pattern below trichobothria eb/esb; 18. fingers of the chela, antiaxial face, partial view, showing trichobothrium sb and sensilla p1 and p2 on movable finger. Roncus cadinensis Zaragoza, 2007, male paratype. SEM photograph: 19. chelal microsetae pattern below trichobothria eb/ esb. Scale bars (mm): 0.05 (Figs 17, 18), 0.10 (Fig. 19), 0.50 (Fig. 16).

opencc-by-4.0Apr 2009View details →
zenodo40/100

Figures 3-9 in Roncus elbulli (Arachnida, Pseudoscorpiones), a new species from Cap de Creus Nature Park (Catalonia, Spain), with a key to the Spanish species of the genus Roncus

Figures 3-9. Roncus elbulli sp. n., male holotype. 3. carapace; 4. anterior margin of carapace, showing epistome; 5. anterior and medial processes of coxa I; 6. left chelicera; 7. fingers of left chelicera, partial view; 8. right leg IV, lateral view; 9. distal end of tarsus and apotele of left leg IV, lateral view. Scale bars (mm): 0.05 (Figs 4, 5, 7, 9), 0.10 (Fig. 6), 0.20 (Figs 3, 8).

opencc-by-4.0Apr 2009View details →
zenodo40/100

Figures 10-15 in Roncus elbulli (Arachnida, Pseudoscorpiones), a new species from Cap de Creus Nature Park (Catalonia, Spain), with a key to the Spanish species of the genus Roncus

Figures 10-15. Roncus elbulli sp. n., male holotype (except where otherwise noted): 10. left palp, without chela, dorsal view; 11. left chela, dorsal view; 12. left chela, lateral view; 13. male paratype, chelal microsetae pattern below trichobothria eb/esb. 14. chelal microsetae pattern below trichobothria eb/esb. Scale bars (mm): 0.05 (Figs 13-14), 0.20 (Figs 10-12). 15. Roncus cadinensis Zaragoza, 2007, male holotype: chelal microsetae pattern below trichobothria eb/esb. Scale bar (mm): 0.05.

opencc-by-4.0Apr 2009View details →
zenodo40/100

Spanish abstracts from PubMed (machine-translated from English)

<p>Spanish abstracts from PubMed (machine-translated from English)</p> <p>A state-of-the-art, domain-specific Neural Machine Translation system has been used to automatically translate a large number of articles (titles and abstracts) from PubMed, from English to Spanish. This dataset, in JSON format, provides MeSH terms and DeCS codes (if available), as well as other fields such as the year of publication for all articles.</p> <p>Copyright (c) 2020 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0May 2020View details →
zenodo40/100

MESINESP: Medical Semantic Indexing in Spanish - Train dataset

<p><em><strong>Please use the <a href="https://doi.org/10.5281/zenodo.4612274">MESINESP2 corpus (the second edition of the shared-task)</a> since it has a higher level of curation, quality and is organized by document type (scientific articles, patents and clinical trials).</strong></em></p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>INTRODUCTION</strong>:</p> <p>The Mesinesp (Spanish BioASQ track, see https://temu.bsc.es/mesinesp) training set has a total of 369,368 records.&nbsp;</p> <p>The training dataset contains all records from LILACS and IBECS databases at the Virtual Health Library (VHL) with a non-empty abstract written in Spanish. The URL used to retrieve records is as follows:<br> http://pesquisa.bvsalud.org/portal/?output=xml&amp;lang=es&amp;sort=YEAR_DESC&amp;format=abstract&amp;filter[db][]=LILACS&amp;filter[db][]=IBECS&amp;q=&amp;index=tw&amp;</p> <p>We have filtered out empty abstracts and non-Spanish abstracts.&nbsp;</p> <p>The training dataset was crawled on 10/22/2019. This means that the data is a snapshot of that moment and that may change over time. In fact, it is very likely that the data will undergo minor changes as the different databases that make up LILACS and IBECS may add or modify the indexes.</p> <p>&nbsp;</p> <p><strong>ZIP STRUCTURE:</strong></p> <p>The training data sets contain 369,368 records from 26,609 different journals. Two different data sets are distributed as described below:</p> <p>&nbsp;- <em>Original Train set</em> with 369,368 records that also include the qualifiers, as retrieved from VHL.&nbsp;<br> &nbsp;- <em>Pre-processed Train set</em><strong> </strong>with the 318,658 records with at least one DeCS code and with no qualifiers.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>STATISTICS</strong>:</p> <p>Abstracts&rsquo; length (measured in characters)<br> Min: 12<br> Avg: 1140.41<br> Median: 1094<br> Max: 9428</p> <p>Number of DeCS codes per file<br> Min: 1<br> Avg: 8.12<br> Median: 7<br> Max: 53</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>CORPUS FORMAT</strong>:</p> <p>The training data sets are distributed as a JSON file with the following format:</p> <pre><code>{   "articles": [     {       "id": "Id of the article",       "title": "Title of the article",       "abstractText": "Content of the abstract",       "journal": "Name of the journal",       "year": 2018,       "db": "Name of the database",       "decsCodes": [         "code1",         "code2",         "code3"       ]     }   ] } </code></pre> <p>Note that the decsCodes field lists the DeCs Ids assigned to a record in the source data. Since the original XML data contain descriptors (no codes), we provide a DeCs conversion table (https://temu.bsc.es/mesinesp/wp-content/uploads/2019/12/DeCS.2019.v5.tsv.zip) with:</p> <p>&nbsp;- DeCs codes<br> &nbsp;- Preferred descriptor (the label used in the European DeCs 2019 set)<br> &nbsp;- List of synonyms (the descriptors and synonyms from both European and Latin Spanish DeCs 2019 data sets, separated by pipes)</p> <p>&nbsp;</p> <p>For more details on the Latin and European Spanish DeCs codes see: http://decs.bvs.br and http://decses.bvsalud.org/ respectively.</p> <p>Please, cite: Krallinger M, Krithara A, Nentidis A, Paliouras G, Villegas M. BioASQ at CLEF2020: Large-Scale Biomedical Semantic Indexing and Question Answering. InEuropean Conference on Information Retrieval 2020 Apr 14 (pp. 550-556). Springer, Cham.</p> <p>&nbsp;</p> <p>Copyright (c) 2020 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0May 2020View details →
zenodo40/100

spanish_used_car_market: Coches de segunda mano a la venta en España

<p>En el siguiente <em>dataset</em> se presentan los datos sobre <strong>coches en venta en el territorio espa&ntilde;ol </strong>obtenidos fruto del desarrollo de la&nbsp;Pr&aacute;ctica 1 de la asignatura &#39;Tipolog&iacute;a y Ciclo de Vida de los Datos&#39; del M&aacute;ster en Ciencia de Datos de la Universitat Oberta de Catalunya, con fin meramente acad&eacute;mico.</p> <p>Los datos han sido recolectados utilizando t&eacute;cnicas de <em>web scraping&nbsp;</em>sobre el sitio web www.milanuncios.com.</p> <p>Los atributos del fichero .CSV se definen a continuaci&oacute;n:</p> <ul> <li><strong>ad_id</strong>: Identificador del anuncio del coche.</li> <li><strong>ad_type</strong>: Tipo de anuncio. En nuestro caso, siempre va a ser Oferta.</li> <li><strong>ad_time</strong>: Tiempo que llevaba publicado el anuncio cuando se recogi&oacute; la informaci&oacute;n, en formato X horas o X d&iacute;as. En el caso en que fuese un anuncio destacado, no tenemos esa informaci&oacute;n.</li> <li><strong>ad_title</strong>: T&iacute;tulo del anuncio de venta, con formato {Marca} &ndash; {Modelo}.</li> <li><strong>car_desc</strong>: Preview de la descripci&oacute;n del anuncio de venta del veh&iacute;culo.</li> <li><strong>car_km</strong>: Kil&oacute;metros que tiene recorridos el coche.</li> <li><strong>car_year</strong>: A&ntilde;o de matriculaci&oacute;n del veh&iacute;culo.</li> <li><strong>car_engine_type</strong>: Tipo de transmisi&oacute;n. Posibles valores: Manual | Autom&aacute;tico.</li> <li><strong>car_door_num</strong>: N&uacute;mero de puertas de las que dispone el coche.</li> <li><strong>car_power</strong>: Potencia del veh&iacute;culo, en formato XXX CV.</li> <li><strong>car_price</strong>: Precio en euros por el que se vende el coche.</li> <li><strong>advertizer_type</strong>: Indica cu&aacute;l es el tipo de vendedor del veh&iacute;culo. Valores posibles: Profesional | Particular.</li> <li><strong>image_url</strong>: Foto principal del anuncio de venta del coche.</li> <li><strong>ts</strong>: Hora en la que se recogi&oacute; la informaci&oacute;n, con formato YYYY-MM-DD hh:mm:ss.ms.</li> <li><strong>region</strong>: Provincia en la que se est&aacute; vendiendo el veh&iacute;culo.</li> </ul>

opencc-by-4.0Nov 2020View details →
zenodo40/100

PharmaCoNER corpus: gold standard annotations of Pharmacological Substances, Compounds and proteins in Spanish clinical case reports

<p><strong>Intro:</strong></p><p>The PharmaCoNER corpus (divided into train, dev and test) is a Gold Standard manually annotated dataset used for the the PharmaCoNER shared task posed at BIONLP-ST (at EMNLP). In addition, we include here the PharmaCoNER background set. It contains the train, development and test sets of the two subtasks (subtask-1 and subtask-2) with Gold Standard annotations. In addition, it contains the documents of the background set, without annotations.</p><p>The PharmaCoNER corpus consists of:</p><ul><li>Manually classified clinical case sections derived from Open access Spanish medical publications, named the Spanish Clinical Case Corpus (SPACCC).</li><li>It was manually selected by a practicing oncologist and revised by a clinical documentalist to assure that records were relevant/representative and resembled structure and content relevant to process clinical records.</li><li>The final corpus: 1000 clinical cases 16,504 sentences</li><li>The corpus contains a total of 396,988 words, with an average of 396.2 words per clinical case.</li><li>It covers a range of medical disciplines including oncology, urology, cardiology, pneumology or infections diseases, etc.</li><li>The corpus has been annotated at the mention level by experts in medicinal chemistry and pharmacology following a granular annotation scheme covering four mention types:<ul><li><i>Entity type 1 (NORMALIZABLES)</i>: mentions of chemicals that can be manually normalized to a unique concept identifier (primarily SNOMED-CT).</li><li><i>Entity type 2 (NO_NORMALIZABLES)</i>: mentions of chemicals that could not be normalized manually to a unique concept identifier.</li><li><i>Entity type 3 (PROTEINAS)</i>: mentions of proteins/genes following an adaptation of the BioCreative GPRO track annotation guidelines (includes peptides, peptide hormones &amp; antibodies).</li><li><i>Entity type 4 (UNCLEAR )</i>: cases of general substance class mentions of clinical relevance, including certain pharmaceutical formulations, general treatments, chemotherapy programs, and vaccines.</li><li>Mentions class "<i>UNCLEAR</i>" (not evaluated for the PharmaCoNER track)</li></ul></li></ul><p>&nbsp;</p><p>&nbsp;</p><p><strong>Please, cite:&nbsp;</strong></p><p>A. G. Agirre, M. Marimon, A. Intxaurrondo, O. Rabal, M. Villegas, M. Krallinger, Pharmaconer: Pharmacological substances, compounds and proteins named entity recognition track, in: Proceedings of The 5th Workshop on BioNLP Open Shared Tasks, 2019, pp. 1–10.</p><p>&nbsp;</p><p><strong>Annotation quality</strong></p><p>Inter-annotator agreement: 93% for annotation, 73% for mapping.</p><p>For more information, see the <a href="https://paperswithcode.com/paper/pharmaconer-pharmacological-substances">paper</a>.</p><p>&nbsp;</p><p><strong>Format</strong></p><p>For subtask 1 annotations are distributed in <i>Brat</i> format. (More info at Brat webpage&nbsp;https://brat.nlplab.org/standoff.html)</p><p>For subtask-2, codes are associated with each document are given in a <i>TSV</i> file with the following columns:&nbsp;</p><blockquote><p>filename&nbsp;&nbsp; &nbsp;code</p></blockquote><p>&nbsp;</p><p><strong>Shared task goal:</strong></p><p>In the two subtasks, the goal is to predict the annotations of the test files (either the ANN files or the TSV with the codes) given only the plain text files.&nbsp;</p><p>&nbsp;</p><p><strong>Resources:</strong></p><ul><li><a href="https://temu.bsc.es/pharmaconer/"><strong>Web</strong></a></li><li><a href="https://www.aclweb.org/anthology/D19-5701.pdf"><strong>Citation</strong></a><strong>:&nbsp;</strong>A. G. Agirre, M. Marimon, A. Intxaurrondo, O. Rabal, M. Villegas, M. Krallinger, Pharmaconer: Pharmacological substances, compounds and proteins named entity recognition track, in: Proceedings of The 5th Workshop on BioNLP Open Shared Tasks, 2019, pp. 1–10.</li><li><a href="https://doi.org/10.5281/zenodo.4271908"><strong>Silver Standard corpus</strong></a></li><li><a href="https://doi.org/10.5281/zenodo.3763276"><strong>Annotation guidelines</strong></a></li><li><a href="https://github.com/TeMU-BSC/PharmaCoNER-Tagger"><strong>PharmaCoNER tagger</strong></a></li><li><a href="https://www.youtube.com/watch?v=B3ZzJl5OMkY"><strong>Youtube video(general setting)</strong></a></li><li><a href="https://www.slideshare.net/MartinKrallinger/pharmaconer-pharmacological-substances-compounds-and-proteins-named-entity-recognition-track-at-bionlpost-workshop-november-4-skycity-rm-2-hong-kong-emnlp2019"><strong>Slides PharmacoNER overview talk at BIONLP-ST / EMNLP&nbsp;</strong></a></li></ul><p>For further information, please visit <a href="https://temu.bsc.es/pharmaconer/">https://temu.bsc.es/pharmaconer/</a> or email us at encargo-pln-life@bsc.es</p><p>Copyright (c) 2018 Secretaría de Estado para el Avance Digital (SEAD)</p><p>&nbsp;</p><p><strong>License</strong></p><p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p><p>&nbsp;</p><p><strong>Contact</strong></p><p>If you have any questions or suggestions, please contact us at:</p><p><br>- Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p><p><strong>Additional resources and corpora</strong></p><p>If you are interested in PharmaCoNER, you might want to check out these corpora and resources:</p><ul><li><a href="https://zenodo.org/records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/8413866">SympTEMIST</a> (Corpus of symptoms, signs and findings mentions and normalization, same document collection)</li><li><a href="https://zenodo.org/records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI), modified synthetic verions of the document collection)</li><li><a href="https://zenodo.org/records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization, different document collection)</li><li><a href="https://zenodo.org/records/3837305">CodiESp</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version, same document collection)</li><li><a href="https://zenodo.org/records/7684093">LivingNER</a> (Corpus of mentions of species, including human/family members, pathogens, food, etc.. and normalization to NCBI Taxonomy, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags, same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries), same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags, same document collection)</li><li><a href="https://zenodo.org/records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts, different document collection)</li></ul>

opencc-by-4.0Nov 2020View details →
zenodo40/100

MEDDOCAN corpus: gold standard annotations for Medical Document Anonymization on Spanish clinical case reports

<p><strong>Intro:</strong></p> <p>Meddocan shared task dataset (divided in train, dev and test). In addition, we include here the Meddocan background set.</p> <p>It contains the training, development and test sets of the Meddocan shared task with Gold Standard annotations.</p> <p>In addition, it contains the documents of the background set, without annotations.</p> <p>&nbsp;</p> <p><strong>Annotation quality</strong></p> <p>Inter-annotator agreement: 98%&nbsp;</p> <p>For more information, see the <a href="http://ceur-ws.org/Vol-2421/MEDDOCAN_overview.pdf">paper</a>.&nbsp;</p> <p>&nbsp;</p> <p><strong>Format:</strong></p> <p>Annotations are distributed in Brat format. See&nbsp;<a href="https://brat.nlplab.org/standoff.html">Brat webpage</a>&nbsp;for more information.</p> <p>In addition, annotations are also distributed in XML format (based on i2b2 XML format).</p> <p>In the <a href="https://temu.bsc.es/meddocan/index.php/resources/">Meddocan webpage</a>, there is a script to convert&nbsp;between MEDDOCAN-Brat, MEDDOCAN-XML, and i2b2 formats.</p> <p>&nbsp;</p> <p><strong>Shared task goal:</strong></p> <p>In the three subtasks, the goal will be to predict the annotations&nbsp;given only the plain text files.&nbsp;</p> <p>&nbsp;</p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/meddocan/">Web</a></strong></li> <li><strong>Citation:&nbsp;</strong>Montserrat Marimon et al. &ldquo;Automatic De-identification of Medical Texts in Spanish: the MEDDOCAN Track, Corpus, Guidelines, Methods and Evaluation of Results.&rdquo; In: IberLEF@ SEPLN. 2019, pp. 618&ndash;638.</li> <li><strong>Silver Standard corpus</strong></li> <li><a href="https://doi.org/10.5281/zenodo.4279337"><strong>Annotation guidelines</strong></a></li> </ul> <p>&nbsp;</p> <p>For further information, please visit&nbsp;<a href="https://temu.bsc.es/meddocan/">https://temu.bsc.es/meddocan/</a>&nbsp;or email us at encargo-pln-life@bsc.es</p> <p>Copyright (c) 2019 Secretar&iacute;a de Estado para el Avance Digital (SEAD)</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Collection of 19th century Spanish-American Novels

<p>This is a text collection prepared for use with the TXM text analysis tool (http://textometrie.ens-lyon.fr/).&nbsp;The collection contains a selection of novels from 1880-1916. There are currently 24 novels with a total of about 1.2 million words. All texts have been tokenised, lemmatised and POS-tagged using TreeTagger.&nbsp;</p>

opencc-zeroMar 2016View details →
zenodo40/100

Datasets for paper: J. A. Cerón-Guzmán and E. León-Guzmán (2016), A Sentiment Analysis System of Spanish Tweets and Its Application in Colombia 2014 Presidential Election. SocialCom 2016. DOI: 10.1109/BDCloud-SocialCom-SustainCom.2016.47

<p>Datasets for paper: J. A. Cerón-Guzmán and E. León-Guzmán (2016), A Sentiment Analysis System of Spanish Tweets and Its Application in Colombia 2014 Presidential Election. The 9th IEEE International Conference on Social Computing and Networking (SocialCom). DOI: 10.1109/BDCloud-SocialCom-SustainCom.2016.47</p>

opencc-by-4.0Jul 2016View details →
zenodo40/100

FIGURE 3 Halosaurus ovenii from the northeastern Atlantic Ocean, MHNUSC 25010 - 2, 578 in Halosaur fishes (Notacanthiformes: Halosauridae) from Atlantic Spanish waters according to integrative taxonomy

FIGURE 3 Halosaurus ovenii from the northeastern Atlantic Ocean, MHNUSC 25010 - 2, 578 mm total length.

opencc-zeroDec 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record