Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
824
datasets available to search
ShareScore release 0.9.0
Dataset results
824 results for “spanish”
TRANSIT long-distance multimodal trips model results - a Spanish case study
<p>The files downloaded present the results per scenario considered obtained with the agent-based model developed for the assessment of the Intermodal Timetable Synchronisation solution proposed in the scope of the TRANSIT project (<a href="https://www.transit-h2020.eu/">https://www.transit-h2020.eu/</a>).</p> <p>An agent-based modelling framework called <a href="https://github.com/StefanoPenazzi/jtap/tree/main">J-TAP</a> has been developed and put at work to implement a Spanish long-distance multimodal trips model. The enhanced version of J-TAP used in this work can be found on a <a href="https://github.com/NommonSolutionsAndTechnologies/jtap">github repository</a>.</p> <p>The case study is focused on modelling the long-distance travel patterns of the residents in the Valencia (Spain) area. The destinations considered include the whole of Spain. The period under study is a full year from March 2019 to February 2020 (both inclusive).The files are structured in three different scenarios:</p> <ol> <li><strong>CS01 - Baseline</strong>. The current state of the network is considered and the actual long-distance travel patterns are obtained.</li> <li><strong>CS02 - HSR connection with Madrid-Barajas airport</strong>. The long-distance travel patterns are modelled with hard measures, the high-speed rail is connected to Madrid-Barajas airport.</li> <li><strong>CS03 - HSR connection with Madrid-Barajas airport and timetable synchronisation</strong>. The effects of the timetable synchronisation are modelled.</li> </ol> <p>Each scenario includes the following files:</p> <ul> <li>ctapModelParameters. A folder containing all the information extracted from the <a href="https://neo4j.com/product/graph-data-science/?utm_program=emea-prospecting&utm_source=google&utm_medium=cpc&utm_campaign=emea-search-offers&utm_adgroup=dynamic&utm_content=dynamic&utm_placement=&utm_network=g&gclid=Cj0KCQiAwJWdBhCYARIsAJc4idAo4CEi9lU8TXwmBym8MNHpEIZHPBs3x_4phxbu76y1XKbYlFoZCjIaAiGhEALw_wcB">neo4j</a> graph database created to model the multimodal network and the agents. J-TAP contains packages that simplify network creation in the graph database. This information is stored in .json files (e.g., "Os2DsTravelCostParameter.json" contains the generalised cost for each OD pair and transport mode, "AttractivenessParameter.json" contains the attractiveness by destination, activity, time of the year and agent, etc.). The solver included in the J-TAP framework uses this information to calculate the agents plans.</li> <li>population.json. The result of the J-TAP optimisation. It includes the fitness value for each agent plan evaluated during the J-TAP execution. The best plan for each agent is selected as the plan performed by the agent. An agent plan includes: <ul> <li>activities - Sequence of activities.</li> <li>locations - Sequence of locations</li> <li>ts - Initial time of the activity</li> <li>te - Final time of the activity</li> </ul> </li> <li>LinkTimeFlow.csv. It is obtained after processing the previous file. It contains the number of agents using each link in the network (i.e., road, rail, air and cross links) in each time interval. The first column represents the link id and the rest of columns indicates the number of agents in each interval.</li> </ul> <p>The J-TAP simulation framework is explained in detail in TRANSIT's deliverable <a href="http://www.nommon-files.es/transit/TRANSIT-D5.1_Modelling_Framework_v02.00.00.pdf">D5.1. TRANSIT Modelling and Simulation Framework</a> and the complete description of the case studies and scenarios tested is included in TRANSIT's deliverable <a href="http://www.nommon-files.es/transit/TRANSIT-D6.1_Assessment_of_Intermodal_Concepts_00.02.00.pdf">D6.1. Impact Assessment of New Intermodal Concepts and Passenger Information Services: Conclusions and Recommendations</a>.</p> <p>Thank you for downloading the dataset! It would be very helpful if you share your view on the data show with us. </p>
Spanish semantic fields
<p>A database with 73.000 frequent spanish words classified in semantic fields in a three level hierarchy</p>
Nos_Machine translation test suite Spanish-Galician
<p>334 spanish-galician parallel sentences for the evaluation of machine translation. Sentences are classified in categories based on challenging linguistic phenomena.</p>
Nos_Machine translation gold standard Spanish-Galician 2
<p>1998 carefully curated Galician-Spanish parallel sentences for the evaluation of automatic machine translation. Sentence pairs are classified according to whether the translations are highly literal or not.</p>
Nos_Machine translation gold standard Spanish-Galician 1
<p>1998 carefully curated Galician-Spanish parallel sentences for the evaluation of automatic machine translation. Sentence pairs are classified according to whether the translations are highly literal or not.</p>
Spanish Workers' Statute Legal Relations and RDF
<p>These datasets result from extracting legal events and relationships from the Spanish Workers' Statutes and structuring the extracted data in an RDF graph. After a set of experiments conducted with GPT-3.5 using scarce annotated data from previous works <a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6432">[1]</a>, a 5-shot learning approach was applied to the full text of the Spanish Workers’ Statute. Approximately 1500 relations were extracted in a JSON format, structured into a dataset, and represented in an RDF graph.</p>
OneNet project - T9.3 - Spanish demo open data
<p>-Anonymized technical data from flexible resources participating in OneNet Spanish demonstration </p> <p>-Market results assessed by the local market platforms considering market bids and DSOs requirement for the OneNet Spanish demonstration</p>
Hydrodynamic and morphological information, and the absolute variations of the vulnerability indices for the period 2000-2015 of the Spanish Iberia Peninsula estuaries.
<p>The dataset included in this repository was obtained during the project entitled 'Sensibilidad física y biotic de los estuarios peninsulares al cambio global (SENSES)' funded by 'Fundación Biodiversidad', PRCV00487. The data were used in the research article 'Sensitivity of Iberian estuaries to changes in sea water temperature, salinity, river-flow, mean sea level, and tidal amplitudes' submitted to <em>Estuarine, Coastal and Shelf Science</em>.</p> <p>Brief description of dataset:</p> <p>For each estuary, the following parameters were calculated</p> <ul> <li>Fachade: the location of the estuary</li> <li>Area (km<sup>2</sup>) </li> <li>D (m): water depth at the mouth of the estuary in 2000 and 2015</li> <li>Tidal Prim (m<sup>3</sup>)</li> <li>Q<sub><em>f</em></sub> (m<sup>3</sup>/s): river flow in 2000 and 2015</li> <li><em>a </em>(m): tidal amplitude of the free surface elevation in 2000 and 2015</li> <li>∆<em>U</em> (m/s): absolute variation of the tidal current amplitude between 2000 and 2015</li> <li>∆<em>E</em> (W/m<sup>2</sup>): absolute variation of the tidal energy flux propagation index between 2000 and 2015</li> <li>∆<em>Ri </em>: absolute variation of the bulk Richardson number index between 2000 and 2015</li> <li>∆<em>SI</em>: absolute variation of the salinity intrusion index between 2000 and 2015</li> </ul> <p>A wide description of the parameters can be found in Serrano, M. A. et al (submitted to <em>Estuarine, Coastal and Shelf Science</em>)</p> <p>Contact person: mserranog@ugr.es</p>
FastText Spanish Medical Embeddings
<p>[Plan TL/medicine/word embeddings] Word embeddings generated from Spanish corpora that include: (a) the full-text in Spanish available in SciELO.org (until December/2018), (b) all articles from the following Wikipedia categories: Pharmacology, Pharmacy, Medicine and Biology (during December/2018) and (c) the concatenation of the previous two corpora.</p> <p>We used fastText to train the word embeddings.</p> <p>For more information, we refer to the corresponding article: <a href="https://www.aclweb.org/anthology/W19-1916/">https://www.aclweb.org/anthology/W19-1916/</a></p> <p>Copyright (c) 2021 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
Figures 16-19 in Roncus elbulli (Arachnida, Pseudoscorpiones), a new species from Cap de Creus Nature Park (Catalonia, Spain), with a key to the Spanish species of the genus Roncus
Figures 16-19. Roncus elbulli sp. n., female paratype, Cala Canadell. SEM photographs: 16. left palp, dorsal view; 17. chelal microsetae pattern below trichobothria eb/esb; 18. fingers of the chela, antiaxial face, partial view, showing trichobothrium sb and sensilla p1 and p2 on movable finger. Roncus cadinensis Zaragoza, 2007, male paratype. SEM photograph: 19. chelal microsetae pattern below trichobothria eb/ esb. Scale bars (mm): 0.05 (Figs 17, 18), 0.10 (Fig. 19), 0.50 (Fig. 16).
Figures 3-9 in Roncus elbulli (Arachnida, Pseudoscorpiones), a new species from Cap de Creus Nature Park (Catalonia, Spain), with a key to the Spanish species of the genus Roncus
Figures 3-9. Roncus elbulli sp. n., male holotype. 3. carapace; 4. anterior margin of carapace, showing epistome; 5. anterior and medial processes of coxa I; 6. left chelicera; 7. fingers of left chelicera, partial view; 8. right leg IV, lateral view; 9. distal end of tarsus and apotele of left leg IV, lateral view. Scale bars (mm): 0.05 (Figs 4, 5, 7, 9), 0.10 (Fig. 6), 0.20 (Figs 3, 8).
Figures 10-15 in Roncus elbulli (Arachnida, Pseudoscorpiones), a new species from Cap de Creus Nature Park (Catalonia, Spain), with a key to the Spanish species of the genus Roncus
Figures 10-15. Roncus elbulli sp. n., male holotype (except where otherwise noted): 10. left palp, without chela, dorsal view; 11. left chela, dorsal view; 12. left chela, lateral view; 13. male paratype, chelal microsetae pattern below trichobothria eb/esb. 14. chelal microsetae pattern below trichobothria eb/esb. Scale bars (mm): 0.05 (Figs 13-14), 0.20 (Figs 10-12). 15. Roncus cadinensis Zaragoza, 2007, male holotype: chelal microsetae pattern below trichobothria eb/esb. Scale bar (mm): 0.05.
Spanish abstracts from PubMed (machine-translated from English)
<p>Spanish abstracts from PubMed (machine-translated from English)</p> <p>A state-of-the-art, domain-specific Neural Machine Translation system has been used to automatically translate a large number of articles (titles and abstracts) from PubMed, from English to Spanish. This dataset, in JSON format, provides MeSH terms and DeCS codes (if available), as well as other fields such as the year of publication for all articles.</p> <p>Copyright (c) 2020 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
MESINESP: Medical Semantic Indexing in Spanish - Train dataset
<p><em><strong>Please use the <a href="https://doi.org/10.5281/zenodo.4612274">MESINESP2 corpus (the second edition of the shared-task)</a> since it has a higher level of curation, quality and is organized by document type (scientific articles, patents and clinical trials).</strong></em></p> <p> </p> <p> </p> <p><strong>INTRODUCTION</strong>:</p> <p>The Mesinesp (Spanish BioASQ track, see https://temu.bsc.es/mesinesp) training set has a total of 369,368 records. </p> <p>The training dataset contains all records from LILACS and IBECS databases at the Virtual Health Library (VHL) with a non-empty abstract written in Spanish. The URL used to retrieve records is as follows:<br> http://pesquisa.bvsalud.org/portal/?output=xml&lang=es&sort=YEAR_DESC&format=abstract&filter[db][]=LILACS&filter[db][]=IBECS&q=&index=tw&</p> <p>We have filtered out empty abstracts and non-Spanish abstracts. </p> <p>The training dataset was crawled on 10/22/2019. This means that the data is a snapshot of that moment and that may change over time. In fact, it is very likely that the data will undergo minor changes as the different databases that make up LILACS and IBECS may add or modify the indexes.</p> <p> </p> <p><strong>ZIP STRUCTURE:</strong></p> <p>The training data sets contain 369,368 records from 26,609 different journals. Two different data sets are distributed as described below:</p> <p> - <em>Original Train set</em> with 369,368 records that also include the qualifiers, as retrieved from VHL. <br> - <em>Pre-processed Train set</em><strong> </strong>with the 318,658 records with at least one DeCS code and with no qualifiers. </p> <p> </p> <p> </p> <p><strong>STATISTICS</strong>:</p> <p>Abstracts’ length (measured in characters)<br> Min: 12<br> Avg: 1140.41<br> Median: 1094<br> Max: 9428</p> <p>Number of DeCS codes per file<br> Min: 1<br> Avg: 8.12<br> Median: 7<br> Max: 53</p> <p> </p> <p> </p> <p><strong>CORPUS FORMAT</strong>:</p> <p>The training data sets are distributed as a JSON file with the following format:</p> <pre><code>{ "articles": [ { "id": "Id of the article", "title": "Title of the article", "abstractText": "Content of the abstract", "journal": "Name of the journal", "year": 2018, "db": "Name of the database", "decsCodes": [ "code1", "code2", "code3" ] } ] } </code></pre> <p>Note that the decsCodes field lists the DeCs Ids assigned to a record in the source data. Since the original XML data contain descriptors (no codes), we provide a DeCs conversion table (https://temu.bsc.es/mesinesp/wp-content/uploads/2019/12/DeCS.2019.v5.tsv.zip) with:</p> <p> - DeCs codes<br> - Preferred descriptor (the label used in the European DeCs 2019 set)<br> - List of synonyms (the descriptors and synonyms from both European and Latin Spanish DeCs 2019 data sets, separated by pipes)</p> <p> </p> <p>For more details on the Latin and European Spanish DeCs codes see: http://decs.bvs.br and http://decses.bvsalud.org/ respectively.</p> <p>Please, cite: Krallinger M, Krithara A, Nentidis A, Paliouras G, Villegas M. BioASQ at CLEF2020: Large-Scale Biomedical Semantic Indexing and Question Answering. InEuropean Conference on Information Retrieval 2020 Apr 14 (pp. 550-556). Springer, Cham.</p> <p> </p> <p>Copyright (c) 2020 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
spanish_used_car_market: Coches de segunda mano a la venta en España
<p>En el siguiente <em>dataset</em> se presentan los datos sobre <strong>coches en venta en el territorio español </strong>obtenidos fruto del desarrollo de la Práctica 1 de la asignatura 'Tipología y Ciclo de Vida de los Datos' del Máster en Ciencia de Datos de la Universitat Oberta de Catalunya, con fin meramente académico.</p> <p>Los datos han sido recolectados utilizando técnicas de <em>web scraping </em>sobre el sitio web www.milanuncios.com.</p> <p>Los atributos del fichero .CSV se definen a continuación:</p> <ul> <li><strong>ad_id</strong>: Identificador del anuncio del coche.</li> <li><strong>ad_type</strong>: Tipo de anuncio. En nuestro caso, siempre va a ser Oferta.</li> <li><strong>ad_time</strong>: Tiempo que llevaba publicado el anuncio cuando se recogió la información, en formato X horas o X días. En el caso en que fuese un anuncio destacado, no tenemos esa información.</li> <li><strong>ad_title</strong>: Título del anuncio de venta, con formato {Marca} – {Modelo}.</li> <li><strong>car_desc</strong>: Preview de la descripción del anuncio de venta del vehículo.</li> <li><strong>car_km</strong>: Kilómetros que tiene recorridos el coche.</li> <li><strong>car_year</strong>: Año de matriculación del vehículo.</li> <li><strong>car_engine_type</strong>: Tipo de transmisión. Posibles valores: Manual | Automático.</li> <li><strong>car_door_num</strong>: Número de puertas de las que dispone el coche.</li> <li><strong>car_power</strong>: Potencia del vehículo, en formato XXX CV.</li> <li><strong>car_price</strong>: Precio en euros por el que se vende el coche.</li> <li><strong>advertizer_type</strong>: Indica cuál es el tipo de vendedor del vehículo. Valores posibles: Profesional | Particular.</li> <li><strong>image_url</strong>: Foto principal del anuncio de venta del coche.</li> <li><strong>ts</strong>: Hora en la que se recogió la información, con formato YYYY-MM-DD hh:mm:ss.ms.</li> <li><strong>region</strong>: Provincia en la que se está vendiendo el vehículo.</li> </ul>
PharmaCoNER corpus: gold standard annotations of Pharmacological Substances, Compounds and proteins in Spanish clinical case reports
<p><strong>Intro:</strong></p><p>The PharmaCoNER corpus (divided into train, dev and test) is a Gold Standard manually annotated dataset used for the the PharmaCoNER shared task posed at BIONLP-ST (at EMNLP). In addition, we include here the PharmaCoNER background set. It contains the train, development and test sets of the two subtasks (subtask-1 and subtask-2) with Gold Standard annotations. In addition, it contains the documents of the background set, without annotations.</p><p>The PharmaCoNER corpus consists of:</p><ul><li>Manually classified clinical case sections derived from Open access Spanish medical publications, named the Spanish Clinical Case Corpus (SPACCC).</li><li>It was manually selected by a practicing oncologist and revised by a clinical documentalist to assure that records were relevant/representative and resembled structure and content relevant to process clinical records.</li><li>The final corpus: 1000 clinical cases 16,504 sentences</li><li>The corpus contains a total of 396,988 words, with an average of 396.2 words per clinical case.</li><li>It covers a range of medical disciplines including oncology, urology, cardiology, pneumology or infections diseases, etc.</li><li>The corpus has been annotated at the mention level by experts in medicinal chemistry and pharmacology following a granular annotation scheme covering four mention types:<ul><li><i>Entity type 1 (NORMALIZABLES)</i>: mentions of chemicals that can be manually normalized to a unique concept identifier (primarily SNOMED-CT).</li><li><i>Entity type 2 (NO_NORMALIZABLES)</i>: mentions of chemicals that could not be normalized manually to a unique concept identifier.</li><li><i>Entity type 3 (PROTEINAS)</i>: mentions of proteins/genes following an adaptation of the BioCreative GPRO track annotation guidelines (includes peptides, peptide hormones & antibodies).</li><li><i>Entity type 4 (UNCLEAR )</i>: cases of general substance class mentions of clinical relevance, including certain pharmaceutical formulations, general treatments, chemotherapy programs, and vaccines.</li><li>Mentions class "<i>UNCLEAR</i>" (not evaluated for the PharmaCoNER track)</li></ul></li></ul><p> </p><p> </p><p><strong>Please, cite: </strong></p><p>A. G. Agirre, M. Marimon, A. Intxaurrondo, O. Rabal, M. Villegas, M. Krallinger, Pharmaconer: Pharmacological substances, compounds and proteins named entity recognition track, in: Proceedings of The 5th Workshop on BioNLP Open Shared Tasks, 2019, pp. 1–10.</p><p> </p><p><strong>Annotation quality</strong></p><p>Inter-annotator agreement: 93% for annotation, 73% for mapping.</p><p>For more information, see the <a href="https://paperswithcode.com/paper/pharmaconer-pharmacological-substances">paper</a>.</p><p> </p><p><strong>Format</strong></p><p>For subtask 1 annotations are distributed in <i>Brat</i> format. (More info at Brat webpage https://brat.nlplab.org/standoff.html)</p><p>For subtask-2, codes are associated with each document are given in a <i>TSV</i> file with the following columns: </p><blockquote><p>filename code</p></blockquote><p> </p><p><strong>Shared task goal:</strong></p><p>In the two subtasks, the goal is to predict the annotations of the test files (either the ANN files or the TSV with the codes) given only the plain text files. </p><p> </p><p><strong>Resources:</strong></p><ul><li><a href="https://temu.bsc.es/pharmaconer/"><strong>Web</strong></a></li><li><a href="https://www.aclweb.org/anthology/D19-5701.pdf"><strong>Citation</strong></a><strong>: </strong>A. G. Agirre, M. Marimon, A. Intxaurrondo, O. Rabal, M. Villegas, M. Krallinger, Pharmaconer: Pharmacological substances, compounds and proteins named entity recognition track, in: Proceedings of The 5th Workshop on BioNLP Open Shared Tasks, 2019, pp. 1–10.</li><li><a href="https://doi.org/10.5281/zenodo.4271908"><strong>Silver Standard corpus</strong></a></li><li><a href="https://doi.org/10.5281/zenodo.3763276"><strong>Annotation guidelines</strong></a></li><li><a href="https://github.com/TeMU-BSC/PharmaCoNER-Tagger"><strong>PharmaCoNER tagger</strong></a></li><li><a href="https://www.youtube.com/watch?v=B3ZzJl5OMkY"><strong>Youtube video(general setting)</strong></a></li><li><a href="https://www.slideshare.net/MartinKrallinger/pharmaconer-pharmacological-substances-compounds-and-proteins-named-entity-recognition-track-at-bionlpost-workshop-november-4-skycity-rm-2-hong-kong-emnlp2019"><strong>Slides PharmacoNER overview talk at BIONLP-ST / EMNLP </strong></a></li></ul><p>For further information, please visit <a href="https://temu.bsc.es/pharmaconer/">https://temu.bsc.es/pharmaconer/</a> or email us at encargo-pln-life@bsc.es</p><p>Copyright (c) 2018 Secretaría de Estado para el Avance Digital (SEAD)</p><p> </p><p><strong>License</strong></p><p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p><p> </p><p><strong>Contact</strong></p><p>If you have any questions or suggestions, please contact us at:</p><p><br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p><p><strong>Additional resources and corpora</strong></p><p>If you are interested in PharmaCoNER, you might want to check out these corpora and resources:</p><ul><li><a href="https://zenodo.org/records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/8413866">SympTEMIST</a> (Corpus of symptoms, signs and findings mentions and normalization, same document collection)</li><li><a href="https://zenodo.org/records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI), modified synthetic verions of the document collection)</li><li><a href="https://zenodo.org/records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization, different document collection)</li><li><a href="https://zenodo.org/records/3837305">CodiESp</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version, same document collection)</li><li><a href="https://zenodo.org/records/7684093">LivingNER</a> (Corpus of mentions of species, including human/family members, pathogens, food, etc.. and normalization to NCBI Taxonomy, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags, same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries), same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags, same document collection)</li><li><a href="https://zenodo.org/records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts, different document collection)</li></ul>
MEDDOCAN corpus: gold standard annotations for Medical Document Anonymization on Spanish clinical case reports
<p><strong>Intro:</strong></p> <p>Meddocan shared task dataset (divided in train, dev and test). In addition, we include here the Meddocan background set.</p> <p>It contains the training, development and test sets of the Meddocan shared task with Gold Standard annotations.</p> <p>In addition, it contains the documents of the background set, without annotations.</p> <p> </p> <p><strong>Annotation quality</strong></p> <p>Inter-annotator agreement: 98% </p> <p>For more information, see the <a href="http://ceur-ws.org/Vol-2421/MEDDOCAN_overview.pdf">paper</a>. </p> <p> </p> <p><strong>Format:</strong></p> <p>Annotations are distributed in Brat format. See <a href="https://brat.nlplab.org/standoff.html">Brat webpage</a> for more information.</p> <p>In addition, annotations are also distributed in XML format (based on i2b2 XML format).</p> <p>In the <a href="https://temu.bsc.es/meddocan/index.php/resources/">Meddocan webpage</a>, there is a script to convert between MEDDOCAN-Brat, MEDDOCAN-XML, and i2b2 formats.</p> <p> </p> <p><strong>Shared task goal:</strong></p> <p>In the three subtasks, the goal will be to predict the annotations given only the plain text files. </p> <p> </p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/meddocan/">Web</a></strong></li> <li><strong>Citation: </strong>Montserrat Marimon et al. “Automatic De-identification of Medical Texts in Spanish: the MEDDOCAN Track, Corpus, Guidelines, Methods and Evaluation of Results.” In: IberLEF@ SEPLN. 2019, pp. 618–638.</li> <li><strong>Silver Standard corpus</strong></li> <li><a href="https://doi.org/10.5281/zenodo.4279337"><strong>Annotation guidelines</strong></a></li> </ul> <p> </p> <p>For further information, please visit <a href="https://temu.bsc.es/meddocan/">https://temu.bsc.es/meddocan/</a> or email us at encargo-pln-life@bsc.es</p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital (SEAD)</p>
Collection of 19th century Spanish-American Novels
<p>This is a text collection prepared for use with the TXM text analysis tool (http://textometrie.ens-lyon.fr/). The collection contains a selection of novels from 1880-1916. There are currently 24 novels with a total of about 1.2 million words. All texts have been tokenised, lemmatised and POS-tagged using TreeTagger. </p>
Datasets for paper: J. A. Cerón-Guzmán and E. León-Guzmán (2016), A Sentiment Analysis System of Spanish Tweets and Its Application in Colombia 2014 Presidential Election. SocialCom 2016. DOI: 10.1109/BDCloud-SocialCom-SustainCom.2016.47
<p>Datasets for paper: J. A. Cerón-Guzmán and E. León-Guzmán (2016), A Sentiment Analysis System of Spanish Tweets and Its Application in Colombia 2014 Presidential Election. The 9th IEEE International Conference on Social Computing and Networking (SocialCom). DOI: 10.1109/BDCloud-SocialCom-SustainCom.2016.47</p>
FIGURE 3 Halosaurus ovenii from the northeastern Atlantic Ocean, MHNUSC 25010 - 2, 578 in Halosaur fishes (Notacanthiformes: Halosauridae) from Atlantic Spanish waters according to integrative taxonomy
FIGURE 3 Halosaurus ovenii from the northeastern Atlantic Ocean, MHNUSC 25010 - 2, 578 mm total length.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.