Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
822
datasets available to search
ShareScore release 0.7.1
Dataset results
822 results for “Spanish”
Database Mobbing-UNIPSICO Scale in Spanish Teachers
<p>This dataset contains data of non-university teachers collected by paper and pencil at the workplace between October 2015 and May 2020. These data were collected by employees working in the INVASSAT (Instituto Valenciano de Seguridad y Salud en el Trabajo, Government of the Valencian Community, Spain). The INVASSAT employees went to all educational center and informed the director, union representative, and teachers at each school of the procedure. Then each teacher filled in the questionnaire individually. The questionnaire was done in the presence of the INVASSAT employees to answer any doubts, and the filled questionnaires were given to the INVASSAT employee.</p><p>The file contains demographic variables, the responses to the 20 items of the Mobbing-UNIPSICO scale questionnaire, the responses to the items on alcohol, tobacco and medication use, and the response to the item regarding the necessity of professional support.</p><p>The name of the variables and the value labels have been written in English to facilitate their understanding.</p><p>Data and codebooks are provided in csv format, following the FAIR principles.</p><p>Three files are provided:</p><p>1. Mobbing database, with the data related to sample characteristics and the answers to the items of the questionnaires and the other items.</p><p>2. Database codebook of variables, with information of the labels of the variables of the Database file.</p><p>3. Variable values codebook, with the labels of the values of the variables in the Database file.</p><p> </p>
SCLabels: Labelled rectified RGB images from the Spanish CoastSnap network
<h1>Training dataset</h1> <p><span>The SCLabels dataset is intended to be used in the exploring and development of Artificial Intelligence (AI) applications aimed at the automation of the shoreline extraction process from rectified images. SCLabels includes rectified RGB images from the Spanish CoastSnap network and their corresponding masks, together with a metadata file and a README file. RGB images encompass variable geographic locations, fields of view, beach types and degrees of occupation, tidal regimes, meteoceanic and lightning conditions, and a variety of environmental characteristics. Masks account for dense pixel labels including 5 categories: i) No data; ii) Not classified; iii) Landwards; iv) Seawards; and v) Shoreline. In the metadata file, images are linked to their corresponding masks, and information about the geographic location of each image, capture characteristics and image source, shoreline position and other auxiliary data are provided. The README file enhances the explainability and comprehension of the dataset, elaborating on the context and contents, and providing detailed explanations of the metadata, potential limitations, technical aspects of the image processing and annotation stages, usage recommendations, and related works. </span></p> <h1>Technical details</h1> <p>The SCLabels dataset version 1.0.0 is packaged in a compressed file (SCLabels_v1.0.0.zip). A total of 1717 RGB images are shared in JPG format, corresponding masks in PNG format, a metadata file in JSON format, and the README file in PDF format.</p> <h2>Data preprocessing</h2> <p><span>To generate the SCLabels masks, rectified RGB images and their corresponding shorelines were used. RGB images were cropped to the minimum and maximum alongshore pixel coordinates of the shoreline (vertical axis) plus 10 additional pixels above and below to preserve contextual information. A grayscale image was then derived from each cropped RGB image for subsequent pixel labelling. First, a binary mask was derived, marking "NoData'' for black and white padded pixels resulting from the registration and rectification steps. Subsequently, the shoreline was densified, ensuring at least one pixel per row was assigned the "Shoreline" label. Next, "Landwards" and "Seawards" labels were assigned to the right and left of the shoreline. Pixels left unlabelled were categorised as "NotClassified". Finally, masks’ values were reclassified to align with the predefined labels, and the grayscale masks were exported. For additional information, please consult the README file. </span></p> <h2>Data splitting</h2> <p><span>Data splitting requirements may vary depending on the chosen AI approach (e.g., splitting by entire images, image patches, or image rows). Researchers should use a consistent data splitting method and document the approach and splits used in publications. This transparency enables reproducible results and facilitates comparisons between studies.</span></p> <h2>Classes, labels and annotations</h2> <p><span>The SCLabels dataset includes one mask per rectified RGB image, sharing the same width and height. These masks are in greyscale and PNG format, and consist of five different labels:</span></p> <table> <tbody> <tr> <td><strong> Mask value</strong></td> <td><strong> Label</strong></td> <td><strong> Description</strong></td> </tr> <tr> <td>0</td> <td>NoData</td> <td>High probability of being black or white padded pixels, used to pad non-rectangular images within the image registration and rectification processes</td> </tr> <tr> <td>25</td> <td>NotClassified</td> <td>Not labeled pixels</td> </tr> <tr> <td>75</td> <td>Landwards</td> <td>All pixels that are towards the landside with respect to the shoreline (row-wise), excluding “NoData” ones</td> </tr> <tr> <td>150</td> <td>Seawards</td> <td>All pixels that are towards the seaside with respect to the shoreline (row-wise), excluding “NoData” ones</td> </tr> <tr> <td>255</td> <td>Shoreline</td> <td>Pixels intersected by the mapped shoreline densified to cover one pixel per row, at least</td> </tr> </tbody> </table> <h2>Parameters</h2> <p><span>RGB values or any transformation in the colour space can be used as parameters.</span><span> </span></p> <h2>Data sources</h2> <p><span>In the CoastSnap initiative, citizens capture images (oblique smartphone photos) from fixed CoastSnap stations and share them with the scientific managers. Images are subjected to a quality control process, spatially registered to a designated target image, and rectified (georeferencing). The shoreline is subsequently digitised from each rectified image.</span><span> </span></p> <h2>Data quality</h2> <p><span>All images included have been supervised by CSs’ scientific managers. However, citizen scientists take images by smartphones (different camera quality) at irregular intervals across various sites with varying weather and illumination conditions. Users of SCLabels dataset must be aware of this variance. </span></p> <h2>Image resolution</h2> <p><span>The resolution of the images depends on the CoastSnap station and the length of the shoreline, ranging from 241x188 pixels to 801x796 pixels.</span></p> <h2>Spatial coverage</h2> <p><span>The SCLabels dataset version 1.0.0 contains data from five Spanish CoastSnap stations, including sandy beaches in the northwest (</span><span>agrelo</span><span>), the Cíes Islands (</span><span>cies</span><span>), the south (</span><span>cadiz</span><span>), and the Balearic Islands (</span><span>samarador </span><span>and </span><span>arenaldentem</span><span>).</span></p> <table> <tbody> <tr> <td><strong> CoastSnap station</strong></td> <td><strong> Longitude</strong></td> <td><strong> Latitude</strong></td> </tr> <tr> <td><em>agrelo</em></td> <td>-8.772</td> <td>42.331</td> </tr> <tr> <td><em>cies</em></td> <td>-8.900</td> <td>42.226</td> </tr> <tr> <td><em>cadiz</em></td> <td>-6.288</td> <td>36.522</td> </tr> <tr> <td><em>samarador</em></td> <td>3.185</td> <td>39.350</td> </tr> <tr> <td><em>arenaldentem</em></td> <td>2.974</td> <td>39.353</td> </tr> </tbody> </table> <h2>Contact information</h2> <p><span>For further technical inquiries or additional information about the annotated dataset, please contact jsoriano@socib.es.</span></p>
Polifonia Corpus - Encyclopedic Module Metadata - Spanish Language
<p>We make available the Metadata related to the Wikipedia pages that constitute the Encyclopedic Module of the Polifonia Textual Corpus. Metadata for this module includes, per each Wikipedia page, its Wikipedia ID, BabelNet ID, gloss, resource type (that can be named entity or concept), Lemmata, Sensekey, WikiData ID.</p> <p>Full description at https://github.com/polifonia-project/Polifonia-Corpus</p>
Polifonia Corpus - Books Module Metadata - Spanish Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
DATASET: characterization of the seed coat extractable phenolic profile and color in 308 common bean lines of the Spanish Diversity Panel
<p>Characterizarion of the seed coat extractable phenolic profile and color in 308 common bean lines of the Spanish Diversity Panel</p>
Multi-head CRF classifier for biomedical multi-class Named Entity Recognition on Spanish clinical notes
<p>This contains the merged dataset as described in the work "<strong>Multi-head CRF classifier for biomedical multi-class Named Entity Recognition on Spanish clinical notes"</strong>.</p> <p>This dataset consists of 4 seperate datasets:</p> <ul> <li><a href="../records/8224056" target="_blank" rel="noopener">MedProcNer</a></li> <li><a href="../records/7614764" target="_blank" rel="noopener">DisTEMIST</a></li> <li><a href="../records/4270158" target="_blank" rel="noopener">PharmaCoNER</a></li> <li><a href="../records/10635215" target="_blank" rel="noopener">SympTEMIST</a></li> </ul> <p>The dataset contains two tasks:</p> <p><strong>Task 1:</strong> This task is related to multi-class Named Entity Recognition. This dataset contains 5 possible classes: SYMPTOM, PROCEDURE, DISEASE, CHEMICAL and PROTEIN.</p> <p><strong>Task 2:</strong> This task is related to Named Entity Linking, where each code corresponds to a code within the SNOMED-CT corpus. The exact corpus used can be obtained <a href="https://download.nlm.nih.gov/umls/kss/IHTSDO20190131/SnomedCT_SpanishRelease-es_PRODUCTION_20190430T120000Z.zip" target="_blank" rel="noopener">here</a>. Further for the MedProcNER, SympTEMIST and DisTEMIST datasets, a gazetteer is provided in the original datasets. </p> <p>For more information on the construction of the dataset, aswell as dataloaders, we refer you to our <a href="https://github.com/ieeta-pt/Multi-Head-CRF" target="_blank" rel="noopener">GitHub repository</a>.<br><br>Further this also contains the embeddings from the <a href="https://huggingface.co/cambridgeltl/SapBERT-UMLS-2020AB-all-lang-from-XLMR-large" target="_blank" rel="noopener">SapBERT</a> model.</p> <p><strong>Please, cite:</strong></p> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <blockquote> <div>@article{jonker2024a, title = {Multi-head {{CRF}} classifier for biomedical multi-class named entity recognition on {{Spanish}} clinical notes}, author = {Jonker, Richard A. A. and Almeida, Tiago and Antunes, Rui and Almeida, Jo{\~a}o R. and Matos, S{\'e}rgio}, year = {2024}, journal = {Database}, publisher = {Oxford University Press} }</div> </blockquote> <div>Jonker, R. A. A., Almeida, T., Antunes, R., Almeida, J. R., & Matos, S. (2024). Multi-head CRF classifier for biomedical multi-class named entity recognition on Spanish clinical notes. (Submitted.) </div> <div> </div> <div> <p><strong>License</strong></p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div>
VIOLENDINGS Violent actions contained in pastoral novels written in Spanish (1559-1633)
<p>This dataset contains a categorization of the violent actions contained in pastoral novels (and in the courtly novels narrated by their characters) written in Spanish between 1559 and 1633 for a diachronic study of the representation of violence in this literary genre and its intersection with other literary traditions. It classifies violent actions by gender and social position of victims and aggressors, relationship between them, motive of aggression, weapon and correspondence with the motives of the Sith Thompson index. In addition, it proposes a categorization for the types of solutions to violent scenes in this literary tradition and the 'distancing devices' used in their representation. The concepts proposed for this categorization are explained in the document INTRO[VIOLENDINGS]20240613_v1. This is the dataset of the research project identified by the acronym VIOLENDINGS —Violence and Happy Endings in the Spanish Golden Age Narrative— (Grant Agreement ID: 101062513),funded by the European Commission’s Marie Skłodowska Curie Actions under Horizon Europe (2021). The project was developed at the Dipartimento di Lingue, Letterature, Culture e Mediazioni of the Università degli Studi di Milano between 2022 and 2024. (2024-06-13) </p> <p> </p> <p> </p>
SCShores: time-series of shorelines from Spanish Sandy beaches from citizen-science monitoring program.
<p>This repository contains 5 years of sandy beaches shorelines deriverd from a citizen-science monitoring program in the Spanish coast. The methodology and the dataset are described in:</p> <p><em><strong>González-Villanueva, R., Soriano-González, J., Alejo, I., Criado-Sudau, F., Plomaritis, T., Fernàndez-Mora, À., Benavente, J., Del Río, L., Nombela, M. Á., and Sánchez-García, E.: SCShores: a comprehensive shoreline dataset of Spanish sandy beaches from a citizen-science monitoring programme, Earth System Science Data. V. 15, 4613-4629 , <a href="https://essd.copernicus.org/articles/15/4613/2023/essd-15-4613-2023.html">https://doi.org/10.5194/essd-15-4613-2023</a>, 2023. </strong></em></p> <p>The shoreline dataset is provided in 1 GEOJSON file: SCShores.geojson. This dataset covers five<strong> </strong>sandy beaches located on the Atlantic and Mediterranean coasts of Spain where CoastSnap stations were available, and it includes a total of 1721 shorelines. The coordinate system for the geospatial layer is WGS84.</p> <ul> <li><strong><em>SCShores.geojson</em></strong>: this layer contains the sandy shorelines . Each feature in this layer is a multipoint with the following attributtes: <ul> <li><strong>site</strong>: CoastSnap station name id, e.g. agrelo, samarador, cadiz, ….</li> <li><strong>date</strong>: date and time of the shoreline, yyyyy-mm-dd hh:mm:ss</li> <li><strong>timezone</strong>: Coordinated Universal Time, UTC</li> <li><strong>timestampQuality</strong>: quality flag indicating the confidence in the date-time indicated by the image provider, e.g. 1, 2</li> <li><strong>imageSource</strong>: source of the original image from which the shoreline has been derived, e.g. Instagram, Twitter, Facebook, Email, CoastSnapApp</li> <li><strong>elevation_m:</strong> same as Z coordinate, defined by the observed tide and the tidal offset, in meters, Tide+tide offset</li> <li><strong>verticalDatum</strong>: mean sea level in Alicante, which is considered the zero topographic reference in the Spanish territory, NMMA</li> <li><strong>geometry</strong>: type of geometry used in the file, MultiPoint</li> <li><strong>coordinates</strong>: Geographic WGS84 coordinates for each point in the geometry, longitude, latitude, Z</li> </ul> </li> </ul> <p> </p>
Effects of fallen Spanish moss (Tillandsia usneoides) on understory plant, invertebrate, and fungi communities
For nearly two years, natural deposition of Spanish moss was excluded from 2m x 2m plots positioned in the understory of a single live oak at each of two field sites. After 29 months, we measured the effects of fallen Spanish moss, relative to unmanipulated control plots that intercepted natural levels of fallen Spanish moss, on understory grass, invertebrate, and fungi communities as well as on litter layer depth, and litter decomposition rates using litter bags. This research was conducted in two fields on Sapelo Island, GA: Long Tabby (LT) and King's Field (KF). Experimental exclusions were maintained by manually removing Spanish moss from experimental plots monthly over the duration of the experiment.
MESINESP: Medical Semantic Indexing in Spanish - Development dataset
<p><em><strong>Please use the <a href="https://doi.org/10.5281/zenodo.4612274">MESINESP2 corpus (the second edition of the shared-task)</a> since it has a higher level of curation, quality and is organized by document type (scientific articles, patents and clinical trials).</strong></em></p> <p> </p> <p> </p> <p><strong>Introduction</strong></p> <p>The Mesinesp (Spanish BioASQ track, see https://temu.bsc.es/mesinesp) development set has a total of 750 records indexed manually by seven experienced medical literature indexers. Indexing is done using <em>DeCS codes, a sort of Spanish equivalent to MeSH terms</em>. Records were distributed in a way that each article was annotated, at least, by two different human indexers.</p> <p>The data annotation process consisted in two steps:</p> <ol> <li>Manual indexing step. DeCS codes were manually assigned to each record following the DeCS manual indexing guidelines.</li> <li>Manual validation and consensus. The joined set of manually indexed DeCS codes generated by both indexers were manually revised and corrections were done.</li> </ol> <p>These annotations were analyzed, resulting in an agreement using the Jaccard index.</p> <p>Records consisted basically in medical literature abstracts and titles from the IBECS and LILACS databases.</p> <p><strong>Zip structure</strong><br> The zip file contains two different development sets:</p> <ul> <li><em>Official development set</em>, which has the union of the annotations, with an agreement of macro = 0.6568 and micro = 0.6819. This set is composed by all the different (unique) DeCS codes that have been added by any annotator for each document; and</li> <li><em>Core-descriptors development set</em>, which has the intersection of the annotations, with an agreement of macro = 1.0 and micro = 1.0. This set is composed of the common DeCS codes that have been added by two or more annotators for each document.</li> </ul> <p><strong>Corpus format</strong></p> <p>Each dataset is a JSON object with one single key named "articles", which contains a list of documents. So, the raw format of the file is one line per document plus two additional lines (the first and the last) to enclose that list of documents and the expected type of data is as follows:</p> <pre><code class="language-json">{"articles":[ {"abstractText":str,"db":str,"decsCodes":list,"id":str,"journal":str,"title":str,"year":int}, ... ]}</code></pre> <p>To clarify, the order of appearance of the fields in each document is as follows (note that this example it is pretty printed for readability purposes):</p> <pre><code class="language-json">{ "articles": [ { "abstractText": "Content of the abstract", "db": "Name of the source database", "decsCodes": [ "code1", "code2", "code3" ], "id": "Id of the document", "journal": "Name of the journal", "title": "Title of the document", "year": 2019 } ] }</code></pre> <p>Note: The fields "db", "journal" and "year" might be null.</p> <p>Copyright (c) 2020 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
CodiEsp corpus: gold standard Spanish clinical cases coded in ICD10 (CIE10) - eHealth CLEF2020
<p><strong>Introduction</strong></p> <p>These are the train, development and test sets of the CodiEsp corpus. Train, development and test have gold standard annotations. In addition, the unannotated background set is also distributed. All documents are released in the context of the CodiEsp track for CLEF ehealth 2020 (<a href="http://temu.bsc.es/codiesp/">http://temu.bsc.es/codiesp/</a>).</p> <p>The CodiEsp corpus contains manually coded clinical cases. All documents are in Spanish language and CIE10 is the coding terminology (it is the Spanish version of ICD10-CM and ICD10-PCS). The CodiEsp corpus has been randomly sampled into three subsets: the train, the development, and the test set. The train set contains 500 clinical cases, and the development and test set 250 clinical cases each. CodiEsp participants must submit predictions for the test and background set, but they will only be evaluated on the test set.</p> <p> </p> <p><strong>Please cite if you use this dataset:</strong></p> <p>Antonio Miranda-Escalada, Aitor Gonzalez-Agirre, Jordi Armengol-Estapé and Martin Krallinger. Overview of automatic clinical coding: annotations, guidelines, and solutions for non-English clinical cases at CodiEsp track of CLEF eHealth 2020. In CLEF (Working Notes). 2020</p> <pre><code>@inproceedings{miranda2020overview, title={Overview of automatic clinical coding: annotations, guidelines, and solutions for non-english clinical cases at codiesp track of CLEF eHealth 2020}, author={Miranda-Escalada, Antonio and Gonzalez-Agirre, Aitor and Armengol-Estap{\'e}, Jordi and Krallinger, Martin}, booktitle={Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings}, year={2020} }</code></pre> <p> </p> <p><strong>Annotation quality</strong></p> <p>Inter-annotator agreement: 88.6% for diagnosis coding, 88.9% for procedure coding and 80.5% for the textual reference annotation. For more information, see the <a href="http://ceur-ws.org/Vol-2696/paper_263.pdf">paper</a>.</p> <p><br> <strong>Zip structure</strong><br> Four folders: train, dev, test and background. Each one of them contains the files for the train, development, test and background corpora, respectively.</p> <ul> <li><strong>train, dev and test</strong> folders have: <ul> <li>3 tab-separated files with the annotation information relevant for each of the 3 sub-tracks of CodiEsp. </li> <li>A subfolder named <em>text_files</em> with the plain text files of the clinical cases.</li> <li>A subfolder named <em>text_files_en</em> with the plain text files machine-translated to English. Due to the translation process, the text files are sentence-splitted.</li> </ul> </li> <li>The <strong>background</strong> folder has only <em>text_files</em> and <em>text_files_en</em> subfolders with the plain text files.</li> </ul> <p><br> <strong>Format</strong><br> The CodiEsp corpus is distributed in plain text in UTF8 encoding, where each clinical case is stored as a single file whose name is the clinical case identifier. Annotations are released in a tab-separated file. Since the CodiEsp track has 3 sub-tracks, every set of documents (train and test) has 3 tab-separated files associated with it. </p> <p>For the sub-tracks CodiEsp-D and CodiEsp-P, the file has the following fields:</p> <pre>articleID ICD10-code </pre> <p>Tab-separated files for the sub-track CodiEsp-X contain extra fields that provide the text-reference and its position:</p> <pre>articleID label ICD10-code text-reference reference-position</pre> <p><br> <strong>Corpus summary statistics</strong><br> The final collection of 1000 clinical cases that make up the corpus had a total of 16504 sentences, with an average of 16.5 sentences per clinical case. It contains a total of 396,988 words, with an average of 396.2 words per clinical case.</p> <p> </p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/codiesp/">Web</a></strong></li> <li><strong><a href="http://ceur-ws.org/Vol-2696/paper_263.pdf">Citation</a>: </strong>Antonio Miranda-Escalada, Aitor Gonzalez-Agirre, Jordi Armengol-Estapé and Martin Krallinger. Overview of automatic clinical coding: annotations, guidelines, and solutions for non-English clinical cases at CodiEsp track of CLEF eHealth 2020. In CLEF (Working Notes). 2020</li> <li><strong><a href="https://doi.org/10.5281/zenodo.3859869">Silver Standard corpus</a></strong></li> <li><strong><a href="https://doi.org/10.5281/zenodo.3730566">Annotation guidelines</a></strong></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhA0crlSVCYMPqMUWd4mXc4x"><strong>YouTube presentations</strong></a></li> <li><a href="https://temu.bsc.es/codiesp/index.php/participants-systems/"><strong>Participant codes</strong></a></li> </ul> <p> </p> <p>For more information, visit the track webpage: http://temu.bsc.es/codiesp/ or email us at encargo-pln-life@bsc.es</p> <p> </p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital</p>
Ontolex-lemon and TIAD versions of Apertium English-Spanish dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2008, Universitat d'Alacant (Transducens group) (c) 2007, Generalitat de Catalunya (dades anglés) (c) 2007, Universitat Pompeu Fabra (IULA) (c) 2005, Universitat Politècnica de Catalunya (c) 2009, Jimmy O'Regan (c) 2009, Paul "greenbreen" Breen
Ontolex-lemon and TIAD versions of Apertium Esperanto-Spanish dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2005--2007, Universitat d'Alacant (Transducens group) (c) 2007, Universitat Pompeu Fabra (c) 2009--, Hèctor Alòs i Font
Ontolex-lemon and TIAD versions of Apertium Spanish-Romanian dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2006--2008, Universitat d'Alacant (Transducens group) (c) 2008-- , Prompsit Language Engineering
Ontolex-lemon and TIAD versions of Apertium Spanish-Galician dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2007 Prompsit Language Engineering (Parts of build system, documentation, and lexical data) (c) 2005-2006, Universidade de Vigo (Seminario de Lingüística Informática, http://sli.uvigo.es) (Galician monolingual data, Galician-Spanish bilingual lexical data, parts of Spanish monolingual data, and parts of Galician-Spanish bilingual structural data) (c) 2005-2006, Universitat d'Alacant (Transducens group, http://transducens.dlsi.ua.es) (Spanish monolingual data and parts of Galician-Spanish bilingual structural data) (c) 2005-2006, Universitat Politècnica de Catalunya (TALP group, http://www.talp.upc.edu) (parts of Spanish monolingual data)
Ontolex-lemon and TIAD versions of Apertium Spanish-Asturian dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2008-2010, Universidá d'Uviéu (Equipu d'Investigación Eslema / Dptos. de Filoloxía Española ya Informática, http://di098.edv.uniovi.es/apertium/comun/nos.php) (c) 2007-2009, Prompsit Language Engineering (Parts of build system, documentation, and lexical data) (c) 2005-2006, Universidade de Vigo (Seminario de Lingüística Informática, http://sli.uvigo.es) (parts of Spanish monolingual data) (c) 2005-2006, Universitat d'Alacant (Transducens group, http://transducens.dlsi.ua.es) (Spanish monolingual data) (c) 2005-2006, Universitat Politècnica de Catalunya (TALP group, http://www.talp.upc.edu) (parts of Spanish monolingual data)
Ontolex-lemon and TIAD versions of Apertium Spanish-Italian dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2007--2008 Prompsit Language Engineering S.L. (http://www.prompsit.com) (c) 2008 Universitat d'Alacant, Grup Transducens (http://transducens.dlsi.ua.es)
MAVIS Twitter dataset: A collection of tweets and sentiment analysis in Spanish about vaccines and diseases during the period 2015-2018
<p>MAVIS dataset comprises a full knowledge base regarding Twitter messages published in Spanish during the period 2015-2018, in the context of sentiment analysis of specific vaccines and their related diseases. Such diseases and vaccines are summarized as follows:</p> <ul> <li>Invasive meningococcal disease (“EMI” in Spanish): Bexsero, Trumenba, Nimenrix</li> <li>Invasive pneumococcal disease (“ENI” in Spanish)</li> <li>Influenza</li> <li>Hepatitis</li> <li>Rotavirus: Rotarix, Rotateq</li> <li>Measles (“Sarampión” in Spanish) and MMR (“Triple vírica” in Spanish)</li> <li>Sepsis</li> <li>Whooping cough (“Tosferina” in Spanish)</li> <li>Chickenpox (“Varicela” in Spanish): Varivax, Varilrix; and Shingles (“Zoster” in Spanish)</li> <li>Human papillomavirus infection (“VPH” in Spanish): Cervarix, Gardasil</li> </ul> <p>Tweets have been manually classified as having a negative or non-negative sentiment by 5 experts. Moreover, an automatic classification has been performed by 3 different tools: IBM Watson (now Watson Tone Analyzer, <a href="https://www.ibm.com/watson/services/tone-analyzer/">https://www.ibm.com/watson/services/tone-analyzer/</a>), Google Cloud Natural Language (<a href="https://cloud.google.com/natural-language">https://cloud.google.com/natural-language</a>), and Meaning Cloud (<a href="https://www.meaningcloud.com/">https://www.meaningcloud.com/</a>). IBM Watson and Google Cloud Natural Language returned a numerical sentiment score ranging from -1 to 1, while Meaning Cloud returned a categorical variable with the values ‘P+’, ‘P’, ‘NEU’, ‘N’ and ‘N+’, which were converted to 1, 2, 3, 4 and 5 respectively.</p> <p>With these variables (IBM Watson, Google Cloud Natural Language, and Meaning Cloud annotations and the experts’ classification as the target label), a machine learning metamodel was developed. Tweets were also annotated with the sentiment output given by this classifier. </p> <p>The provided data includes intrinsic tweets information, intrinsic information regarding the users that posted the tweets, the keywords mentioned in each tweet, and the annotations that the experts, the tools, and the model gave to each tweet.</p> <p><strong>Funding</strong>: This dataset was obtained with funding from MSD, Spain under MAVIS Study (VEAP ID: 7789).</p> <p><strong>Current studies using this dataset at the moment of the publication</strong>:</p> <ul> <li>Rodríguez-González et al., “Creating a metamodel based on machine learning to identify the sentiment of vaccine and disease-related messages in Twitter: the MAVIS study” in 2020 IEEE 33st International Symposium on Computer-Based Medical Systems (CBMS), Jul. 2020, p. 6. DOI: 10.1109/CBMS49503.2020.00053</li> <li>Rodríguez-González et al., "Identifying Polarity in Tweets from an Imbalanced Dataset about Diseases and Vaccines Using a Meta-Model Based on Machine Learning Techniques" in Applied Sciences, 2020, 10. DOI: 10.3390/app10249019</li> </ul>
Photonics4All Bookmark Bubble (Spanish)
<p>The purpose of the bookmarks for the project Photonics4All is to increase the public awareness of photonics and especially of the technological advances of photonics which have changed and improved everyday life (basic technology introduction).<br> <br> Why do soap bubbles have colour?<br> <br> Light reflects off both the inner and outer surfaces of a soap bubble. As the bubble dries out it changes thickness and the light waves reflecting off both surfaces have to travel different distances. White light is made up of all different colours – or waves of different lengths and - when light waves meet – or overlap - they create different colours. Because reflected light travels different distances due to the different film thicknesses we see iridescence in soap bubbles. This phenomenon is used in photonics to provide anti-reflection coating on your glasses for example.</p> <p> </p>
Photonics4All Bookmark LED (Spanish)
<p>The purpose of the bookmarks for the project Photonics4All is to increase the public awareness of photonics and especially of the technological advances of photonics which have changed and improved everyday life (basic technology introduction).<br> <br> How can Light Emitting Diodes (LEDs) transform local food production?<br> <br> Because LEDs emit pure and specific colours they can be used to make plants grow faster and larger. LEDs can replace sunlight or costly greenhouse lamps to grow crops in cold climates or during off-season periods. Growing food locally reduces the need for long-distance transport and lessens the environmental impact used to produce the food. All thanks to Photonics!</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.