Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4,725

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4,725 results for “Normalization”

Learn how ShareScore rates datasets ↗
zenodo40/100

Text-fig. 3. Correlation model of the studied sections of Late Miocene deposits. a. Tuapse highway bridge. b. Gaverdovsky. c. Volchaya Balka. 1. reversed polarity; 2. normal polarity; 3. unstudied interval. in Late Miocene (Early Turolian) Vertebrate Faunas And Associated Biotic Record Of The Northern Caucasus: Geology, Taxonomy, Palaeoenvironment, Biochronology

Text-fig. 3. Correlation model of the studied sections of Late Miocene deposits. a. Tuapse highway bridge. b. Gaverdovsky. c. Volchaya Balka. 1. reversed polarity; 2. normal polarity; 3. unstudied interval.

opencc-by-4.0Dec 2017View details →
zenodo40/100

Fig. 16. A. Vertebral line normally separated from the parietal region. B in Phylogenetic relationships based on morphological data and taxonomy of the genus Salvadora Baird & Girard, 1853 (Reptilia, Colubridae)

Fig. 16. A. Vertebral line normally separated from the parietal region. B. Vertebral line reaching the parietal region.

opencc-by-4.0Aug 2021View details →
zenodo40/100

Five vital signs of normal people

<p>The csv file published contains vital signs dataset of people whom are&nbsp;not suffering from any diseases. Each row in the dataset related to one person, the readings measured from human body three times a day for three days ( total of nine readings per vital sign per subject ). Each time the doctor measures the 5 vital signs (Body temperature, heart rate, oxygen saturation, systolic blood pressure, diastolic blood pressure).&nbsp;This dataset is published under CC0 license.</p>

opencc-zeroOct 2021View details →
zenodo40/100

The dataset for the submitted paper " Time Series Analysis of Normal Mode Energetics for Rossby Wave Breaking and Saturation using a Simple Barotropic Model".

<p>These files are the data of the result in the submitted paper, titled &quot;Time Series Analysis of Normal Mode Energetics for Rossby Wave Breaking and Saturation using a Simple Barotropic Model&quot;.</p> <ul> <li>File Description</li> </ul> <p>pv13.data&nbsp;&nbsp; : Exp. 1<br> pv17.data&nbsp;&nbsp; : Exp. 2</p> <p>The raw potential vorticity (PV) data for the Exp.1 and Exp.2, respectively, used in drawing the Fig.1, 2, and the supplemental movie 1 and 2.<br> These are the grid point value files, 72 levels for the zonal direction, 30 levels for meridional direction.&nbsp; More details are described in the next ctl files.</p> <p>&nbsp;</p> <p>pv13.ctl<br> pv17.ctl</p> <p>Description files for pv13.data and pv17.data. This will be called from grads_pv13.gs and grads_pv17.data, respectively.</p> <p>grads_pv13.gs<br> grads_pv17.gs</p> <p>GrADS script for mapping the PV.</p> <p>&nbsp;</p> <p>energy17.txt&nbsp; : Exp.2</p> <p>The time series table of energy values for exp.2.<br> One raw is identified by combination of the TIME in the experiment and zonal wave number N.</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

Normalized Difference Vegetation Index (NDVI) data

<p>&nbsp;</p> <p>The raster data&nbsp;&#39;viCH_Landsat_NAfill.tif&#39; represents&nbsp;Normalized Difference Vegetation Index (NDVI) &nbsp;data of Switzerland. Raster pixels have&nbsp;250x250 m&nbsp;resolution and contain&nbsp;median values of the years 2013 to 2017. Data was collected&nbsp;by the Landsat&nbsp;satellite and was downloaded via Google Earth Engine. Missing values were closed with a local neighborhood average.</p> <p>The raster data&nbsp;&#39;viNE_Modis2018_NAfill.tif&#39; represents NDVI&nbsp;data of North&nbsp;Eurasia,&nbsp;at 1x1 km resolution, and averaged&nbsp;over the growing season 2018. Data was collected by the Modis sensor and was downloaded via Google Earth Engine. Missing values were closed with a local neighborhood average.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Correlation of the kinetics of viral antigen and genomic RNA with restoration of normal cell homeostasis

<p>The&nbsp;main objective of the studies within COCID Work Package 6 is to understand basic&nbsp;mechanisms by which viral replication&nbsp;machineries are removed from cells after&nbsp;pharmacological interruption of viral replication using functional as well as imaging&nbsp;techniques, including soft X-ray tomography.</p> <p>In this&nbsp;document, we highlight the progress in establishing the HCV replication models&nbsp;(replicons), the antiviral treatment&nbsp;chosen to eliminate viral structures from&nbsp;the host cells and the correlation between elimination of the viral replication&nbsp;machinery&nbsp;and restoration of a normal host cell homeostasis, using markers of&nbsp;HCV-induced stress. These data are essential to frame the&nbsp;experimental setup&nbsp;chosen for the imaging process required in subsequent stps of the project.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Normalized daily average Schumann resonance (SR) intensity data from 13 to 31 January, 2019

<p>Data are provided for the whole globe (T), South America (SA), Africa (AF) and Asia (AS).</p> <p>File prepared by Tam&aacute;s Boz&oacute;ki (20 January, 2023)<br> Contact: bozoki.tamas@epss.hu</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

DisTEMIST corpus: detection and normalization of disease mentions in spanish clinical cases

<p><strong>DisTEMIST corpus:&nbsp;training set&nbsp;+ MULTILINGUAL RESOURCES + CROSSMAPPINGS + test set + background set</strong></p><p>&nbsp;</p><p><strong>Please cite if you use this dataset:</strong></p><p>Miranda-Escalada, A., Gascó, L., Lima-López, S., Farré-Maduell, E., Estrada, D., Nentidis, A., Krithara, A., Katsimpras, G., Paliouras, G., &amp; Krallinger, M. (2022). Overview of DisTEMIST at BioASQ: Automatic detection and normalization of diseases from clinical texts: results, methods, evaluation and multilingual resources.&nbsp;<i>Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings</i></p><p>@article{miranda2022overview, title={Overview of DisTEMIST at BioASQ: Automatic detection and normalization of diseases from clinical texts: results, methods, evaluation and multilingual resources}, author={Miranda-Escalada, Antonio and Gascó, Luis and Lima-López, Salvador and Farré-Maduell, Eulàlia and Estrada, Darryl and Nentidis, Anastasios and Krithara, Anastasia and Katsimpras, Georgios and Paliouras, Georgios and Krallinger, Martin}, booktitle={Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings}, year={2022} }</p><p>&nbsp;</p><p><strong>Introduction</strong></p><p>The DisTEMIST corpus is a collection of 1000 clinical cases with <strong>disease mention annotations manually mapped to</strong> <strong>Snomed-CT concepts</strong>. All documents&nbsp;are released in the context of the BioASQ DisTEMIST track for CLEF 2022.&nbsp;For more information about the track and its schedule, please <a href="https://temu.bsc.es/distemist/">visit the website</a>.</p><p>&nbsp;</p><p><strong>File structure:</strong></p><p>The DisTEMIST corpus has been randomly divided into a training set, containing 750 clinical cases, and a test set (584 in the case of subtrack2), consisting of 250 additional cases. Participants must train their systems using the train set and submit predictions for the test set, on which they will be evaluated.&nbsp;&nbsp;The file structure of the corpus is as follows:</p><ul><li><strong>train_set:</strong><ul><li><strong>text_files</strong>: Folder with plain text files of the clinical cases</li><li><strong>subtrack1_entities</strong>: It contains annotations in a tab-separated file (TSV) with the following columns:<ul><li><i>filename</i>: document name</li><li><i>mark</i>: identifier mention id</li><li><i>label</i>: mentions type (ENFERMEDAD)</li><li><i>off0</i>: starting position of the mention in the document</li><li><i>off1</i>: ending position of the mention in the document</li><li><i>span</i>:&nbsp; text span</li></ul></li><li><strong>subtrack2_linking:&nbsp;</strong>It contains annotations in a tab-separated file (TSV) with the following columns:<ul><li><i>filename</i>: document name</li><li><i>mark</i>: identifier mention id</li><li><i>label</i>: mentions type (ENFERMEDAD)</li><li><i>off0</i>: starting position of the mention in the document</li><li><i>off1</i>: ending position of the mention in the document</li><li><i>span</i>:&nbsp; text span</li><li><i>codes</i>: List of Snomed-CT concept codes linked to the mention. If there is more than one code associated with a mention, they will be concatenated by the symbol "+".</li><li><i>semantic relation</i>: the relationship between the assigned code and the mention. It can be EXACT, when the code corresponds exactly with the mention, or NARROW, when the mention corresponds to a narrower concept than the Snomed-CT code. For instance, the concept "Chorioretinal lacunae" does not exist in Snomed-CT. Then, it is normalized to the Snomed-CT ID 302893000 ("Chorioretinal disorder").</li><li>(Note: the training for entity linking were released in two parts, you must join them)</li></ul></li></ul></li></ul><p>&nbsp;</p><ul><li><strong>test_annotated:</strong><ul><li><strong>text_files</strong>: Folder with plain text files of the clinical cases</li><li><strong>brat</strong>: Folder with text files and their annotations in brat's .ann format</li><li><strong>subtrack1_entities: </strong>It contains annotations in a tab-separated file (TSV) with the same columns as the train set</li><li><strong>subtrack2_linking: </strong>It contains annotations in a tab-separated file (TSV) with the same columns as the train set</li></ul></li></ul><p>&nbsp;</p><ul><li><strong>test_background_unannotated/text_files:&nbsp;</strong>3000 clinical cases (test + background). In the the original task, participants had to make predictions for these 3000 files and they were evaluated on a subset of them.</li></ul><p>&nbsp;</p><ul><li><strong>multilingual-resources</strong>: we have generated the annotated training and validation sets in <strong>6 languages</strong>: <i><strong>English</strong></i>, <i><strong>Portuguese</strong></i>, <i><strong>Catalan</strong></i>, <i><strong>Italian</strong></i>, <i><strong>French</strong></i> and <i><strong>Romanian</strong></i>.&nbsp;The process was:<ol><li>The&nbsp;text files were translated with a neural machine translation system.</li><li>The annotations were translated with the same&nbsp;neural machine translation system.</li><li>The translated annotations were transferred to the translated&nbsp;text files using an annotation transfer technology.</li><li>The text files are stored in the multilingual_resources/training-text-files&nbsp;subfolder.</li><li>The annotated TSV files are stored in the multilingual_resources/lang&nbsp;subfolders.</li><li>If you want to visualize the multilingual resources, check out this Brat server:&nbsp;<a href="https://temu.bsc.es/mDistemist/#/translations/">https://temu.bsc.es/mDistemist/#/translations/</a><br>For instance, you can see the parallel annotations&nbsp;in <a href="http://temu.bsc.es/mDistemist/diff.xhtml#/translations/fr/train/es-S0004-06142008000100011-1?diff=/translations/en/train/">English vs&nbsp;in French</a>, or in&nbsp;<a href="https://temu.bsc.es/mDistemist/diff.xhtml#/translations/cat/train/S0004-06142005000500011-1?diff=/gold-standard/train/">Spanish (the gold standard) vs in Catalan</a>.</li></ol></li></ul><p>&nbsp;</p><ul><li><strong>cross-mappings.&nbsp;</strong>We include the same entities as in DISTEMIST-linking but mapped to Snomed-CT, <a href="https://www.ncbi.nlm.nih.gov/mesh/"><strong>MeSH</strong></a>, <a href="https://icd.who.int/browse10/2019/en#/"><strong>ICD-10</strong></a>, <a href="https://hpo.jax.org/"><strong>HPO</strong></a>, and <a href="https://www.omim.org/"><strong>OMIM</strong></a>. The original mappings are manual and to Snomed-CT. The mapping to the other terminologies&nbsp;was done through the UMLS Metathesaurus.</li></ul><p>&nbsp;</p><p><strong>Resources</strong>&nbsp;</p><ul><li><a href="https://temu.bsc.es/distemist/">Task Web</a></li><li><a href="https://github.com/TeMU-BSC/distemist_evaluation_library">Evaluation Library</a></li><li><a href="https://doi.org/10.5281/zenodo.6458114"><strong>DISTEMIST&nbsp;gazetteer</strong></a></li><li><a href="https://doi.org/10.5281/zenodo.6458078"><strong>DISTEMIST guidelines</strong></a></li><li><strong>Citation: </strong>Miranda-Escalada, A., Gascó, L., Lima-López, S., Farré-Maduell, E., Estrada, D., Nentidis, A., Krithara, A., Katsimpras, G., Paliouras, G., &amp; Krallinger, M. (2022). Overview of DisTEMIST at BioASQ: Automatic detection and normalization of diseases from clinical texts: results, methods, evaluation and multilingual resources.&nbsp;<i>Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings</i></li><li><a href="https://ceur-ws.org/Vol-3180/paper-11.pdf"><strong>Overview paper</strong></a></li><li><a href="https://ceur-ws.org/Vol-3180/"><strong>CLEF/BioASQ Workshop proceedings</strong></a></li><li><a href="https://www.youtube.com/watch?v=KvvhogUMjl8&amp;list=PL5uSCzf1azhC_Sg9leXavmcA4nJ4HQa_n"><strong>Youtube videos (overview &amp; teams)</strong></a></li><li><a href="https://www.slideshare.net/MartinKrallinger/overview-of-distemist-at-bioasq-automatic-detection-and-normalization-of-diseases-from-clinical-texts-results-methods-evaluation-and-multilingual-resources-at-bioasq-workshop-of-clef-conference"><strong>DISTEMIST overview talk slides</strong></a></li></ul><p>&nbsp;</p><p><strong>License</strong></p><p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p><p>&nbsp;</p><p><strong>Additional resources and corpora</strong></p><p>If you are interested in DisTEMIST, you might want to check out these corpora and resources:</p><ul><li><a href="https://zenodo.org/records/8413866">SymTEMIST</a> (Corpus of symptoms, signs and findings mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/4270158">PharmaCoNER</a> (Corpus of medications, drugs, chemical substances, genes, proteins and vaccine mentions and normalization, same document collection)</li><li><a href="https://zenodo.org/records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI), modified synthetic verions of the document collection)</li><li><a href="https://zenodo.org/records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization, different document collection)</li><li><a href="https://zenodo.org/records/3837305">CodiESp</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version, same document collection)</li><li><a href="https://zenodo.org/records/7684093">LivingNER</a> (Corpus of mentions of species, including human/family members, pathogens, food, etc.. and normalization to NCBI Taxonomy, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags, same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries), same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags, same document collection)</li><li><a href="https://zenodo.org/records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts, different document collection)</li></ul><p><strong>Contact</strong></p><p>If you have any questions or suggestions, please contact us at:</p><p><br>- Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

LivingNER corpus: Named entity recognition, normalization & classification of species, pathogens and food

<p><strong>LivingNER Gold Standard corpus (includes training, validation, test and background&nbsp;sets + MULTILINGUAL RESOURCES</strong>)</p><p>&nbsp;</p><p><strong>Please cite if you use this dataset:</strong></p><p>A. Miranda-Escalada, E. Farré-Maduell, S. Lima-López, D. Estrada, L. Gascó, M. Krallinger, Mention detection, normalization &amp; classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources, <i>Procesamiento del Lenguaje Natural</i> (2022)</p><p>@article{amiranda2022nlp, title={Mention detection, normalization \&amp; classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources}, author={Miranda-Escalada, Antonio and Farr{\'e}-Maduell, Eul{`a}lia and Lima-L{\'o}pez, Salvador and Estrada, Darryl and Gasc{\'o}, Luis and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, year={2022} }</p><p>&nbsp;</p><p><i><strong>1. Introduction</strong></i></p><p>The LivingNER Gold Standard corpus is a collection of<strong> 2000 clinical case reports</strong> covering a <strong>broad range of medical specialities</strong>, i.e. infectious diseases (including Covid-19 cases), cardiology, neurology, oncology, dentistry, pediatrics, endocrinology, primary care, allergology, radiology, psychiatry, ophthalmology, urology, internal medicine, emergency and intensive care medicine, tropical medicine, and dermatology <strong>annotated with species</strong> [SPECIES] (including <strong>living organisms</strong> and <strong>microorganisms</strong>) and <strong>infectious diseases</strong> [ENFERMEDAD] mentions. Species mentions include many <strong>pathogens</strong> and infectious agents, but also <strong>food</strong>, allergens, <strong>pets</strong> or other species, taxonomic groups and organisms of clinical relevance.&nbsp;</p><p>The &nbsp;LivingNER corpus has also annotations of mentions of <strong>humans</strong> (tag HUMAN), including the patients itself, <strong>family members</strong>, healhcare professionals or other persons mentioned in the case reports. Thus it can be useful to extract family history information of patients or information about the social and healthcare personal environment and interactions.</p><p>All mentions have been exhaustively manually mapped by experts to their corresponding <a href="https://www.ncbi.nlm.nih.gov/taxonomy"><strong>NCBI Taxonomy</strong></a> identifiers.&nbsp;</p><p>It was used for the&nbsp;<a href="https://temu.bsc.es/livingner/">LivingNER</a>&nbsp;Shared Task on pathogens and living beings detection and normalization in Spanish medical documents, which was celebrated as part of IberLEF 2022.</p><p>&nbsp;</p><p><i><strong>2. Training, validation, test and background sets</strong></i></p><p>The training set is composed of 1000 clinical case reports. The validation set includes 500 clinical case reports with the same characteristics and the test set includes 485. The background set is a collection of around 13k unannotated case reports that were originally added to prevent manual annotations in the test set during the competition and to create a Silver Standard.</p><p><i><strong>2.1 Annotations format</strong></i></p><p>Annotations and text files are distributed separately. The texts are in plain text (.txt in UTF-8) format, while the annotations are are distributed in a tab-separated file&nbsp;(.tsv) file with one row per annotation:</p><p>- For&nbsp;<strong>subtask 1 (LivingNER-Species NER track)</strong>, the .tsv file has the following columns:</p><ul><li>filename: document name</li><li>mark: identifier mention mark</li><li>label: mention type (SPECIES or HUMAN)</li><li>off0: starting&nbsp;position of the mention in the document</li><li>off1: ending position of the mention in the document</li><li>span: textual span</li></ul><p>&nbsp;- For <strong>subtask 2 (LivingNER-Species Norm track)</strong>, the .tsv file has the same columns as the previous one, plus:</p><ul><li>isH: whether the span is narrower than the&nbsp;NCBITax assigned code</li><li>isN: whether the mention corresponds to a nosocomial infection</li><li>iscomplex: whether the span has assigned a combination of NCBITax&nbsp;codes</li><li>NCBITax: mention code in the&nbsp;NCBI Taxonomy</li></ul><p>- For&nbsp;<strong>subtask 3 (LivingNER-Clinical IMPACT track)</strong>,&nbsp;the .tsv file has the following columns:</p><ul><li>filename</li><li>isPet (Yes/No)</li><li>PetIDs (NCBITaxonomy codes of pet &amp; farm animals present in document)</li><li>isAnimalInjury (Yes/No)</li><li>AnimalInjuryIDs (NCBITaxonomy codes of animals causing injuries present in document)</li><li>IsFood (Yes/No)</li><li>FoodIDs (NCBITaxonomy codes of food mentions present in document)</li><li>isNosocomial (Yes/No)</li><li>NosocomialIDs (NCBITaxonomy codes of nosocomial species mentions present in document)</li></ul><p><i><strong>2.2 Important notes about subtask 3 (LivingNER-Clinical IMPACT track):</strong></i></p><ul><li><strong>Less clinical case reports</strong>. Subtask&nbsp;3 (LivingNER-Clinical IMPACT track) contains half of the clinical case reports&nbsp;(500 in the training partition, 250 in the validation partition). The list of valid&nbsp;clinical case reports for task 3 is included in the data&nbsp;(train_files_task3.txt and validation_files_task3.txt)</li><li><strong>Enriched dataset.</strong>&nbsp;The GS format is the one described above (a TSV with one line per clinical case report). However, we believe participants may find useful and&nbsp;<strong>enriched dataset.&nbsp;</strong>Then, we provide an additional dataset, with the mentions of the NER track classified in the 4 Clinical impact categories (food, pet&amp;farm animals, animals causing injuries and nosocomial). It is a TSV file with one row per annotation, and with the following columns:&nbsp;filename,&nbsp;mark,&nbsp;label,&nbsp;off0,&nbsp;off1,&nbsp;span,&nbsp;isPet,&nbsp;isAnimalInjury,&nbsp;isFood,&nbsp;isNosocomial,&nbsp;isH,&nbsp;iscomplex,&nbsp;code</li></ul><p>&nbsp;</p><p><i><strong>3. Multilingual resources</strong></i></p><p>We have generated the annotated training and validation sets in <strong>7 languages</strong>:</p><ul><li><i><strong>English</strong></i></li><li><i><strong>Portuguese</strong></i></li><li><i><strong>Catalan</strong></i></li><li><i><strong>Galician</strong></i></li><li><i><strong>Italian</strong></i></li><li><i><strong>French</strong></i></li><li><i><strong>Romanian</strong></i></li></ul><p>&nbsp;</p><p>The process was:</p><ol><li>The&nbsp;text files were translated with a neural machine translation system.</li><li>The annotations were translated with the same&nbsp;neural machine translation system.</li><li>The translated annotations were transferred to the translated&nbsp;text files using an annotation transfer technology.</li></ol><p>The&nbsp;text files are stored in the&nbsp;multilingual_resources/<strong>training-text-files</strong> and&nbsp;multilingual_resources/<strong>validation-text-files </strong>subfolders.</p><p>The annotated TSV files are stored in the&nbsp;multilingual_resources/<strong>annotation_transfer </strong>subfolder.</p><p>For the sake of comparison, we incorporate as well the annotations that resulted from the <a href="https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-11-85">LINNAEUS tool</a>&nbsp;in the&nbsp;multilingual_resources/<strong>linneaus</strong>&nbsp;subfolder.</p><p>If you want to visualize the multilingual resources, check out this Brat server:&nbsp;<a href="https://temu.bsc.es/mLivingNER/#/translations/">https://temu.bsc.es/mLivingNER/#/translations/</a></p><p>For instance, you can see the parallel annotations in <a href="https://temu.bsc.es/mLivingNER/diff.xhtml#/translations/en/annotation_transfer/train/casos_clinicos_cardiologia34?diff=/translations/fr/annotation_transfer/train/">English vs&nbsp;in French</a>, or <a href="https://temu.bsc.es/mLivingNER/diff.xhtml#/translations/cat/annotation_transfer/train/casos_clinicos_cardiologia35?diff=/gold-standard/train/">in Spanish (the gold standard) vs in Catalan.</a></p><p>&nbsp;</p><p><strong>Resources</strong></p><ul><li><a href="https://temu.bsc.es/livingner/"><strong>Task Web</strong></a></li><li><strong>Citation:&nbsp;</strong>A. Miranda-Escalada, E. Farré-Maduell, S. Lima-López, D. Estrada, L. Gascó, M. Krallinger, Mention detection, normalization &amp; classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources, <i>Procesamiento del Lenguaje Natural</i> (2022)</li><li><a href="https://doi.org/10.5281/zenodo.6385162"><strong>Annotation guidelines</strong></a></li><li><a href="https://github.com/tonifuc3m/livingner-evaluation-library"><strong>Evaluation library</strong></a></li><li><a href="https://doi.org/10.5281/zenodo.6390506">LivingNER terminology</a></li><li><a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6444"><strong>Overview paper</strong></a></li><li><a href="https://ceur-ws.org/Vol-3202/"><strong>Proceedings participant papers</strong></a></li><li><a href="https://www.youtube.com/watch?v=8VcZw8ywyJY&amp;list=PL5uSCzf1azhA_gMLC3DBZe6NvmMJiggTg"><strong>Youtube videos</strong></a></li><li><a href="https://www.slideshare.net/MartinKrallinger/mention-detection-normalization-classification-of-species-pathogens-humans-and-food-in-clinical-documents-overview-of-the-livingner-shared-task-and-resources-talk-at-iberlef-sepln-2022"><strong>LivingNER overview talk sides at IberLEF/SEPLN</strong></a></li></ul><p>&nbsp;</p><p><strong>License</strong></p><p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p><p><strong>Contact</strong></p><p>If you have any questions or suggestions, please contact us at:</p><p><br>- Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p><p><strong>Additional resources and corpora</strong></p><p>If you are interested in SympTEMIST, you might want to check out these corpora and resources:</p><ul><li><a href="https://zenodo.org/records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/8413866">SympTEMIST</a> (Corpus of symptoms, sign and findings mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/4270158">PharmaCoNER</a> (Corpus of medications, drugs, chemical substances, genes, proteins and vaccine mentions and normalization, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization, different document collection)</li><li><a href="https://zenodo.org/records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization, different document collection)</li><li><a href="https://zenodo.org/records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI), differentdocument collection)</li><li><a href="https://zenodo.org/records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization, different document collection)</li><li><a href="https://zenodo.org/records/3837305">CodiESp</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries), different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, different document collection, some overlapping documents)</li></ul>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Gyroscope sensor data of shank motion during normal and barefoot walking

<p>The measurements were performed in closed and disturbance free space, where an unobstructed 10 m walkway was arranged. All tests were performed on a hard floor surface first with shoes that were adapt for walking (indoor sports shoes, sneakers etc.) and afterwards walking barefooted the same protocol. The subjects had clothing which did not restrict lower limb movement. In each test, the 10 m walking was repeated three times. Prior to testing, the procedure was demonstrated, and the sensors were carefully positioned to correct locations. The subjects were instructed to walk with their own natural walking velocity and to begin each 10 m walk from a completely stationary position.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

The Chilean Waiting List sub-Corpus with medical entities normalized to UMLS terminology

<p>A collection of 2000 medical referrals from the Chilean Waiting List Corpus, manually annotated with six entity types (Finding, Procedure, Disease, Family Member, Body Part, and Medication) and manually normalized to the Unified Medical Language System (UMLS).</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Figure 4. A normal female and a in Dichromothrips smithi (Zimmermann), a New Thrips Species Infesting Bamboo Orchids Arundina graminifolia (D. Don) Hochr. and Commercially Grown Orchids in Hawaii

Figure 4. A normal female and a teneral (newly molted, light colored) female D. smithi after storage in 70% ethanol (4a). Male and female D. smithi in ethanol (4b). The two males are lighter in color. Note that females stored in ethanol appear lighter brown with abdominal segments stretched out in comparison to the live female on blossom tissue pictured in 4c.

opencc-by-4.0Dec 2012View details →
dryad40/100

Cacao gene atlas: Additional file 13 (CPM normalized read counts) and Additional file 14 (fractional counts)

<p class="MsoNormal">A large dataset of replicated transcriptomes was developed to accelerate<span class="MsoCommentReference"> <em>T</em></span><em>heobroma cocoa</em> genomics research with the long-term goal of progressing breeding towards developing high-yielding elite varieties of cacao. RNAs were extracted and transcriptomes were sequenced from 123 different tissues and stages of development representing major organs and developmental stages of the cacao lifecycle. In addition, several experimental treatments and time courses were performed to measure gene expression in tissues responding to biotic and abiotic stressors. Samples were collected in replicates (3-5) to enable statistical analysis of gene expression levels for a total of 390 transcriptomes.  We describe the creation of the atlas,and its global characterization and define sets of genes co-regulated in highly organ- and temporally-specific manners. To promote wider use of these data, all raw sequencing data, expression read mapping matrices, scripts, and other information used to create the resource are freely available online. A gene expression browser with a graphical user interface was developed to display gene expression patterns and to provide easy access of raw data and statistical analyses.</p>

opencc-zeroAug 2023View details →
zenodo40/100

Hybrid Formats - The New Normal?!

<p>How can hybrid meetings, workshops and trainings be conceptualized successfully?</p> <p>A hybrid workshop hosted by Instituto Mora, Mexico City.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Datasets and Trained Models of the paper: The NFLikelihood: an unsupervised DNNLikelihood from Normalizing Flows

<p><strong>Training Data and Trained models corresponding to the publication: 'The NFLikelihood: an unsupervised DNNLikelihood from Normalizing Flows' (<a href="https://arxiv.org/abs/2309.09743">arXiv:2309.09743</a>).</strong></p> <p>The files cointain the relevant resources for 3 trained Likelihood functions: The Toy-Likelihood, the EW-Likelihood and the Flavor-Likelihood. In each corresponding directory, the training data is found in \data. The trained model and generated samples are found in \NFmodel.</p> <p>To reproduce the published results, git clone&nbsp;https://github.com/NF4HEP/NFLikelihoods,&nbsp; and plug in the provided resources in the corresponding directories of the code.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
dryad40/100

Supplemental files: Q-Q plots and Shapiro-Wilk tests of normality

<p><em>Symphyotrichum</em> (Asteraceae) is a well-circumscribed genus on the basis of morphological and molecular data, but species boundaries remain poorly understood. Here, the species delimitation of the contentious <em>Symphyotrichum</em> <em>subulatum</em> group (<em>Symphyotrichum</em> subg. <em>Astropolium</em>) is illuminated using morphometric, phylogenomic, and geographical analyses. <em>Symphyotrichum mexicanum</em> sp. nov., a new species endemic to central Mexico, is described and distinguished from <em>Symphyotrichum</em> <em>expansum</em> based on its morphometric attributes, phylogenetic placement, geographic range, and ecological specialization.</p>

opencc-zeroSep 2023View details →
ClinicalTrials.gov40/100

A Study of Long-Term Safety and Efficacy of Lanadelumab for Prevention of Acute Attacks of Non-histaminergic Angioedema With Normal C1-Inhibitor

ClinicalTrials.gov study NCT04444895. IPD Sharing: YES. Countries: 10. Publications: 1.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov40/100

A Study of Lanadelumab in Teenagers and Adults to Prevent Acute Attacks of Non-histaminergic Angioedema With Normal C1-Inhibitor (C1-INH)

ClinicalTrials.gov study NCT04206605. IPD Sharing: YES. Countries: 10. Publications: 1.

controlledIPD-YESFeb 2026View details →
dryad40/100

The new normal? Redaction bias in biomedical science

Open the record for dataset details and reuse information.

publicDec 2021View details →
dryad40/100

The phageome in normal and inflamed human skin

Open the record for dataset details and reuse information.

publicJan 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record