Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,725
datasets available to search
ShareScore release 0.9.0
Dataset results
4,725 results for “Normalization”
Text-fig. 3. Correlation model of the studied sections of Late Miocene deposits. a. Tuapse highway bridge. b. Gaverdovsky. c. Volchaya Balka. 1. reversed polarity; 2. normal polarity; 3. unstudied interval. in Late Miocene (Early Turolian) Vertebrate Faunas And Associated Biotic Record Of The Northern Caucasus: Geology, Taxonomy, Palaeoenvironment, Biochronology
Text-fig. 3. Correlation model of the studied sections of Late Miocene deposits. a. Tuapse highway bridge. b. Gaverdovsky. c. Volchaya Balka. 1. reversed polarity; 2. normal polarity; 3. unstudied interval.
Fig. 16. A. Vertebral line normally separated from the parietal region. B in Phylogenetic relationships based on morphological data and taxonomy of the genus Salvadora Baird & Girard, 1853 (Reptilia, Colubridae)
Fig. 16. A. Vertebral line normally separated from the parietal region. B. Vertebral line reaching the parietal region.
Five vital signs of normal people
<p>The csv file published contains vital signs dataset of people whom are not suffering from any diseases. Each row in the dataset related to one person, the readings measured from human body three times a day for three days ( total of nine readings per vital sign per subject ). Each time the doctor measures the 5 vital signs (Body temperature, heart rate, oxygen saturation, systolic blood pressure, diastolic blood pressure). This dataset is published under CC0 license.</p>
The dataset for the submitted paper " Time Series Analysis of Normal Mode Energetics for Rossby Wave Breaking and Saturation using a Simple Barotropic Model".
<p>These files are the data of the result in the submitted paper, titled "Time Series Analysis of Normal Mode Energetics for Rossby Wave Breaking and Saturation using a Simple Barotropic Model".</p> <ul> <li>File Description</li> </ul> <p>pv13.data : Exp. 1<br> pv17.data : Exp. 2</p> <p>The raw potential vorticity (PV) data for the Exp.1 and Exp.2, respectively, used in drawing the Fig.1, 2, and the supplemental movie 1 and 2.<br> These are the grid point value files, 72 levels for the zonal direction, 30 levels for meridional direction. More details are described in the next ctl files.</p> <p> </p> <p>pv13.ctl<br> pv17.ctl</p> <p>Description files for pv13.data and pv17.data. This will be called from grads_pv13.gs and grads_pv17.data, respectively.</p> <p>grads_pv13.gs<br> grads_pv17.gs</p> <p>GrADS script for mapping the PV.</p> <p> </p> <p>energy17.txt : Exp.2</p> <p>The time series table of energy values for exp.2.<br> One raw is identified by combination of the TIME in the experiment and zonal wave number N.</p>
Normalized Difference Vegetation Index (NDVI) data
<p> </p> <p>The raster data 'viCH_Landsat_NAfill.tif' represents Normalized Difference Vegetation Index (NDVI) data of Switzerland. Raster pixels have 250x250 m resolution and contain median values of the years 2013 to 2017. Data was collected by the Landsat satellite and was downloaded via Google Earth Engine. Missing values were closed with a local neighborhood average.</p> <p>The raster data 'viNE_Modis2018_NAfill.tif' represents NDVI data of North Eurasia, at 1x1 km resolution, and averaged over the growing season 2018. Data was collected by the Modis sensor and was downloaded via Google Earth Engine. Missing values were closed with a local neighborhood average.</p>
Correlation of the kinetics of viral antigen and genomic RNA with restoration of normal cell homeostasis
<p>The main objective of the studies within COCID Work Package 6 is to understand basic mechanisms by which viral replication machineries are removed from cells after pharmacological interruption of viral replication using functional as well as imaging techniques, including soft X-ray tomography.</p> <p>In this document, we highlight the progress in establishing the HCV replication models (replicons), the antiviral treatment chosen to eliminate viral structures from the host cells and the correlation between elimination of the viral replication machinery and restoration of a normal host cell homeostasis, using markers of HCV-induced stress. These data are essential to frame the experimental setup chosen for the imaging process required in subsequent stps of the project.</p>
Normalized daily average Schumann resonance (SR) intensity data from 13 to 31 January, 2019
<p>Data are provided for the whole globe (T), South America (SA), Africa (AF) and Asia (AS).</p> <p>File prepared by Tamás Bozóki (20 January, 2023)<br> Contact: bozoki.tamas@epss.hu</p>
DisTEMIST corpus: detection and normalization of disease mentions in spanish clinical cases
<p><strong>DisTEMIST corpus: training set + MULTILINGUAL RESOURCES + CROSSMAPPINGS + test set + background set</strong></p><p> </p><p><strong>Please cite if you use this dataset:</strong></p><p>Miranda-Escalada, A., Gascó, L., Lima-López, S., Farré-Maduell, E., Estrada, D., Nentidis, A., Krithara, A., Katsimpras, G., Paliouras, G., & Krallinger, M. (2022). Overview of DisTEMIST at BioASQ: Automatic detection and normalization of diseases from clinical texts: results, methods, evaluation and multilingual resources. <i>Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings</i></p><p>@article{miranda2022overview, title={Overview of DisTEMIST at BioASQ: Automatic detection and normalization of diseases from clinical texts: results, methods, evaluation and multilingual resources}, author={Miranda-Escalada, Antonio and Gascó, Luis and Lima-López, Salvador and Farré-Maduell, Eulàlia and Estrada, Darryl and Nentidis, Anastasios and Krithara, Anastasia and Katsimpras, Georgios and Paliouras, Georgios and Krallinger, Martin}, booktitle={Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings}, year={2022} }</p><p> </p><p><strong>Introduction</strong></p><p>The DisTEMIST corpus is a collection of 1000 clinical cases with <strong>disease mention annotations manually mapped to</strong> <strong>Snomed-CT concepts</strong>. All documents are released in the context of the BioASQ DisTEMIST track for CLEF 2022. For more information about the track and its schedule, please <a href="https://temu.bsc.es/distemist/">visit the website</a>.</p><p> </p><p><strong>File structure:</strong></p><p>The DisTEMIST corpus has been randomly divided into a training set, containing 750 clinical cases, and a test set (584 in the case of subtrack2), consisting of 250 additional cases. Participants must train their systems using the train set and submit predictions for the test set, on which they will be evaluated. The file structure of the corpus is as follows:</p><ul><li><strong>train_set:</strong><ul><li><strong>text_files</strong>: Folder with plain text files of the clinical cases</li><li><strong>subtrack1_entities</strong>: It contains annotations in a tab-separated file (TSV) with the following columns:<ul><li><i>filename</i>: document name</li><li><i>mark</i>: identifier mention id</li><li><i>label</i>: mentions type (ENFERMEDAD)</li><li><i>off0</i>: starting position of the mention in the document</li><li><i>off1</i>: ending position of the mention in the document</li><li><i>span</i>: text span</li></ul></li><li><strong>subtrack2_linking: </strong>It contains annotations in a tab-separated file (TSV) with the following columns:<ul><li><i>filename</i>: document name</li><li><i>mark</i>: identifier mention id</li><li><i>label</i>: mentions type (ENFERMEDAD)</li><li><i>off0</i>: starting position of the mention in the document</li><li><i>off1</i>: ending position of the mention in the document</li><li><i>span</i>: text span</li><li><i>codes</i>: List of Snomed-CT concept codes linked to the mention. If there is more than one code associated with a mention, they will be concatenated by the symbol "+".</li><li><i>semantic relation</i>: the relationship between the assigned code and the mention. It can be EXACT, when the code corresponds exactly with the mention, or NARROW, when the mention corresponds to a narrower concept than the Snomed-CT code. For instance, the concept "Chorioretinal lacunae" does not exist in Snomed-CT. Then, it is normalized to the Snomed-CT ID 302893000 ("Chorioretinal disorder").</li><li>(Note: the training for entity linking were released in two parts, you must join them)</li></ul></li></ul></li></ul><p> </p><ul><li><strong>test_annotated:</strong><ul><li><strong>text_files</strong>: Folder with plain text files of the clinical cases</li><li><strong>brat</strong>: Folder with text files and their annotations in brat's .ann format</li><li><strong>subtrack1_entities: </strong>It contains annotations in a tab-separated file (TSV) with the same columns as the train set</li><li><strong>subtrack2_linking: </strong>It contains annotations in a tab-separated file (TSV) with the same columns as the train set</li></ul></li></ul><p> </p><ul><li><strong>test_background_unannotated/text_files: </strong>3000 clinical cases (test + background). In the the original task, participants had to make predictions for these 3000 files and they were evaluated on a subset of them.</li></ul><p> </p><ul><li><strong>multilingual-resources</strong>: we have generated the annotated training and validation sets in <strong>6 languages</strong>: <i><strong>English</strong></i>, <i><strong>Portuguese</strong></i>, <i><strong>Catalan</strong></i>, <i><strong>Italian</strong></i>, <i><strong>French</strong></i> and <i><strong>Romanian</strong></i>. The process was:<ol><li>The text files were translated with a neural machine translation system.</li><li>The annotations were translated with the same neural machine translation system.</li><li>The translated annotations were transferred to the translated text files using an annotation transfer technology.</li><li>The text files are stored in the multilingual_resources/training-text-files subfolder.</li><li>The annotated TSV files are stored in the multilingual_resources/lang subfolders.</li><li>If you want to visualize the multilingual resources, check out this Brat server: <a href="https://temu.bsc.es/mDistemist/#/translations/">https://temu.bsc.es/mDistemist/#/translations/</a><br>For instance, you can see the parallel annotations in <a href="http://temu.bsc.es/mDistemist/diff.xhtml#/translations/fr/train/es-S0004-06142008000100011-1?diff=/translations/en/train/">English vs in French</a>, or in <a href="https://temu.bsc.es/mDistemist/diff.xhtml#/translations/cat/train/S0004-06142005000500011-1?diff=/gold-standard/train/">Spanish (the gold standard) vs in Catalan</a>.</li></ol></li></ul><p> </p><ul><li><strong>cross-mappings. </strong>We include the same entities as in DISTEMIST-linking but mapped to Snomed-CT, <a href="https://www.ncbi.nlm.nih.gov/mesh/"><strong>MeSH</strong></a>, <a href="https://icd.who.int/browse10/2019/en#/"><strong>ICD-10</strong></a>, <a href="https://hpo.jax.org/"><strong>HPO</strong></a>, and <a href="https://www.omim.org/"><strong>OMIM</strong></a>. The original mappings are manual and to Snomed-CT. The mapping to the other terminologies was done through the UMLS Metathesaurus.</li></ul><p> </p><p><strong>Resources</strong> </p><ul><li><a href="https://temu.bsc.es/distemist/">Task Web</a></li><li><a href="https://github.com/TeMU-BSC/distemist_evaluation_library">Evaluation Library</a></li><li><a href="https://doi.org/10.5281/zenodo.6458114"><strong>DISTEMIST gazetteer</strong></a></li><li><a href="https://doi.org/10.5281/zenodo.6458078"><strong>DISTEMIST guidelines</strong></a></li><li><strong>Citation: </strong>Miranda-Escalada, A., Gascó, L., Lima-López, S., Farré-Maduell, E., Estrada, D., Nentidis, A., Krithara, A., Katsimpras, G., Paliouras, G., & Krallinger, M. (2022). Overview of DisTEMIST at BioASQ: Automatic detection and normalization of diseases from clinical texts: results, methods, evaluation and multilingual resources. <i>Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings</i></li><li><a href="https://ceur-ws.org/Vol-3180/paper-11.pdf"><strong>Overview paper</strong></a></li><li><a href="https://ceur-ws.org/Vol-3180/"><strong>CLEF/BioASQ Workshop proceedings</strong></a></li><li><a href="https://www.youtube.com/watch?v=KvvhogUMjl8&list=PL5uSCzf1azhC_Sg9leXavmcA4nJ4HQa_n"><strong>Youtube videos (overview & teams)</strong></a></li><li><a href="https://www.slideshare.net/MartinKrallinger/overview-of-distemist-at-bioasq-automatic-detection-and-normalization-of-diseases-from-clinical-texts-results-methods-evaluation-and-multilingual-resources-at-bioasq-workshop-of-clef-conference"><strong>DISTEMIST overview talk slides</strong></a></li></ul><p> </p><p><strong>License</strong></p><p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p><p> </p><p><strong>Additional resources and corpora</strong></p><p>If you are interested in DisTEMIST, you might want to check out these corpora and resources:</p><ul><li><a href="https://zenodo.org/records/8413866">SymTEMIST</a> (Corpus of symptoms, signs and findings mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/4270158">PharmaCoNER</a> (Corpus of medications, drugs, chemical substances, genes, proteins and vaccine mentions and normalization, same document collection)</li><li><a href="https://zenodo.org/records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI), modified synthetic verions of the document collection)</li><li><a href="https://zenodo.org/records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization, different document collection)</li><li><a href="https://zenodo.org/records/3837305">CodiESp</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version, same document collection)</li><li><a href="https://zenodo.org/records/7684093">LivingNER</a> (Corpus of mentions of species, including human/family members, pathogens, food, etc.. and normalization to NCBI Taxonomy, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags, same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries), same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags, same document collection)</li><li><a href="https://zenodo.org/records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts, different document collection)</li></ul><p><strong>Contact</strong></p><p>If you have any questions or suggestions, please contact us at:</p><p><br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p>
LivingNER corpus: Named entity recognition, normalization & classification of species, pathogens and food
<p><strong>LivingNER Gold Standard corpus (includes training, validation, test and background sets + MULTILINGUAL RESOURCES</strong>)</p><p> </p><p><strong>Please cite if you use this dataset:</strong></p><p>A. Miranda-Escalada, E. Farré-Maduell, S. Lima-López, D. Estrada, L. Gascó, M. Krallinger, Mention detection, normalization & classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources, <i>Procesamiento del Lenguaje Natural</i> (2022)</p><p>@article{amiranda2022nlp, title={Mention detection, normalization \& classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources}, author={Miranda-Escalada, Antonio and Farr{\'e}-Maduell, Eul{`a}lia and Lima-L{\'o}pez, Salvador and Estrada, Darryl and Gasc{\'o}, Luis and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, year={2022} }</p><p> </p><p><i><strong>1. Introduction</strong></i></p><p>The LivingNER Gold Standard corpus is a collection of<strong> 2000 clinical case reports</strong> covering a <strong>broad range of medical specialities</strong>, i.e. infectious diseases (including Covid-19 cases), cardiology, neurology, oncology, dentistry, pediatrics, endocrinology, primary care, allergology, radiology, psychiatry, ophthalmology, urology, internal medicine, emergency and intensive care medicine, tropical medicine, and dermatology <strong>annotated with species</strong> [SPECIES] (including <strong>living organisms</strong> and <strong>microorganisms</strong>) and <strong>infectious diseases</strong> [ENFERMEDAD] mentions. Species mentions include many <strong>pathogens</strong> and infectious agents, but also <strong>food</strong>, allergens, <strong>pets</strong> or other species, taxonomic groups and organisms of clinical relevance. </p><p>The LivingNER corpus has also annotations of mentions of <strong>humans</strong> (tag HUMAN), including the patients itself, <strong>family members</strong>, healhcare professionals or other persons mentioned in the case reports. Thus it can be useful to extract family history information of patients or information about the social and healthcare personal environment and interactions.</p><p>All mentions have been exhaustively manually mapped by experts to their corresponding <a href="https://www.ncbi.nlm.nih.gov/taxonomy"><strong>NCBI Taxonomy</strong></a> identifiers. </p><p>It was used for the <a href="https://temu.bsc.es/livingner/">LivingNER</a> Shared Task on pathogens and living beings detection and normalization in Spanish medical documents, which was celebrated as part of IberLEF 2022.</p><p> </p><p><i><strong>2. Training, validation, test and background sets</strong></i></p><p>The training set is composed of 1000 clinical case reports. The validation set includes 500 clinical case reports with the same characteristics and the test set includes 485. The background set is a collection of around 13k unannotated case reports that were originally added to prevent manual annotations in the test set during the competition and to create a Silver Standard.</p><p><i><strong>2.1 Annotations format</strong></i></p><p>Annotations and text files are distributed separately. The texts are in plain text (.txt in UTF-8) format, while the annotations are are distributed in a tab-separated file (.tsv) file with one row per annotation:</p><p>- For <strong>subtask 1 (LivingNER-Species NER track)</strong>, the .tsv file has the following columns:</p><ul><li>filename: document name</li><li>mark: identifier mention mark</li><li>label: mention type (SPECIES or HUMAN)</li><li>off0: starting position of the mention in the document</li><li>off1: ending position of the mention in the document</li><li>span: textual span</li></ul><p> - For <strong>subtask 2 (LivingNER-Species Norm track)</strong>, the .tsv file has the same columns as the previous one, plus:</p><ul><li>isH: whether the span is narrower than the NCBITax assigned code</li><li>isN: whether the mention corresponds to a nosocomial infection</li><li>iscomplex: whether the span has assigned a combination of NCBITax codes</li><li>NCBITax: mention code in the NCBI Taxonomy</li></ul><p>- For <strong>subtask 3 (LivingNER-Clinical IMPACT track)</strong>, the .tsv file has the following columns:</p><ul><li>filename</li><li>isPet (Yes/No)</li><li>PetIDs (NCBITaxonomy codes of pet & farm animals present in document)</li><li>isAnimalInjury (Yes/No)</li><li>AnimalInjuryIDs (NCBITaxonomy codes of animals causing injuries present in document)</li><li>IsFood (Yes/No)</li><li>FoodIDs (NCBITaxonomy codes of food mentions present in document)</li><li>isNosocomial (Yes/No)</li><li>NosocomialIDs (NCBITaxonomy codes of nosocomial species mentions present in document)</li></ul><p><i><strong>2.2 Important notes about subtask 3 (LivingNER-Clinical IMPACT track):</strong></i></p><ul><li><strong>Less clinical case reports</strong>. Subtask 3 (LivingNER-Clinical IMPACT track) contains half of the clinical case reports (500 in the training partition, 250 in the validation partition). The list of valid clinical case reports for task 3 is included in the data (train_files_task3.txt and validation_files_task3.txt)</li><li><strong>Enriched dataset.</strong> The GS format is the one described above (a TSV with one line per clinical case report). However, we believe participants may find useful and <strong>enriched dataset. </strong>Then, we provide an additional dataset, with the mentions of the NER track classified in the 4 Clinical impact categories (food, pet&farm animals, animals causing injuries and nosocomial). It is a TSV file with one row per annotation, and with the following columns: filename, mark, label, off0, off1, span, isPet, isAnimalInjury, isFood, isNosocomial, isH, iscomplex, code</li></ul><p> </p><p><i><strong>3. Multilingual resources</strong></i></p><p>We have generated the annotated training and validation sets in <strong>7 languages</strong>:</p><ul><li><i><strong>English</strong></i></li><li><i><strong>Portuguese</strong></i></li><li><i><strong>Catalan</strong></i></li><li><i><strong>Galician</strong></i></li><li><i><strong>Italian</strong></i></li><li><i><strong>French</strong></i></li><li><i><strong>Romanian</strong></i></li></ul><p> </p><p>The process was:</p><ol><li>The text files were translated with a neural machine translation system.</li><li>The annotations were translated with the same neural machine translation system.</li><li>The translated annotations were transferred to the translated text files using an annotation transfer technology.</li></ol><p>The text files are stored in the multilingual_resources/<strong>training-text-files</strong> and multilingual_resources/<strong>validation-text-files </strong>subfolders.</p><p>The annotated TSV files are stored in the multilingual_resources/<strong>annotation_transfer </strong>subfolder.</p><p>For the sake of comparison, we incorporate as well the annotations that resulted from the <a href="https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-11-85">LINNAEUS tool</a> in the multilingual_resources/<strong>linneaus</strong> subfolder.</p><p>If you want to visualize the multilingual resources, check out this Brat server: <a href="https://temu.bsc.es/mLivingNER/#/translations/">https://temu.bsc.es/mLivingNER/#/translations/</a></p><p>For instance, you can see the parallel annotations in <a href="https://temu.bsc.es/mLivingNER/diff.xhtml#/translations/en/annotation_transfer/train/casos_clinicos_cardiologia34?diff=/translations/fr/annotation_transfer/train/">English vs in French</a>, or <a href="https://temu.bsc.es/mLivingNER/diff.xhtml#/translations/cat/annotation_transfer/train/casos_clinicos_cardiologia35?diff=/gold-standard/train/">in Spanish (the gold standard) vs in Catalan.</a></p><p> </p><p><strong>Resources</strong></p><ul><li><a href="https://temu.bsc.es/livingner/"><strong>Task Web</strong></a></li><li><strong>Citation: </strong>A. Miranda-Escalada, E. Farré-Maduell, S. Lima-López, D. Estrada, L. Gascó, M. Krallinger, Mention detection, normalization & classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources, <i>Procesamiento del Lenguaje Natural</i> (2022)</li><li><a href="https://doi.org/10.5281/zenodo.6385162"><strong>Annotation guidelines</strong></a></li><li><a href="https://github.com/tonifuc3m/livingner-evaluation-library"><strong>Evaluation library</strong></a></li><li><a href="https://doi.org/10.5281/zenodo.6390506">LivingNER terminology</a></li><li><a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6444"><strong>Overview paper</strong></a></li><li><a href="https://ceur-ws.org/Vol-3202/"><strong>Proceedings participant papers</strong></a></li><li><a href="https://www.youtube.com/watch?v=8VcZw8ywyJY&list=PL5uSCzf1azhA_gMLC3DBZe6NvmMJiggTg"><strong>Youtube videos</strong></a></li><li><a href="https://www.slideshare.net/MartinKrallinger/mention-detection-normalization-classification-of-species-pathogens-humans-and-food-in-clinical-documents-overview-of-the-livingner-shared-task-and-resources-talk-at-iberlef-sepln-2022"><strong>LivingNER overview talk sides at IberLEF/SEPLN</strong></a></li></ul><p> </p><p><strong>License</strong></p><p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p><p><strong>Contact</strong></p><p>If you have any questions or suggestions, please contact us at:</p><p><br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p><p><strong>Additional resources and corpora</strong></p><p>If you are interested in SympTEMIST, you might want to check out these corpora and resources:</p><ul><li><a href="https://zenodo.org/records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/8413866">SympTEMIST</a> (Corpus of symptoms, sign and findings mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/4270158">PharmaCoNER</a> (Corpus of medications, drugs, chemical substances, genes, proteins and vaccine mentions and normalization, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization, different document collection)</li><li><a href="https://zenodo.org/records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization, different document collection)</li><li><a href="https://zenodo.org/records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI), differentdocument collection)</li><li><a href="https://zenodo.org/records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization, different document collection)</li><li><a href="https://zenodo.org/records/3837305">CodiESp</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries), different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, different document collection, some overlapping documents)</li></ul>
Gyroscope sensor data of shank motion during normal and barefoot walking
<p>The measurements were performed in closed and disturbance free space, where an unobstructed 10 m walkway was arranged. All tests were performed on a hard floor surface first with shoes that were adapt for walking (indoor sports shoes, sneakers etc.) and afterwards walking barefooted the same protocol. The subjects had clothing which did not restrict lower limb movement. In each test, the 10 m walking was repeated three times. Prior to testing, the procedure was demonstrated, and the sensors were carefully positioned to correct locations. The subjects were instructed to walk with their own natural walking velocity and to begin each 10 m walk from a completely stationary position.</p>
The Chilean Waiting List sub-Corpus with medical entities normalized to UMLS terminology
<p>A collection of 2000 medical referrals from the Chilean Waiting List Corpus, manually annotated with six entity types (Finding, Procedure, Disease, Family Member, Body Part, and Medication) and manually normalized to the Unified Medical Language System (UMLS).</p>
Figure 4. A normal female and a in Dichromothrips smithi (Zimmermann), a New Thrips Species Infesting Bamboo Orchids Arundina graminifolia (D. Don) Hochr. and Commercially Grown Orchids in Hawaii
Figure 4. A normal female and a teneral (newly molted, light colored) female D. smithi after storage in 70% ethanol (4a). Male and female D. smithi in ethanol (4b). The two males are lighter in color. Note that females stored in ethanol appear lighter brown with abdominal segments stretched out in comparison to the live female on blossom tissue pictured in 4c.
Cacao gene atlas: Additional file 13 (CPM normalized read counts) and Additional file 14 (fractional counts)
<p class="MsoNormal">A large dataset of replicated transcriptomes was developed to accelerate<span class="MsoCommentReference"> <em>T</em></span><em>heobroma cocoa</em> genomics research with the long-term goal of progressing breeding towards developing high-yielding elite varieties of cacao. RNAs were extracted and transcriptomes were sequenced from 123 different tissues and stages of development representing major organs and developmental stages of the cacao lifecycle. In addition, several experimental treatments and time courses were performed to measure gene expression in tissues responding to biotic and abiotic stressors. Samples were collected in replicates (3-5) to enable statistical analysis of gene expression levels for a total of 390 transcriptomes. We describe the creation of the atlas,and its global characterization and define sets of genes co-regulated in highly organ- and temporally-specific manners. To promote wider use of these data, all raw sequencing data, expression read mapping matrices, scripts, and other information used to create the resource are freely available online. A gene expression browser with a graphical user interface was developed to display gene expression patterns and to provide easy access of raw data and statistical analyses.</p>
Hybrid Formats - The New Normal?!
<p>How can hybrid meetings, workshops and trainings be conceptualized successfully?</p> <p>A hybrid workshop hosted by Instituto Mora, Mexico City.</p>
Datasets and Trained Models of the paper: The NFLikelihood: an unsupervised DNNLikelihood from Normalizing Flows
<p><strong>Training Data and Trained models corresponding to the publication: 'The NFLikelihood: an unsupervised DNNLikelihood from Normalizing Flows' (<a href="https://arxiv.org/abs/2309.09743">arXiv:2309.09743</a>).</strong></p> <p>The files cointain the relevant resources for 3 trained Likelihood functions: The Toy-Likelihood, the EW-Likelihood and the Flavor-Likelihood. In each corresponding directory, the training data is found in \data. The trained model and generated samples are found in \NFmodel.</p> <p>To reproduce the published results, git clone https://github.com/NF4HEP/NFLikelihoods, and plug in the provided resources in the corresponding directories of the code.</p> <p> </p> <p> </p>
Supplemental files: Q-Q plots and Shapiro-Wilk tests of normality
<p><em>Symphyotrichum</em> (Asteraceae) is a well-circumscribed genus on the basis of morphological and molecular data, but species boundaries remain poorly understood. Here, the species delimitation of the contentious <em>Symphyotrichum</em> <em>subulatum</em> group (<em>Symphyotrichum</em> subg. <em>Astropolium</em>) is illuminated using morphometric, phylogenomic, and geographical analyses. <em>Symphyotrichum mexicanum</em> sp. nov., a new species endemic to central Mexico, is described and distinguished from <em>Symphyotrichum</em> <em>expansum</em> based on its morphometric attributes, phylogenetic placement, geographic range, and ecological specialization.</p>
A Study of Long-Term Safety and Efficacy of Lanadelumab for Prevention of Acute Attacks of Non-histaminergic Angioedema With Normal C1-Inhibitor
ClinicalTrials.gov study NCT04444895. IPD Sharing: YES. Countries: 10. Publications: 1.
A Study of Lanadelumab in Teenagers and Adults to Prevent Acute Attacks of Non-histaminergic Angioedema With Normal C1-Inhibitor (C1-INH)
ClinicalTrials.gov study NCT04206605. IPD Sharing: YES. Countries: 10. Publications: 1.
The new normal? Redaction bias in biomedical science
Open the record for dataset details and reuse information.
The phageome in normal and inflamed human skin
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.