Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
447
datasets available to search
ShareScore release 0.9.0
Dataset results
447 results for “Model validation”
Data from: A probabilistic metric for the validation of computational models
Open the record for dataset details and reuse information.
External validation of EPIC’s Risk of Unplanned Readmission model, the LACE+ index and SQLape® as predictors of unplanned hospital readmissions: A monocentric, retrospective, diagnostic cohort study in Switzerland
Open the record for dataset details and reuse information.
Data from: Screening of health care workers for tuberculosis: development and validation of a new health economic model to inform practice
Open the record for dataset details and reuse information.
Data from: The importance of validating the demethylating effect of 5-aza-dC in model species
Open the record for dataset details and reuse information.
Data from: Validating two-dimensional leadership models on three-dimensionally structured fish schools
Open the record for dataset details and reuse information.
Data from: Validation of a rodent model of source memory
Open the record for dataset details and reuse information.
Periodic and heterogeneous solid and velocity data used to train and validate CNN models
Open the record for dataset details and reuse information.
Synthetic and reticulated foam solid and velocity data used to train and validate CNN models
Open the record for dataset details and reuse information.
Validation of noise models for single-cell transcriptomics
GEO Series GSE54695. Mus musculus. 16 samples. Type: Expression profiling by high throughput sequencing.
Comparative Validation of the D. melanogaster Encyclopedia of DNA Elements Transcript Models
GEO Series GSE44612. Drosophila mojavensis; Drosophila yakuba; Drosophila ananassae; Drosophila simulans; Drosophila virilis; Drosophila elegans; Drosophila biarmipes; Drosophila melanogaster; Drosophila pseudoobscura; Drosophila eugracilis; Drosophila takahashii; Drosophila ficusphila; Drosophila kikkawai; Drosophila bipectinata; Drosophila rhopaloa. 93 samples. Type: Expression profiling by high throughput sequencing.
Design of a gene signature to analyze the proinflammatory and the profibrotic polarization of murine fibroblasts and validation in an experimental murine model of scleroderma.
GEO Series GSE191223. Mus musculus. 15 samples. Type: Expression profiling by array.
CNA data from The human glioblastoma cell culture resource: validated cell models representing all molecular subtypes
GEO Series GSE72209. Homo sapiens. 48 samples. Type: Genome variation profiling by SNP array.
Validation of Baboon Pluripotent Cells as a Model for Translational Stem Cell Research [RNA-Seq]
GEO Series GSE165184. Papio anubis. 18 samples. Type: Expression profiling by high throughput sequencing.
DNA methylation analysis validates organoids as a viable model for studying human intestinal aging (methylation)
GEO Series GSE141254. Homo sapiens. 84 samples. Type: Methylation profiling by array.
Training and validation of a novel 4-miRNA ratio model (MiCaP) for prediction of post-operative outcome in prostate cancer patients
GEO Series GSE115402. Homo sapiens. 136 samples. Type: Expression profiling by RT-PCR.
Based on RNA-SEQ and experimental validation to explore the similarities and differences in catabolism and proliferation of chondrocyte inflammatory models between rats and mice.
GEO Series GSE163080. Mus musculus; Rattus norvegicus. 12 samples. Type: Expression profiling by high throughput sequencing.
Development and validation of a gene-based classification model for pN2 lung adenocarcinoma
GEO Series GSE282774. Homo sapiens. 58 samples. Type: Expression profiling by high throughput sequencing.
Creation and validation of models to predict response to chemotherapy before treatment in serous ovarian cancer
GEO Series GSE156699. Homo sapiens. 88 samples. Type: Expression profiling by high throughput sequencing.
Development, evaluation, and validation of machine learning models for COVID-19 detection based on routine blood tests
<p>The .xlsx dataset includes all patients used for training, internal-external and external validation: these can be distinguished by looking at the ID (first column) in the dataset: those in format Axxxx-<Date> are the data used for the training, those in the format 20xx are the data used for the internal-external validation, while the remaining data were used for external validation.</p> <p>As regards the features: for the Target feature the value 1 stands for "Positive to COVID-19" while the value 0 stands for "Negative to COVID-19"; while for the Sex feature the value 1 stands for "Male" while the value 0 stands for "Female".</p> <p>The full article is available at: https://www.degruyter.com/view/journals/cclm/ahead-of-print/article-10.1515-cclm-2020-1294/article-10.1515-cclm-2020-1294.xml.</p> <p>A pre-print version of the article is also available on MedrXiv: https://www.medrxiv.org/content/10.1101/2020.10.02.20205070v1</p> <p><strong>ABSTRACT</strong></p> <p><strong>Background</strong> The rRT-PCR test, the current gold standard for the detection of coronavirus disease (COVID-19), presents with known shortcomings, such as long turnaround time, potential shortage of reagents, false-negative rates around 15–20%, and expensive equipment. The hematochemical values of routine blood exams could represent a faster and less expensive alternative. </p> <p><strong>Methods</strong> Three different training data set of hematochemical values from 1,624 patients (52% COVID-19 positive), admitted at San Raphael Hospital (OSR) from February to May 2020, were used for developing machine learning (ML) models: the complete OSR dataset (72 features: complete blood count (CBC), biochemical, coagulation, hemogasanalysis and CO-Oxymetry values, age, sex and specific symptoms at triage) and two sub datasets (COVID-specific and CBC dataset, 32 and 21 features respectively). 58 cases (50% COVID-19 positive) from another hospital, and 54 negative patients collected in 2018 at OSR, were used for internal-external and external validation.</p> <p><strong>Results</strong> We developed five ML models: for the complete OSR dataset, the area under the receiver operating characteristic curve (AUC) for the algorithms ranged from 0.83 to 0.90; for the COVID-specific dataset from 0.83 15 to 0.87; and for the CBC dataset from 0.74 to 0.86. The validations also achieved good results: respectively, AUC 16 from 0.75 to 0.78; and specificity from 0.92 to 0.96. </p> <p><strong>Conclusions</strong> ML can be applied to blood tests as both an adjunct and alternative method to rRT-PCR for the fast and cost-effective identification of COVID-19-positive patients. This is especially useful in developing countries, or in countries facing an increase in contagions.</p>
Validation of the Bond et. al. (2010) SDSS-derived kinematic models for the Milky Way's disk and halo stars with Gaia Data Release 3 proper motion and radial velocity data
<p>We validate the Bond et. al. (2010) kinematic models for the Milky Way's disk and halo stars with Gaia Data Release 3 data. Bond et al. constructed models for stellar velocity distributions using stellar radial velocities measured by the Sloan Digital Sky Survey (SDSS) and stellar proper motions derived from SDSS and the Palomar Observatory Sky Survey astrometric measurements. These models describe velocity distributions as functions of position in the Galaxy, with separate models for disk and halo stars that were labeled using SDSS photometric and spectroscopic metallicity measurements. We find that the Bond et al. model predictions are in good agreement with recent measurements of stellar radial velocities and proper motions by the Gaia survey. In particular, the model accurately predicts the skewed non-Gaussian distribution of rotational velocity for disk stars and its vertical gradient, as well as the dispersions for all three velocity components. Additionally, the spatial invariance of velocity ellipsoid for halo stars when expressed in spherical coordinates is also confirmed by Gaia data at galacto-centric radial distances of up to 15 kpc.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.