Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

470

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

470 results for “prediction of disease”

Learn how ShareScore rates datasets ↗
ClinicalTrials.gov36/100

Using Clinical Prediction Models to Improve Treatment for Patients With Chronic Obstructive Pulmonary Disease (COPD)

ClinicalTrials.gov study NCT05309356. IPD Sharing: YES. Countries: 1. Publications: 1.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Coronary Imaging and Metabolic Indicators-Based Risk Prediction Model for Coronary Artery Disease(CMI-RiskCAD)

ClinicalTrials.gov study NCT07353762. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

CloudConnect: Predictive And Retrospective Clinical Decision Support For Chronic Disease Management

ClinicalTrials.gov study NCT03676465. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
dryad36/100

Host traits and temperature predict biogeographic variation in seagrass disease prevalence

Open the record for dataset details and reuse information.

publicJan 2025View details →
dryad36/100

Data from: Activity patterns during the mating season predict sex-biased infections in an emerging fungal disease

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad36/100

Data from: Predicting disease risk areas through co-production of spatial models: the example of Kyasanur Forest Disease in India’s forest landscapes

Open the record for dataset details and reuse information.

publicMar 2020View details →
dryad36/100

Predicting spatio-temporal population patterns of Borrelia burgdorferi, the Lyme disease pathogen

Open the record for dataset details and reuse information.

publicAug 2022View details →
dryad36/100

Plasma neurofilament light for prediction of disease progression in familial frontotemporal lobar degeneration

Open the record for dataset details and reuse information.

publicFeb 2022View details →
zenodo32/100

Predicting disease in transition dairy cattle based on behaviors measured before calving

<p>Data set and code for the article &quot;Predicting disease in transition dairy cattle based on behaviors measured before calving&quot;</p>

opencc-by-4.0May 2020View details →
dryad32/100

Experimental evidence of warming-induced disease emergence and its prediction by a trait-based mechanistic model

<p>Predicting the effects of seasonality and climate change on the emergence and spread of infectious disease remains difficult, in part because of poorly understood connections between warming and the mechanisms driving disease. Trait-based mechanistic models combined with thermal performance curves arising from the Metabolic Theory of Ecology (MTE) have been highlighted as a promising approach going forward; however, this framework has not been tested under controlled experimental conditions that isolate the role of gradual temporal warming on disease dynamics and emergence. Here, we provide experimental evidence that a slowly warming host – parasite system can be pushed through a critical transition into an epidemic state. We then show that a trait-based mechanistic model with MTE functional forms can predict the critical temperature for disease emergence, subsequent disease dynamics through time, and final infection prevalence in an experimentally warmed system of <i>Daphnia </i>and a microsporidian parasite. Our results serve as a proof of principle that trait-based mechanistic models using MTE sub-functions can predict warming-induced disease emergence in data-rich systems – a critical step towards generalizing the approach to other systems.</p>

opencc-zeroSep 2020View details →
zenodo32/100

Prediction of cardiovascular diseases by integrating multi-modal features with machine learning methods

<p>Electrocardiogram (ECG) and Phonocardiogram (PCG) play important roles in early prevention and diagnosis of cardiovascular diseases. As the development of machine learning technique, detection of cardiovascular diseases from ECG and PCG has been attracted much attention.&nbsp;However, current available methods are mostly based on single data resource. It is desirable to develop efficient multi-modal machine learning methods to predict and diagnose cardiovascular diseases. In this study, we propose a novel multi-modal method for predicting cardiovascular diseases based on ECG and PCG features. By building up conventional neural networks, we extract ECG and PCG deep coding features respectively. The genetic algorithm is used to screen the combined features and obtain the best feature subset. Then support vector machine makes classification decision. Experimental results show that compared with using single-modal features ECG and PCG, the performance of this method reaches an AUC value of 0.936 when using multi-modal data resources.</p> <p>This&nbsp;dataset is&nbsp;developed&nbsp;from a real-world dataset which was assembled by PhysioNet/CinC Challenge&nbsp;in 2016. The original dataset can be downloaded from website (<a href="http://www.physionet.org/challenge/2016/">http://www.physionet.org/challenge/2016/</a>).</p>

opencc-by-4.0Nov 2020View details →
dryad32/100

Data from: Accurate genomic predictions for chronic wasting disease in U.S. white-tailed deer

<p>The geographic expansion of chronic wasting disease (CWD) in U.S. white-tailed deer (<em>Odocoileus virginianus</em>) has been largely unabated by best management practices, diagnostic surveillance, and depopulation of positive herds. Using a custom Affymetrix Axiom® single nucleotide polymorphism (SNP) array, we demonstrate that both differential susceptibility to CWD, and natural variation in disease progression, are moderately to highly heritable ( among farmed U.S. white-tailed deer, and that loci other than <em>PRNP</em> are involved. Genome-wide association analyses using 123,987 quality filtered SNPs for a geographically diverse cohort of 807 farmed U.S. white-tailed deer (n = 284 CWD positive; n = 523 CWD non-detect) confirmed the prion gene (<em>PRNP</em>; G96S) as a large-effect risk locus (<em>P</em>-value &lt; 6.3E-11), as evidenced by the estimated proportion of phenotypic variance explained (PVE ≥ 0.05), but also demonstrated that more phenotypic variance was collectively explained by loci other than <em>PRNP</em>.<em> </em>Genomic best linear unbiased prediction (GBLUP; n = 123,987 SNPs) with <em>k</em>-fold cross validation (<em>k</em> = 3; <em>k</em> = 5) and random sampling (n = 50 iterations) for the same cohort of 807 farmed U.S. white-tailed deer produced mean genomic prediction accuracies ≥ 0.81; thereby providing the necessary foundation for exploring a genomically-estimated CWD eradication program.</p>

opencc-zeroMar 2020View details →
dryad32/100

Data from: Detection error influences both temporal seroprevalence predictions and risk factors associations in wildlife disease models

Understanding the prevalence of pathogens in invasive species is essential to guide efforts to prevent transmission to agricultural animals, wildlife, and humans. Pathogen prevalence can be difficult to estimate for wild species due to imperfect sampling and testing (pathogens may not be detected in infected individuals and erroneously detected in individuals that are not infected). The invasive wild pig (Sus scrofa, also referred to as wild boar and feral swine) is one of the most widespread hosts of domestic animal and human pathogens in North America. We developed hierarchical Bayesian models that account for imperfect detection to estimate the seroprevalence of five pathogens (porcine reproductive and respiratory syndrome virus, pseudorabies virus, Influenza A virus in swine, Hepatitis E virus, and Brucella spp.) in wild pigs in the United States using a dataset of over 50,000 samples across nine years. To assess the effect of incorporating detection error in models, we also evaluated models that ignored detection error. Both sets of models included effects of demographic parameters on seroprevalence. We compared our predictions of seroprevalence to 40 published studies, only one of which accounted for imperfect detection. We found a range of seroprevalence among the pathogens with a high seroprevalence of pseudorabies virus, indicating significant risk to livestock and wildlife. Demographics had mostly weak effects, indicating that other variables may have greater effects in predicting seroprevalence. Models that ignored detection error led to different predictions of seroprevalence as well as different inferences on the effects of demographic parameters. Our results highlight the importance of incorporating detection error in models of seroprevalence and demonstrate that ignoring such error may lead to erroneous conclusions about the risk associated with pathogen transmission. When using opportunistic sampling data to model seroprevalence and evaluate risk factors, detection error should be included.

opencc-zeroAug 2019View details →
zenodo32/100

Supplemental Materials - Performance Figures for "A Model for Predicting the (re)-occurrence of a ≥40% eGFR Decline in a large Population-based cohort of Persons with or At-Risk of Chronic Kidney Disease " paper

<p>The zip file contains performance metrics figures for each dynamic Bayesian Network (DBN) model, stratified by comorbidities, race, CKD stages, and ethnicity.</p> <p>Contains:</p> <ul> <li>Stratified: Bootstrapping of 1000 iterations and 1000 samples with stratified proportions (as in the original population of the test set) of rapid eGFR decliners and non-decliners.</li> </ul> <p>&nbsp;</p> <p>Second zip contains DBN structures as matrices for 2 periods study entry to entry period and entry period to year 1 for all sites in 2 excel files.</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

TRACE Dataset: Predicting molecular mechanisms of hereditary diseases by using their tissue-selective manifestation

<p>Features dataset, as described in Simonovsky, Eyal, et al. "Predicting molecular mechanisms of hereditary diseases by using their tissue‐selective manifestation." <i>Molecular Systems Biology</i> (2023): e11407.</p><p>Article: https://doi.org/10.15252/msb.202211407</p><p>Code: https://github.com/eyalsim/trace</p>

openMay 2023View details →
zenodo32/100

Demo datasets for the protocol to identify shared transcriptional risks between diseases and compounds predicted to result in mutual benefit

<p>We present a computational protocol (https://github.com/ghbore/protocol-cancer-cvd-similarity), implemented as a Snakemake workflow, that was used in previous works (Gao et al., 2022; Baylis et al., 2023). This protocol allows researchers to identify shared transcriptional processes that drive disease and to screen existing compounds for mutual benefit. The protocol also includes a description of the pharmacovigilance study design used to validate the effect of novel compounds using electronic health records, where applicable. This repository bundles the datasets used in previous works as an example to run through the Snakemake workflow. These datasets include the TCGA cancer dataset, the STARNET and BiKE CVD datasets, and other dependent resources.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Disease trajectories in hospitalized COVID-19 patients are predicted by clinical and peripheral blood signatures representing distinct lung pathologies

<p><span>COVID-19 is characterized by a broad range of symptoms and disease trajectories. Understanding the correlation between clinical biomarkers and lung pathology over the course of acute COVID-19 is necessary to understand its diverse pathogenesis and inform more precise and effective treatments. Here, we present an integrated analysis of longitudinal clinical parameters, peripheral blood biomarkers, and lung pathology in COVID-19 patients from the Brazilian Amazon. We identified core clinical and peripheral blood signatures differentiating disease progression between recovered patients from severe disease and fatal cases. Signatures were heterogenous among fatal cases yet clustered into two patient groups: &ldquo;early death&rdquo; (&lt; 15 days of disease until death) and &ldquo;late death&rdquo; (&gt; 15 days). Progression to early death was characterized systemically and in lung histopathology by rapid, intense endothelial and myeloid activation/chemoattraction and presence of thrombi, associated with SARS-CoV-2<sup>+</sup> macrophages. In contrast, progression to late death was associated with fibrosis, apoptosis and abundant SARS-CoV-2<sup>+</sup> epithelial cells in post-mortem lung, with cytotoxicity, interferon and Th17 signatures only detectable in the peripheral blood 2 weeks into hospitalization. Progression to recovery was associated with higher lymphocyte counts, Th2 and anti-inflammatory-mediated responses. By integrating ante-mortem longitudinal systemic and spatial single-cell lung signatures, we defined an enhanced set of prognostic clinical parameters predicting disease outcome for guiding more precise and optimal treatments.</span><span> Finally, this study represents a major advance in the investigation of acute respiratory infections by integrating serial clinical data and peripheral blood samples with histopathological and </span><span>spatially-resolved single-cell </span><span>analyses of post-mortem lung samples.</span></p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Prediction of individual disease progression including parameter uncertainty in rare neurodegenerative diseases: the example of Autosomal-Recessive Spastic Ataxia Charlevoix Saguenay (ARSACS) - code and data sets

<p>This repository contains the scripts for the paper in revision to the AAPS J: Prediction of individual disease progression including parameter uncertainty in rare neurodegenerative diseases: the example of Autosomal-Recessive Spastic Ataxia Charlevoix Saguenay (ARSACS)&nbsp;</p> <p>Authors: Niels Hendrickx, MSc, France Mentr&eacute;, MD, PhD, Andreas Trasch&uuml;tz, MD, PhD, Cynthia Gagnon, PhD, Rebecca Sch&uuml;le, MD, ARCA Study Group, EVIDENCE-RND consortium, Matthis Synofzik, MD, Emmanuelle Comets, PhD</p> <p>A simulated dataset (<strong>simulated_arsacs.csv</strong>) has been included in the repository to make the code executable as a standalone. Four main scripts have been provided in addition with the present Readme describing the files. The repository also includes 3 R objects and 2 folders which will be overwritten when the scripts are run, and are included as examples of the expected outputs. The main scripts are:</p> <p>- <strong>Script_imputation_selection.R</strong>: runs the covariate selection method. It uses a simulated dataset provided in the depot. The multiple imputation model is hardcoded as an input to the mice package to generate 10 imputed datasets, saved in current_directory/imputed_data_sets/df_arsacs_mi_i.csv. The script then runs the covariate selection method. The script prints out the list of selected covariates and returns a saemixObject containing the fit of the selected covariate model.<br>&nbsp;After the script executes, a list will be saved with the name of the selected covariates in the current directory (an example is included under the name "cov_matrix_model.RData" in the repository), the output of the selection, containing the whole history of runs will be saved under "final_covariate_model.RData", the list of selected covariate names will be saved under "list_covariates.RData".</p> <p>- <strong>source_mi.R</strong>: contains the functions used by Script_imputation_selection.R</p> <p>- <strong>script_bootstrap_indfit.R</strong>: This script loads "cov_matrix_model.RData" containing the matrix of covariate effects (used by saemix) and "list_covariates.RData", the list of covariates included, fits the model on the imputed data sets and computes its bootstrap distribution for each imputed data set (in the script, using only 20 samples for computation time, saved in current_directory/bootstrap/boot.arsacs.case.mi.i). It then computes the mean parameter and relative standard error of each parameter. It then computes the conditional distribution of each patient in each bootstrap samples and returns a data frame of individual predictions. The script will then plot 4 indivudal predictions.&nbsp;</p> <p>-<strong> source_bootstrap.R</strong>: contains the functions used by script_bootstrap_indfit.R</p> <p>Both scripts need the saemix package to run, which we haven&rsquo;t included in the repository as it is freely available on the CRAN (https://cran.r-project.org/web/packages/saemix/index.html). Additional libraries we make use of in the code (MICE, tidyverse, ggplot2) also need to be installed prior to execution.&nbsp;<br>The R code provided can be further customised to be adapted to different scenarios.</p> <p>For the code to run, it is preferable to unzip the whole folder and set the working directory to the source file location as the script uses the "bootstrap" and "imputed_data_sets" sub-folders</p> <p>To execute this code, assuming the required libraries are available in the local R installation, please open an R session and run:<br>source("Script_imputation_selection.R") # for the covariate selection method (runtime: 3h on a &nbsp;i7-8565U laptop)<br>source("script_bootstrap_indfit.R") # to obtain individual trajectories (runtime: 1h on a &nbsp;i7-8565U laptop)</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Predicting wildlife susceptibility to infectious diseases atglobal scales

<p>Dataset included as supplementary material of &nbsp;the paper entitled https://doi.org/10.5281/zenodo.4914750. &nbsp;It contains phylogenetic, geographical and environmental distance for birds and bats, counts of incidence of &nbsp;<em>Plasmodium relictum</em> on birds, counts of incidence of West Nile Virus on birds and counts of incidence of coronavirus in bats, and susceptibility calculated by the random forest algorithm. Also we include the r scripts to run the models, and the outputs of the models after 1000 runs.&nbsp;</p>

opencc-by-4.0Oct 2021View details →
dryad32/100

Modeling management strategies for chronic disease in wildlife: predictions for the control of respiratory disease in bighorn sheep

<p>1. Controlling persistent infectious disease in wildlife populations is an on-going challenge for wildlife managers and conservationists worldwide.</p> <p>2. Here, we develop a dynamic pathogen transmission model capturing key features of M. ovipneumoniae infection, a major cause of population declines in North American bighorn sheep (Ovis canadensis). We explore the effects of model assumptions and parameter values on disease dynamics, including density versus frequency dependent transmission, the inclusion of a carrier class versus a longer infectious period, host survival rates, disease-induced mortality and recovery rates, and the epidemic growth rate.</p> <p>3. We compare the effectiveness of a suite of management actions following an epidemic, including test-and-remove, depopulation-and-reintroduction, range expansion, herd augmentation, and density reduction.</p> <p>4. Our results suggest that test-and-remove, depopulation-and-reintroduction, and range expansion have the potential to facilitate recovery of persistently infected bighorn sheep herds post-epidemic. By contrast, augmentation could lead to worse outcomes than those expected in the absence of management. Management that improves host survival or reduces disease-induced mortality are also likely to improve population size and persistence of chronically infected herds.</p> <p>5. Dynamic transmission models like the one employed here offer a structured, logical approach towards exploring hypotheses and can serve as a basis for planning field experiments and adaptive management. Models should be used iteratively with the field empirical approaches to triangulate on better approaches to wildlife management.</p>

opencc-zeroFeb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record