Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

470

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

470 results for “prediction of disease”

Learn how ShareScore rates datasets ↗
zenodo48/100

An integrated polygenic tool substantially enhances coronary artery disease prediction

<p>Summary-level CAD GWAS data generated by Genomics plc as presented in:</p> <p>Riveros-Mckay F. et al. An integrated polygenic tool substantially enhances coronary artery disease prediction. Circulation: Genomics and Precision Medicine (in press).&nbsp;</p> <p>If you have any questions or comments regarding these files, please contact Genomics plc at research@genomicsplc.com</p> <p>&nbsp;</p> <p>NOTES<br> -----------------------------<br> These analyses were carried out using the full UK Biobank imputation data release (v3b). Analyses were restricted to a subset of UK Biobank, described as &ldquo;Group I&rdquo; in the published paper.&nbsp; Group I, &ldquo;no PCE/QRISK3 available&rdquo;, included 114,196 European-ancestry individuals with missing data that prevented PCE or QRISK3 calculation.</p> <p>CAD case phenotypes were defined as described in the &ldquo;Phenotype definitions&rdquo; section of the paper&rsquo;s Supplementary Materials, using both prevalent (pre-baseline) and incident (post-baseline) events.</p> <p>All analyses included Age at assessment, sex, genotyping chip, and 10 principal components as covariates.&nbsp;</p> <p>We used plink2.0 logistic regression. For chromosome X variants males were treated as having 0 or 2 alternative alleles.&nbsp;</p> <p>The results are not adjusted for genomic control.</p> <p>&nbsp;</p> <p>DATA FILE CONTENT DESCRIPTION<br> -----------------------------<br> cpra Variant ID in &lsquo;CPRA&rsquo; format. Position reflects position in b37.&nbsp;<br> chrom Chromosome<br> pos Position in base pairs (b37, 1-based)<br> alt Alternative allele (effect allele)<br> beta Effect size (log odds ratio)<br> standard_error Standard error of beta&nbsp;<br> minus_log10_p Minus log(base 10) of P-value<br> ref Reference allele (non-effect allele)<br> ncase Number of cases<br> ncontrol Number of controls</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Weather-related Disease Prediction Dataset

<p>This dataset integrates medical symptoms and weather conditions to facilitate research into the prediction of diseases influenced by meteorological factors. The data spans a range of weather parameters alongside reported medical symptoms from a sizable number of anonymized individuals.</p> <p>&nbsp;</p> <p><strong>Variables:</strong></p> <p>Age: Age of the patient.</p> <p>Gender: Gender of the patient (encoded numerically).</p> <p>Temperature (C): Daily average temperature in Celsius.</p> <p>Humidity: Daily average humidity percentage.</p> <p>Wind Speed (km/h): Daily average wind speed in kilometers per hour.</p> <p>Symptoms: Various symptoms such as nausea, joint pain, abdominal pain, high fever, chills, fatigue, runny nose, pain behind the eyes, etc., encoded as binary values (1 for present, 0 for absent).</p> <p>Pre-existing Conditions: Conditions like asthma history, high cholesterol, diabetes, obesity, HIV/AIDS, nasal polyps, high blood pressure, encoded as binary values.</p> <p><strong>Data Collection Method:</strong></p> <p>The data was collected from anonymous medical records and corresponding local weather stations. All personal identifiers have been removed to ensure patient confidentiality and data privacy, adhering to ethical standards for medical data handling.</p> <p><strong>Usage Notes:</strong></p> <p>This dataset is intended for use in academic and research settings, especially in studies focused on the impact of weather on human health. It could be particularly useful for developing machine learning models to predict the likelihood of disease outbreaks based on weather patterns.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

# Single-cell network biology characterizes cell type gene regulation for drug repurposing and phenotype prediction in Alzheimer's disease

<p>Dysregulation of gene expression in Alzheimer&rsquo;s disease (AD) remains elusive, especially at the cell type level. Gene regulatory network, a key molecular mechanism linking transcription factors (TFs) and regulatory elements to govern target gene expression, can change across cell types in the human brain and thus serve as a model for studying gene dysregulation in AD. However, it is still challenging to understand how cell type networks work abnormally under AD. To address this, we integrated single-cell multi-omics data and predicted the gene regulatory networks in AD and control for four major cell types, excitatory and inhibitory neurons, microglia and oligodendrocytes. Importantly, we applied network biology approaches to analyze the changes of network characteristics across these cell types, and between AD and control. For instance, many hub TFs target different genes between AD and control (rewiring). Also, these networks show strong hierarchical structures in which top TFs (master regulators) are largely common across cell types, whereas different TFs operate at the middle levels in some cell types (e.g., microglia). The regulatory logics of enriched network motifs (e.g., feed-forward loops) further uncover cell type-specific TF-TF cooperativities in gene regulation. The cell type networks are highly modular and several network modules with cell-type-specific expression changes in AD pathology are enriched with AD-risk genes and putative targets of approved and pending AD drugs, suggesting possible cell-type genomic medicine in AD. Finally, using the cell type gene regulatory networks, we developed machine learning models to classify and prioritize additional AD genes. We found that top prioritized genes predict clinical phenotypes (e.g., cognitive impairment) with reasonable accuracy. Overall, this single-cell network biology analysis provides a comprehensive map linking genes, regulatory networks, cell types and drug targets and reveals dysregulated cell type gene dysregulatory mechanisms in AD.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Global predictions for the risk of establishment of Pierce's disease of grapevines

<p>Example to compute the risk of Pierce's Disease establishment from ERA5-Land temperature data. Related to the publication&nbsp;<em><a href="https://www.nature.com/articles/s42003-022-04358-w" target="_blank" rel="noopener">Global predictions for the risk of establishment of Pierce&rsquo;s disease of grapevines</a></em></p> <p>See <a href="https://github.com/agimenezromero/PierceDisease-GlobalRisk-Predictions" target="_blank" rel="noopener">GitHub</a> repo.</p>

openother-openMar 2022View details →
zenodo40/100

Development of a machine learning model to predict non- durable response to anti-TNF therapy in Crohn's disease using transcriptome imputed from genotypes

<p>This is the expression value predicted using PrediXcan version 7 to find a gene feature that can distinguish between patients with and without effect on infliximab.</p> <p>Among the various tissue models provided by PrediXcan v7, three models were selected and used: whole blood, Colon&nbsp;transverse, and terminal ileum of small intestine, and the predicted gene counts of each model were 6,294, 5,612 and 3,107.</p> <p>For each of the three models, predicted gene expression values and phenotype information per sample were submitted.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

BRAIN Journal-Prediction of Thyroid Disease Using Data Mining Techniques-Figure 2. Attributes of the classification models used in the experiments

<p>The authors used for their experiments a data set (UCI, 2016) containing 756 records about persons with thyroid dysfunctions. The classification model has 22 attributes; the class attribute is the target and it has three possible values: hypothyroidism, hyperthyroidism and normal. The current data set was extracted and preprocessed from the original file. A description of the attributes used in the experiments is given in Figure 2 (an extract from thyroid.arff test file).&nbsp;</p>

opencc-by-4.0Jun 2016View details →
zenodo40/100

BRAIN Journal-Prediction of Thyroid Disease Using Data Mining Techniques-Figure 3. KNIME Diagram

<p>The proposed KNIME diagram representing the data mining models is given in Figure 3. The nodes that constitute the model diagram are: ARFF Reader &ndash; the input node used to load the data set in arff format, Partitioning &ndash; the node with the role of data set partition (for training and for the validation of the classification model), Naive Bayes Learner and Decision Tree Learner &ndash; the nodes used to build the classification model, Naive Bayes Predictor and Decision Tree Predictor &ndash; the nodes used to validate the model, Scorer &ndash; the node reports a confusion matrix and the accompanying quality measures in its view, Normalizer &ndash; the data set are normalized to be able to apply the neural network models, Multilayer Perceptron and RBFNetwork &ndash; the nodes corresponding to the neural network classification models, Weka Predictor &ndash; a node implemented in Weka to validate the models.&nbsp;&nbsp;</p>

opencc-by-4.0Jun 2016View details →
zenodo40/100

BRAIN Journal-Prediction of Thyroid Disease Using Data Mining Techniques-Figure 1. Factors that Affect Thyroid Function (The Institute for Functional Medicine, 2014)

<p>&nbsp;In Figure 1 are presented the main factors that affect the thyroid function. It is obvious that factors such as stress, infection, toxins, trauma and certain medication are directly responsible for the improper production of thyroid hormones. Symptoms identification and the early detection of abnormal values of thyroid hormones after clinical investigation will help in establishing the proper diagnostic and to prescribe the right medication. The patient must periodically evaluate his clinical state in order to receive the treatment as long as he needs it.&nbsp;&nbsp;</p>

opencc-by-4.0Jun 2016View details →
zenodo40/100

iDPP@CLEF 2022 - Participants' repositories for the Intelligent Disease Prediction Progression Challenge

<p><a href="https://brainteaser.health/open-evaluation-challenges/idpp-2022/">iDPP@CLEF 2022</a> (Intelligent Disease Progression Prediction at CLEF) is a challenge organised by the&nbsp;<a href="https://brainteaser.health/">BRAINTEASER</a>&nbsp;Horizon 2020 project and co-located with&nbsp;<a href="https://clef2022.clef-initiative.eu/">CLEF 2022</a>&nbsp;(Conference and Labs of the Evaluation Forum).&nbsp;</p> <p>BRAINTEASER is a data science project that seeks to exploit the value of big data, including those related to health, lifestyle habits, and environment, to support patients with amyotrophic lateral sclerosis (ALS) and multiple sclerosis (MS) and their clinicians. Taking advantage of cost-efficient sensors and apps, BRAINTEASER will integrate large, clinical datasets that host both patient-generated and environmental data.</p> <p>The goal of iDPP@CLEF is to design and develop an evaluation infrastructure for AI algorithms able to:</p> <ul> <li> <p>Better describe<strong>&nbsp;disease mechanisms</strong>.</p> </li> <li> <p><strong>Stratify patients&nbsp;</strong>according to their phenotype assessed all over the disease evolution.</p> </li> <li> <p><strong>Predict disease progression</strong>&nbsp;in a probabilistic, time dependent fashion.</p> </li> </ul> <p>iDPP@CLEF 2022 offered the following tasks:</p> <ul> <li> <p><strong>Pilot Task 1 &ndash; Ranking Risk of Impairment</strong>: It focuses on ranking of patients based on the risk of impairment in specific domains. More in detail, we will use the ALSFRS-R scale to monitor speech, swallowing, handwriting, dressing/hygiene, walking and respiratory ability in time and will ask participants to rank patients based on time to event risk of experiencing impairment in each specific domain.</p> </li> <li> <p><strong>Pilot Task 2 &ndash; Predicting Time of Impairment</strong>: It refines Task 1 asking participants to predict when specific impairments will occur (i.e. in the correct time-window). In this regard, we assess model calibration in terms of the ability of the proposed algorithms to estimate a probability of an event close to the true probability within a specified time-window.</p> </li> <li> <p><strong>Position Papers Task 3 &ndash; Explainability of AI algorithms</strong>: We call for proposals of different visualization frameworks able to show the multivariate nature of the data and the model predictions in an explainable, possibly interactive, way.</p> </li> </ul> <p>&nbsp;</p> <p>This dataset contains the repositories of the participants to iDPP@CLEF 2022. These repositories contain the output, i.e. the predictions, produced by the participating systems as well as the performance scores for those systems.</p> <p>For additional information about iDPP@CLEF 2022, please see:</p> <ul> <li> <p>Guazzo, A., Trescato, I., Longato, E., Hazizaj, E., Dosso, D., Faggioli, G., Di Nunzio, G. M., Silvello, G., Vettoretti, M., Tavazzi, E., Roversi, C., Fariselli, P., Madeira, S. C., de Carvalho, M., Gromicho, M., Chi&ograve;, A., Manera, U., Dagliati, A., Birolo, G., Aidos, H., Di Camillo, B., and Ferro, N. (2022). Intelligent Disease Progression Prediction: Overview of iDPP@CLEF 2022. In Barr ́on-Cedeno, A., Da San Martino, G., Degli Es- posti, M., Sebastiani, F., Macdonald, C., Pasi, G., Hanbury, A., Potthast, M., Faggioli, G., and Ferro, N., editors, <em>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Thirteenth International Conference of the CLEF Association (CLEF 2022)</em>, pages 395&ndash;422. Lecture Notes in Computer Science (LNCS) 13390, Springer, Heidelberg, Germany.</p> </li> <li> <p>Guazzo, A., Trescato, I., Longato, E., Hazizaj, E., Dosso, D., Faggioli, G., Di Nunzio, G. M., Silvello, G., Vettoretti, M., Tavazzi, E., Roversi, C., Fariselli, P., Madeira, S. C., de Carvalho, M., Gromicho, M., Chi&ograve;, A., Manera, U., Dagliati, A., Birolo, G., Aidos, H., Di Camillo, B., and Ferro, N. (2022). Overview of iDPP@CLEF 2022: The Intelligent Disease Progression Prediction Challenge. In Faggioli, G., Ferro, N., Hanbury, A., and Potthast, M., editors, <em>CLEF 2022 Working Notes</em>, pages 1130&ndash; 1210. CEUR Workshop Proceedings (CEUR-WS.org), ISSN 1613-0073. <a href="https://ceur-ws.org/Vol-3180/paper-88.pdf">http://ceur-ws.org/Vol-3180/</a>.</p> </li> </ul>

opencc-by-4.0Dec 2022View details →
dryad40/100

Data from: Host disease tolerance predicts transmission probability for a songbird pathogen

Open the record for dataset details and reuse information.

publicAug 2024View details →
dryad40/100

Improving genomic prediction for plant disease using environmental covariates

Open the record for dataset details and reuse information.

publicAug 2025View details →
dryad36/100

Predicting spatio-temporal population patterns of Borrelia burgdorferi, the Lyme disease pathogen

<p>The causative bacterium of Lyme disease, <em>Borrelia burgdorferi</em>, expanded from an undetected human pathogen into the etiologic agent of the most common vector-borne disease in the United States over the last several decades. Systematic field collections of the tick vector reveal increases in the geographic range and population size of <em>B. burgdorferi</em> that coincided with increases in human Lyme disease incidence across New York State. Here we investigate the impact of environmental features on the population dynamics of <em>B. burgdorferi</em>. Analytical models developed using field collections of nearly 19,000 nymphal <em>Ixodes scapularis </em>and spatially- and temporally-explicit environmental features accurately explained the variation of <em>B. burgdorferi </em>population sizes across space and time. Importantly, the model identified environmental features that can be used to predict the biogeographical patterns of <em>B. burgdorferi-</em>infected ticks into future years and in previously unsampled areas. Forecasting the distribution and abundance of a pathogen at fine geographic scales offers a powerful strategy to mitigate a serious public health threat.</p>

opencc-zeroAug 2022View details →
zenodo36/100

Evolution of retinal degeneration and prediction of disease activity in relapsing and progressive multiple sclerosis

<p><span>Retinal optical coherence tomography has been identified as biomarker for disease progression in relapsing-remitting multiple sclerosis (RRMS), while the dynamics of retinal atrophy in progressive MS are less clear. We investigated retinal layer thickness changes in RRMS, </span><span>primary and secondary progressive MS (PPMS, SPMS)</span><span>, and their prognostic value for disease activity. Here, we analyzed 2651 OCT measurements of 195 RRMS, 87 SPMS, 125 PPMS patients, and 98 controls from five German MS centers after quality control. Peripapillary and macular retinal nerve fiber layer (pRNFL, mRNFL) thickness </span><span>predicted</span><span> future relapses in all MS and RRMS patients while mRNFL</span><span> and </span><span>ganglion cell-inner plexiform layer (GCIPL) </span><span>thickness predicted </span><span>future </span><span>MRI activity </span><span>in RRMS (mRNFL, GCIPL) and PPMS (GCIPL). mRNFL thickness </span><span>predicted </span><span>future disability progression </span><span>in PPMS.</span><span> </span><span>However, thickness change rates were subject to considerable amounts of measurement variability. In conclusion, retinal degeneration, most pronounced of pRNFL and GCIPL, occurs in all subtypes. Using the current state of technology, longitudinal assessments of retinal thickness may not be suitable on a single patient level.</span></p>

opencc-by-4.0May 2024View details →
zenodo36/100

Sway frequencies may predict postural instability in Parkinson's disease: Data

<p>Dataset with raw Center of Pressure (COP) and Center of Mass (COM) time series along with respective wavelet spectrograms.</p> <p>Recorded during 30 seconds of quiet stance. 10 trials per participant. Sampled at 50 Hz.</p> <p>18 individuals with Parkinson's disease, 15 healthy controls.</p> <p>Detailed data description can be found in the word file.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Speos: An ensemble graph representation framework to predict core genes for complex diseases (Datasets)

<p>the &quot;data.tar.gz&quot; tarball contains the unprocessed or minimally processed data used by Speos. If you intend to use the framework or want to inspect the data, download this part of the dataset.</p> <p>The &quot;final_datasets.tar.gz&quot; tarball contains tsv-formatted, processed data matrices which are directly used as input features for the ensemble models.&nbsp;</p> <p>There are two tsv-files&nbsp;per disease, one labeled &quot;normal&quot;, which contains the p input features alongside the gene identifiers and a column which indicates if the gene is labeled as Mendelian or not, and another file labeled &quot;with_n2v_vectors&quot;, which also contains the 100-dimensional vectors produced by Node2Vec so the N2V+MLP method can be reproduced with the exact same parameters.&nbsp;</p> <p>All files contain a header row which describes the column and no index column.</p> <p>The &quot;model_parameters.tar.gz&quot; tarball contains the model parameters for all ensemble models used to produce the candidate genes.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Vector species richness predicts local mortality rates by Chagas disease

<p>Vector species richness may drive the prevalence of vector-borne diseases by influencing pathogen transmission rates. The dilution effect hypothesis predicts that higher biodiversity reduces disease prevalence, but with inconclusive evidence. In contrast, the amplification effect hypothesis suggests that higher vector diversity may result in greater disease transmission by increasing and diversifying the transmission pathways. The relationship between vector diversity and pathogen transmission remains unclear and requires further study. Chagas disease is a vector-borne disease most prevalent in Brazil and transmitted by multiple species of Triatominae insect vectors, yet the drivers of spatial variation in its impact on human populations remain unresolved. We tested whether triatomine species richness, latitude, bioclimatic variables, human host population density, and socioeconomic variables predict Chagas disease mortality rates across over 5000 spatial grid cells covering all of Brazil. Results show that species richness of triatomine vectors is a good predictor of mortality rates caused by Chagas disease, which supports the amplification effect hypothesis. Vector richness and the impact of Chagas disease may also be driven by latitudinal components of climate and human socioeconomic factors. We provide evidence that vector diversity is a strong predictor of disease prevalence and give support to the amplification effect hypothesis.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Monitoring of postpartum body condition at the cow and herd levels: assessing explanatory and predictive power of disease risk models

<p>Objectives</p> <p>1- To define the herd threshold for cows with poor body condition based on its predictive capacity for disease risk at the herd level, and</p> <p>2- to estimate the impact measures on disease rates due to body condition indicators in transition period.</p> <p>Two commercial grazing dairy herds (Herd A=5.034 and herd B=7.965 lactations) from Argentinean Pampa region were used to perform a longitudinal retrospective study during a 4-year period (2014 &ndash;2017).Health, reproductive and body condition score (BCS) records were gathered. The BCS (5-point scale) was performed at calving and at the time of reproductive release. The difference between both measures of BCS was used to assess the body condition loss (∆BCS). All the cows not bred by 70 DIM were checked for anestrus.Calving cohorts of 21-day were defined at each herd and parity group through the entire study period. The frequency of cows with BCS&lt;3 or ∆BC&gt;-0.5 at each cohort were calculated and used to define quartiles through whole study period. Quartiles were used, one at a time, as threshold to dichotomize the cohorts to predict the risk that a cohort has a frequency of anestrus over the median.The higher AUC was used as selection criterium to determine the herd level threshold at each HERD and PARITY level.&nbsp;The population attributable fraction (AFP) of anestrus rate to body condition indicators at each cohort was calculated, for every HERD and PARITY level.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
ClinicalTrials.gov36/100

Predicting Disease Progression and/or Recurrence in Cancer

ClinicalTrials.gov study NCT04776837. IPD Sharing: YES. Countries: 1. Publications: 0.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Longitudinal Follow up to Assess Biomarkers Predictive of Emphysema Progression in Patients With COPD (Chronic Obstructive Pulmonary Disease)

ClinicalTrials.gov study NCT02719184. IPD Sharing: YES. Countries: 12. Publications: 2.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Predictive Tools for the Presence of Significant Coronary Artery Disease Requiring Intervention in Chronic Hemodialysis Patients

ClinicalTrials.gov study NCT06661941. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record