Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

109

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

109 results for “prediction accuracy”

Learn how ShareScore rates datasets ↗
zenodo44/100

ZIRFs: zero-inflated random forests for estimating gene regulatory networks from single cell RNA-seq data (assessment of predictive accuracy and VIM stability)

<p>We developed a zero-inflated random forests (ZIRFs) algorithm to produce a metric of connection strength&nbsp;between regulator genes and target genes. This file contains SCENIC results for the aorta and diaphragm tissue data sets from the Tabula Muris Consortium results. SCENIC is a genetic regulatory network analysis published by Aibar et al. (2017). The purpose of the data sets and R source code are described by README files in each directory.</p>

opencc-by-3.0-usJul 2021View details →
zenodo40/100

Dataset: Accuracy of Motor Error Predictions for Different Sensory Signals

<p>Supplementary Data for&nbsp;<em><strong>Accuracy of Motor Error Predictions for Different Sensory Signals&nbsp;</strong></em>article</p> <p>Dataset associated with the following publication:</p> <p>Joch, M., Hegele, M., Maurer, H., M&uuml;ller, H., &amp; Maurer, L. K. (2018). Accuracy of Motor Error Predictions for Different Sensory Signals.&nbsp;<em>Frontiers in Psychology</em>,&nbsp;<em>9&nbsp;</em>(August), 1&ndash;13. https://doi.org/10.3389/fpsyg.2018.01376</p>

opencc-by-4.0Jun 2020View details →
dryad40/100

Simulated Herbarium data for testing the accuracy with which specimen data can predict the timing and duration of population-level flowering displays

<p>This dataset provides code and example data for simulating specimen collections of flowering plants across North America, and for developing phenological predictions of population-level flowering onset and termination for these data.  It further presents code for assessing the accuracy of these predictions relaticve to known (simulated) population-level flowering dates at the location of each collection.</p>

opencc-zeroMar 2024View details →
zenodo40/100

Data for "High-accuracy determination of Paul-trap stability parameters for electric-quadrupole-shift prediction", J. Appl. Phys. 132, 124401 (2022)

<p>Data required to reproduce the key results in: &quot;High-accuracy determination of Paul-trap stability parameters for electric-quadrupole-shift prediction&quot;, J. Appl. Phys. <strong>132</strong>, 124401 (2022). <a href="https://doi.org/10.1063/5.0106633">https://doi.org/10.1063/5.0106633</a></p> <ul> <li>The file &quot;sec_freq.dat&quot; contains measured secular frequencies, the rf frequency, the applied bias voltages, and the MJD of the measurement: <ul> <li>Figures 4-5 use rows 31-33 of this data.</li> <li>Figure 6 uses all data corresponding to -1.1e-3 &lt; <em>a</em><sub>x</sub> &lt; -0.6e-3.</li> <li>Figure 7(a) uses all the data.</li> </ul> </li> <li>The file &quot;data2021-12-21_MJD.txt&quot; contains the secular frequency data used to derive Eq. (16) and plot Figure 8.</li> <li>The file &quot;RF_monitor_rectifier_1d.txt&quot; contains the rectified monitor voltage used in Figure 8.</li> <li>The file &quot;Temperature_108_1d.txt&quot; contains the helical-resonator temperature used in Figure 8.</li> <li>The file &quot;Fig9_EQS.dat&quot; contains the measured electric quadrupole shift (EQS) used for Figure 9.</li> <li>The file &quot;interleavedEQS.dat&quot; contains the data from the interleaved EQS measurement used to determine a lower value of 1070 for the cancellation factor.</li> </ul>

opencc-by-4.0Oct 2022View details →
zenodo40/100

MACHINE LEARNING APPROACHES FOR DEMAND FORECASTING: THE IMPACT OF CUSTOMER SATISFACTION ON PREDICTION ACCURACY

<p><span>This study investigates the effectiveness of various machine learning models in predicting product demand based on customer satisfaction data. Four models&mdash;Linear Regression, Random Forest, Gradient Boosting, and Support Vector Machine (SVM)&mdash;were evaluated using performance metrics, including Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R&sup2; score. The results indicate that Gradient Boosting achieved the highest accuracy, with an MAE of 2.56, MSE of 12.75, RMSE of 3.57, and R&sup2; score of 0.82, effectively capturing the complex, non-linear relationships inherent in customer satisfaction factors. Random Forest also demonstrated strong performance, while Linear Regression and SVM showed limitations in handling intricate datasets. These findings underscore the importance of utilizing advanced machine learning techniques for accurate demand forecasting, highlighting the critical role of customer satisfaction data in enhancing predictive capabilities. The insights gained from this research can guide organizations in optimizing inventory management and improving customer satisfaction in a rapidly evolving market.</span></p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

BRCA1-specific machine learning model predicts variant pathogenicity with high accuracy - Supplementary material

<p>Figure S1: Distribution of the reviewed 141&nbsp;<em>BRCA1</em>&nbsp;missense variants; Figure S2:&nbsp;The Shapely values for the&nbsp;<em>BRCA1</em>&nbsp;XGBoost models; Figure S3: The Shapely values of the&nbsp;<em>BRCA1</em>&nbsp;XGBoost model used to predict the functional assays&rsquo; results for variants of uncertain significance;&nbsp;Table S1: The receiver operating characteristic (ROC) curve analysis for the different in silico predictions; Table S2: Cross validation of the BRCA1 model in 5 different random training and test samples; Table S3: Pathogenicity prediction and prioritization of the 31,058 unreviewed BRCA1 variants from the BRCA Exchange database.</p>

opencc-by-4.0Apr 2023View details →
dryad40/100

Simulated Herbarium data for testing the accuracy with which specimen data can predict the timing and duration of population-level flowering displays

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad36/100

Improving genome-wide association discovery and genomic prediction accuracy in biobank data

<p>Genetically informed, deep-phenotyped biobanks are an important research resource and it is imperative that the most powerful, versatile, and efficient analysis approaches are used. Here, we apply our recently developed Bayesian grouped mixture of regressions model (GMRM) in the UK and Estonian Biobanks and obtain the highest genomic prediction accuracy reported to date across 21 heritable traits. When compared to other approaches, GMRM accuracy was greater than annotation prediction models run in the LDAK or LDPred-funct software by 15% (SE 7%) and 14% (SE 2%), respectively, and was 18% (SE 3%) greater than a baseline BayesR model without single-nucleotide polymorphism (SNP) markers grouped into minor allele frequency–linkage disequilibrium (MAF-LD) annotation categories. For height, the prediction accuracy R 2 was 47% in a UK Biobank holdout sample, which was 76% of the estimated h SNP 2 . We then extend our GMRM prediction model to provide mixed-linear model association (MLMA) SNP marker estimates for genome-wide association (GWAS) discovery, which increased the independent loci detected to 16,162 in unrelated UK Biobank individuals, compared to 10,550 from BoltLMM and 10,095 from Regenie, a 62 and 65% increase, respectively. The average χ<sup>2</sup> value of the leading markers increased by 15.24 (SE 0.41) for every 1% increase in prediction accuracy gained over a baseline BayesR model across the traits. Thus, we show that modeling genetic associations accounting for MAF and LD differences among SNP markers, and incorporating prior knowledge of genomic function, is important for both genomic prediction and discovery in large-scale individual-level studies.</p>

opencc-zeroSep 2022View details →
zenodo36/100

Mapping the relative accuracy of cross-ancestry prediction

<p>GWAS for six traits for SNPs from the UK Biobank arrays set, six traits for SNPs from the HapMap set, and three traits for SNPs from the ARIC study set.</p>

opencc-by-4.0Sep 2024View details →
dryad36/100

Variable prediction accuracy of polygenic scores within an ancestry group

<p class="formanuscript">Fields as diverse as human genetics and sociology are increasingly using polygenic scores based on genome-wide association studies (GWAS) for phenotypic prediction. However, recent work has shown that polygenic scores have limited portability across groups of different genetic ancestries, restricting the contexts in which they can be used reliably and potentially creating serious inequities in future clinical applications. Using the UK Biobank data, we demonstrate that even within a single ancestry group (i.e., when there are negligible differences in linkage disequilibrium or in causal alleles frequencies), the prediction accuracy of polygenic scores can depend on characteristics such as the socio-economic status, age or sex of the individuals in which the GWAS and the prediction were conducted, as well as on the GWAS design. Our findings highlight both the complexities of interpreting polygenic scores and underappreciated obstacles to their broad use.</p>

opencc-zeroFeb 2020View details →
zenodo36/100

Comparative Accuracies of Models for Drag Prediction During Geomagnetically Disturbed Periods: A First Principles Model versus Empirical Models

<p>This dataset contains observational and TIEGCM simulation data used for a manuscript that is being submitted to the <em>Space Weather</em> journal. The abstract for the study follows: We examine the accuracy of density prediction by the first principals model Thermosphere Ionsosphere Electrodynamics General Circulation Model (TIEGCM) developed by the National Center for Atmospheric Research and compare it to the accuracy of three empirical models: Jacchia 71, the Naval Research Laboratory Mass Spectrometer Incoherent Scatter Extended 2000 (NRLMSIS), Jacchia 1971 and Jacchia-Bowman 2008. Comparisons are made for three large storms: the October 2003 storm, the March 2013 storm, and the March 2015 storm. To evaluate the accuracy of these models we use tracking data for nine space objects in low earth orbit (three for each storm). Additionally, and evaluate the accuracy of the TIEGCM and NRLMSIS with data from high precision accelerometers on the Challenging Minisatellite Payload (CHAMP) and Gravity field and Circulation Explorer (GOCE) satellites. The goal is to assess the use of a first principles model as a potential tool for forecasting satellite drag during large magnetic storms. We find that the TIEGCM accuracy is substantially better than for the Jacchia 71 and NRLMSIS models. The accuracies of the TIEGCM and JB2008 models are similar, but overall the TIEGCM is more accurate. We found smaller mean percentage differences for TIEGCM versus CHAMP than for NRLMIS for the Halloween Storm and smaller differences than results published for JB2008 and the assimilative model HASDM. The empirical models are at present more practical for operational purposes, but the first principles TIEGCM was developed as a research model and with a greater focus on operational use offers the potential for improved utility during stressing conditions.</p>

opencc-by-4.0Oct 2022View details →
ClinicalTrials.gov36/100

Correlation of Predictive Accuracy of PREDICT Version 2.2 of Indian Women With Operable Breast Cancer

ClinicalTrials.gov study NCT04985253. IPD Sharing: NO. Countries: 1. Publications: 17.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Accuracy of Venous Excess Ultrasound Score at Hospital Admission to Predict Acute Kidney Injury

ClinicalTrials.gov study NCT07040696. IPD Sharing: NO. Countries: 1. Publications: 4.

closedIPD-NOFeb 2026View details →
dryad36/100

Variable prediction accuracy of polygenic scores within an ancestry group

Open the record for dataset details and reuse information.

publicFeb 2020View details →
dryad36/100

Improving genome-wide association discovery and genomic prediction accuracy in biobank data

Open the record for dataset details and reuse information.

publicSep 2022View details →
dryad36/100

Data from: Hybrid machine learning approach to zero-inflated data improves accuracy of dengue prediction

Open the record for dataset details and reuse information.

publicDec 2025View details →
dryad32/100

Data from: Genomic prediction accuracies in space and time for height and wood density of Douglas-fir using exome capture as the genotyping platform

Background Genomic selection (GS) can offer unprecedented gains, in terms of cost efficiency and generation turnover, to forest tree selective breeding; especially for late expressing and low heritability traits. Here, we used: 1) exome capture as a genotyping platform for 1372 Douglas-fir trees representing 37 full-sib families growing on three sites in British Columbia, Canada and 2) height growth and wood density (EBVs), and deregressed estimated breeding values (DEBVs) as phenotypes. Representing models with (EBVs) and without (DEBVs) pedigree structure. Ridge regression best linear unbiased predictor (RR-BLUP) and generalized ridge regression (GRR) were used to assess their predictive accuracies over space (within site, cross-sites, multi-site, and multi-site to single site) and time (age-age/ trait-trait). Results The RR-BLUP and GRR models produced similar predictive accuracies across the studied traits. Within-site GS prediction accuracies with models trained on EBVs were high (RR-BLUP: 0.79–0.91 and GRR: 0.80–0.91), and were generally similar to the multi-site (RR-BLUP: 0.83–0.91, GRR: 0.83–0.91) and multi-site to single-site predictive accuracies (RR-BLUP: 0.79–0.92, GRR: 0.79–0.92). Cross-site predictions were surprisingly high, with predictive accuracies within a similar range (RR-BLUP: 0.79–0.92, GRR: 0.78–0.91). Height at 12 years was deemed the earliest acceptable age at which accurate predictions can be made concerning future height (age-age) and wood density (trait-trait). Using DEBVs reduced the accuracies of all cross-validation procedures dramatically, indicating that the models were tracking pedigree (family means), rather than marker-QTL LD. Conclusions While GS models' prediction accuracies were high, the main driving force was the pedigree tracking rather than LD. It is likely that many more markers are needed to increase the chance of capturing the LD between causal genes and markers.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Diagnostic accuracy of presepsin in predicting bacteremia in elderly patients admitted to the emergency department: prospective study in Japan

<p class="CxSpFirst"><span><b>Objective</b></span></p> <p class="CxSpMiddle"><span><a name="_Hlk1399924">Early detection of bacteremia in the elderly is needed in the emergency department (ED). </a></span></p> <p class="CxSpMiddle"><span><a name="_Hlk11878757"><b>Design, Setting, and Participants</b></a></span></p> <p class="CxSpMiddle"><span>Prospective study in Japan, single center trial in patients who satisfied the sepsis criteria was conducted between September 2014 and March 2016. Forty-six elderly patients aged ≥70 years were included.</span></p> <p class="CxSpMiddle"><span><b>Interventions</b></span></p> <p class="CxSpMiddle"><span>Blood sampling to evaluate C-reactive protein (CRP), procalcitonin (PCT) and presepsin plasma levels; two sets of blood sampling for bacterial cultures; and evaluations of the Sequential Organ Failure Assessment (SOFA) and Acute Physiology and Chronic Health Evaluation (APACHE II) scores were performed upon arrival at the ED. The results were compared between patients with bacteremia and those without bacteremia.</span></p> <p class="CxSpMiddle"><span><a name="_Hlk11878769"><b>Main Outcome Measure</b></a></span></p> <p class="CxSpMiddle"><span>The accuracy of detecting bacteremia</span></p> <p class="CxSpMiddle"><span><b>Results</b></span></p> <p class="CxSpMiddle"><span>The presepsin value was significantly higher in the bacteremia group than in the non-bacteremia group (866.6 ± 184.6 pg/mL vs. 639.9 ± 137.1 pg/mL, p = 0.03). The PCT and CRP did not significantly differ between the groups. The area under the receiver-operating-characteristic curve (AUC) values were not significantly different among presepsin (0.69), PCT (0.61), and CRP (0.53). Multivariate analysis showed that presepsin was independently associated with bacteremia (odds ratio, 8.84; 95% confidence interval, 0.95–81.79; p = 0.02).</span></p> <p class="CxSpMiddle"><span><b>Conclusion<a name="_Hlk1399945"></a></b></span></p> <p class="CxSpMiddle"><span>Presepsin could be a good biomarker to predict bacteremia in elderly patients with sepsis criteria admitted to the ED.</span></p>

opencc-zeroNov 2019View details →
dryad32/100

Data from: Combining high-throughput phenotyping and genomic information to increase prediction and selection accuracy in wheat breeding

Genomics and phenomics have promised to revolutionize the field of plant breeding. The integration of these two fields has just begun and is being driven through big data by advances in next-generation sequencing and developments of field-based high-throughput phenotyping (HTP) platforms. Each year the International Maize and Wheat Improvement Center (CIMMYT) evaluates tens-of-thousands of advanced lines for grain yield across multiple environments. To evaluate how CIMMYT may utilize dynamic HTP data for genomic selection (GS), we evaluated 1170 of these advanced lines in two environments, drought (2014, 2015) and heat (2015). A portable phenotyping system called 'Phenocart' was used to measure normalized difference vegetation index and canopy temperature simultaneously while tagging each data point with precise GPS coordinates. For genomic profiling, genotyping-by-sequencing (GBS) was used for marker discovery and genotyping. Several GS models were evaluated utilizing the 2254 GBS markers along with over 1.1 million phenotypic observations. The physiological measurements collected by HTP, whether used as a response in multivariate models or as a covariate in univariate models, resulted in a range of 33% below to 7% above the standard univariate model. Continued advances in yield prediction models as well as increasing data generating capabilities for both genomic and phenomic data will make these selection strategies tractable for plant breeders to implement increasing the rate of genetic gain.

opencc-zeroDec 2017View details →
dryad32/100

Validation of the predictive accuracy of health-state utility values based on the Lloyd model for metastatic or recurrent breast cancer in Japan

<p>Although there is a lack of data on health-state utility values (HSUVs) for calculating quality-adjusted life years in Japan, Cost-utility analysis has been introduced by the Japanese government to inform decision-making in the medical field since 2016. This study aimed to determine whether the Lloyd model which was a predictive model of HSUVs for metastatic breast cancer (MBC) patients in the United Kingdom can accurately predict actual HSUVs for Japanese patients with MBC. The prospective observational study, followed by the validation study of the clinical predictive model.<b> </b>Forty-four Japanese patients with MBC were studied at 336 survey points. This study consisted of two phases. In the first phase, we constructed a database of clinical data prospectively and HSUVs for Japanese patients with MBC to evaluate the predictive accuracy of HSUVs calculated using the Lloyd model. In the second phase, Bland-Altman analysis was used to determine how accurately predicted HSUVs (based on the Lloyd model) correlated with actual HSUVs obtained using the EuroQol 5-Dimension 5-Level questionnaire, a preference-based measure of HSUVs in patients with MBC. In the Bland-Altman analysis, the mean difference between HSUVs estimated by the Lloyd model and actual HSUVs, or systematic error, was -0.106. The precision was 0.165. The 95% limits of agreement ranged from -0.436 to 0.225. The t value was 4.6972, which was greater than the t value with 2 degrees of freedom at the 5% significance level (p=0.425). There were acceptable degrees of fixed and proportional errors associated with the prediction of HSUVs based on the Lloyd model for Japanese patients with MBC. We recommend that sensitivity analysis be performed when conducting cost-effectiveness analyses with HSUVs calculated using the Lloyd model.</p>

opencc-zeroNov 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record