Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
289
datasets available to search
ShareScore release 0.9.0
Dataset results
289 results for “Genomic prediction”
Data from: Multi-generation genomic prediction of maize yield using parametric and non-parametric sparse selection indices
<p>Genomic prediction models are often calibrated using multi-generation data. Over time, as data accumulates, training data sets become increasingly heterogeneous. Differences in allele frequency and linkage disequilibrium patterns between the training and prediction genotypes may limit prediction accuracy. This leads to the question of whether all available data or a subset of it should be used to calibrate genomic prediction models. Previous research on training set optimization has focused on identifying a subset of the available data that is optimal for a given prediction set. However, this approach does not contemplate the possibility that different training sets may be optimal for different prediction genotypes. To address this problem, we recently introduced a sparse selection index (SSI) that identifies an optimal training set for each individual in a prediction set. Using additive genomic relationships, the SSI can provide increased accuracy relative to genomic-BLUP (GBLUP). Non-parametric genomic models using Gaussian kernels (KBLUP) have, in some cases, yielded higher prediction accuracies than standard additive models. Therefore, here we studied whether combining SSIs and kernel methods could further improve prediction accuracy when training genomic models using multi-generation data. Using four years of doubled haploid maize data from the International Maize and Wheat Improvement Center (CIMMYT), we found that when predicting grain yield the KBLUP outperformed the GBLUP, and that using SSI with additive relationships (GSSI) lead to 5-17% increases in accuracy, relative to the GBLUP. However, differences in prediction accuracy between the KBLUP and the kernel-based SSI were smaller and not always significant.</p>
Incorporation of soil-derived covariates in progeny testing and line selection to enhance genomic prediction accuracy in soybean breeding
<p>The availability of high-dimensional molecular markers has allowed plant breeding programs to maximize their efficiency through the genomic prediction of a phenotype of interest. Yield is a highly complex and quantitative trait whose expression is sensitive to environmental stimuli. In this research, we investigated the potential of incorporating soil texture and its interaction with molecular markers through covariance structures to enhance predictive ability. A total of 797 advanced soybean breeding lines derived from 367 unique bi-parental populations were genotyped using the Illumina Infinium BARCSoySNP6K BeadChip and tested for yield for five years in Tiptonville silt loam, Sharkey clay, and Malden fine sand environments. Four statistical models were considered, including a default GBLUP model (M1), a reaction norm model (M2) accounting for the interaction between molecular markers and the environment (GE), an expansion of M2 including soil type (S), and the interaction between soil type and molecular markers (GS) (M3), and an alternative version of M3 without the GE term. Four cross-validation scenarios simulating progeny testing and line selection were implemented (CV2, CV1, CV0, and CV00). Across environments, the addition of GS in M3 decreased the amount of variability captured by both the environment (-30.4%) and residual (-39.2%) terms as compared to M1. Within environments, the GS term in M3 reduced the variability captured by the residual term by roughly 60% and 30% when compared to M1 and M2, respectively. M3 outperformed all models in CV2 (0.577), CV1 (0.480), and CV0 (0.488). The addition of soil texture seems to structure the environment term revealing its components that could enhance or hinder the predictability of a model. The availability of soil texture before the growing season may maximize the functionality of covariance structures, particularly in scenarios with untested genotypes in untested environments. Genomic selection can optimize the efficiency of a soybean breeding program by allowing the reconsideration of field experimental design, allocation of resources, reduction of preliminary trials, and shortening of the breeding cycle.</p>
Oopsacas minuta alternative masked genomes and predicted proteomes
<p>Oopsacas minuta alternative versions.</p> <p>The genome was masked using the Dfam TE tool container (https://github.com/Dfam-consortium/TETools).</p> <p>oopsacas_minuta_hardmasked.fna : genome hardmasked (with N). Low complexity regions are also masked.</p> <p>oopsacas_minuta_hardmasked_nolow.fna: genome hardmasked (with N). Low complexity regions are not masked.</p> <p>oopsacas_minuta_softmasked.fna: genome softmasked (lower-case letters). Low complexity regions are also masked.</p> <p>oopsacas_minuta_softmasked_nolow.fna: genome softmasked (lower-case letters). Low complexity regions are not masked.</p> <p>Proteins were predicted using Braker v1.9.</p> <p>oopsacas_minuta_hardmasked.faa : from oopsacas_minuta_hardmasked.fna</p> <p>oopsacas_minuta_hardmasked_nolow.faa: fromoopsacas_minuta_hardmasked_nolow.fna</p> <p>oopsacas_minuta_softmasked.faa: from oopsacas_minuta_softmasked.fna</p> <p>oopsacas_minuta_softmasked_nolow.faa: from oopsacas_minuta_softmasked_nolow.fna</p> <p> </p> <p> </p>
Genome-wide association and genomic prediction for a reproductive index summarizing fertility outcomes in U.S. Holsteins
<p>Subfertility represents one major challenge to enhancing dairy production and efficiency. Herein, we use a reproductive index (RI) expressing the predicted probability of pregnancy following artificial insemination with Illumina 778K genotypes to perform single and multi-locus genome-wide association analyses (GWAA) on 2,448 geographically diverse U.S. Holstein cows and produce genomic heritability estimates. Moreover, we use genomic best linear unbiased prediction (GBLUP) to investigate the potential utility of the RI by performing genomic predictions with cross-validation. Notably, genomic heritability estimates for the U.S. Holstein RI were moderate ( 0.1654± 0.0317 – 0.2550 ± 0.0348), while single and multi-locus GWAA revealed overlapping quantitative trait loci (QTL) on BTA6 and BTA29, including known QTL for daughter pregnancy rate (DPR) and cow conception rate (CCR). Multi-locus GWAA revealed seven additional QTL, including one on BTA7 (60 Mb) which is adjacent to a known heifer conception rate (HCR) QTL (59 Mb). Positional candidate genes for the detected QTL included male and female fertility loci (i.e., spermatogenesis, oogenesis), meiotic and mitotic regulators, and genes associated with immune response, milk yield, enhanced pregnancy rates, and the reproductive-longevity pathway. Based on the proportion of phenotypic variance explained (PVE), all detected QTL (n = 13; P ≤ 5e<sup>-05</sup>) were estimated to have moderate (1.0% < PVE ≤ 2.0%) or small effects (PVE ≤ 1.0%) on the predicted probability of pregnancy. Genomic prediction using GBLUP with cross-validation (<em>k</em> = 3) produced mean predictive abilities (0.1692–0.2301) and mean genomic prediction accuracies (0.4119–0.4557) that were similar to bovine health and production traits previously investigated.</p>
Predicted protein functions for 2031 Saccharomyces cerevisiae genome assemblies
<p>Predicted protein functions for 2031 Saccharomyces cerevisiae genome assemblies</p>
Genomic Analysis to Identify a Predictive Biomarker for Immunotherapy
ClinicalTrials.gov study NCT03578185. IPD Sharing: Not stated. Countries: 1. Publications: 5.
Exploration and Determination of Genomic Markers Predictive of Uterine Atony
ClinicalTrials.gov study NCT03413917. IPD Sharing: NO. Countries: 1. Publications: 1.
Analysis of Prognostic and Predictive Genomic Signatures Using Archival Paraffin-embedded Breast Tumor - a Pilot Study
ClinicalTrials.gov study NCT01247467. IPD Sharing: Not stated. Countries: 1. Publications: 2.
MyLeukoMAP™ Genomic Survival Prediction Assay Pivotal Clinical Study
ClinicalTrials.gov study NCT05258942. IPD Sharing: NO. Countries: 1. Publications: 3.
Analysis of Prognostic and Predictive Genomic Signatures Using Archival Paraffin-embedded Tumor Specimens in Breast Cancer
ClinicalTrials.gov study NCT01247480. IPD Sharing: Not stated. Countries: 1. Publications: 2.
Genomic Signatures to Predict Treatment Response
ClinicalTrials.gov study NCT02032745. IPD Sharing: Not stated. Countries: 1. Publications: 2.
A Validation Study of Relationships Among Genomic Gene Expression Profile, Prognosis and Prediction of Adjuvant Chemotherapy Benefit With Capecitabine and Oxaliplatin in Gastric Cancer Stage II and II
ClinicalTrials.gov study NCT03403296. IPD Sharing: NO. Countries: 1. Publications: 5.
Metabolism Imaging-genomics for Predicting the Surgical Outcomes of Colorectal Cancer
ClinicalTrials.gov study NCT06614660. IPD Sharing: UNDECIDED. Countries: 1. Publications: 3.
Genome-wide association and genomic prediction for a reproductive index summarizing fertility outcomes in U.S. Holsteins
Open the record for dataset details and reuse information.
Data for: Machine learning for genomic and pedigree prediction in sugarcane
Open the record for dataset details and reuse information.
Data from: Accurate genomic prediction of Coffea canephora in multiple environments using whole-genome statistical models
Open the record for dataset details and reuse information.
Predicting genome-wide tissue-specific enhancers via combinatorial transcription factor genomic occupancy analysis
Open the record for dataset details and reuse information.
Haplotype-based genome-wide association increases the predictability of leaf rust (Puccinia triticina) resistance in wheat
Open the record for dataset details and reuse information.
Data from: Predictable genome-wide sorting of standing genetic variation during parallel adaptation to basic versus acidic environments in stickleback fish
Open the record for dataset details and reuse information.
Predictability and parallelism in the contemporary evolution of hybrid genomes
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.