Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
289
datasets available to search
ShareScore release 0.9.0
Dataset results
289 results for “Genomic prediction”
Predicted genome-wide chromatin contact differences among 71 bonobos and chimpanzees
Open the record for dataset details and reuse information.
Genomic demography predicts community dynamics in a temperate Montane forest
Open the record for dataset details and reuse information.
Data from: Do genomics and sex predict migration in a partially migratory salmonid fish, Oncorhynchus mykiss?
Open the record for dataset details and reuse information.
Data from: How accurate is genomic prediction across wild populations?
Open the record for dataset details and reuse information.
SNP genotype and hyperspectral reflectance data from: Ensembles of genomic and hyperspectral imaging-based prediction enable selection for reduced deoxynivalenol content in wheat grains
Open the record for dataset details and reuse information.
Data from: Genomic signals of selection predict climate-driven population declines in a migratory bird
Open the record for dataset details and reuse information.
Improving genome-wide association discovery and genomic prediction accuracy in biobank data
Open the record for dataset details and reuse information.
Genomic prediction in the wild: a case study in Soay sheep
Open the record for dataset details and reuse information.
Data from: Genotyping by sequencing and genome–environment associations in wild common bean predict widespread divergent adaptation to drought
Open the record for dataset details and reuse information.
Genomic prediction for growth using a low-density SNP panel in dromedary camels
Open the record for dataset details and reuse information.
Data from: Genomic analysis and prediction within a US public collaborative winter wheat regional testing nursery
Open the record for dataset details and reuse information.
Data from: A mix of old British and modern European breeds: Genomic prediction of breed composition of smallholder pigs in Uganda
Open the record for dataset details and reuse information.
How useful is genomic data for predicting maladaptation to future climate?
Open the record for dataset details and reuse information.
Genomics of new ciliate lineages provides insight into the evolution of obligate anaerobiosis - single gene datasets for phylogenomic analysis of anaerobic ciliates (SAL, Ciliophora), protein datasets for mitochondrial pathways prediction, and mitochondrial genomes
Open the record for dataset details and reuse information.
Data for: Multi-trait/environment sparse genomic prediction using the SFSI R-package
Open the record for dataset details and reuse information.
Genomic prediction enables rapid selection of high-performing genets in an intermediate wheatgrass (Thinopyrum intermedium) breeding program
Open the record for dataset details and reuse information.
Splice altering variant predictions in four archaic hominin genomes
Open the record for dataset details and reuse information.
Predicting Phenotype from Multi-Scale Genomic and Environment Data using Neural Networks and Knowledge Graphs
<p><strong>Background: To mitigate the effects of climate change on public health and conservation, we need to better understand the dynamic interplay between biological processes and environmental effects. Machine learning (ML) methods in general, and Deep Learning (DL) methods in particular, are a potential way forward because they are able to cope with the nonlinearity of natural systems. However, there are several barriers that exist, including the absence of ML-ready data. We propose to develop a machine learning framework capable of predicting phenotypes based on multi-scale data about genes and environments. A critical part of this framework are data transformation methods that map the heterogeneous input data into formats that are consumable by the ML techniques. The central hypothesis of this research is that deep learning algorithms and biological knowledge graphs will predict phenotypes more accurately across more taxa and more ecosystems than do current numerical and traditional statistical modeling methods. Our long term goal is to develop predictive analytics for organismal response to environmental perturbations using innovative data science approaches. This pilot project on predicting emergent properties of complex systems and multidimensional interactions is funded by the NSF (Award # 1939945, 1940059, 1940062, 1940330). </strong></p> <p> </p> <p><strong>Results: We have established shared project governance, communication channels, project timeline, and data and computing environment across four universities. We have successfully reached out to three other projects for broader collaboration.</strong></p>
Efficient weighting methods for genomic best linear unbiased prediction (BLUP) adaption to the genetic architectures of quantitative traits
<p><a name="_Hlk19877414"></a>Genomic best linear unbiased prediction (GBLUP) assumes equal variance for all marker effects, which is suitable for traits that conform to the infinitesimal model. For traits controlled by major genes, Bayesian methods with shrinkage priors or genome-wide association study (GWAS) methods can be used to identify <a name="_Hlk24974556">causal variants</a> effectively. The information from Bayesian/GWAS methods can be used to construct the weighted genomic relationship matrix (<b>G</b>). However, it remains unclear which methods perform best for traits varying in genetic architecture. Therefore, we developed several methods to <a name="_Hlk23592218">optimize</a> the performance of weighted GBLUP and compare them with other available methods using simulated and real datasets. First, two types of methods (marker effects with local-shrinkage or normal prior) were used to obtain test statistics and estimates for each marker effect. Second, three weighted <b>G</b> matrices were constructed based on the marker information from the first step: (1) the genomic-feature weighted <b>G</b> (GFWG), (2) the estimated marker-variance weighted <b>G</b> (EVWG), and (3) the absolute value of estimated marker-effect weighted <b>G</b> (AEWG). Following the above process, six different weighted GBLUP methods (local-shrinkage/normal prior GF/EV/AE-WGBLUP) were proposed for genomic prediction. Analyses with both simulated and real data demonstrated that these options offer flexibility for optimizing the weighted GBLUP for traits with a broad spectrum of genetic architectures. The advantage of weighting methods over GBLUP in terms of accuracy were trait dependent, ranging from 14.8% to marginal for simulated traits and from 44% to marginal for real traits. Local-shrinkage prior EVWGBLUP is superior for traits mainly controlled by loci of large effect. Normal prior AEWGBLUP performs well for traits mainly controlled by loci of moderate effect. For traits controlled by some loci with large effects (<a name="_Hlk49869847">explain 25%~50% genetic variance</a>) and a range of loci with small effects, GFWGBLUP has advantages. In conclusion, the optimal weighted GBLUP method for genomic selection should take both the genetic architecture and number of QTLs of traits into consideration carefully.</p>
Using genomic prediction to detect microevolutionary change of a quantitative trait
<p>Detecting microevolutionary responses to natural selection by observing temporal changes in individual breeding values is challenging. The collection of suitable datasets can take many years and disentangling the contributions of the environment and genetics to phenotypic change is not trivial. Furthermore, pedigree-based methods of obtaining individual breeding values have known biases. Here, we apply a genomic prediction approach to estimate breeding values of adult weight in a 35-year dataset of Soay sheep (<i>Ovis aries)</i>. Comparisons are made with a traditional pedigree-based approach. During the study period adult body weight decreased, but the underlying genetic component of body weight increased, at a rate that is unlikely to be attributable to genetic drift. Thus cryptic microevolution of greater adult body weight has probably occurred. Genomic and pedigree-based approaches gave largely consistent results. Thus, using genomic prediction to study microevolution in wild populations can remove the requirement for pedigree data, potentially opening up new study systems for similar research.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.