Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
289
datasets available to search
ShareScore release 0.9.0
Dataset results
289 results for “Genomic prediction”
Data from: Genomic BLUP decoded: a look into the black box of genomic prediction
Genomic best linear unbiased prediction (BLUP) is a statistical method that uses relationships between individuals calculated from single-nucleotide polymorphisms (SNPs) to capture relationships at quantitative trait loci (QTL). We show that genomic BLUP exploits not only linkage disequilibrium (LD) and additive-genetic relationships, but also cosegregation to capture relationships at QTL. Simulations were used to study the contributions of those types of information to accuracy of genomic estimated breeding values (GEBVs), their persistence over generations without retraining, and their effect on the correlation of GEBVs within families. We show that accuracy of GEBVs based on additive-genetic relationships can decline with increasing training data size and speculate that modeling polygenic effects via pedigree relationships jointly with genomic breeding values using Bayesian methods may prevent that decline. Cosegregation information from half sibs contributes little to accuracy of GEBVs in current dairy cattle breeding schemes but from full sibs it contributes considerably to accuracy within family in corn breeding. Cosegregation information also declines with increasing training data size, and its persistence over generations is lower than that of LD, suggesting the need to model LD and cosegregation explicitly. The correlation between GEBVs within families depends largely on additive-genetic relationship information, which is determined by the effective number of SNPs and training data size. As genomic BLUP cannot capture short-range LD information well, we recommend Bayesian methods with t-distributed priors.
Genomically predicted theoretical protein mass database for mass spectrometry (GPMsDB) R01-RS95
<p>GPMsDB-tk/GPMsDB-dbtk are software toolkits for assigning taxonomic identification to user-provided MALDI-TOF mass spectrometry profiles obtained from bacterial and archaeal cultured isolates. This is a database (release R01-RS95) used for the toolkits. This database is released under the <a href="https://creativecommons.org/licenses/by-sa/4.0/">Creative Commons Attribution-ShareAlike 4.0 International License</a>.</p> <p>GPMsDB-tk/GPMsDB-dbtk are available at <a href="https://github.com/ysekig">https://github.com/ysekig</a>. Details are shown at the GitHub sites.</p>
Data from: Ribosomal DNA sequence heterogeneity reflects intra-species phylogenies and predicts genome structure in two contrasting yeast species
Open the record for dataset details and reuse information.
Data from: An equation to predict the accuracy of genomic values by combining data from multiple traits, populations, or environments
Open the record for dataset details and reuse information.
Data from: Shared patterns of genome-wide differentiation are more strongly predicted by geography than by ecology.
Open the record for dataset details and reuse information.
Phased, chromosome-scale genome assemblies of tetraploid potato reveals a complex genome, transcriptome, and predicted proteome landscape underpinning genetic diversity
Open the record for dataset details and reuse information.
Data from: The impact of variable degrees of freedom and scale parameters in Bayesian methods for genomic prediction in Chinese Simmental beef cattle
Open the record for dataset details and reuse information.
Data from: Incorporating single-step strategy into random regression model to enhance genomic prediction of longitudinal trait
Open the record for dataset details and reuse information.
Optimizing whole-genomic prediction for autotetraploid blueberry breeding
Open the record for dataset details and reuse information.
Data from: Predicted input of uncultured fungal symbionts to a lichen symbiosis from metagenome-assembled genomes
Open the record for dataset details and reuse information.
Data from: Gene prediction and annotation in Penstemon (Plantaginaceae): a workflow for marker development from extremely low-coverage genome sequencing
Open the record for dataset details and reuse information.
The predicted haploid gene set of the genome of Nitzschia putrida
Open the record for dataset details and reuse information.
Data from: Genome-environment associations in sorghum landraces predict adaptive traits
Open the record for dataset details and reuse information.
Data from: Genome-wide association and genomic prediction models of tocochromanols in fresh sweet corn kernels
Open the record for dataset details and reuse information.
Data from: Efficiency of genomic prediction of nonassessed testcrosses
Open the record for dataset details and reuse information.
Genotyping of marine sticklebacks - Predicting future from past: The genomic basis of recurrent and rapid stickleback evolution
Open the record for dataset details and reuse information.
Data from: Genome-wide prediction of bacterial effector candidates across six secretion system types using a feature-based statistical framework
Open the record for dataset details and reuse information.
Data from: Genomic BLUP decoded: a look into the black box of genomic prediction
Open the record for dataset details and reuse information.
Accelerating wheat breeding for end-use quality through association mapping and multivariate genomic prediction
Open the record for dataset details and reuse information.
Data from: Across population genomic prediction scenarios in which Bayesian variable selection outperforms GBLUP
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.