Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

289

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

289 results for “Genomic prediction”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Genomic BLUP decoded: a look into the black box of genomic prediction

Genomic best linear unbiased prediction (BLUP) is a statistical method that uses relationships between individuals calculated from single-nucleotide polymorphisms (SNPs) to capture relationships at quantitative trait loci (QTL). We show that genomic BLUP exploits not only linkage disequilibrium (LD) and additive-genetic relationships, but also cosegregation to capture relationships at QTL. Simulations were used to study the contributions of those types of information to accuracy of genomic estimated breeding values (GEBVs), their persistence over generations without retraining, and their effect on the correlation of GEBVs within families. We show that accuracy of GEBVs based on additive-genetic relationships can decline with increasing training data size and speculate that modeling polygenic effects via pedigree relationships jointly with genomic breeding values using Bayesian methods may prevent that decline. Cosegregation information from half sibs contributes little to accuracy of GEBVs in current dairy cattle breeding schemes but from full sibs it contributes considerably to accuracy within family in corn breeding. Cosegregation information also declines with increasing training data size, and its persistence over generations is lower than that of LD, suggesting the need to model LD and cosegregation explicitly. The correlation between GEBVs within families depends largely on additive-genetic relationship information, which is determined by the effective number of SNPs and training data size. As genomic BLUP cannot capture short-range LD information well, we recommend Bayesian methods with t-distributed priors.

opencc-zeroDec 2012View details →
zenodo28/100

Genomically predicted theoretical protein mass database for mass spectrometry (GPMsDB) R01-RS95

<p>GPMsDB-tk/GPMsDB-dbtk are software toolkits for assigning taxonomic identification to user-provided MALDI-TOF mass spectrometry profiles obtained from bacterial and archaeal cultured isolates. This is a database&nbsp;(release R01-RS95) used for the toolkits.&nbsp;This&nbsp;database&nbsp;is released under the <a href="https://creativecommons.org/licenses/by-sa/4.0/">Creative Commons Attribution-ShareAlike 4.0 International License</a>.</p> <p>GPMsDB-tk/GPMsDB-dbtk&nbsp;are available at&nbsp;<a href="https://github.com/ysekig">https://github.com/ysekig</a>. Details are shown at the GitHub sites.</p>

openother-atMar 2023View details →
dryad28/100

Data from: Ribosomal DNA sequence heterogeneity reflects intra-species phylogenies and predicts genome structure in two contrasting yeast species

Open the record for dataset details and reuse information.

publicMar 2014View details →
dryad28/100

Data from: An equation to predict the accuracy of genomic values by combining data from multiple traits, populations, or environments

Open the record for dataset details and reuse information.

publicNov 2016View details →
dryad28/100

Data from: Shared patterns of genome-wide differentiation are more strongly predicted by geography than by ecology.

Open the record for dataset details and reuse information.

publicSep 2020View details →
dryad28/100

Phased, chromosome-scale genome assemblies of tetraploid potato reveals a complex genome, transcriptome, and predicted proteome landscape underpinning genetic diversity

Open the record for dataset details and reuse information.

publicJan 2022View details →
dryad28/100

Data from: The impact of variable degrees of freedom and scale parameters in Bayesian methods for genomic prediction in Chinese Simmental beef cattle

Open the record for dataset details and reuse information.

publicApr 2017View details →
dryad28/100

Data from: Incorporating single-step strategy into random regression model to enhance genomic prediction of longitudinal trait

Open the record for dataset details and reuse information.

publicAug 2016View details →
dryad28/100

Optimizing whole-genomic prediction for autotetraploid blueberry breeding

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad28/100

Data from: Predicted input of uncultured fungal symbionts to a lichen symbiosis from metagenome-assembled genomes

Open the record for dataset details and reuse information.

publicMar 2021View details →
dryad28/100

Data from: Gene prediction and annotation in Penstemon (Plantaginaceae): a workflow for marker development from extremely low-coverage genome sequencing

Open the record for dataset details and reuse information.

publicOct 2015View details →
dryad28/100

The predicted haploid gene set of the genome of Nitzschia putrida

Open the record for dataset details and reuse information.

publicMar 2022View details →
dryad28/100

Data from: Genome-environment associations in sorghum landraces predict adaptive traits

Open the record for dataset details and reuse information.

publicJun 2016View details →
dryad28/100

Data from: Genome-wide association and genomic prediction models of tocochromanols in fresh sweet corn kernels

Open the record for dataset details and reuse information.

publicAug 2019View details →
dryad28/100

Data from: Efficiency of genomic prediction of nonassessed testcrosses

Open the record for dataset details and reuse information.

publicAug 2019View details →
dryad28/100

Genotyping of marine sticklebacks - Predicting future from past: The genomic basis of recurrent and rapid stickleback evolution

Open the record for dataset details and reuse information.

publicMay 2021View details →
dryad28/100

Data from: Genome-wide prediction of bacterial effector candidates across six secretion system types using a feature-based statistical framework

Open the record for dataset details and reuse information.

publicSep 2017View details →
dryad28/100

Data from: Genomic BLUP decoded: a look into the black box of genomic prediction

Open the record for dataset details and reuse information.

publicMay 2013View details →
dryad28/100

Accelerating wheat breeding for end-use quality through association mapping and multivariate genomic prediction

Open the record for dataset details and reuse information.

publicSep 2022View details →
dryad28/100

Data from: Across population genomic prediction scenarios in which Bayesian variable selection outperforms GBLUP

Open the record for dataset details and reuse information.

publicDec 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record