Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

289

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

289 results for “Genomic prediction”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Accurate genomic predictions for chronic wasting disease in U.S. white-tailed deer

Open the record for dataset details and reuse information.

publicMar 2020View details →
dryad32/100

Benchmarking parametric and machine learning models for genomic prediction of complex traits

Open the record for dataset details and reuse information.

publicOct 2019View details →
dryad32/100

Incorporation of soil-derived covariates in progeny testing and line selection to enhance genomic prediction accuracy in soybean breeding

Open the record for dataset details and reuse information.

publicOct 2022View details →
zenodo28/100

Genomes and predicted proteins of Naegleria fowleri strains CDC:V212, 986, and ATCC30863

<p>Genomes and predicted proteins for three strains of&nbsp;<em>Naegleria fowleri</em>: strain CDC:V212, 986, and ATCC30863.</p> <p>&nbsp;</p> <p>These data have been made available for the publication of the manuscript <strong>&quot;A comparative &lsquo;omics approach to candidate pathogenicity factor discovery in the brain-eating amoeba&nbsp;<em>Naegleria&nbsp;fowleri</em>&quot;. </strong>See&nbsp;https://doi.org/10.1101/2020.01.16.908186&nbsp;for more information.</p>

opencc-by-4.0Jul 2021View details →
dryad28/100

Optimizing whole-genomic prediction for autotetraploid blueberry breeding

<p><span><span><span><span><span><span><span><span><span><span><span><span><span>Blueberry (<em>Vaccinium</em> spp.) is an important autopolyploid crop with significant benefits for human health. Apart from its genetic complexity, the feasibility of genomic prediction has been proven for blueberry, </span></span><span><span>enabling a reduction in the breeding cycle time and increasing genetic gain. </span></span>However, as for other polyploid crops<span><span>, </span></span>sequencing <span><span>costs still hinder the implementation of genome-based breeding methods for blueberry.</span></span> This motivated us to evaluate the effect of training population sizes and composition, as well as the impact of marker density and sequencing depth on phenotype prediction for the species. For this, data from a large real breeding population of 1 804 individuals was used. Genotypic data from 86 930 markers and three traits with different genetic architecture (fruit firmness, fruit weight, and total yield) were evaluated. Herein, we suggested that marker density, sequencing depth, and training population size can be substantially reduced with no significant impact on model accuracy. Our results can help guide decisions towards resource allocation (e.g., genotyping and phenotyping) in order to maximize prediction accuracy. These findings have the potential to allow for a faster and more accurate release of varieties with a substantial reduction of resources for the application of genomic prediction in blueberry. We anticipate that the benefits and pipeline described in our study <span><span>can be applied to optimize genomic prediction for other diploid and polyploid species.</span></span></span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroAug 2020View details →
dryad28/100

Data from: Shared patterns of genome-wide differentiation are more strongly predicted by geography than by ecology.

Closely related populations often display similar patterns of genomic differentiation, yet it remains an open question which ecological and evolutionary forces generate these patterns. The leading hypothesis is that this similarity in divergence is driven by parallel natural selection. However, several recent studies have suggested that these patterns may instead be a product of the depletion of genetic variation that occurs as result of background selection (i.e. linked negative selection). To date, there have been few direct tests of these competing hypotheses. To determine the relative contributions of background selection and parallel selection to patterns of repeated differentiation, we examined 24 independently derived populations of freshwater stickleback occupying a variety of niches and estimated genomic patterns of differentiation in each relative to their common marine ancestor. Patterns of genetic differentiation were strongly correlated across pairs of freshwater populations adapting to the same ecological niche, supporting a role for parallel natural selection. In contrast to other recent work, by examining populations adapting to the same niche we did not find evidence that similar patterns of genomic differentiation are generated by background selection. We also found that overall patterns of genetic differentiation were considerably more similar for populations found in closer geographic proximity. In fact, the effect of geography on the repeatability of differentiation was greater than that of parallel selection. Our results suggest that shared selective landscapes and ancestral variation are the key drivers of repeated patterns of differentiation in systems that have recently colonized novel environments.

opencc-zeroSep 2020View details →
dryad28/100

Data from: Integration of genomics and transcriptomics predicts diabetic retinopathy susceptibility genes

<p class="Normal1">We determined differential gene expression in response to high glucose in lymphoblastoid cell lines derived from matched individuals with type 1 diabetes with and without retinopathy. Those genes exhibiting the largest difference in glucose response were assessed for association to diabetic retinopathy in a genome-wide association study meta-analysis. Expression Quantitative Trait Loci (eQTLs) of the glucose response genes were tested for association with diabetic retinopathy. We detected an enrichment of the eQTLs from the glucose response genes among small association p-values and identified <i>FLCN</i> as a susceptibility gene for diabetic retinopathy. Expression of <i>FLCN </i>in response to glucose was greater in individuals with diabetic retinopathy. Independent cohorts of individuals with diabetes revealed an association of <i>FLCN</i> eQTLs to diabetic retinopathy. Mendelian randomization confirmed a direct positive effect of increased <i>FLCN</i> expression on retinopathy. Integrating genetic association with gene expression implicated <i>FLCN </i>as a disease gene for diabetic retinopathy.</p>

opencc-zeroNov 2020View details →
dryad28/100

Data from: Efficiency of genomic prediction of nonassessed testcrosses

In plant breeding, genomic selection has been mainly used to predict untested single crosses and testcrosses. The objectives of this study were to assess the efficiency of prediction of untested testcrosses and the significance of the factors affecting the prediction. We simulated 20,000 testcrosses from two groups of 5000 related doubled haploid lines (DHs) and two unrelated elite inbred lines. The DHs were genotyped for 20,000 single nucleotide polymorphisms (SNPs) and the testcrosses were phenotyped for grain yield. The average SNP density was 0.1 cM. We assumed genetic control by 400 quantitative trait loci (QTLs) and heritability of 25, 50, and 75% for the assessed testcrosses. The training set sizes were 10 and 30% of the available DHs. The process of random sampling of the field assessed and predicted testcrosses were replicated 50 times. We computed the prediction accuracy and the coincidence index, a measure of the selection efficacy of the nonassessed testcrosses. The results evidenced that genomic selection is an efficient process for selecting superior nonassessed testcrosses, if there is sufficient relatedness between the available DHs, linkage disequilibrium in the DHs and genotypic variance between testcrosses, a training set size of at least 10% of the available DHs, and a SNP density of at least 1 cM. The efficacy of selecting the superior nonassessed testcrosses ranged between 0.1 and 0.9, proportional to the heritability of the assessed testcrosses (0.25–1.00). It is important to highlight that the parametric coincidence ranged from 0.1 to 0.8. Furthermore, genomic prediction is much more efficient than pedigree-based best linear unbiased prediction.

opencc-zeroAug 2019View details →
dryad28/100

Data from: The impact of variable degrees of freedom and scale parameters in Bayesian methods for genomic prediction in Chinese Simmental beef cattle

Three conventional Bayesian approaches (BayesA, BayesB and BayesCπ) have been demonstrated to be powerful in predicting genomic merit for complex traits in livestock. A priori, these Bayesian models assume that the non-zero SNP effects (marginally) follow a t-distribution depending on two fixed hyperparameters, degrees of freedom and scale parameters. In this study, we performed genomic prediction in Chinese Simmental beef cattle and treated degrees of freedom and scale parameters as unknown with inappropriate priors. Furthermore, we compared the modified methods (BayesFA, BayesFB and BayesFCπ) with their corresponding counterparts using simulation datasets. We found that the modified methods with distribution assumed to the two hyperparameters were beneficial for improving the predictive accuracy. Our results showed that the predictive accuracies of the modified methods were slightly higher than those of their counterparts especially for traits with low heritability and a small number of QTLs. Moreover, cross-validation analysis for three traits, namely carcass weight, live weight and tenderloin weight, in 1136 Simmental beef cattle suggested that predictive accuracy of BayesFCπ noticeably outperformed BayesCπ with the highest increase (3.8%) for live weight using the cohort masking cross-validation.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Genome-wide prediction of bacterial effector candidates across six secretion system types using a feature-based statistical framework

Gram-negative bacteria are responsible for hundreds of millions infections worldwide, including the emerging hospital-acquired infections and neglected tropical diseases in the third-world countries. Finding a fast and cheap way to understand the molecular mechanisms behind the bacterial infections is critical for efficient diagnostics and treatment. An important step towards understanding these mechanisms is the discovery of bacterial effectors, the proteins secreted into the host through one of the six common secretion system types. Unfortunately, current prediction methods are designed to specifically target one of three secretion systems, and no accurate "secretion system-agnostic" method is available. Here, we present PREFFECTOR, a computational feature-based approach to discover effector candidates in Gram-negative bacteria, without prior knowledge on bacterial secretion system(s) or cryptic secretion signals. Our approach was first evaluated using several assessment protocols on a manually curated, balanced dataset of experimentally determined effectors across all six secretion systems, as well as non-effector proteins. The evaluation revealed high accuracy of the top performing classifiers in PREFFECTOR, with the small false positive discovery rate across all six secretion systems. Our method was also applied to six bacteria that had limited knowledge on virulence factors or secreted effectors. PREFFECTOR web-server is freely available at: http://korkinlab.org/preffector.

opencc-zeroSep 2017View details →
zenodo28/100

Multimodal Features Integration: Genomics and Histopathological images for colon cancer stage prediction and survival stratification

<p>File Description</p> <p>clinical_patient_coad.xls - Clinical data for colon cancer patients<br>COAD__geneExp.xls - RNA data for colon cancer patients<br>COAD__methylation_450__TSS200-TSS1500.xls - DNA methylation data for colon cancer patients<br>COAD__miRNAExp__ReadCount.xls - miRNA data for colon cancer patients<br>COAD_DX_STG1.tar.gz - Image tiles for stage 1 colon cancer patients<br>COAD_DX_STG2.tar.gz - Image tiles for stage 2 colon cancer patients<br>COAD_DX_STG3.tar.gz - Image tiles for stage 3 colon cancer patients<br>COAD_DX_STG4.tar.gz - Image tiles for stage 4 colon cancer patients</p>

opencc-by-4.0Mar 2024View details →
dryad28/100

Phased, chromosome-scale genome assemblies of tetraploid potato reveals a complex genome, transcriptome, and predicted proteome landscape underpinning genetic diversity

<p>Hoopes G., Meng X., Hamilton J.P., Achakkagari S.R., de Alves Freitas Guesdes F., Bolger M.E., Coombs J.J., Esselink D., Kaiser N.R., Kodde L., Kyriakidou M., Lavrijssen B., van Lieshout N., Shereda R., Tuttle H.K., Vaillancourt B., Wood J.C., de Boer J.M., Bornowski N., Bourke P., Douches D., van Eck H.J., Ellis D., Feldman M.J., Gardner K.M., Hopman J.C.P., Jiang J., De Jong W.S., Kuhl J.C., Novy R.G., Oome S., Sathuvalli V., Tan E.H., Ursum R.A., Vales M.I., Vining K., Visser R.G.F., Vossen J., Yencho G.C., Anglin N.L., Bachem C.W.B., Endelman J.B., Shannon L.M., Strömvik M.V., Tai H.H., Usadel B., Buell C.R., and Finkers R. (2022). Phased, chromosome-scale genome assemblies of tetraploid potato reveals a complex genome, transcriptome, and predicted proteome landscape underpinning genetic diversity. Mol. Plant. doi: https://doi.org/10.1016/j.molp.2022.01.003.</p> <p>Cultivated potato is a clonally propagated autotetraploid species with a highly heterogeneous genome. Phased assemblies of six cultivars including two chromosome-scale phased genome assemblies revealed extensive allelic diversity including altered coding and transcript sequences, preferential allele expression, and structural variation that collectively result in a highly complex transcriptome and predicted proteome which are distributed across the homologous chromosomes. Wild species contribute to the extensive allelic diversity in tetraploid cultivars, demonstrating ancestral introgressions predating modern breeding efforts. As a clonally propagated autotetraploid that undergoes limited meiosis, dysfunctional and deleterious alleles are not purged in tetraploid potato. Nearly a quarter of the loci bore mutations predicted to have a high negative impact on protein function, complicating breeder's efforts to reduce genetic load. The <em>StCDF1</em> locus controls maturity and analysis of six tetraploid genomes revealed 12 allelic variants correlated with maturity in a dosage dependent manner. Knowledge of the complexity of the tetraploid potato genome with its rampant structural variation and embedded deleterious and dysfunctional alleles will be key not only to implementing precision breeding of tetraploid cultivars but also to the construction of homozygous, diploid potato germplasm containing favorable alleles to capitalize on heterosis in F1 hybrids.</p>

opencc-zeroDec 2021View details →
dryad28/100

The predicted haploid gene set of the genome of Nitzschia putrida

<p>Secondary loss of photosynthesis is observed across almost all plastid-bearing branches of the eukaryotic tree of life. However, genome-based insights into the transition from a phototroph into a secondary heterotroph have so far only been revealed for parasitic species. Free-living organisms can yield unique insights into the evolutionary consequence of the loss of photosynthesis, as the parasitic lifestyle requires specific adaptations to host environments. Here we report on the diploid genome of the free-living diatom <i>Nitzschia putrida </i>(35 Mbp), a non-photosynthetic osmotroph whose photosynthetic relatives contribute ca. 40% of net oceanic primary production. Comparative analyses with photosynthetic diatoms and heterotrophic algae with parasitic lifestyle revealed that a combination of gene loss, the accumulation of genes involved in organic carbon degradation, a unique secretome and the rapid divergence of conserved gene families involved in cell wall and extracellular metabolism appear to have facilitated the lifestyle of a free-living secondary heterotroph.</p>

opencc-zeroMar 2022View details →
zenodo28/100

The Cytoscape session file for the network of siderophore BGCs predicted from Streptomyces genomes

<p>The Cytoscape session file for the network of siderophore BGCs predicted from Streptomyces genomes</p>

opencc-by-4.0May 2022View details →
zenodo28/100

The Cytoscape session file for the network of indole BGCs predicted from Streptomyces genomes

<p>The Cytoscape session file for the network of indole BGCs predicted from Streptomyces genomes</p>

opencc-by-4.0May 2022View details →
dryad28/100

Data from: Incorporating single-step strategy into random regression model to enhance genomic prediction of longitudinal trait

In prediction of genomic values, single-step method has been demonstrated to outperform multi-step methods. In statistical analyses of longitudinal traits, random regression test-day model (RR-TDM) has clear advantages over other models. Our goal in this study was to evaluate the performance of the model integrating both single-step and RR-TDM prediction methods, called single-step random regression test-day model (SS RR-TDM), in comparison with the pedigree-based RR-TDM and genomic best linear unbiased prediction (GBLUP) model. We performed extensive simulations to exploit potential advantages of SS RR-TDM over the other two models under various scenarios with different level of heritability, the number of QTL as well as the selection scheme. SS RR-TDM was found to achieve the highest accuracy and unbiasedness under all scenarios, exhibiting robust prediction ability in longitudinal trait analyses. Moreover, SS RR-TDM showed better persistency of accuracy over generations than GBLUP model. In addition, we also found that the SS RR-TDM had advantages over RR-TDM and GBLUP in terms of a real dataset of human contributed by the GAW18 workshop. The findings in our study firstly proved the feasibility and advantages of the SS RR-TDM, and further enhanced strategies for the genomic prediction of longitudinal traits in the future.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Ribosomal DNA sequence heterogeneity reflects intra-species phylogenies and predicts genome structure in two contrasting yeast species

The ribosomal RNA encapsulates a wealth of evolutionary information, including genetic variation that can be used to discriminate between organisms at a wide range of taxonomic levels. For example, the prokaryotic 16S rDNA sequence is very widely used both in phylogenetic studies and as a marker in metagenomic surveys and the ITS region, frequently used in plant phylogenetics, is now recognised as a fungal DNA barcode. However, this widespread use does not escape criticism, principally due to issues such as difficulties in classification of paralogous versus orthologous rDNA units and intragenomic variation, both of which may be significant barriers to accurate phylogenetic inference. We recently analysed datasets from the Saccharomyces Genome Resequencing Project, characterising rDNA sequence variation within multiple strains of the baker's yeast <i>Saccharomyces cerevisiae</i> and its nearest wild relative <i>Saccharomyces paradoxus</i> in unprecedented detail. Notably, both species possess single locus rDNA systems. Here, we use these new variation datasets to assess whether a more detailed characterisation of the rDNA locus can alleviate the second of these phylogenetic issues, sequence heterogeneity, while controlling for the first. We demonstrate that a strong phylogenetic signal exists within both datasets and illustrate how they can be used, with existing methodology, to estimate intra-species phylogenies of yeast strains consistent with those derived from whole-genome approaches. We also describe the use of partial Single Nucleotide Polymorphisms, a type of sequence variation found only in repetitive genomic regions, in identifying key evolutionary features such as genome hybridisation events and show their consistency with whole-genome Structure analyses. We conclude that our approach can transform rDNA sequence heterogeneity from a problem to a useful source of evolutionary information, enabling the estimation of highly accurate phylogenies of closely related organisms, and discuss how it could be extended to future studies of multi-locus rDNA systems.

opencc-zeroDec 2013View details →
dryad28/100

Accelerating wheat breeding for end-use quality through association mapping and multivariate genomic prediction

<p>In hard winter wheat breeding, the evaluation of end-use quality is expensive and time-consuming, being relegated to the final stages of the breeding program after selection for many traits including disease resistance, agronomic performance and grain yield. In this study, our objectives were to identify genetic variants underlying baking quality traits through genome-wide association mapping (GWAS) and develop improved genomic selection (GS) models for the quality traits in hard winter wheat.  Advanced breeding lines (n=462) from 2015-2017 were genotyped using genotyping-by-sequencing (GBS) and evaluated for baking quality.  Significant associations were detected for mixograph mixing time and bake mixing time; most of which were within or in tight linkage to glutenin and gliadin loci, and could be suitable for marker-assisted breeding.  Candidate genes for newly associated loci are phosphate-dependent decarboxylase and lipid transfer protein genes, which are believed to affect nitrogen metabolism and dough development, respectively.  The use of GS can both shorten the breeding cycle time and significantly increase the number of lines that could be selected for quality traits; thus we evaluated various GS models for end-use quality traits.  As a baseline, univariate GS models had 0.25 to 0.55 prediction accuracy in cross-validation and from 0 to 0.41 in forward-prediction.  By including secondary traits as additional predictor variables (univariate GS with covariates) or correlated response variables (multivariate GS), the prediction accuracies were increased relative to the univariate model using only genomic information.  The improved genomic prediction models have great potential to further accelerate wheat breeding for end-use quality.</p>

opencc-zeroSep 2022View details →
zenodo28/100

The effect of different statistical methods on the accuracy of predicting genomic selection in beef cattle

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

Article Multivariate Adaptive Regression Splines enhances Genomic Prediction of non-additive traits

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record