Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
289
datasets available to search
ShareScore release 0.9.0
Dataset results
289 results for “Genomic prediction”
Swordtail fish hybrids reveal that genome evolution is surprisingly predictable after initial hybridization
Open the record for dataset details and reuse information.
Data from: Integrating genomic data and simulations to evaluate alternative species distribution models and improve predictions of glacial refugia and future responses to climate change
Open the record for dataset details and reuse information.
Improving genomic prediction for plant disease using environmental covariates
Open the record for dataset details and reuse information.
Source code for StrVCTVRE: a supervised learning method to predict the pathogenicity of human genome structural variants
Open the record for dataset details and reuse information.
Data from: Genome assembly of the ragweed leaf beetle, a step forward to better predict rapid evolution of a weed biocontrol agent to environmental novelties
<p><span>Rapid evolution of weed biological control agents (BCAs) to new biotic and abiotic conditions is poorly understood and so far, only little considered both in pre-release and post-release studies, despite potential major negative or positive implications for risks of non-targeted attacks or for colonizing yet unsuitable habitats, respectively. Provision of genetic resources, such as assembled and annotated genomes, is essential to assess potential adaptive processes by identifying underlying genetic mechanisms. Here, we provide the first sequenced genome of a phytophagous insect used as a BCA, <i>i.e.</i> the leaf beetle <i>Ophraella communa</i>, a promising BCA of common ragweed, recently and accidentally introduced into Europe. A total 33.98 Gb of raw DNA sequences, representing c. 43-fold coverage, were obtained using the PacBio SMRT-Cell sequencing approach. Among the five different assemblers tested, the SMARTdenovo assembly displaying the best scores was then corrected with Illumina short reads. A final genome of 774 Mb containing 7,003 scaffolds was obtained. The reliability of the final assembly was then assessed by benchmarking universal single-copy orthologous genes (> 96.0% of the 1,658 expected insect genes) and by remapping tests of Illumina short reads (average of 98.6% ± 0.7% without filtering). The number of protein-coding genes of 75,642, representing 82% of the published antennal transcriptome, and the phylogenetic analyses based on 825 orthologous genes placing <i>O. communa </i>in the monophyletic group of Chrysomelidae, confirm the relevance of our genome assembly. Overall, the genome provides a valuable resource for studying potential risks and benefits of this BCA facing environmental novelties.</span></p>
Data from: Do genomics and sex predict migration in a partially migratory salmonid fish, Oncorhynchus mykiss?
Partial migration is a common phenomenon wherein populations include migratory and resident individuals. Whether an individual migrates or not has important ecological and management implications, particularly within protected populations. Within partially migratory populations of O. mykiss, migration is highly correlated with a specific genomic region, but it is unclear how well this region predicts migration at the individual level. Here, we relate sex and life history genotype, determined using >400 SNPs on the migratory-linked genomic region, to life history expression of marked juvenile O. mykiss from two tributaries to the South Fork Eel River, northern California. Most resident fish were resident-genotypes (57% resident, 37% heterozygous, 6% migratory genotype) and male (78%). Most migratory fish were female (62%), but were a mixture of genotypes (30% resident, 45% heterozygous, 25% migratory genotype). Sex was more strongly correlated with life history expression than genotype, but the best-supported model included both. Resident genotypes regularly migrated, highlighting the importance of conserving the full suite of life history and genetic diversity in partially migratory populations.
Data from: Genomic analysis and prediction within a US public collaborative winter wheat regional testing nursery
The development of inexpensive, whole-genome profiling enables a transition to allele-based breeding using genomic prediction models. These models consider alleles shared between lines to predict phenotypes and select new lines based on estimated breeding values. This approach can leverage highly-unbalanced datasets common to breeding programs. The Southern Regional Performance Nursery (SRPN) is a public nursery established by the USDA-ARS in 1931 to characterize performance and quality of near-release wheat varieties from breeding programs in the US Central Plains. New entries are submitted annually and can be reentered only once. The trial is grown at more than 30 locations each year and lines are evaluated for grain yield, disease resistance, and agronomic traits. Overall genetic gain is measured across years by including common check cultivars for comparison. We have generated whole-genome profiles via genotyping-by-sequencing for 939 SPRN entries dating back to 1992. We measured the diversity within the nursery and have explored its potential use as a GS training population. GS prediction models across years (average r= 0.33) outperformed year-to-year phenotypic correlation for yield (r=0.27) for a majority of the years evaluated, suggesting that genomic selection has the potential to outperform low heritability selection on yield in these highly variable environments. We also examined the predictability of programs using both program-specific and whole-set training populations. Generally, the predictability of a program was similar with both approaches. These results suggest that wheat breeding programs can collaboratively leverage the immense datasets that are generated from regional testing networks.
Data from: Genomic signals of selection predict climate-driven population declines in a migratory bird
The ongoing loss of biodiversity caused by rapid climatic shifts requires accurate models for predicting species' responses. Despite evidence that evolutionary adaptation could mitigate climate change impacts, evolution is rarely integrated into predictive models. Integrating population genomics and environmental data, we identified genomic variation associated with climate across the breeding range of the migratory songbird, yellow warbler (Setophaga petechia). Populations requiring the greatest shifts in allele frequencies to keep pace with future climate change have experienced the largest population declines, suggesting that failure to adapt may have already negatively affected populations. Broadly, our study suggests that the integration of genomic adaptation can increase the accuracy of future species distribution models and ultimately guide more effective mitigation efforts.
Data from: Genotyping by sequencing and genome–environment associations in wild common bean predict widespread divergent adaptation to drought
Drought will reduce global crop production by >10% in 2050 substantially worsening global malnutrition. Breeding for resistance to drought will require accessing crop genetic diversity found in the wild accessions from the driest high stress ecosystems. Genome–environment associations in crop wild relatives reveal natural adaptation, and therefore can be used to identify adaptive variation. We explored this approach in the food crop Phaseolus vulgaris L., characterizing 86 geo-referenced wild accessions using Genotyping by Sequencing (GBS) to discover single-nucleotide-polymorphisms (SNPs). The wild beans represented Mesoamerica, Guatemala, Colombia, Ecuador/Northern Peru and Andean groupings. We found high polymorphism with a total of 22,845 SNPs across the 86 accessions loci that confirmed genetic relationships for the groups. As a second objective, we quantified allelic associations with a bioclimatic-based drought index using 10 different statistical models that accounted for population structure. Based on the optimum model, 115 SNPs in 90 regions, widespread in all 11 common bean chromosomes, were associated with the bioclimatic-based drought index. A gene coding for an Ankyrin repeat-containing protein and a phototropic-responsive NPH3 gene were identified as potential candidates. Genomic windows of 1Mb containing associated SNPs had more positive Tajima's D scores than windows without associated markers. This indicates that adaptation to drought, as estimated by bioclimatic variables, has been under natural divergent selection, suggesting that drought tolerance may be favorable under dry conditions but harmful in humid conditions. Our work exemplifies that genomic signatures of adaptation are useful for germplasm characterization, potentially enhancing future marker-assisted selection and crop improvement.
Predicted genome-wide chromatin contact differences among 71 bonobos and chimpanzees
<p>This file contains predicted chromatin contact differences in HFF cells using Akita among pairs of 71 bonobos and chimpanzees at 4,420 ~ 1 Mb genomic windows in the panTro6 genome. Each entry corresponds to a pairwise comparison at a given window. Data per comparison includes the individual IDs in the pairwise comparison, lineages represented, chromosome, position, window ID, mean squared error, Spearman correlation, divergence (1 - Spearman correlation), and the number of nucleotide differences for the pair at the given window.</p>
The Role of Genomic Data in Stratifying Patients within Predictive Models for Breast Cancer Survival Outcome
<p>Data associated with my PhD thesis titled "The Role of Genomic Data in Stratifying Patients within Predictive Models for Breast Cancer Survival Outcome".</p>
Genomic prediction in the wild: a case study in Soay sheep
<p>Genomic prediction, the technique whereby an individual's genetic component of their phenotype is estimated from its genome, has revolutionised animal and plant breeding and medical genetics. However, despite being first introduced nearly two decades ago, it has hardly been adopted by the evolutionary genetics community studying wild organisms. Here, genomic prediction is performed on eight traits in a wild population of Soay sheep. The population has been the focus of a >30 year evolutionary ecology study and there is already considerable understanding of the genetic architecture of the focal Mendelian and quantitative traits. We show that the accuracy of genomic prediction is high for all traits, but especially those with loci of large effect segregating. Five different methods are compared, and the two methods that can accommodate zero-effect and large-effect loci in the same model tend to perform best. If the accuracy of genomic prediction is similar in other wild populations, then there is a real opportunity for pedigree-free molecular quantitative genetics research to be enabled in many more wild populations; currently the literature is dominated by studies that have required decades of field data collection to generate sufficiently deep pedigrees. Finally, some of the potential applications of genomic prediction in wild populations are discussed.</p>
Combining climatic and genomic data improves range-wide tree height growth prediction in a forest tree
<p>Population response functions based on climatic and phenotypic data from common gardens have long been the gold standard for predicting quantitative trait variation in new environments. However, prediction accuracy might be enhanced by incorporating genomic information that captures the neutral and adaptive processes behind intra-population genetic variation. We used five clonal common gardens containing 34 provenances (523 genotypes) of maritime pine (<em>Pinus pinaster</em> Aiton) to determine whether models combining climatic and genomic data capture the underlying drivers of height-growth variation, and thus improve predictions at large geographical scales. The plastic component explained most of the height-growth variation, probably resulting from population responses to multiple environmental factors. The genetic component stemmed mainly from climate adaptation, and the distinct demographic and selective histories of the different maritime pine gene pools. Models combining climate-of-origin and gene pool of the provenances, and positive-effect height-associated alleles (PEAs) captured most of the genetic component of height-growth and better predicted new provenances compared to the climate-based population response functions. Regionally-selected PEAs were better predictors than globally-selected PEAs, showing high predictive ability in some environments, even when included alone in the models. These results are therefore promising for the future use of genome-based prediction of quantitative traits.</p>
Hydractinia strain 236-21 genome assembly and Alr domain predictions
<p>This dataset is related to the preprint "A family of unusual A family of unusual immunoglobulin superfamily genes in an invertebrate histocompatibility complex" (<a href="https://www.biorxiv.org/content/10.1101/2022.03.04.482883v2">https://www.biorxiv.org/content/10.1101/2022.03.04.482883v2</a>).</p> <p><strong>Preprint Abstract:</strong></p> <p>Most colonial marine invertebrates are capable of allorecognition, the ability to distinguish between themselves and conspecifics. One long-standing question is whether invertebrate allorecognition genes are homologous to vertebrate histocompatibility genes. In the cnidarian <em>Hydractinia symbiolongicarpus, </em>allorecognition is controlled by at least two genes, <em>Allorecognition 1</em> (<em>Alr1</em>) and <em>Allorecognition 2 </em>(<em>Alr2</em>), which encode highly polymorphic cell surface proteins that serve as markers of self. Here, we show that <em>Alr1</em> and <em>Alr2</em> are part of a family of 41 <em>Alr </em>genes, all of which reside a single genomic interval called the Allorecognition Complex (ARC). Using sensitive homology searches and highly accurate structural predictions, we demonstrate that the Alr proteins are members of the immunoglobulin superfamily (IgSF) with V-set and I-set Ig domains unlike any previously identified in animals. Specifically, their primary amino acid sequences lack many of the motifs considered diagnostic for V-set and I-set domains, yet they adopt secondary and tertiary structures nearly identical to canonical Ig domains. Thus, the V-set domain, which played a central role in the evolution of vertebrate adaptive immunity, was present in the last common ancestor of cnidarians and bilaterians. Unexpectedly, several Alr proteins also have immunoreceptor tyrosine-based activation motifs (ITAMs) and immunoreceptor tyrosine-based inhibitory motifs (ITIMs) in their cytoplasmic tails, suggesting they could participate in pathways homologous to those that regulate immunity in humans and flies. This work expands our definition of the IgSF with the addition of a family of unusual members, several of which play a role in invertebrate histocompatibility.</p> <p><strong>This dataset contains:</strong></p> <ol> <li><strong>Hsym-236-21-genome-assembly.fa.gz</strong>: A gzip-compressed FASTA-formatted file of the genome assembly generated in the paper. </li> <li><strong>Alr-domain-structure-predictions.zip:</strong> a zip-compressed file with structural predictions produced with Colabfold for all domains of the Alr proteins described in that manuscript.</li> </ol>
Kinship matrices for deep learning for whole-genome predictions
<p>Kinship matrices to be used for the deep learning model implementation used in Tensorflow/Keras for whole-genome predictions (see Github repo at https://github.com/filippob/paper_deep_learning_vs_gblup/)</p>
Switchgrass flowering time measurements for genomic prediction
<p>The seasonal timing of the transition from vegetative to reproductive growth has a major impact on biomass accumulation in switchgrass. Late-flowering switchgrass cultivars produce greater biomass, a critical trait for sustainable bioenergy production. Genomic prediction (GP) may allow rapid selection of late-flowering individuals with reduced time and expense for field evaluations. To evaluate GP, two flowering time traits (heading date and anthesis date) were collected on 1,532 genotypes from four breeding populations: Midwest, Gulf, Atlantic, and Hybrid. These were sequenced using genotype-by-sequencing (530,792 SNPs). Predictive ability of single-trait and multi-trait models were evaluated by cross-validation, by prediction of a progeny trial (n=122), and through prediction of yield performance in a parallel experiment (n=52). Predictive ability was not improved by sharing information among breeding groups. Overall, multi-trait models provided an advantage during cross-validation, but a smaller advantage during progeny prediction. Within populations, GP resulted in lower per-cycle progress than previously reported field evaluations (3.1 vs 5.0 day<sup>-1</sup> cycle<sup>-1</sup>). However, GP cycles are potentially much faster than field evaluations. When directly predicting biomass yield, the Hybrid training population had a predictive ability of 0.54-0.63. This reinforces the strong linkage between biomass yields in swards and flowering time. These results highlight the value of GP for rapid yield improvement in switchgrass, particularly in a breeding program designed to share information between biomass yield trials and low-cost flowering time evaluations.</p>
Improving genome-wide association discovery and genomic prediction accuracy in biobank data
<p>Genetically informed, deep-phenotyped biobanks are an important research resource and it is imperative that the most powerful, versatile, and efficient analysis approaches are used. Here, we apply our recently developed Bayesian grouped mixture of regressions model (GMRM) in the UK and Estonian Biobanks and obtain the highest genomic prediction accuracy reported to date across 21 heritable traits. When compared to other approaches, GMRM accuracy was greater than annotation prediction models run in the LDAK or LDPred-funct software by 15% (SE 7%) and 14% (SE 2%), respectively, and was 18% (SE 3%) greater than a baseline BayesR model without single-nucleotide polymorphism (SNP) markers grouped into minor allele frequency–linkage disequilibrium (MAF-LD) annotation categories. For height, the prediction accuracy R 2 was 47% in a UK Biobank holdout sample, which was 76% of the estimated h SNP 2 . We then extend our GMRM prediction model to provide mixed-linear model association (MLMA) SNP marker estimates for genome-wide association (GWAS) discovery, which increased the independent loci detected to 16,162 in unrelated UK Biobank individuals, compared to 10,550 from BoltLMM and 10,095 from Regenie, a 62 and 65% increase, respectively. The average χ<sup>2</sup> value of the leading markers increased by 15.24 (SE 0.41) for every 1% increase in prediction accuracy gained over a baseline BayesR model across the traits. Thus, we show that modeling genetic associations accounting for MAF and LD differences among SNP markers, and incorporating prior knowledge of genomic function, is important for both genomic prediction and discovery in large-scale individual-level studies.</p>
Genomic Prediction for the Germplasm Enhancement of Maize Project
<p>The Germplasm Enhancement of Maize (GEM) project was initiated in 1993 as a cooperative effort of public and private sector maize breeders to enhance the genetic diversity of the U.S. maize crop. The GEM project selects progeny lines with high topcross yield potential from crosses between elite temperate lines and exotic parents. The GEM project has released hundreds of useful breeding lines based on phenotypic selection within selfing generations and multi-environment yield evaluations of GEM line topcrosses to elite adapted testers. Developing genomic selection (GS) models for the GEM project may contribute to increases in the rate of genetic gain. Here we evaluated the prediction ability of GS models trained on six years of topcross evaluations from the two GEM programs in Raleigh, NC and Ames, IA, documenting prediction abilities ranging from 0.36 to 0.75 for grain yield and from 0.78 to 0.96 for grain moisture when models were cross-validated within program and heterotic group. Predicted genetic gain from genomic selection ranged from 0.95 to 2.58 times the gain from phenotypic selection. Prediction ability across program and heterotic group was generally poorer than within groups. Based on observed genomic relationships between GEM breeding lines and their tropical ancestors, GS for either yield or moisture would reduce recovery of exotic germplasm only slightly. Using GS models trained within-program, the GEM programs should be able to more effectively deliver on its mission to broaden the genetic base of U.S. germplasm. </p>
Predicted genes annotated with Prokka for the 623 archaeal UBA genomes
<p>Gene prediction and annotation for the 623 archaeal UBA metagenome-assembled genomes using Prokka v1.12 with Pfam v31 and UniProt databases created on April 17, 2017 according to the Prokka instructions.</p> <p>These genomes are described in:</p> <p>Parks DH, et al. 2017. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol, doi:10.1038/s41564-017-0012-7/ .</p> <p>https://www.nature.com/articles/s41564-017-0012-7</p>
Amino acid sequences of the proteins predicted from the whole genome of hilsa shad (Tenualosa ilisha) of the Bay of Bengal
<p>Gene prediction was performed by AUGUSTUS (Stanke et al., 2006) from the whole genome sequence of <em>T. ilisha</em> (<a href="https://www.ncbi.nlm.nih.gov/bioproject/400122">PRJNA400122</a>). The data contain amino acid sequences of 37,450 predicted protein coding genes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.