Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
448
datasets available to search
ShareScore release 0.9.0
Dataset results
448 results for “Genomic selection”
Dataset from: Selecting deep neural networks that yield consistent attribution-based interpretations for genomics
<p>Deep neural networks (DNNs) have demonstrated great promise at taking DNA sequences as input and predicting a wide variety of functional activity. Post hoc attribution analysis has been employed to provide insights into the features learned by DNNs, often revealing patterns such as known motifs. However, attribution maps are noisy in practice to an extent that varies from model to model, even across DNNs that yield similar generalization performance. This makes it challenging to identify which high-performing DNN will provide trustworthy explanations. Here we propose a summary statistic that characterizes the consistency of learned features across a population of attribution maps which can be utilized as an additional criterion for model selection. We demonstrate the efficacy of this approach quantitatively using synthetic data and qualitatively with chromatin accessibility data. Together, this work advances our ability to select optimal DNNs that not only yield high generalization performance but also reliable attribution maps that will, in turn, accelerate scientific discovery in genomics.</p>
Genomic Selection Paves Way for the Identification of Rust Disease Resistant Genotypes in Bread Wheat (Triticum aestivum).
<p><span>In the last two decades, genomic prediction (GP) or Genomic Selection (GS) methods have been widely adopted in various plant and animal breeding programs globally. GP/GS is a promising method that employs genomic markers to calculate genomic-estimated breeding values (GEBVs) to select best individuals. To evaluate the performance of different genomic selection (GS) models, we examined six different models namely, ridge regression (RR), least absolute shrinkage and selection operator (LASSO), genomic best linear unbiased prediction (GBLUP), elastic net (EN), reproducing kernel Hilbert spacing (RKHS), and random forest (RF) models, for seedling and adult plant resistance to leaf, stem and stripe rust of wheat using a panel of 347 wheat germplasm accessions. The GBLUP and RF models performed noticeably better than the other GS models, with mean predictive abilities of 0.5 and 0.4 for seedling resistance and 0.4 and 0.3 for adult plant resistance (APR) for leaf and stem rust, respectively. Unfortunately, except for a few environments, the performance of GP models in the current study is quite low for stripe rust for both seedling and APR. The outcomes of this study revealed the capability of GP to be applied for breeding initiatives aimed at developing wheat varieties resistant to rust diseases. </span><span>Moreover, based on favorable allele analysis we also identified a total of 2 lines (CRP-165/42, HGP1-470) that showed resistance to most of the pathotypes at seedling and adult plant stage to all three rusts.<strong><span> </span></strong>These lines can serve as valuable resources for future breeding programs focused on rust resistance.</span></p> <p><strong><span>Keywords: </span></strong><span>GS;</span><strong><span> </span></strong><span>GEBVs; leaf rust; stem rust; stripe rust; seedling resistance; APR</span></p>
The phylogeny of Triticeae Dumort. (Poaceae): resolution and reticulation based on a genome-wide selection of nuclear loci.
<p>Chloroplast-genome and nuclear-locus phylogenetic datasets for the wheat tribe Triticeae.</p>
FIGURE 10 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 10. Phylogeny of the liparid clade Aenigmoliparia from the majority rule (50%) consensus tree from the Bayesian inference of a 490 bp alignment of 270 cytochrome c oxidase subunit one gene (COI) sequences. Nodal values represent Bayesian posterior probabilities and bootstrap values from the maximum likelihood analysis (above and below branches, respectively). Species names are followed by a catalog number or BOLD "Sequence ID" number when represented by a sequence from a single specimen in our dataset. N indicates number of sequences, when multiple sequences support a branch tip. Boldface species names indicate species placed in different positions in COI and RADseq trees. Only unique sequences were subjected to the analyses (Appendix Table 1); other identical sequences surveyed are listed in Appendix Table 2.
FIGURE 7 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 7. Majority-rule (50%) consensus phylogenetic tree of Shen et al. (2017, after fig. S6), derived from a Bayesian inference of a 440 bp alignment of cytochrome c oxidase subunit 1 gene (COI) sequences for 84 samples of 83 liparid species. Bayesian posterior probabilities are above branches. Tree is rooted with species of the Cyclopteridae. Corrected identifications based on our study are in parentheses.
FIGURE 9 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 9. Phylogeny of the genus Liparis, excluding L. fucensis depicted in Figure 8, from the majority rule (50%) consensus tree from the Bayesian inference of a 490 bp alignment of 270 cytochrome c oxidase subunit one gene (COI) sequences. Nodal values represent Bayesian posterior probabilities and bootstrap values from the maximum likelihood analysis (above and below, respectively). Species names are followed by a catalog number or BOLD "Sequence ID" number when represented by a sequence from a single specimen in our dataset. N indicates number of sequences, when multiple sequences support a branch tip. Only unique sequences were subjected to the analyses (Appendix Table 1); other identical sequences surveyed are listed in Appendix Table 2.
FIGURE 6 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 6. Majority rule (50%) consensus phylogenetic tree of Gardner et al. (2016, after fig. 4), derived from Bayesian inference and maximum parsimony analysis of a 492 bp alignment of cytochrome c oxidase subunit 1 gene (COI) sequences of 492 bp for 128 samples of 23 liparid species. Bootstrap values are above and Bayesian posterior probabilities are below branches that lead to multiple species. Tree is rooted with Liparis gibbus. Corrected identifications based on our study are in parentheses.
FIGURE 12 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 12. Phylogeny of selected eastern North Pacific liparids inferred using genome-wide restriction-site associated DNA sequences (RADseq; –p 28, –r 0.5) with maximum likelihood and Bayesian methods. Majority rule (50%) consensus tree of individual sequences. Nodal values represent Bayesian posterior probabilities and bootstrap values from the maximum likelihood analysis (above and below branches, respectively); double asterisks denote Bayesian posterior probabilities of 1 and bootstrap support of 100%. Species names are followed by the University of Washington Fish Collection catalog number for the specimen. Boldface species names indicate species placed in different positions in COI and RADseq trees.
FIGURE 4 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 4. Unrooted neighbor-joining tree of Steinke et al. (2009, after fig. 4), derived from cluster analysis of a 650 bp alignment of cytochrome c oxidase subunit 1 gene (COI) sequences for 78 samples of 19 liparid species. Bootstrap values>80 are above branches leading to multiple species. Corrected identifications based on our study are in parentheses.
FIGURE 2 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 2. Phylogenetic hypothesis of Balushkin (1996, after fig. 4), derived from a manual cladistic analysis of morphological data, including seven osteological and external characters, for 26 liparid genera.
FIGURE 1 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 1. Phylogenetic hypothesis of Kido (1988, after fig. 20), derived from a maximum parsimony analysis of morphological data, using 34 osteological and external characters, for 60 liparid species.
FIGURE 3 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 3. Majority-rule (50%) consensus tree of Knudsen et al. (2007, after fig. 3), derived from a Bayesian analysis of three combined datasets composed of mitochondrial DNA (16S and cytochrome b) and morphological data for 24 liparid species. Tree is rooted with species of the Cyclopteridae. Posterior probabilities are above branches.
FIGURE 5 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 5. Consensus phylogenetic tree of Duhamel et al. (2010, after fig. 3), derived from Bayesian and maximum parsimony analyses of a 668 bp alignment of cytochrome c oxidase subunit 1 gene (COI) sequences for 157 samples of 46 liparid species. Bayesian posterior probabilities are above branches that lead to multiple species. Tree is rooted with species of the Cyclopteridae and Zoarcidae. Corrected identifications based on our study are in parentheses.
FIGURE 8 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 8. Phylogeny of the Liparidae. Majority rule (50%) consensus tree from the Bayesian inference of a 490 bp alignment of 270 cytochrome c oxidase subunit one gene (COI) sequences. Nodal values represent Bayesian posterior probabilities and bootstrap values from the maximum likelihood analysis (above and below, respectively). Species names are followed by a catalog number or BOLD "Sequence ID" number when represented by a sequence from a single specimen in our dataset. N indicates number of sequences, when multiple sequences support a branch tip. Only unique sequences were subjected to the analyses (Appendix Table 1); other identical sequences surveyed are listed in Appendix Table 2. Clades Liparis, Aenigmoliparia, and Paraliparia are depicted in Figures 9, 10, and 11, respectively.
FIGURE 13 H–O in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 13 H–O. Pectoral girdles of selected species of the Cyclopteridae and Liparidae: H) Prognatholiparis ptychomandibularis, UW 156749; I) Acantholiparis opercularis, UW 118624; J) Careproctus sp. cf. melanurus, UW 119240; K) Paraliparis dactylosus, UW 116232; L) Rhinoliparis attenuatus, UW 113736; M) P. cephalus, UW 117527; N) Paraliparis pectoralis, UW 117515; O) P. ulochir, UW 150802.
FIGURE 13 A–G in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 13 A–G. Pectoral girdles of selected species of the Cyclopteridae and Liparidae: A) Eumicrotremus orbis, UW 111284; B) Nectoliparis pelagicus, UW 119455; C) Liparis gibbus, UW 119092; D) Crystallichthys cyclospilus, UW 47840; E) Careproctus macrodiscus, FAKU 137835; F) Careproctus marginatus, FAKU 144616; G) Careproctus roseofuscus, FAKU 144615
FIGURE 11 in Molecular phylogenetics of snailfishes (Cottoidei: Liparidae) based on MtDNA and RADseq genomic analyses, with comments on selected morphological characters
FIGURE 11. Phylogeny of the liparid clade Paraliparia from the majority rule (50%) consensus tree from the Bayesian inference of a 490 bp alignment of 270 cytochrome c oxidase subunit one gene (COI) sequences. Nodal values represent Bayesian posterior probabilities and bootstrap values from the maximum likelihood analysis (above and below, respectively). Species names are followed by a catalog number or BOLD "Sequence ID" number when represented by a sequence from a single specimen in our dataset. N indicates number of sequences, when multiple sequences support a branch tip. Only unique sequences were subjected to the analyses (Appendix Table 1); other identical sequences surveyed are listed in Appendix Table 2.
Widespread intersex differentiation across the stickleback genome – the signature of sexually antagonistic selection?
<p>Females and males within a species commonly have distinct reproductive roles, and the associated traits may be under perpetual divergent natural selection between the sexes if their sex-specific control has not yet evolved. We here explore whether such sexually antagonistic selection can be detected based on the magnitude of differentiation between the sexes across genome-wide genetic polymorphisms by whole-genome sequencing of large pools of female and male threespine stickleback fish. We find numerous autosomal genome regions exhibiting intersex allele frequency differences beyond the range plausible under pure sampling stochasticity. Alternative sequence alignment strategies rule out that these high-differentiation regions represent sex chromosome segments misassembled into the autosomes. Instead, comparing allele frequencies and sequence read depth between the sexes reveals that regions of high intersex differentiation arise because autosomal chromosome segments got copied into the male-specific sex chromosome (Y), where they acquired new mutations. Because the Y chromosome is missing in the stickleback reference genome, sequence reads from derived DNA copies on the Y chromosome still align to the original homologous regions on the autosomes. We argue that this phenomenon hampers the identification of sexually antagonistic selection within a genome, and can lead to spurious conclusions from population genomic analyses when the underlying samples differ in sex ratios. Because the hemizygous sex chromosome sequence (Y or W) is not represented in most reference genomes, these problems may apply broadly.</p>
Data from: Effects of gene action, marker density, and time since selection on the performance of landscape genomic scans of local adaptation
Genomic "scans" to identify loci that contribute to local adaptation are becoming increasingly common. Many methods used for such studies have assumed that local adaptation is created by loci experiencing antagonistic pleiotropy and that the selected locus itself is assayed, and few consider how signals of selection change through time. However, most empirical data sets have marker density too low to assume that a selected locus itself is assayed, researchers seldom know when selection was first imposed, and many locally adapted loci likely experience not antagonistic pleiotropy but conditional neutrality. We simulated data to evaluate how these factors affect the performance of tests for genotype-environment association. We found that three types of regression-based analyses (linear models, mixed linear models, and latent factor mixed models) and an implementation of BayEnv all performed well, with high rates of true positives and low rates of false positives, when the selected locus experienced antagonistic pleiotropy, and when the selected locus was assayed directly. However, all tests had reduced power to detect loci experiencing conditional neutrality, and the probability of detecting associations was sharply reduced when physically linked rather than causative loci were sampled. Antagonistic pleiotropy also maintained detectable genotype-environment associations much longer than conditional neutrality. Our analyses suggest that if local adaptation is often driven by loci experiencing conditional neutrality, genome-scan methods will have limited capacity to find loci responsible for local adaptation.
Data from: Whole-genome resequencing uncovers molecular signatures of natural and sexual selection in wild bighorn sheep
The identification of genes influencing fitness is central to our understanding of the genetic basis of adaptation and how it shapes phenotypic variation in wild populations. Here, we used whole-genome resequencing of wild Rocky Mountain bighorn sheep (Ovis canadensis) to >50-fold coverage to identify 2.8 million single nucleotide polymorphisms (SNPs) and genomic regions bearing signatures of directional selection (i.e. selective sweeps). A comparison of SNP diversity between the X chromosome and the autosomes indicated that bighorn males had a dramatically reduced long-term effective population size compared to females. This probably reflects a long history of intense sexual selection mediated by male–male competition for mates. Selective sweep scans based on heterozygosity and nucleotide diversity revealed evidence for a selective sweep shared across multiple populations at RXFP2, a gene that strongly affects horn size in domestic ungulates. The massive horns carried by bighorn rams appear to have evolved in part via strong positive selection at RXFP2. We identified evidence for selection within individual populations at genes affecting early body growth and cellular response to hypoxia; however, these must be interpreted more cautiously as genetic drift is strong within local populations and may have caused false positives. These results represent a rare example of strong genomic signatures of selection identified at genes with known function in wild populations of a nonmodel species. Our results also showcase the value of reference genome assemblies from agricultural or model species for studies of the genomic basis of adaptation in closely related wild taxa.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.