Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
448
datasets available to search
ShareScore release 0.9.0
Dataset results
448 results for “Genomic selection”
Data from: Landscape genomics of Populus trichocarpa: the role of hybridization, limited gene flow and natural selection in shaping patterns of population structure
Populus trichocarpa is an ecologically important tree across western North America. We used a large population sample of 498 accessions over a wide geographical area genotyped with a 34K Populus SNP array to quantify geographical patterns of genetic variation in this species (landscape genomics). We present evidence that three processes contribute to the observed patterns: (1) introgression from the sister species P. balsamifera (2) isolation-by-distance and (3) natural selection. Introgression was detected only at the margins of the species' distribution. Isolation-by-distance was significant across the sampled area as a whole, but no evidence of restricted gene flow was detected in a core of drainages from southern British Columbia. We identified a large number of FST outliers. GO analyses revealed that FST outliers are overrepresented in genes involved in circadian rhythm and response to red/far-red light when the entire dataset is considered, while in southern British Columbia heat response genes are overrepresented. We also identified strong correlations between geoclimate variables and allele frequencies at FST outlier loci that provide clues regarding the selective pressures acting at these loci.
Data from: Genetic diversity, linkage disequilibrium and selection signatures in Chinese and Western pigs revealed by genome-wide SNP markers
To investigate population structure, linkage disequilibrium (LD) pattern and selection signature at the genome level in Chinese and Western pigs, we genotyped 304 unrelated animals from 18 diverse populations using porcine 60 K SNP chips. We confirmed the divergent evolution between Chinese and Western pigs and showed distinct topological structures of the tested populations. We acquired the evidence for the introgression of Western pigs into two Chinese pig breeds. Analysis of runs of homozygosity revealed that historical inbreeding reduced genetic variability in several Chinese breeds. We found that intrapopulation LD extents are roughly comparable between Chinese and Western pigs. However, interpopulation LD is much longer in Western pigs compared with Chinese pigs with average r20.3 values of 125 kb for Western pigs and only 10.5 kb for Chinese pigs. The finding indicates that higher-density markers are required to capture LD with causal variants in genome-wide association studies and genomic selection on Chinese pigs. Further, we looked across the genome to identify candidate loci under selection using FST outlier tests on two contrast samples: Tibetan pigs versus lowland pigs and belted pigs against non-belted pigs. Interestingly, we highlighted several genes including ADAMTS12, SIM1 and NOS1 that show signatures of natural selection in Tibetan pigs and are likely important for genetic adaptation to high altitude. Comparison of our findings with previous reports indicates that the underlying genetic basis for high-altitude adaptation in Tibetan pigs, Tibetan peoples and yaks is likely distinct from one another. Moreover, we identified the strongest signal of directional selection at the EDNRB loci in Chinese belted pigs, supporting EDNRB as a promising candidate gene for the white belt coat color in Chinese pigs. Altogether, our findings advance the understanding of the genome biology of Chinese and Western pigs.
Data from: Rapid evolution and the genomic consequences of selection against interspecific mating
While few species introduced into a new environment become invasive, those that do provide critical information on ecological mechanisms that determine invasions success and the evolutionary responses that follow invasion. Aedes albopictus (the Asian tiger mosquito) was introduced into the naturalized range of Aedes aegypti (the yellow fever mosquito) in the USA in the mid-1980s, resulting in the displacement of A. aegypti in much of the southeastern USA. The rapid displacement was likely due to the superior competitive ability of A. albopictus as larvae and asymmetric mating interference competition, in which male A. albopictus mate with and sterilize A. aegypti females, a process called "satyrization". The goal of this study was to examine the genomic responses of a resident species to an invasive species in which the mechanism of character displacement is understood. We used double-digest restriction enzyme DNA sequencing (ddRADseq) to analyze outlier loci between selected and control lines of laboratory-reared A. aegypti females from two populations (Tucson, AZ and Key West, Florida, USA), and individual females classified as either "resisted" or "mated with" A. albopictus males via mating trials of wild-derived females from four populations in Florida. We found significant outlier loci in comparing selected and control lines and between mated and non-mated A. aegypti females in the laboratory and wild-derived populations, respectively. We found overlap in specific outlier loci between different source populations that support consistent genomic signatures of selection within A. aegypti. Our results point to regions of the A. aegypti genome and potential candidate genes that may be involved in mating behavior, and specifically in avoiding interspecific mating choices.
Data from: Genomic analysis of a migratory divide reveals candidate genes for migration and implicates selective sweeps in generating islands of differentiation
Differential gene flow, reductions in diversity following linked selection and/or features of the genome can structure patterns of genomic differentiation during the process of speciation. Possible sources of reproductive isolation are well studied between coastal and inland subspecies groups of Swainson's thrushes, with differences in seasonal migratory behaviour likely playing a key role in reducing hybrid fitness. We assembled and annotated a draft reference genome for this species and generated whole-genome shotgun sequence data for populations adjacent to the hybrid zone between these groups. We documented substantial genomewide heterogeneity in relative estimates of genetic differentiation between the groups. Within population diversity was lower in areas of high relative differentiation, supporting a role for selective sweeps in generating this pattern. Absolute genetic differentiation was reduced in these areas, further suggesting that recurrent selective sweeps in the ancestral population and/or between divergent populations following secondary contact likely occurred. Relative genetic differentiation was also higher near centromeres and on the Z chromosome, suggesting that features of the genome also contribute to genomewide heterogeneity. Genes linked to migratory traits were concentrated in islands of differentiation, supporting previous suggestions that seasonal migration is under divergent selection between Swainson's thrushes. Differences in migratory behaviour likely play a central role in the speciation of many taxa; we developed the infrastructure here to permit future investigations into the role several candidate genes play in reducing gene flow between not only Swainson's thrushes but other species as well.
Data from: Whole genome sequencing of two North American Drosophila melanogaster populations reveals genetic differentiation and positive selection
The prevailing demographic model for Drosophila melanogaster suggests that the colonization of North America occurred very recently from a subset of European flies that rapidly expanded across the continent. This model implies a sudden population growth and range expansion consistent with very low or no population subdivision. As flies adapt to new environments, local adaptation events may be expected. To describe demographic and selective events during North American colonization, we have generated a data set of 35 individual whole-genome sequences from inbred lines of D. melanogaster from a west coast US population (Winters, California, USA) and compared them with a public genome data set from Raleigh (Raleigh, North Carolina, USA). We analysed nuclear and mitochondrial genomes and described levels of variation and divergence within and between these two North American D. melanogaster populations. Both populations exhibit negative values of Tajima's D across the genome, a common signature of demographic expansion. We also detected a low but significant level of genome-wide differentiation between the two populations, as well as multiple allele surfing events, which can be the result of gene drift in local subpopulations on the edge of an expansion wave. In contrast to this genome-wide pattern, we uncovered a 50-kilobase segment in chromosome arm 3L that showed all the hallmarks of a soft selective sweep in both populations. A comparison of allele frequencies within this divergent region among six populations from three continents allowed us to cluster these populations in two differentiated groups, providing evidence for the action of natural selection on a global scale.
Demographic History and Genomic Targets of Positive Selection in Giant Gough Mice
<p>A key challenge in understanding how natural selection operates is to identify the mutations and genes that make it possible. Positive selection on beneficial mutations distorts linked variation by altering the site frequency spectrum, the configuration of haplotypes, and population differentiation. By comparing patterns of sequence variation to neutral predictions across genomes, the targets of positive selection can be located. We applied this logic to an unusual population of house mice that shows phenotypic and ecological hallmarks of selection. Mice from Gough Island are twice the body size of mainland mice, eat live seabirds, maintain a very high population density, and inhabit an environment without predators or humans. We used massively parallel short-read sequencing to survey the genomes of 14 Gough Island mice. We computed a set of summary statistics to capture diverse aspects of variation across these genome sequences, used approximate Bayesian computation to reconstruct a null demographic model, and then applied machine learning to estimate the posterior probability of positive selection in each region of the genome. We conducted parallel analyses on genome sequences from 8 mice from Germany, treating them as representatives of a mainland reference population. A few thousand 5kb windows show strong evidence for positive selection in Gough Island mice but not in German mice. Genic regions and the X chromosome contain disproportionate shares of these selection windows. Over-represented gene ontologies in selection windows emphasize neurological themes. Inspection of genomic regions harboring many selection windows with high posterior probabilities pointed to genes with known effects on exploratory behavior and body size as potential targets. Some genes in these regions have missense mutations and/or putative regulatory mutations with large differences between Gough Island mice and German/French mice in the frequency of the derived allele; these are candidates for adaptive variants. Our results provide a genomic portrait of adaptation to island conditions and position Gough Island mice as a powerful system for understanding the genetic component of natural selection.</p>
Data from: Multi-generation genomic prediction of maize yield using parametric and non-parametric sparse selection indices
<p>Genomic prediction models are often calibrated using multi-generation data. Over time, as data accumulates, training data sets become increasingly heterogeneous. Differences in allele frequency and linkage disequilibrium patterns between the training and prediction genotypes may limit prediction accuracy. This leads to the question of whether all available data or a subset of it should be used to calibrate genomic prediction models. Previous research on training set optimization has focused on identifying a subset of the available data that is optimal for a given prediction set. However, this approach does not contemplate the possibility that different training sets may be optimal for different prediction genotypes. To address this problem, we recently introduced a sparse selection index (SSI) that identifies an optimal training set for each individual in a prediction set. Using additive genomic relationships, the SSI can provide increased accuracy relative to genomic-BLUP (GBLUP). Non-parametric genomic models using Gaussian kernels (KBLUP) have, in some cases, yielded higher prediction accuracies than standard additive models. Therefore, here we studied whether combining SSIs and kernel methods could further improve prediction accuracy when training genomic models using multi-generation data. Using four years of doubled haploid maize data from the International Maize and Wheat Improvement Center (CIMMYT), we found that when predicting grain yield the KBLUP outperformed the GBLUP, and that using SSI with additive relationships (GSSI) lead to 5-17% increases in accuracy, relative to the GBLUP. However, differences in prediction accuracy between the KBLUP and the kernel-based SSI were smaller and not always significant.</p>
Incorporation of soil-derived covariates in progeny testing and line selection to enhance genomic prediction accuracy in soybean breeding
<p>The availability of high-dimensional molecular markers has allowed plant breeding programs to maximize their efficiency through the genomic prediction of a phenotype of interest. Yield is a highly complex and quantitative trait whose expression is sensitive to environmental stimuli. In this research, we investigated the potential of incorporating soil texture and its interaction with molecular markers through covariance structures to enhance predictive ability. A total of 797 advanced soybean breeding lines derived from 367 unique bi-parental populations were genotyped using the Illumina Infinium BARCSoySNP6K BeadChip and tested for yield for five years in Tiptonville silt loam, Sharkey clay, and Malden fine sand environments. Four statistical models were considered, including a default GBLUP model (M1), a reaction norm model (M2) accounting for the interaction between molecular markers and the environment (GE), an expansion of M2 including soil type (S), and the interaction between soil type and molecular markers (GS) (M3), and an alternative version of M3 without the GE term. Four cross-validation scenarios simulating progeny testing and line selection were implemented (CV2, CV1, CV0, and CV00). Across environments, the addition of GS in M3 decreased the amount of variability captured by both the environment (-30.4%) and residual (-39.2%) terms as compared to M1. Within environments, the GS term in M3 reduced the variability captured by the residual term by roughly 60% and 30% when compared to M1 and M2, respectively. M3 outperformed all models in CV2 (0.577), CV1 (0.480), and CV0 (0.488). The addition of soil texture seems to structure the environment term revealing its components that could enhance or hinder the predictability of a model. The availability of soil texture before the growing season may maximize the functionality of covariance structures, particularly in scenarios with untested genotypes in untested environments. Genomic selection can optimize the efficiency of a soybean breeding program by allowing the reconsideration of field experimental design, allocation of resources, reduction of preliminary trials, and shortening of the breeding cycle.</p>
Data from: Genomic signatures of artificial selection during early domestication of a wood crop
<p>To determine how a century of artificial selection has changed the genome of <i>E. grandis</i>, we generated SNP genotypes for 1080 individuals from three advanced South African breeding programmes using the EUChip60K chip, and investigated population structure and genome-wide differentiation patterns relative to wild progenitors.</p>
Liquid Biopsy-informed Precision Oncology Study to Evaluate Utility of Plasma Genomic Profiling for Therapy Selection
ClinicalTrials.gov study NCT05585684. IPD Sharing: NO. Countries: 1. Publications: 0.
Genomic Opioid Optimization of Dosing and Selections (GOODS) Study
ClinicalTrials.gov study NCT03579121. IPD Sharing: NO. Countries: 1. Publications: 1.
Accuracy of genomic selection and long-term genetic gain for resistance to Verticillium wilt in a genetically diverse strawberry population
Open the record for dataset details and reuse information.
Data from: AFLP genome scan in the black rat (Rattus rattus) from Madagascar: detecting genetic markers undergoing plague-mediated selection
Open the record for dataset details and reuse information.
Data from: Genomic response to selection for predatory behavior in a mammalian model of adaptive radiation
Open the record for dataset details and reuse information.
Data from: Genomic sequencing reveals historical, demographic and selective factors associated with the diversification of the fire-associated fungus Neurospora discreta
Open the record for dataset details and reuse information.
Data from: Genome-wide differentiation in closely related populations: the roles of selection and geographic isolation
Open the record for dataset details and reuse information.
Data from: Signatures of selection in the Iberian honey bee (Apis mellifera iberiensis) revealed by a genome scan analysis of single nucleotide polymorphisms
Open the record for dataset details and reuse information.
Data from: AFLP genome scans suggest divergent selection on colour patterning in allopatric colour morphs of a cichlid fish
Open the record for dataset details and reuse information.
Data from: Experimental evidence for ecological selection on genome variation in the wild
Open the record for dataset details and reuse information.
Data from: Genomic data illuminates demography, genetic structure and selection of a popular dog breed
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.