Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
153
datasets available to search
ShareScore release 0.9.0
Dataset results
153 results for “Selective signature”
Data from: Signatures of selection for bonamiosis resistance in European flat oyster (Ostrea edulis): new genomic tools for breeding programs and management of natural resources
The European flat oyster (Ostrea edulis) is a highly appreciated mollusk with an important aquaculture production throughout the 20th century, in addition to playing an important role on coastal ecosystems. Overexploitation of natural beds, habitat degradation, introduction of non-native species and epidemic outbreaks have severely affected this important resource, particularly, the protozoan parasite Bonamia ostreae, which is the main concern affecting its production and conservation. In order to identify genomic regions and markers potentially associated with bonamiosis resistance, six oyster beds distributed throughout the European Atlantic coast were sampled. Three of them have been exposed to this parasite since the early 1980's and showed some degree of innate resistance (long-term affected group, LTA), while the other three were free of B. ostreae at least until sampling date (naïve group, NV). A total of 14,065 SNPs were analyzed, including 37 markers from candidate genes and 14,028 from a medium density SNP array. Gene diversity was similar between LTA and NV groups suggesting no genetic erosion due to long term exposure to the parasite, and three population clusters were detected using the whole dataset. Tests for divergent selection between NV and LTA groups detected the presence of a very consistent set of 22 markers, located within a putative single genomic region, which suggests the presence of a major quantitative trait locus associated with B. ostreae resistance. Moreover, 324 outlier loci associated with factors other than bonamiosis were identified allowing fully discrimination of all the oyster beds. A practical tool which included the 84 highest discriminative markers for tracing O. edulis populations was developed and tested with empirical data. Results reported herein could assist the production of stocks with improved resistance to bonamiosis, and facilitate the management of oyster beds for recovery production and ecosystem services provided by this species.
Data from: Detecting genomic signatures of natural selection with principal component analysis: application to the 1000 Genomes data
To characterize natural selection, various analytical methods for detecting candidate genomic regions have been developed. We propose to perform genome-wide scans of natural selection using principal component analysis (PCA). We show that the common FST index of genetic differentiation between populations can be viewed as the proportion of variance explained by the principal components. Considering the correlations between genetic variants and each principal component provides a conceptual framework to detect genetic variants involved in local adaptation without any prior definition of populations. To validate the PCA-based approach, we consider the 1000 Genomes data (phase 1) considering 850 individuals coming from Africa, Asia, and Europe. The number of genetic variants is of the order of 36 millions obtained with a low-coverage sequencing depth (3×). The correlations between genetic variation and each principal component provide well-known targets for positive selection (EDAR, SLC24A5, SLC45A2, DARC), and also new candidate genes (APPBPP2, TP1A1, RTTN, KCNMA, MYO5C) and noncoding RNAs. In addition to identifying genes involved in biological adaptation, we identify two biological pathways involved in polygenic adaptation that are related to the innate immune system (beta defensins) and to lipid metabolism (fatty acid omega oxidation). An additional analysis of European data shows that a genome scan based on PCA retrieves classical examples of local adaptation even when there are no well-defined populations. PCA-based statistics, implemented in the PCAdapt R package and the PCAdapt fast open-source software, retrieve well-known signals of human adaptation, which is encouraging for future whole-genome sequencing project, especially when defining populations is difficult.
Data from: Signatures of selection in the Iberian honey bee (Apis mellifera iberiensis) revealed by a genome scan analysis of single nucleotide polymorphisms
Understanding the genetic mechanisms of adaptive population divergence is one of the most fundamental endeavours in evolutionary biology and is becoming increasingly important as it will allow predictions about how organisms will respond to global environmental crisis. This is particularly important for the honey bee, a species of unquestionable ecological and economical importance that has been exposed to increasing human-mediated selection pressures. Here, we conducted a single nucleotide polymorphism (SNP)-based genome scan in honey bees collected across an environmental gradient in Iberia and used four FST-based outlier tests to identify genomic regions exhibiting signatures of selection. Additionally, we analysed associations between genetic and environmental data for the identification of factors that might be correlated or act as selective pressures. With these approaches, 4.4% (17 of 383) of outlier loci were cross-validated by four FST-based methods, and 8.9% (34 of 383) were cross-validated by at least three methods. Of the 34 outliers, 15 were found to be strongly associated with one or more environmental variables. Further support for selection, provided by functional genomic information, was particularly compelling for SNP outliers mapped to different genes putatively involved in the same function such as vision, xenobiotic detoxification and innate immune response. This study enabled a more rigorous consideration of selection as the underlying cause of diversity patterns in Iberian honey bees, representing an important first step towards the identification of polymorphisms implicated in local adaptation and possibly in response to recent human-mediated environmental changes.
Data from: Signatures of positive selection in African Butana and Kenana dairy zebu cattle
Butana and Kenana are two types of zebu cattle found in Sudan. They are unique amongst African indigenous zebu cattle because of their high milk production. Aiming to understand their genome structure, we genotyped 25 individuals from each breed using the Illumina BovineHD Genotyping BeadChip. Genetic structure analysis shows that both breeds have an admixed genome composed of an even proportion of indicine (0.75 ± 0.03 in Butana, 0.76 ± 0.006 in Kenana) and taurine (0.23 ± 0.009 in Butana, 0.24 ± 0.006 in Kenana) ancestries. We also observe a proportion of 0.02 to 0.12 of European taurine ancestry in ten individuals of Butana that were sampled from cattle herds in Tamboul area suggesting local crossbreeding with exotic breeds. Signatures of selection analyses (iHS and Rsb) reveal 87 and 61 candidate positive selection regions in Butana and Kenana, respectively. These regions span genes and quantitative trait loci (QTL) associated with biological pathways that are important for adaptation to marginal environments (e.g., immunity, reproduction and heat tolerance). Trypanotolerance QTL are intersecting candidate regions in Kenana cattle indicating selection pressure acting on them, which might be associated with an unexplored level of trypanotolerance in this cattle breed. Several dairy traits QTL are overlapping the identified candidate regions in these two zebu cattle breeds. Our findings underline the potential to improve dairy production in the semi-arid pastoral areas of Africa through breeding improvement strategy of indigenous local breeds.
Data from: Genetic signatures of natural selection in response to air pollution in red spruce (Picea rubens, Pinaceae)
One of the most important drivers of local adaptation for forest trees is climate. Coupled to these patterns, however, are human-induced disturbances through habitat modification and pollution. The confounded effects of climate and disturbance have rarely been investigated with regard to selective pressure on forest trees. Here, we have developed and used a population genetic approach to search for signals of selection within a set of 36 candidate genes chosen for their putative effects on adaptation to climate and human-induced air pollution within five populations of red spruce (Picea rubens Sarg.), distributed across its natural range and air pollution gradient in eastern North America. Specifically, we used FST outlier and environmental correlation analyses to highlight a set of seven single nucleotide polymorphisms (SNPs) that were overly correlated with climate and levels of sulphate pollution after correcting for the confounding effects of population history. Use of three age cohorts within each population allowed the effects of climate and pollution to be separated temporally, as climate-related SNPs (n = 7) showed the strongest signals in the oldest cohort, while pollution-related SNPs (n = 3) showed the strongest signals in the youngest cohorts. These results highlight the usefulness of population genetic scans for the identification of putatively nonneutral evolution within genomes of nonmodel forest tree species, but also highlight the need for the development and application of robust methodologies to deal with the inherent multivariate nature of the genetic and ecological data used in these types of analyses.
Data from: Comparative transcriptomics uncovers alternative splicing changes and signatures of selection from maize improvement
Background: Alternative splicing (AS) is an important regulatory mechanism that greatly contributes to eukaryotic transcriptome diversity. A substantial amount of evidence has demonstrated that AS complexity is relevant to eukaryotic evolution, development, adaptation, and complexity. In this study, six teosinte and ten maize transcriptomes were sequenced to analyze AS changes and signatures of selection in maize domestication and improvement. Results In maize and teosinte, 13,593 highly conserved genes, including 12,030 multiexonic genes, were detected. By identifying AS isoforms from mutliexonic genes, we found that AS types were not significantly different between maize and teosinte. In addition, the two main AS types (intron retention and alternative acceptor) contributed to more than 60% of the AS events in the two species, but the average unique AS events per each alternatively spliced gene in maize (4.12) was higher than that in teosinte (2.26). Moreover, 94 genes generating 98 retained introns with transposable element (TE) sequences were detected in maize, which is far more than 9 retained introns with TEs detected in teosinte. This indicates that TE insertion might be an important mechanism for intron retention in maize. Additionally, the AS levels of 3864 genes were significantly different between maize and teosinte. Of these, 151 AS level-altered genes that are involved in transcriptional regulation and in stress responses are located in regions that have been targets of selection during maize improvement. These genes were inferred to be putatively improved genes. Conclusions We suggest that both maize and teosinte share similar AS mechanisms, but more genes have increased AS complexity during domestication from teosinte to maize. Importantly, a subset of AS level-increased genes that encode transcription factors and stress-responsive proteins may have been selected during maize improvement.
Data from: Scans for signatures of selection in Russian cattle breed genomes reveal new candidate genes for environmental adaptation and acclimation
Domestication and selective breeding has resulted in over 1000 extant cattle breeds. Many of these breeds do not excel in important traits but are adapted to local environments. These adaptations are a valuable source of genetic material for efforts to improve commercial breeds. As a step toward this goal we identified candidate regions to be under selection in genomes of nine Russian native cattle breeds adapted to survive in harsh climates. After comparing our data to other breeds of European and Asian origins we found known and novel candidate genes that could potentially be related to domestication, economically important traits and environmental adaptations in cattle. The Russian cattle breed genomes contained regions under putative selection with genes that may be related to adaptations to harsh environments (e.g., AQP5, RAD50, and RETREG1). We found genomic signatures of selective sweeps near key genes related to economically important traits, such as the milk production (e.g., DGAT1, ABCG2), growth (e.g., XKR4), and reproduction (e.g., CSF2). Our data point to candidate genes which should be included in future studies attempting to identify genes to improve the extant breeds and facilitate generation of commercial breeds that fit better into the environments of Russia and other countries with similar climates.
Data from: The population genomic signature of environmental selection in the widespread insect-pollinated tree species Frangula alnus at different geographical scales
The evaluation of the molecular signatures of selection in species lacking an available closely related reference genome remains challenging, yet it may provide valuable fundamental insights into the capacity of populations to respond to environmental cues. We screened 25 native populations of the tree species Frangula alnus subsp. alnus (Rhamnaceae), covering three different geographical scales, for 183 annotated single-nucleotide polymorphisms (SNPs). Standard population genomic outlier screens were combined with individual-based and multivariate landscape genomic approaches to examine the strength of selection relative to neutral processes in shaping genomic variation, and to identify the main environmental agents driving selection. Our results demonstrate a more distinct signature of selection with increasing geographical distance, as indicated by the proportion of SNPs (i) showing exceptional patterns of genetic diversity and differentiation (outliers) and (ii) associated with climate. Both temperature and precipitation have an important role as selective agents in shaping adaptive genomic differentiation in F. alnus subsp. alnus, although their relative importance differed among spatial scales. At the 'intermediate' and 'regional' scales, where limited genetic clustering and high population diversity were observed, some indications of natural selection may suggest a major role for gene flow in safeguarding adaptability. High genetic diversity at loci under selection in particular, indicated considerable adaptive potential, which may nevertheless be compromised by the combined effects of climate change and habitat fragmentation.
Phylogeny and disparate selection signatures suggest two genetically independent domestication events of pea (Pisum L.)
<p>Domestication is considered a model of adaptation that can be used to draw conclusions about the <em>modus operandi</em> of selection in natural systems. Investigating domestication may give insights into how plants react to different intensities of human manipulation, which has direct implication for the continuing efforts of crop improvement. Therefore, scientists of various disciplines study domestication-related questions to understand the biological and cultural bases of the domestication process. We employed restriction site-associated DNA sequencing (RAD-seq) of 494 <em>Pisum sativum</em> (pea) samples from all wild and domesticated groups to analyze the genetic structure of the collection. Patterns of ancient admixture were investigated by analysis of admixture graphs. We used two complementary approaches, one diversity based and one based on differentiation, to detect the selection signatures putatively associated with domestication. An analysis of the subpopulation structure of wild <em>P. sativum</em> revealed five distinct groups with a notable geographic pattern. <em>Pisum abyssinicum</em> clustered unequivocally within the <em>P. sativum</em> complex, without any indication of hybrid origin. We detected 32 genomic regions putatively subjected to selection: 29 in <em>P. sativum</em> ssp. <em>sativum</em> and three in <em>P. abyssinicum</em>. The two domesticated groups did not share regions under selection and did not display similar haplotype patterns within those regions. Wild <em>P. sativum</em> is structured into well-diverged subgroups. Although <em>Pisum sativum</em> ssp<em>. humile</em> is not supported as a taxonomic entity, the so-called 'southern <em>humile</em>' is a genuine wild group. Introgression did not shape the variation observed within the sampled germplasm. The two domesticated pea groups display distinct genetic bases of domestication, suggesting two genetically independent domestication events.</p>
Selective and non-selective evolutionary signatures found in the simplest replicative biological entities
Open the record for dataset details and reuse information.
Widespread intersex differentiation across the stickleback genome – the signature of sexually antagonistic selection?
<p>Females and males within a species commonly have distinct reproductive roles, and the associated traits may be under perpetual divergent natural selection between the sexes if their sex-specific control has not yet evolved. We here explore whether such sexually antagonistic selection can be detected based on the magnitude of differentiation between the sexes across genome-wide genetic polymorphisms by whole-genome sequencing of large pools of female and male threespine stickleback fish. We find numerous autosomal genome regions exhibiting intersex allele frequency differences beyond the range plausible under pure sampling stochasticity. Alternative sequence alignment strategies rule out that these high-differentiation regions represent sex chromosome segments misassembled into the autosomes. Instead, comparing allele frequencies and sequence read depth between the sexes reveals that regions of high intersex differentiation arise because autosomal chromosome segments got copied into the male-specific sex chromosome (Y), where they acquired new mutations. Because the Y chromosome is missing in the stickleback reference genome, sequence reads from derived DNA copies on the Y chromosome still align to the original homologous regions on the autosomes. We argue that this phenomenon hampers the identification of sexually antagonistic selection within a genome, and can lead to spurious conclusions from population genomic analyses when the underlying samples differ in sex ratios. Because the hemizygous sex chromosome sequence (Y or W) is not represented in most reference genomes, these problems may apply broadly.</p>
Data from: Whole-genome resequencing uncovers molecular signatures of natural and sexual selection in wild bighorn sheep
The identification of genes influencing fitness is central to our understanding of the genetic basis of adaptation and how it shapes phenotypic variation in wild populations. Here, we used whole-genome resequencing of wild Rocky Mountain bighorn sheep (Ovis canadensis) to >50-fold coverage to identify 2.8 million single nucleotide polymorphisms (SNPs) and genomic regions bearing signatures of directional selection (i.e. selective sweeps). A comparison of SNP diversity between the X chromosome and the autosomes indicated that bighorn males had a dramatically reduced long-term effective population size compared to females. This probably reflects a long history of intense sexual selection mediated by male–male competition for mates. Selective sweep scans based on heterozygosity and nucleotide diversity revealed evidence for a selective sweep shared across multiple populations at RXFP2, a gene that strongly affects horn size in domestic ungulates. The massive horns carried by bighorn rams appear to have evolved in part via strong positive selection at RXFP2. We identified evidence for selection within individual populations at genes affecting early body growth and cellular response to hypoxia; however, these must be interpreted more cautiously as genetic drift is strong within local populations and may have caused false positives. These results represent a rare example of strong genomic signatures of selection identified at genes with known function in wild populations of a nonmodel species. Our results also showcase the value of reference genome assemblies from agricultural or model species for studies of the genomic basis of adaptation in closely related wild taxa.
Data from: Genetic diversity, linkage disequilibrium and selection signatures in Chinese and Western pigs revealed by genome-wide SNP markers
To investigate population structure, linkage disequilibrium (LD) pattern and selection signature at the genome level in Chinese and Western pigs, we genotyped 304 unrelated animals from 18 diverse populations using porcine 60 K SNP chips. We confirmed the divergent evolution between Chinese and Western pigs and showed distinct topological structures of the tested populations. We acquired the evidence for the introgression of Western pigs into two Chinese pig breeds. Analysis of runs of homozygosity revealed that historical inbreeding reduced genetic variability in several Chinese breeds. We found that intrapopulation LD extents are roughly comparable between Chinese and Western pigs. However, interpopulation LD is much longer in Western pigs compared with Chinese pigs with average r20.3 values of 125 kb for Western pigs and only 10.5 kb for Chinese pigs. The finding indicates that higher-density markers are required to capture LD with causal variants in genome-wide association studies and genomic selection on Chinese pigs. Further, we looked across the genome to identify candidate loci under selection using FST outlier tests on two contrast samples: Tibetan pigs versus lowland pigs and belted pigs against non-belted pigs. Interestingly, we highlighted several genes including ADAMTS12, SIM1 and NOS1 that show signatures of natural selection in Tibetan pigs and are likely important for genetic adaptation to high altitude. Comparison of our findings with previous reports indicates that the underlying genetic basis for high-altitude adaptation in Tibetan pigs, Tibetan peoples and yaks is likely distinct from one another. Moreover, we identified the strongest signal of directional selection at the EDNRB loci in Chinese belted pigs, supporting EDNRB as a promising candidate gene for the white belt coat color in Chinese pigs. Altogether, our findings advance the understanding of the genome biology of Chinese and Western pigs.
Data from: Parallel signatures of selection in temporally-isolated lineages of pink salmon
Studying the effect of similar environments on diverse genetic backgrounds has long been a goal of evolutionary biologists with studies typically relying on experimental approaches. Pink salmon, a highly-abundant and widely-ranging salmonid, provide a naturally-occurring opportunity to study the effects of similar environments on divergent genetic backgrounds due to a strict two-year semelparous life-history. The species is composed of two reproductively-isolated lineages with overlapping ranges that share the same spawning and rearing environments in alternate years. We used restriction site-associated DNA (RAD) sequencing to discover and genotype approximately 8,000 SNP loci in three population pairs of even- and odd-year pink salmon along a latitudinal gradient in North America. We found greater differentiation within the odd-year than the even-year lineage and greater differentiation in the southern pair from Puget Sound than in the northern Alaskan population pairs. We identified 15 SNPs reflecting signatures of parallel selection using both a differentiation-based method (BAYESCAN) and an environmental correlation method (BAYENV). These SNPs represent genomic regions that may be particularly informative in understanding adaptive evolution in pink salmon and exploring how differing genetic backgrounds within a species respond to selection from the same natural environment.
Data from: Genomic signatures of artificial selection during early domestication of a wood crop
<p>To determine how a century of artificial selection has changed the genome of <i>E. grandis</i>, we generated SNP genotypes for 1080 individuals from three advanced South African breeding programmes using the EUChip60K chip, and investigated population structure and genome-wide differentiation patterns relative to wild progenitors.</p>
Signatures of selection in four indigenous horse breeds of Iran
<p>Indigenous Iranian horse breeds were evolutionarily affected by natural and artificial selection in distinct phylogeographic clades which shaped their genomes in several unique ways. The aims of this study were to evaluate genetic diversity and genome-wide selection signatures in four indigenous Iranian horse breeds. We evaluated 169 horses from Caspian (n = 21), Turkmen (n = 29), Kurdish (n = 67), and Persian Arabian (n = 52) populations, using genome-wide genotyping data. The contemporary effective population sizes were 59, 98, 102, and 113 for Turkmen, Caspian, Persian Arabian, and Kurdish, respectively. Analysis of population genetic structure classified the north breeds (Caspian and Turkmen) and west/southwest breeds (Persian Arabian and Kurdish) into two phylogeographic clades reflecting their geographic origin. Using a de-correlated composite of multiple selection signals statistics based on pairwise comparisons, we detected a different number of significant SNPs under putative selection from 13 to 28 for the six pairwise comparisons (FDR < 0.05). The identified SNPs under putative selection coincided with genes previously associated with known QTL for morphological, adaptation, and fitness traits. Our results showed <em>HMGA2</em> and <em>LLPH</em> as strong candidate genes for height variation between Caspian with a small size and the other studied breeds with a medium size. Using results of studies for human height retrieved from the GWAS catalog, we suggested 38 new putative candidate genes under selection. These results provide a genome-wide map of selection signatures in the studied breeds, which represent valuable information for formulating genetic conservation and improved breeding strategies for the breeds.</p>
Data from: Genetic signatures of natural selection in response to air pollution in red spruce (Picea rubens, Pinaceae)
Open the record for dataset details and reuse information.
Data from: Trait variation in response to varying winter temperatures, diversity patterns and signatures of selection along the latitudinal distribution of the widespread grassland plant Arrhenatherum elatius
Open the record for dataset details and reuse information.
Data from: Signatures of selection in the Iberian honey bee (Apis mellifera iberiensis) revealed by a genome scan analysis of single nucleotide polymorphisms
Open the record for dataset details and reuse information.
Signatures of selection in four indigenous horse breeds of Iran
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.