Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

448

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

448 results for “Genomic selection”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: AFLP genome scans suggest divergent selection on colour patterning in allopatric colour morphs of a cichlid fish

Genome scan-based tests for selection are directly applicable to natural populations to study the genetic and evolutionary mechanisms behind phenotypic differentiation. We conducted AFLP genome scans in three distinct geographic colour morphs of the cichlid fish Tropheus moorii to assess whether the extant, allopatric colour pattern differentiation can be explained by drift and to identify markers mapping to genomic regions possibly involved in colour patterning. The tested morphs occupy adjacent shore sections in southern Lake Tanganyika and are separated from each other by major habitat barriers. The genome scans revealed significant genetic structure between morphs, but a very low proportion of loci fixed for alternative AFLP alleles in different morphs. This high level of polymorphism within morphs suggested that colour pattern differentiation did not result exclusively from neutral processes. Outlier detection methods identified six loci with excess differentiation in the comparison between a bluish and a yellow-blotch morph and five different outlier loci in comparisons of each of these morphs with a red morph. As population expansions and the genetic structure of Tropheus make the outlier approach prone to false-positive signals of selection, we examined the correlation between outlier locus alleles and colour phenotypes in a genetic and phenotypic cline between two morphs. Distributions of allele frequencies at one outlier locus were indeed consistent with linkage to a colour locus. Despite the challenges posed by population structure and demography, our results encourage the cautious application of genome scans to studies of divergent selection in subdivided and recently expanded populations.

opencc-zeroDec 2011View details →
dryad32/100

Data from: The impact of selection, gene flow and demographic history on heterogeneous genomic divergence: threespine sticklebacks in divergent environments

Heterogeneous genomic divergence between populations may reflect selection, but should also be seen in conjunction with gene flow and drift, particularly population bottlenecks. Marine and freshwater threespine stickleback (Gasterosteus aculeatus) populations often exhibit different lateral armor plate morphs. Moreover, strikingly parallel genomic footprints across different marine-freshwater population pairs are interpreted as parallel evolution and gene reuse. Nevertheless, in some geographic regions like the North Sea and Baltic Sea different patterns are observed. Freshwater populations in coastal regions are often dominated by marine morphs, suggesting that gene flow overwhelms selection, and genomic parallelism may also be less pronounced. We used RAD sequencing for analyzing 28,888 SNPs in two marine and seven freshwater populations in Denmark, Europe. Freshwater populations represented a variety of environments: river populations accessible to gene flow from marine sticklebacks and large and small isolated lakes with and without fish predators. Sticklebacks in an accessible river environment showed minimal morphological and genome-wide divergence from marine populations, supporting the hypothesis of gene flow overriding selection. Allele frequency spectra suggested bottlenecks in all freshwater populations, and particularly two small lake populations. However, genomic footprints ascribed to selection could nevertheless be identified. No genomic regions were consistent freshwater-marine outliers, and parallelism was much lower than in other comparable studies. Two genomic regions previously described to be under divergent selection in freshwater and marine populations were outliers between different freshwater populations. We ascribe these patterns to stronger environmental heterogeneity among freshwater populations in our study as compared to most other studies, although the demographic history involving bottlenecks should also be considered in the interpretation of results.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Assessing the expected response to genomic selection of individuals and families in Eucalyptus breeding with an additive-dominant model

We report a genomic selection (GS) study of growth and wood quality traits in an outbred F2 hybrid Eucalyptus population (n=768) using high-density single-nucleotide polymorphism (SNP) genotyping. Going beyond previous reports in forest trees, models were developed for different selection targets, namely, families, individuals within families and individuals across the entire population using a genomic model including dominance. To provide a more breeder-intelligible assessment of the performance of GS we calculated the expected response as the percentage gain over the population average expected genetic value (EGV) for different proportions of genomically selected individuals, using a rigorous cross-validation (CV) scheme that removed relatedness between training and validation sets. Predictive abilities (PAs) were 0.40–0.57 for individual selection and 0.56–0.75 for family selection. PAs under an additive+dominance model improved predictions by 5 to 14% for growth depending on the selection target, but no improvement was seen for wood traits. The good performance of GS with no relatedness in CV suggested that our average SNP density (~25 kb) captured some short-range linkage disequilibrium. Truncation GS successfully selected individuals with an average EGV significantly higher than the population average. Response to GS on a per year basis was ~100% more efficient than by phenotypic selection and more so with higher selection intensities. These results contribute further experimental data supporting the positive prospects of GS in forest trees. Because generation times are long, traits are complex and costs of DNA genotyping are plummeting, genomic prediction has good perspectives of adoption in tree breeding practice.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Recurrent selection explains parallel evolution of genomic regions of high relative but low absolute differentiation in a ring species

Recent technological developments allow investigation of the repeatability of evolution at the genomic level. Such investigation is particularly powerful when applied to a ring species, in which spatial variation represents changes during the evolution of two species from one. We examined genomic variation among three subspecies of the greenish warbler ring species, using genotypes at 13 013 950 nucleotide sites along a new greenish warbler consensus genome assembly. Genomic regions of low within-group variation are remarkably consistent between the three populations. These regions show high relative differentiation but low absolute differentiation between populations. Comparisons with outgroup species show the locations of these peaks of relative differentiation are not well explained by phylogenetically conserved variation in recombination rates or selection. These patterns are consistent with a model in which selection in an ancestral form has reduced variation at some parts of the genome, and those same regions experience recurrent selection that subsequently reduces variation within each subspecies. The degree of heterogeneity in nucleotide diversity is greater than explained by models of background selection, but is consistent with selective sweeps. Given the evidence that greenish warblers have had both population differentiation for a long period of time and periods of gene flow between those populations, we propose that some genomic regions underwent selective sweeps over a broad geographic area followed by within-population selection-induced reductions in variation. An important implication of this 'sweep-before-differentiation' model is that genomic regions of high relative differentiation may have moved among populations more recently than other genomic regions.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Genome-wide differentiation in closely related populations: the roles of selection and geographic isolation

Population divergence in geographic isolation is due to a combination of factors. Natural and sexual selection may be important in shaping patterns of population differentiation, a pattern referred to as 'isolation by adaptation' (IBA). IBA can be complementary to the well-known pattern of 'isolation by distance' (IBD), in which the divergence of closely related populations (via any evolutionary process) is associated with geographic isolation. The barn swallow Hirundo rustica complex comprises six closely related subspecies, where divergent sexual selection is associated with phenotypic differentiation among allopatric populations. To investigate the relative contributions of selection and geographic distance to genome-wide differentiation, we compared genotypic and phenotypic variation from 350 barn swallows sampled across eight populations (28 pairwise comparisons) from four different subspecies. We report a draft whole-genome sequence for H. rustica, to which we aligned a set of 9493 single nucleotide polymorphisms (SNPs). Using statistical approaches to control for spatial autocorrelation of phenotypic variables and geographic distance, we find that divergence in traits related to migratory behaviour and sexual signalling, as well as geographic distance, together explain over 70% of genome-wide divergence among populations. Controlling for IBD, we find 42% of genomewide divergence is attributable to IBA through pairwise differences in traits related to migratory behaviour and sexual signalling alone. By (i) combining these results with prior studies of how selection shapes morphological differentiation and (ii) accounting for spatial autocorrelation, we infer that morphological adaptation plays a large role in shaping population-level differentiation in this group of closely related populations.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Combining high-throughput phenotyping and genomic information to increase prediction and selection accuracy in wheat breeding

Genomics and phenomics have promised to revolutionize the field of plant breeding. The integration of these two fields has just begun and is being driven through big data by advances in next-generation sequencing and developments of field-based high-throughput phenotyping (HTP) platforms. Each year the International Maize and Wheat Improvement Center (CIMMYT) evaluates tens-of-thousands of advanced lines for grain yield across multiple environments. To evaluate how CIMMYT may utilize dynamic HTP data for genomic selection (GS), we evaluated 1170 of these advanced lines in two environments, drought (2014, 2015) and heat (2015). A portable phenotyping system called 'Phenocart' was used to measure normalized difference vegetation index and canopy temperature simultaneously while tagging each data point with precise GPS coordinates. For genomic profiling, genotyping-by-sequencing (GBS) was used for marker discovery and genotyping. Several GS models were evaluated utilizing the 2254 GBS markers along with over 1.1 million phenotypic observations. The physiological measurements collected by HTP, whether used as a response in multivariate models or as a covariate in univariate models, resulted in a range of 33% below to 7% above the standard univariate model. Continued advances in yield prediction models as well as increasing data generating capabilities for both genomic and phenomic data will make these selection strategies tractable for plant breeders to implement increasing the rate of genetic gain.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Experimental evidence for ecological selection on genome variation in the wild

Understanding natural selection's effect on genetic variation is a major goal in biology, but the genome-scale consequences of contemporary selection are not well known. In a release and recapture field experiment we transplanted stick insects to native and novel host plants and directly measured allele frequency changes within a generation at 186 576 genetic loci. We observed substantial, genome-wide allele frequency changes during the experiment, most of which could be attributed to random mortality (genetic drift). However, we also documented that selection affected multiple genetic loci distributed across the genome, particularly in transplants to the novel host. Host-associated selection affecting the genome acted on both a known colour-pattern trait as well as other (unmeasured) phenotypes. We also found evidence that selection associated with elevation affected genome variation, although our experiment was not designed to test this. Our results illustrate how genomic data can identify previously underappreciated ecological sources and phenotypic targets of selection.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Genome-wide analysis of colonization history and concomitant selection in Arabidopsis lyrata

The high climatic variability in the past hundred thousand years has affected the demographic and adaptive processes in many species, especially in boreal and temperate regions undergoing glacial cycles. This has also influenced the patterns of genome-wide nucleotide variation, but the details of these effects are largely unknown. Here we study the patterns of genome-wide variation to infer colonization history and patterns of selection of the perennial herb species Arabidopsis lyrata, in locally adapted populations from different parts of its distribution range (Germany, UK, Norway, Sweden, and USA) representing different environmental conditions. Using site frequency spectra based demographic modelling we found strong reduction in the effective population size of the species in general within the past 100 000 years, with more pronounced effects in the colonizing populations. We further found that the northwestern European A. lyrata populations (UK and Scandinavian) are more closely related to each other than with the Central European populations, and coalescent based population split modelling suggests that western European and Scandinavian populations became isolated relatively recently after the glacial retreat. We also highlighted loci showing evidence for local selection associated with the Scandinavian colonization. The results presented here give new insights into post-glacial Scandinavian colonization history and its genome-wide effects.

opencc-zeroDec 2016View details →
dryad32/100

Data from: High genomic diversity and candidate genes under selection associated with range expansion in eastern coyote (Canis latrans) populations

Range expansion is a widespread biological process, with well described theoretical expectations for the genomic outcomes accompanying the colonization of novel ranges. However, comparatively few empirical studies address the genome-wide consequences associated with the range expansion process, particularly in recent or on-going expansions. Here, we assess two recent and distinct eastward expansion fronts of a highly mobile carnivore, the coyote (Canis latrans), to investigate patterns of genomic diversity and identify variants that may have been under selection during range expansion. Using a restriction enzyme assisted sequencing approach (RADseq), we genotyped 394 coyotes at 22,935 SNPs and found that overall population structure corresponded to their 19th century historical range and two distinct populations that expanded during the 20th century. Counter to theoretical expectations for populations to bottleneck during range expansions, we observed minimal evidence for decreased genomic diversity across coyotes sampled along either expansion front, which is likely due to hybridization with other Canis species. Furthermore, we identified 12 SNPs, located either within genes or putative regulatory regions, that were consistently associated with range expansion. Of these 12 genes, three (CACNA1C, ALK, and EPHA6) have putative functions related to dispersal, including habituation to novel environments and spatial learning, consistent with the expectations for traits under selection during range expansion. Although coyote colonization of eastern North America is well-publicized, this study provides novel insights by identifying genes associated with dispersal capabilities in coyotes on the two eastern expansion fronts.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Population genomic footprints of selection and associations with climate in natural populations of Arabidopsis halleri from the Alps

Natural genetic variation is essential for the adaptation of organisms to their local environment and to changing environmental conditions. Here we examine genome-wide patterns of nucleotide variation in natural populations of the outcrossing herb Arabidopsis halleri and associations with climatic variation among populations in the Alps. Using a pooled population sequencing (Pool-Seq) approach, we discovered more than two million SNPs in five natural populations and identified highly differentiated genomic regions and SNPs using FST–based analyses. We tested only the most strongly differentiated SNPs for associations with a non-redundant set of environmental factors using partial Mantel tests to identify topo-climatic factors that may underlie the observed footprints of selection. Possible functions of genes showing signatures of selection were identified by Gene Ontology analysis. We found 175 genes to be highly associated with one or more of the five tested topo-climatic factors. Of these, 23.4% had unknown functions. Genetic variation in four candidate genes was strongly associated with site water balance and solar radiation, and functional annotations were congruent with these environmental factors. Our results provide a genome-wide perspective on the distribution of adaptive genetic variation in natural plant populations from a highly diverse and heterogeneous alpine environment.

opencc-zeroDec 2012View details →
dryad32/100

Data from: The role of parasite-driven selection in shaping landscape genomic structure in red grouse (Lagopus lagopus scotica)

Landscape genomics promises to provide novel insights into how neutral and adaptive processes shape genome-wide variation within and among populations. However, there has been little emphasis on examining whether individual-based phenotype-genotype relationships derived from approaches such as genome-wide association (GWAS) manifest themselves as a population-level signature of selection in a landscape context. The two may prove irreconcilable as individual-level patterns may become diluted by high levels of gene flow and complex phenotypic or environmental heterogeneity. We illustrate this issue with a case study that examines the role of the highly prevalent gastrointestinal nematode Trichostrongylus tenuis in shaping genomic signatures of selection in red grouse (Lagopus lagopus scotica). Individual-level GWAS involving 384 SNPs has previously identified five SNPs that explain variation in T. tenuis burden. Here, we examine whether these same SNPs display population-level relationships between T. tenuis burden and genetic structure across a small-scale landscape of 21 sites with heterogeneous parasite pressure. Moreover, we identify adaptive SNPs showing signatures of directional selection using FST outlier analysis and relate population- and individual-level patterns of multi-locus neutral and adaptive genetic structure to T. tenuis burden. The five candidate SNPs for parasite-driven selection were neither associated with T. tenuis burden on a population level, nor under directional selection. Similarly, there was no evidence of parasite-driven selection in SNPs identified as FST outliers. We discuss these results in the context of red grouse ecology and highlight the broader consequences for the utility of landscape genomics approaches for identifying signatures of selection.

opencc-zeroDec 2014View details →
dryad32/100

Data from: The role of selection in driving landscape genomic structure of the waterflea Daphnia magna

The combined analysis of neutral and adaptive genetic variation is crucial to reconstruct the processes driving population genetic structure of natural populations. However, such combined analysis is challenging because of the complex interaction among neutral and selective processes in the landscape. Overcoming this level of complexity requires an unbiased search for the evidence of selection in the genomes of populations sampled from their natural habitats and the identification of demographic processes that lead to present-day populations genetic structure. Ecological model species with a suite of genomic tools and well-understood ecologies are best suited to resolve this complexity and elucidate the role of selective and demographic processes in the landscape genomic structure of natural populations. Here we investigate the water flea Daphnia magna, an emerging model system in genomics and a renowned ecological model system. We infer past and recent demographic processes by contrasting patterns of local and regional neutral genetic diversity at markers with different mutation rates. We assess the role of the environment in driving genetic variation in our study system by identifying correlates between biotic and abiotic variables naturally occurring in the landscape and patterns of neutral and adaptive genetic variation. Our results indicate that selection plays a major role in determining the population genomic structure of D. magna. First, environmental selection directly impacts genetic variation at loci hitchhiking with genes under selection. Secondly, priority effects enhanced by local genetic adaptation (cf. monopolization) affect neutral genetic variation by reducing gene flow among populations and genetic diversity within populations.

opencc-zeroDec 2011View details →
dryad32/100

Data from: Scans for signatures of selection in Russian cattle breed genomes reveal new candidate genes for environmental adaptation and acclimation

Domestication and selective breeding has resulted in over 1000 extant cattle breeds. Many of these breeds do not excel in important traits but are adapted to local environments. These adaptations are a valuable source of genetic material for efforts to improve commercial breeds. As a step toward this goal we identified candidate regions to be under selection in genomes of nine Russian native cattle breeds adapted to survive in harsh climates. After comparing our data to other breeds of European and Asian origins we found known and novel candidate genes that could potentially be related to domestication, economically important traits and environmental adaptations in cattle. The Russian cattle breed genomes contained regions under putative selection with genes that may be related to adaptations to harsh environments (e.g., AQP5, RAD50, and RETREG1). We found genomic signatures of selective sweeps near key genes related to economically important traits, such as the milk production (e.g., DGAT1, ABCG2), growth (e.g., XKR4), and reproduction (e.g., CSF2). Our data point to candidate genes which should be included in future studies attempting to identify genes to improve the extant breeds and facilitate generation of commercial breeds that fit better into the environments of Russia and other countries with similar climates.

opencc-zeroDec 2017View details →
dryad32/100

Data from: A comparison of genomic selection models across time in interior spruce (Picea engelmannii × glauca) using unordered SNP imputation methods

Genomic selection (GS) potentially offers an unparalleled advantage over traditional pedigree-based selection (TS) methods by reducing the time commitment required to carry out a single cycle of tree improvement. This quality is particularly appealing to tree breeders, where lengthy improvement cycles are the norm. We explored the prospect of implementing GS for interior spruce (Picea engelmannii × glauca) utilizing a genotyped population of 769 trees belonging to 25 open-pollinated families. A series of repeated tree height measurements through ages 3–40 years permitted the testing of GS methods temporally. The genotyping-by-sequencing (GBS) platform was used for single nucleotide polymorphism (SNP) discovery in conjunction with three unordered imputation methods applied to a data set with 60% missing information. Further, three diverse GS models were evaluated based on predictive accuracy (PA), and their marker effects. Moderate levels of PA (0.31–0.55) were observed and were of sufficient capacity to deliver improved selection response over TS. Additionally, PA varied substantially through time accordingly with spatial competition among trees. As expected, temporal PA was well correlated with age-age genetic correlation (r=0.99), and decreased substantially with increasing difference in age between the training and validation populations (0.04–0.47). Moreover, our imputation comparisons indicate that k-nearest neighbor and singular value decomposition yielded a greater number of SNPs and gave higher predictive accuracies than imputing with the mean. Furthermore, the ridge regression (rrBLUP) and BayesCπ (BCπ) models both yielded equal, and better PA than the generalized ridge regression heteroscedastic effect model for the traits evaluated.

opencc-zeroDec 2014View details →
dryad32/100

Data from: The population genomic signature of environmental selection in the widespread insect-pollinated tree species Frangula alnus at different geographical scales

The evaluation of the molecular signatures of selection in species lacking an available closely related reference genome remains challenging, yet it may provide valuable fundamental insights into the capacity of populations to respond to environmental cues. We screened 25 native populations of the tree species Frangula alnus subsp. alnus (Rhamnaceae), covering three different geographical scales, for 183 annotated single-nucleotide polymorphisms (SNPs). Standard population genomic outlier screens were combined with individual-based and multivariate landscape genomic approaches to examine the strength of selection relative to neutral processes in shaping genomic variation, and to identify the main environmental agents driving selection. Our results demonstrate a more distinct signature of selection with increasing geographical distance, as indicated by the proportion of SNPs (i) showing exceptional patterns of genetic diversity and differentiation (outliers) and (ii) associated with climate. Both temperature and precipitation have an important role as selective agents in shaping adaptive genomic differentiation in F. alnus subsp. alnus, although their relative importance differed among spatial scales. At the 'intermediate' and 'regional' scales, where limited genetic clustering and high population diversity were observed, some indications of natural selection may suggest a major role for gene flow in safeguarding adaptability. High genetic diversity at loci under selection in particular, indicated considerable adaptive potential, which may nevertheless be compromised by the combined effects of climate change and habitat fragmentation.

opencc-zeroDec 2014View details →
dryad32/100

The roles of recombination and selection in shaping genomic divergence in an incipient ecological species complex

<p>Speciation genomic studies have revealed that genomes of diverging lineages are shaped jointly by the actions of gene flow and selection. These evolutionary forces acting in concert with processes such as recombination and genome features such as gene density shape a mosaic landscape of divergence. We investigated the roles of recombination and gene density in shaping the patterns of differentiation and divergence between the cyclically parthenogenetic ecological sister-taxa, <i>Daphnia pulicaria </i>and <i>Daphnia pulex. </i>First, we assembled a phased chromosome-scale genome assembly using trio-binning for <i>D. pulicaria</i> and constructed a genetic map using an F2-intercross panel to understand sex-specific recombination rate heterogeneity.<i> </i>Finally, we used a ddRADseq dataset with broad geographic sampling of <i>D. pulicaria, D. pulex, </i>and their hybrids to understand the patterns of genome-scale divergence and demographic parameters. Our study provides the first sex-specific estimates of recombination rates for a cyclical parthenogen, and unlike other eukaryotic species, we observed male-biased heterochiasmy in <i>D. pulicaria</i>, which may be related to this somewhat unique breeding mode. Additionally, regions of high gene density and recombination are generally more divergent than regions of suppressed recombination. Outlier analysis indicated that divergent genomic regions are likely driven by selection on <i>D. pulicaria</i>, the derived lineage colonizing a novel lake habitat. Together, our study supports a scenario of selection acting on genes related to local adaptation shaping genome-wide patterns of differentiation despite high local recombination rates in this species complex. Finally, we discuss the limitations of our data in light of demographic uncertainty.</p>

opencc-zeroJan 2022View details →
dryad32/100

Data from: Population genomic evidence of selection on structural variants in a natural hybrid zone

<p><span>Structural variants (SVs) can promote speciation by directly causing reproductive isolation or by suppressing recombination across large genomic regions. Whereas examples of each mechanism have been documented, systematic tests of the role of SVs in speciation are lacking. Here, we take advantage of long-read (Oxford nanopore) whole-genome sequencing and a hybrid zone between two </span><em>Lycaeides</em> butterfly taxa (<em>L. melissa</em> and Jackson Hole <em>Lycaeides</em>) to comprehensively evaluate genome-wide patterns of introgression for SVs and relate these patterns to hypotheses about speciation. We found &gt;100,000 SVs segregating within or between the two hybridizing species. SVs and SNPs exhibited similar levels of genetic differentiation between species, with the exception of inversions, which were more differentiated. We detected credible variation in patterns of introgression among SV loci in the hybrid zone, with 562 of 1419 ancestry-informative SVs exhibiting genomic clines that deviated from null expectations based on genome-average ancestry. Overall, hybrids exhibited a directional shift towards Jackson Hole <em>Lycaeides</em> ancestry at SV loci, consistent with the hypothesis that these loci experienced more selection on average than SNP loci. Surprisingly, we found that deletions, rather than inversions, showed the highest skew towards excess ancestry from Jackson Hole <em>Lycaeides</em>. Excess Jackson Hole <em>Lycaeides</em> ancestry in hybrids was also especially pronounced for Z-linked SVs and inversions containing many genes. In conclusion, our results show that SVs are ubiquitous and suggest that SVs in general, but especially deletions, might disproportionately affect hybrid fitness and thus contribute to reproductive isolation.</p>

opencc-zeroApr 2022View details →
zenodo32/100

Western redcedar single nucleotide polymorphism (SNP) genotyping data for genomic selection and population genetics

<p>Western redcedar (<em>Thuja plicata</em>) Single Nucleotide Polymorphism (SNP) data in Variant Call Format (VCF) for genomic selection training and target populations, genomic selection parents, and self-fertilized (selfing) lines, comprising 4,833 trees.</p> <p>Targeted sequencing-based genotyping was done by Capture-Seq methodology at Rapid Genomics (Neves est al. 2013). A set of 57,000 probes as designed for initial marker discovery, from which a panel of 20,858 probes was selected for genotyping. A set of transcriptomes (Shalev et al. 2018) (PRJNA704616) was aligned to the reference genome to identify SNPs. Candidate probes (120 nt) were initially designed in silico and 57,000 selected by removing candidates with poor base composition for hybridization (GC content &lt;0.2 and &gt;0.6, high G content &gt;0.2 and long homopolymers &gt;7), followed by removing probes aligning to more than one position on the reference genome (&ge;90% identity and length). The 57,000 probes represent 14,517 scaffolds (average 3.9 probes/scaffold), with 37,275 targeting at least one SNP and 19,725 mapping to intergenic regions not containing pre-identified SNPs. A set of 128 individuals were selected to validate the 57,000 probe panel and associated polymorphisms. Genomic DNA (0.5 ug) was fragmented (mean size 300 bp), followed by repair of ends, phosphorylation, adenylation, ligation of Illumina compatible adapters containing 8bp indexes and 5&rsquo; T-overhang, and 10 cycles PCR amplification with universal primers to produce sequencing-ready libraries. Libraries were quantified using PicoGreen. Libraries from 16 samples were pooled, hybridized to the 120 nt RNA probes following Agilent&rsquo;s SureSelect Target Enrichment System (Agilent Technologies) and sequenced on an Illumina HiSeq X machine with paired-end 150bp cycle for an average sequencing depth per sample of 15X. Sequence data were aligned to the reference genome with BWA-MEM (http://arxiv.org/abs/1303.3997) and sets of four samples were combined to increase sequencing depth for identifying markers. Putative SNPs were identified using Freebayes (http://arxiv.org/abs/1207.3907) in 150bp on either side of the 57,000 probes and filtered probes that had more than 17 SNPs per 420 bp target region (150bp + 120bp + 150bp). The sequencing depth of the probes was used to select the final set of 20,885 probes, removing probes on both sides of the distribution (low and high sequencing depth), for Capture-Seq on the remainder of the samples.</p>

opencc-by-4.0Sep 2022View details →
dryad32/100

Data from: Improvement of genomic predictions in small breeds by construction of genomic relationship matrix through variable selection

<p>Genomic selection has been increasingly implemented in the animal breeding industry, and it is becoming a routine method in many livestock breeding contexts. However, its use is still limited in several small-population local breeds, which are, nonetheless, an important source of genetic variability of great economic value. A major roadblock for their genomic selection is accuracy when population size is limited: to improve breeding value accuracy, variable selection models that assume heterogenous variance have been proposed over the last few years. However, while these models might outperform traditional and genomic predictions in terms of accuracy, they also carry a proportional increase of breeding value bias and dispersion. These mutual increases are especially striking when genomic selection is performed with a low number of phenotypes and high shrinkage value—which is precisely the situation that happens with small local breeds. In our study, we tested several alternative methods to improve the accuracy of genomic selection in a small population. First, we investigated the impact of using only a subset of informative markers regarding prediction accuracy, bias, and dispersion. We used different algorithms to select them, such as recursive feature eliminations, penalized regression, and XGBoost. We compared our results with the predictions of pedigree-based BLUP, single-step genomic BLUP, and weighted single-step genomic BLUP in different simulated populations obtained by combining various parameters in terms of number of QTLs and effective population size. We also investigated these approaches on a real data set belonging to the small local Rendena breed. Our results show that the accuracy of GBLUP in small-sized populations increased when performed with SNPs selected via variable selection methods both in simulated and real data sets. In addition, the use of variable selection models—especially those using XGBoost—in our real data set did not impact bias and the dispersion of estimated breeding values. We have discussed possible explanations for our results and how our study can help estimate breeding values for future genomic selection in small breeds.</p>

opencc-zeroAug 2022View details →
dryad32/100

Variant Call File (VCF) for Genome-wide polymorphism and genic selection in feral and domesticated lineages of Cannabis sativa

<p>A comprehensive understanding of the degree to which genomic variation is maintained by selection versus drift and gene flow is lacking in many important species such as <em>Cannabis</em> <em>sativa </em>(<em>C. sativa</em>), one of the oldest known crops to be cultivated by humans worldwide. We generated whole genome resequencing data across diverse samples of feralized (escaped domesticated lineages) and domesticated lineages of <em>C. sativa</em>. We performed analyses to examine population structure, and genome wide scans for FST, balancing selection, and positive selection. Our analyses identified evidence for sub-population structure and further support the Asian origin hypothesis of this species. Feral plants sourced from the U.S. exhibited broad regions on chromosomes 4 and 10 with high <span>𝐹̅</span>ST which may indicate chromosomal inversions maintained at high frequency in this sub-population. Both our balancing and positive selection analyses identified loci that may reflect differential selection for traits favored by natural selection and artificial selection in feral versus domesticated sub-populations. In the U.S. feral sub-population, we found six loci related to stress response under balancing selection and one gene involved in disease resistance under positive selection, suggesting local adaptation to new climates and biotic interactions. In the marijuana sub-population, we identified the gene <em>SMALLER TRICHOMES</em> <em>WITH VARIABLE BRANCHES 2 </em>to be under positive selection which suggests artificial selection for increased tetrahydrocannabinol yield. Overall the data generated, and results obtained from our study help to form a better understanding of the evolutionary history in <em>C. sativa</em>.</p>

opencc-zeroAug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record