Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
174
datasets available to search
ShareScore release 0.9.0
Dataset results
174 results for “gene architecture”
Identification of novel genes involved in phosphate accumulation in Lotus japonicus through Genome Wide Association mapping of root system architecture and anion content
<p>130 Lotus japonicus accessions were used. The names and accession numbers are<br> listed in S6 Table. Seeds were scarified with sandpaper and then sterilized 14 minutes in 0.05%<br> sodium hypochlorite. Subsequently, seeds were rinsed and washed 5 times in sterile distilled<br> water. For the germination, seeds were positioned in imbibed filter paper, in sterile Petri dishes,<br> and wrapped in aluminium foil. After 3 days at 21°C, young seedling were transferred to square<br> plates (12 x 12 cm) containing growth medium. Both media used in this<br> study were based on Long-Ashton solution (with two levels of phosphate concentration -20 or<br> 750 μM, LP or HP, respectively) with 0.8% MES buffer (Duchefa Biochemie,<br> Haarlem, The Netherlands), 0.8% agarose (to minimize phosphate contamination), and adjusted<br> to pH 5.7 with 1M KOH. After adding the medium, plates were dried, closed, overnight in a<br> sterile laminar flow hood. Two accessions, with four replicates per each accession, were placed<br> on each plate. Each plate was replicated, with mirrored position of each accession to minimize<br> any positional growth effects. Plates were placed vertically, and plants grown under long-day<br> conditions (21°C, 16 h light/8 h dark cycle) with white light bulbs emitting 50 μmol/m 2 /s and<br> roots were exposed to light. Every day at the same time, the racks were transported to the image<br> acquisition room where images of each plate were acquired with eight Epson V600 CCD flatbed<br> color image scanners (Seiko Epson) and then immediately returned to the growth chamber.</p>
Gene flow influences the genomic architecture of local adaptation in six riverine fish species
<p>Understanding how gene flow influences adaptive divergence is important for predicting adaptive responses. Theoretical studies suggest that when gene flow is high, clustering of adaptive genes in fewer genomic regions would protect adaptive alleles from recombination and thus be selected for, but few studies have tested it with empirical data. Here, we used RADseq to generate genomic data for six fish species with contrasting life histories from six reaches of the Upper Mississippi River System, USA. We used four differentiation-based outlier tests and three genotype-environment association analyses to define neutral SNPs and outlier SNPs that were putatively under selection. We then examined the distribution of outlier SNPs along the genome and investigated whether these SNPs were found in genomic islands of differentiation and inversions. We found that gene flow varied among species, and outlier SNPs were clustered more tightly in species with higher gene flow. The two species with the highest overall <em>F</em><sub>ST</sub> (0.0303 - 0.0720) and therefore lowest gene flow showed little evidence of clusters of outlier SNPs, with outlier SNPs in these species spreading uniformly across the genome. In contrast, nearly all outlier SNPs in the species with the lowest <em>F</em><sub>ST</sub> (0.0003) were found in a single large putative inversion. Two other species with intermediate gene flow (<em>F</em><sub>ST</sub> ~ 0.0025 - 0.0050) also showed clustered genomic architectures, with most islands of differentiation clustered on a few chromosomes. Our results provide important empirical evidence to support the hypothesis that increasingly clustered architectures of local adaptation are associated with high gene flow. </p>
Intronization signatures in coding exons reveal the evolutionary fluidity of eukaryotic gene architecture
<p>This dataset accompanies the manuscript "Intronization signatures in coding exons reveal the evolutionary fluidity of eukaryotic gene architecture". It contains the relevant tables, scripts and figures used for and created during data analysis.</p>
Data from: A comparative analysis of stably expressed genes across diverse angiosperms exposes flexibility in underlying promoter architecture
<p><span>Promoters regulate both the amplitude and pattern of gene expression—key factors needed for optimization of many synthetic biology applications. Previous work in <em>Arabidopsis</em> found that promoters that contain a TATA-box element tend to be expressed only under specific conditions or in particular tissues, while promoters which lack any known promoter elements, thus designated as Coreless, tend to be expressed more ubiquitously. To test whether this trend represents a conserved promoter design rule, we identified stably expressed genes across multiple angiosperm species using publicly available RNA-seq data. Comparisons between core promoter architectures and gene expression stability revealed differences in core promoter usage in monocots and eudicots. Furthermore, when tracing the evolution of a given promoter across species, we found that core promoter type was not a strong predictor of expression stability. Our analysis suggests that core promoter types are correlative rather than causative in promoter expression patterns and highlights the challenges in finding or building constitutive promoters that will work across diverse plant species.</span></p>
Data from: A comparative analysis of stably expressed genes across diverse angiosperms exposes flexibility in underlying promoter architecture
Open the record for dataset details and reuse information.
Gene flow influences the genomic architecture of local adaptation in six riverine fish species
Open the record for dataset details and reuse information.
Data from: A genomic assessment of population structure and gene flow in an aquatic salamander identifies the roles of spatial scale, barriers, and river architecture
Open the record for dataset details and reuse information.
Insights from the timber rattlesnake (<em>Crotalus horridus</em>) genome for MHC gene architecture and evolution in threatened rattlesnakes
Open the record for dataset details and reuse information.
Data from: Genetic architecture of a hormonal response to gene knockdown in honey bees
Variation in endocrine signaling is proposed to underlie the evolution and regulation of social life histories, but the genetic architecture of endocrine signaling is still poorly understood. An excellent example of a hormonally influenced set of social traits is found in the honey bee (Apis mellifera): a dynamic and mutually suppressive relationship between juvenile hormone (JH) and the yolk precursor protein vitellogenin (Vg) regulates behavioral maturation and foraging of workers. Several other traits cosegregate with these behavioral phenotypes, comprising the pollen hoarding syndrome (PHS) one of the best-described animal behavioral syndromes. Genotype differences in responsiveness of JH to Vg are a potential mechanistic basis for the PHS. Here, we reduced Vg expression via RNA interference in progeny from a backcross between 2 selected lines of honey bees that differ in JH responsiveness to Vg reduction and measured JH response and ovary size, which represents another key aspect of the PHS. Genetic mapping based on restriction site-associated DNA tag sequencing identified suggestive quantitative trait loci (QTL) for ovary size and JH responsiveness. We confirmed genetic effects on both traits near many QTL that had been identified previously for their effect on various PHS traits. Thus, our results support a role for endocrine control of complex traits at a genetic level. Furthermore, this first example of a genetic map of a hormonal response to gene knockdown in a social insect helps to refine the genetic understanding of complex behaviors and the physiology that may underlie behavioral control in general.
Data from: Genes and QTLs controlling inflorescence and stem branch architecture in Leymus (Poaceae: Triticeae) wildrye
Grass inflorescence and stem branches show recognizable architectural differences among species. The inflorescence branches of Triticeae cereals and grasses, including wheat, barley, and 400–500 wild species, are usually contracted into a spike formation, with the number of flowering branches (spikelets) per node conserved within species and genera. Perennial Triticeae grasses of genus Leymus are unusual in that the number of spikelets per node varies, inflorescences may have panicle branches, and vegetative stems may form subterranean rhizomes. Leymus cinereus and L. triticoides show discrete differences in inflorescence length, branching architecture, node number, and density; number of spikelets per node and florets per spikelet; culm length and width; and perimeter of rhizomatous spreading. Quantitative trait loci controlling these traits were detected in 2 pseudo-backcross populations derived from the interspecific hybrids using a linkage map with 360 expressed gene sequence markers from Leymus tiller and rhizome branch meristems. Alignments of genes, mutations, and quantitative trait loci controlling similar traits in other grass species were identified using the Brachypodium genome reference sequence. Evidence suggests that loci controlling inflorescence and stem branch architecture in Leymus are conserved among the grasses, are governed by natural selection, and can serve as possible gene targets for improving seed, forage, and grain production.
Data from: Genome divergence and the genetic architecture of barriers to gene flow between Lycaeides idas and L. melissa
Genome divergence during speciation is a dynamic process that is affected by various factors, including the genetic architecture of barriers to gene flow. Herein we quantitatively describe aspects of the genetic architecture of two sets of traits, male genitalic morphology and oviposition preference, that putatively function as barriers to gene flow between the butterfly species Lycaeides idas and L. melissa. Our analyses are based on unmapped DNA sequence data and a recently developed Bayesian regression approach that includes variable selection and explicit parameters for the genetic architecture of traits. A modest number of nucleotide polymorphisms explained a small to large proportion of the variation in each trait, and average genetic variant effects were non-negligible. Several genetic regions were associated with variation in multiple traits or with trait variation within- and among-populations. In some instances genetic regions associated with trait variation also exhibited exceptional genetic differentiation between speices or exceptional introgression in hybrids. These results are consistent with the hypothesis that divergent selection on male genitalia has contributed to heterogeneous genetic differentiation, and that both sets of traits affect fitness in hybrids. Although these results are encouraging, we highlight several difficulties related to understanding the genetics of speciation.
Data from: The genetic architecture of reproductive isolation during speciation-with-gene-flow in lake whitefish species pairs assessed by RAD sequencing
During speciation-with-gene-flow, effective migration varies across the genome as a function of several factors, including proximity of selected loci, recombination rate, strength of selection, and number of selected loci. Genome scans may provide better empirical understanding of the genome-wide patterns of genetic differentiation, especially if the variance due to the previously mentioned factors is partitioned. In North American lake whitefish (Coregonus clupeaformis), glacial lineages that diverged in allopatry about 60,000 years ago and came into contact 12,000 years ago have independently evolved in several lakes into two sympatric species pairs (a normal benthic and a dwarf limnetic). Variable degrees of reproductive isolation between species pairs across lakes offer a continuum of genetic and phenotypic divergence associated with adaptation to distinct ecological niches. To disentangle the complex array of genetically based barriers that locally reduce the effective migration rate between whitefish species pairs, we compared genome-wide patterns of divergence across five lakes distributed along this divergence continuum. Using restriction site associated DNA (RAD) sequencing, we combined genetic mapping and population genetics approaches to identify genomic regions resistant to introgression and derive empirical measures of the barrier strength as a function of recombination distance. We found that the size of the genomic islands of differentiation was influenced by the joint effects of linkage disequilibrium maintained by selection on many loci, the strength of ecological niche divergence, as well as demographic characteristics unique to each lake. Partial parallelism in divergent genomic regions likely reflected the combined effects of polygenic adaptation from standing variation and independent changes in the genetic architecture of postzygotic isolation. This study illustrates how integrating genetic mapping and population genomics of multiple sympatric species pairs provide a window on the speciation-with-gene-flow mechanism.
Data from: Clines on the seashore: the genomic architecture underlying rapid divergence in the face of gene flow
Adaptive divergence and speciation may happen despite opposition by gene flow. Identifying the genomic basis underlying divergence with gene flow is a major task in evolutionary genomics. Most approaches (e.g. outlier scans) focus on genomic regions of high differentiation. However, not all genomic architectures potentially underlying divergence are expected to show extreme differentiation. Here, we develop an approach that combines hybrid zone analysis (i.e. focuses on spatial patterns of allele frequency change) with system-specific simulations to identify loci inconsistent with neutral evolution. We apply this to a genome-wide SNP set from an ideally-suited study organism, the intertidal snail Littorina saxatilis, which shows primary divergence between ecotypes associated with different shore habitats. We detect many SNPs with clinal patterns, most of which are consistent with neutrality. Among non-neutral SNPs, most are located within three large putative inversions differentiating ecotypes. Many non-neutral SNPs show relatively low levels of differentiation. We discuss potential reasons for this pattern, including loose linkage to selected variants, polygenic adaptation and a component of balancing selection within populations (which may be expected for inversions). Our work is in line with theory predicting a role for inversions in divergence, and emphasises that genomic regions contributing to divergence may not always be accessible with methods purely based on allele frequency differences. These conclusions call for approaches that take spatial patterns of allele frequency change into account in other systems.
Data from: Deciphering the genomic architecture of the stickleback brain with a novel multi-locus gene-mapping approach
Quantitative traits important to organismal function and fitness, such as brain size, are presumably controlled by many small-effect loci. Deciphering the genetic architecture of such traits with traditional quantitative trait locus (QTL) mapping methods is challenging. Here, we investigated the genetic architecture of brain size (and the size of five different brain parts) in nine-spined sticklebacks (Pungitius pungitius) with the aid of novel multi-locus QTL mapping approaches based on a de-biased LASSO method. Apart from having more statistical power to detect QTL and reduced rate of false positives than conventional QTL mapping approaches, the developed methods can handle large marker panels and provide estimates of genomic heritability. Single-locus analyses of an F2-interpopulation cross with 239 individuals and 15 198 fully informative single nucleotide polymorphisms (SNPs) uncovered 79 QTL associated with variation in stickleback brain size traits. Many of these loci were in strong linkage disequilibrium (LD) with each other, and consequently, a multi-locus mapping of individual SNPs, accounting for LD structure in the data, recovered only four significant QTL. However, a multi-locus mapping of SNPs grouped by linkage group (LG) identified 14 LGs (1-6 depending on the trait) that influence variation in brain traits. For instance, 17.6% of the variation in relative brain size was explainable by cumulative effects of SNPs distributed over six LGs, whereas 42% of the variation was accounted for by all 21 LGs. Hence, the results suggest that variation in stickleback brain traits is influenced by many small-effect loci. Apart from suggesting moderately heritable (h2 ≈ 0.15-0.42) multifactorial genetic architecture of brain traits, the results highlight the challenges in identifying the loci contributing to variation in quantitative traits. Nevertheless, the results demonstrate that the novel QTL mapping approach developed here has distinctive advantages over the traditional QTL mapping methods in analyses of dense marker panels.
Data from: Genetic architecture and genomic patterns of gene flow between hybridizing species of Picea
Hybrid zones provide an opportunity to study the effects of selection and gene flow in natural settings. We employed nuclear microsatellites (single sequence repeat (SSR)) and candidate gene single-nucleotide polymorphism markers (SNPs) to characterize the genetic architecture and patterns of interspecific gene flow in the Picea glauca × P. engelmannii hybrid zone across a broad latitudinal (40–60 degrees) and elevational (350–3500 m) range in western North America. Our results revealed a wide and complex hybrid zone with broad ancestry levels and low interspecific heterozygosity, shaped by asymmetric advanced-generation introgression, and low reproductive barriers between parental species. The clinal variation based on geographic variables, lack of concordance in clines among loci and the width of the hybrid zone points towards the maintenance of species integrity through environmental selection. Congruency between geographic and genomic clines suggests that loci with narrow clines are under strong selection, favoring either one parental species (directional selection) or their hybrids (overdominance) as a result of strong associations with climatic variables such as precipitation as snow and mean annual temperature. Cline movement due to past demographic events (evidenced by allelic richness and heterozygosity shifts from the average cline center) may explain the asymmetry in introgression and predominance of P. engelmannii found in this study. These results provide insights into the genetic architecture and fine-scale patterns of admixture, and identify loci that may be involved in reproductive barriers between the species.
The genomic architecture of the passerine MHC region: high repeat content and contrasting evolutionary histories of single copy and tandemly duplicated MHC genes
<p><span>The Major Histocompatibility Complex (MHC) is of central importance to the immune system, and an optimal MHC diversity is believed to maximize pathogen elimination. Birds show substantial variation in MHC diversity, ranging from few genes in most bird orders to very many genes in passerines. Our understanding of the evolutionary trajectories of the MHC in passerines is hampered by lack of data on genomic organization. Therefore, we assemble and annotate the MHC genomic region of the great reed warbler (<em>Acrocephalus arundinaceus</em>), using long-read sequencing and optical mapping. The MHC region is large (>5.5Mb), characterized by structural changes compared to hitherto investigated bird orders and shows higher repeat content</span><span> than the genome average. These features were supported by analyses in three additional passerines. MHC genes in passerines are found in two different chromosomal arrangements, either as single copy MHC genes located among non-MHC genes, or as tandemly duplicated tightly linked MHC genes. Some single copy MHC genes are old and putative orthologs among species. In contrast tandemly duplicated MHC genes are monophyletic within species and have evolved by simultaneous gene duplication of several MHC genes. Structural differences in the MHC genomic region among bird orders seem substantial compared to mammals and have possibly been fuelled by clade-specific immune system adaptations. Our study provides methodological guidance in characterizing complex genomic regions, constitutes a resource for MHC research in birds, and calls for a revision of the general belief that avian MHC has a conserved gene order and small size compared to mammals.</span></p>
Clines on the seashore: The genomic architecture underlying rapid divergence in the face of gene flow
<p>Adaptive divergence and speciation may happen despite opposition by gene flow. Identifying the genomic basis underlying divergence with gene flow is a major task in evolutionary genomics. Most approaches (e.g., outlier scans) focus on genomic regions of high differentiation. However, not all genomic architectures potentially underlying divergence are expected to show extreme differentiation. Here, we develop an approach that combines hybrid zone analysis (i.e., focuses on spatial patterns of allele frequency change) with system-specific simulations to identify loci inconsistent with neutral evolution. We apply this to a genome-wide SNP set from an ideally suited study organism, the intertidal snail <em>Littorina saxatilis</em>, which shows primary divergence between ecotypes associated with different shore habitats. We detect many SNPs with clinal patterns, most of which are consistent with neutrality. Among non-neutral SNPs, most are located within three large putative inversions differentiating ecotypes. Many non-neutral SNPs show relatively low levels of differentiation. We discuss potential reasons for this pattern, including loose linkage to selected variants, polygenic adaptation and a component of balancing selection within populations (which may be expected for inversions). Our work is in line with theory predicting a role for inversions in divergence, and emphasizes that genomic regions contributing to divergence may not always be accessible with methods purely based on allele frequency differences. These conclusions call for approaches that take spatial patterns of allele frequency change into account in other systems.</p>
Data from: Genetic architecture of a hormonal response to gene knockdown in honey bees
Open the record for dataset details and reuse information.
Data from: Root architecture shaping by the environment is orchestrated by dynamic gene expression in space and time
Open the record for dataset details and reuse information.
Data from: Genetic architecture and genomic patterns of gene flow between hybridizing species of Picea
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.