Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
314
datasets available to search
ShareScore release 0.7.1
Dataset results
314 results for “genome architecture”
Data from: Genome-wide association mapping within a local Arabidopsis thaliana population more fully reveals the genetic architecture for defensive metabolite diversity
<p>A paradoxical finding from genome-wide association studies (GWAS) in plants is that variation in metabolite profiles typically maps to a small number of loci, despite the complexity of underlying biosynthetic pathways. This discrepancy may partially arise from limitations presented by geographically diverse mapping panels. Properties of metabolic pathways that impede GWAS by diluting the additive effect of a causal variant, such as allelic and genic heterogeneity and epistasis, would be expected to increase in severity with the geographic range of the mapping panel. We hypothesized that a population from a single locality would reveal an expanded set of associated loci. We tested this in a French <em>Arabidopsis thaliana</em> population (< 1 km transect) by profiling and conducting GWAS for glucosinolates, a suite of defensive metabolites that have been studied in depth through functional and genetic mapping approaches. For two distinct classes of glucosinolates, we discovered more associations at biosynthetic loci than previous GWAS with continental-scale mapping panels. Candidate genes underlying novel associations were supported by concordance between their observed effects in the TOU-A population and previous functional genetic and biochemical characterization. Local populations complement geographically diverse mapping panels to reveal a more complete genetic architecture for metabolic traits.</p>
Identification of novel genes involved in phosphate accumulation in Lotus japonicus through Genome Wide Association mapping of root system architecture and anion content
<p>130 Lotus japonicus accessions were used. The names and accession numbers are<br> listed in S6 Table. Seeds were scarified with sandpaper and then sterilized 14 minutes in 0.05%<br> sodium hypochlorite. Subsequently, seeds were rinsed and washed 5 times in sterile distilled<br> water. For the germination, seeds were positioned in imbibed filter paper, in sterile Petri dishes,<br> and wrapped in aluminium foil. After 3 days at 21°C, young seedling were transferred to square<br> plates (12 x 12 cm) containing growth medium. Both media used in this<br> study were based on Long-Ashton solution (with two levels of phosphate concentration -20 or<br> 750 μM, LP or HP, respectively) with 0.8% MES buffer (Duchefa Biochemie,<br> Haarlem, The Netherlands), 0.8% agarose (to minimize phosphate contamination), and adjusted<br> to pH 5.7 with 1M KOH. After adding the medium, plates were dried, closed, overnight in a<br> sterile laminar flow hood. Two accessions, with four replicates per each accession, were placed<br> on each plate. Each plate was replicated, with mirrored position of each accession to minimize<br> any positional growth effects. Plates were placed vertically, and plants grown under long-day<br> conditions (21°C, 16 h light/8 h dark cycle) with white light bulbs emitting 50 μmol/m 2 /s and<br> roots were exposed to light. Every day at the same time, the racks were transported to the image<br> acquisition room where images of each plate were acquired with eight Epson V600 CCD flatbed<br> color image scanners (Seiko Epson) and then immediately returned to the growth chamber.</p>
Genome sequencing of 2,000 canids advances the understanding of demography, genome function and architecture
<p><strong>Background: </strong>The international Dog10K project aims to sequence and analyze several thousand canine genomes. Incorporating 20x data from 1,987 individuals, including 1,611 dogs (321 breeds), 309 village dogs, 63 wolves and four coyotes, we identify genomic variation across the canid family, setting the stage for detailed studies of domestication, behavior, morphology, disease susceptibility and genome architecture and function.</p> <p><strong>Results: </strong>We report the analysis of >48M single nucleotide, indel, and structural variants spanning the autosomes, X chromosome and mitochondria. We discover more than 75% of variation for 239 sampled breeds. Allele sharing analysis indicates that 94.9% of breeds form monophyletic clusters and 25 major clades. German Shepherd Dogs and related breeds show the highest allele sharing with independent breeds from multiple clades. On average, each breed dog differs from the UU_Cfam_GSD_1.0 reference at 26,960 deletions and 14,034 insertions greater than 50bp, with wolves having 14% more variants. Discovered variants include retrogene insertions from 926 parent genes. To aid functional prioritization, single nucleotide variants were annotated with SnpEff and Zoonomia phyloP constraint scores. Constrained positions were negatively correlated with allele frequency. Finally, the utility of the Dog10K data as an imputation reference panel is assessed, generating high confidence calls across varied genotyping platform densities including for breeds not included in the Dog10K collection.</p> <p><strong>Conclusions:</strong> We have developed a dense dataset of 1,987 sequenced canids that reveals patterns of allele sharing, identifies likely functional variants, informs breed structure, and enables accurate imputation. Dog10K data are publicly available</p>
Local adaptation and the evolution of genome architecture in threespine stickleback
<p class="MsoNormal"><span>Theory predicts that local adaptation should favour the evolution of a concentrated genetic architecture, where the alleles driving adaptive divergence are tightly clustered on chromosomes. Adaptation to marine vs. freshwater environments in threespine stickleback has resulted in an architecture that seems consistent with this prediction: divergence among populations is mainly driven by a few genomic regions harbouring multiple quantitative trait loci (QTL) for environmentally adapted traits, as well as candidate genes with well-established phenotypic effects. One theory for the evolution of these "genomic islands" is that rearrangements remodel the genome to bring causal loci into tight proximity, but this has not been studied explicitly. We tested this theory using synteny analysis to identify micro- and macro-rearrangements in the stickleback genome and assess their potential involvement in the evolution of genomic islands. To identify rearrangements, we conducted a <em>de novo</em> assembly of the closely-related tubesnout (<em>Aulorhyncus flavidus</em>) genome and compared this to the genomes of threespine stickleback and two other closely related species. We found that small rearrangements, within-chromosome duplications, and Lineage-Specific Genes (LSGs) were enriched around genomic islands, and that all three chromosomes harbouring large genomic islands have experienced macro-rearrangements. We also found that duplicates and micro-rearrangements are 9.9x and 2.9x more likely to involve genes differentially expressed between marine and freshwater genotypes. While not conclusive, these results are consistent with the explanation that strong divergent selection on candidate genes drove the recruitment of rearrangements to yield clusters of locally adaptive loci.</span></p>
Collection of runs with VIA genomics workload on RISC-V architectures
<p>This is a repository of results obtained experimenting with VIA genomics workload on RISC-V architectures. The current repository will be moved into a proper web-site with tools to properly visualize the data gathered. However, for now, we provide the results as txt files with the output of the workload for the several cases we examined so far. Within the dataset there is a README explaining the nomenclature used to save the results.</p>
Local adaptation and the evolution of genome architecture in threespine stickleback
Open the record for dataset details and reuse information.
Data from: Genome-wide association mapping within a local Arabidopsis thaliana population more fully reveals the genetic architecture for defensive metabolite diversity
Open the record for dataset details and reuse information.
Input data of manuscript "CACTUS: integrating clonal architecture with genomic clustering and transcriptome profiling of single tumor cells"
<p>This is the directory containing input data necessary to reproduce analyses presented in the manuscript:</p> <blockquote> <p><strong>CACTUS: integrating clonal architecture with genomic clustering and transcriptome profiling of single tumor cells</strong><br> Shadi Darvish Shafighi, Szymon M Kiełbasa, Julieta Sepúlveda Yáñez, Ramin Monajemi, Davy Cats, Hailiang Mei, Roberta Menafra, Susan Kloet, Hendrik Veelken, Cornelis A.M. van Bergen, Ewa Szczurek</p> </blockquote>
Data from: Cryptic species in the mountaintops: species delimitation and taxonomy of the Bembidion breve species group (Coleoptera: Carabidae) aided by genomic architecture of a century-old type specimen
The breve species group includes closely related Bembidion Latreille ground beetles commonly found at high elevation in the mountains of western North America. For several decades, the group has been considered to consist of two species. Here, we present evidence from morphological, molecular and geographic data that the group contains nine species: Bembidion ampliatum, B. breve, B. geopearlis, B. laxatum, B. lividulum, B. oromaia, B. saturatum, B. testatum and B. vulcanix. We describe three species (B. geopearlis, B. oromaia and B. vulcanix) as new and resurrect four previously synonymized names (B. ampliatum, B. lividulum, B. saturatum and B. testatum). Species diversity is highest throughout the Cascades in Oregon and Washington, and Sierra Nevada of California, where up to seven species can occur in sympatry. We resolved challenging nomenclatural issues through analysis of sequences obtained from century-old type specimens by using a novel application of rDNA copy number analysis – an approach that may prove useful for other historical specimens.
Data from: Pan-evolutionary and regulatory genome architecture delineated by integrated macro- and microsynteny approach
<p>Based on the published algorithms or tools developed by our and other groups, we introduce a detailed protocol for the most comprehensive and up-to-date genome synteny pipeline (called PanSyn) and provides step-by-step instructions as well as application examples for demonstrating how to use it. PanSyn pipeline includes three major modules (microsynteny analysis, macrosynteny analysis, and integrated micro & macro analysis). PanSyn not only fills a gap of lacking a user-friendly, highly-customized tool for genome macrosynteny analysis but also allows for integrated pan-evolutionary and regulatory analysis of genome microsyntenty and macrosynteny which are not yet available in any public synteny software or tools. PanSyn has been tested under Linux system. PanSyn has multiple subroutines. Users only need to simply modify the configuration file and corresponding command parameters to execute them. Outputs include vector diagrams that are suitable for custom modification. <br> </p>
Dissecting the genetic architecture of quantitative traits using genome-wide identity-by-descent sharing
<p>Additive and dominance genetic variances underlying the expression of quantitative traits are important quantities for predicting short-term responses to selection, but they are notoriously challenging to estimate in most non-model wild populations. Specifically, large-sized or panmictic populations may be characterized by low variance in genetic relatedness among individuals which in turn, can prevent accurate estimation of quantitative genetic parameters. We used estimates of genome-wide identity-by-descent (IBD) sharing from autosomal SNP loci to estimate quantitative genetic parameters for ecologically important traits in nine-spined sticklebacks (<em>Pungitius pungitius</em>) from a large, outbred population. Using empirical and simulated datasets, with varying sample sizes and pedigree complexity, we assessed the performance of different crossing schemes in estimating additive genetic variance and heritability for all traits. We found that low variance in relatedness characteristic of wild outbred populations with high migration rate can impair the estimation of quantitative genetic parameters and bias heritability estimates downwards. On the other hand, the use of a half-sib/full-sib design allowed precise estimation of genetic variance components, and revealed significant additive variance and heritability for all measured traits, with negligible dominance contributions. Genome-partitioning and QTL mapping analyses revealed that most traits had a polygenic basis and were controlled by genes at multiple chromosomes. Furthermore, different QTL contributed to variation in the same traits in different populations suggesting heterogenous underpinnings of parallel evolution at the phenotypic level. Our results provide important guidelines for future studies aimed at estimating adaptive potential in the wild, particularly for those conducted in outbred large-sized populations.</p>
Gene flow influences the genomic architecture of local adaptation in six riverine fish species
<p>Understanding how gene flow influences adaptive divergence is important for predicting adaptive responses. Theoretical studies suggest that when gene flow is high, clustering of adaptive genes in fewer genomic regions would protect adaptive alleles from recombination and thus be selected for, but few studies have tested it with empirical data. Here, we used RADseq to generate genomic data for six fish species with contrasting life histories from six reaches of the Upper Mississippi River System, USA. We used four differentiation-based outlier tests and three genotype-environment association analyses to define neutral SNPs and outlier SNPs that were putatively under selection. We then examined the distribution of outlier SNPs along the genome and investigated whether these SNPs were found in genomic islands of differentiation and inversions. We found that gene flow varied among species, and outlier SNPs were clustered more tightly in species with higher gene flow. The two species with the highest overall <em>F</em><sub>ST</sub> (0.0303 - 0.0720) and therefore lowest gene flow showed little evidence of clusters of outlier SNPs, with outlier SNPs in these species spreading uniformly across the genome. In contrast, nearly all outlier SNPs in the species with the lowest <em>F</em><sub>ST</sub> (0.0003) were found in a single large putative inversion. Two other species with intermediate gene flow (<em>F</em><sub>ST</sub> ~ 0.0025 - 0.0050) also showed clustered genomic architectures, with most islands of differentiation clustered on a few chromosomes. Our results provide important empirical evidence to support the hypothesis that increasingly clustered architectures of local adaptation are associated with high gene flow. </p>
Genome-wide association and multi-trait analyses characterize the common genetic architecture of heart failure
<p>Genome-wide association study summary statistics.</p>
Data from: Genomic architecture and introgression shape a butterfly radiation
We probe the history of rapidly radiating Heliconius butterflies by means of 20 new genome assemblies and employ them to investigate the genomic architecture of gene flow among lineages. By developing a test to distinguish incomplete lineage sorting from introgression, we demonstrate that histories of loci that differ from the species tree arose mostly through introgression. Moreover, these loci are underrepresented in low recombination and gene-rich regions, consistent with the purging of introgressed alleles tightly linked with incompatibility loci. Additionally, our analysis identifies an inversion that captures a color pattern switch locus which was transferred between lineages via introgression and is convergent with a similar rearrangement in another part of the genus. This analysis of multiple de novo genome sequences enables an improved understanding of the importance of introgression and selective processes in adaptive radiation.
Data from: Genome-wide association studies across environmental and genetic contexts reveal complex genetic architecture of symbiotic extended phenotypes
<p>A goal of modern biology is to develop the genotype-phenotype (G→P) map, a predictive understanding of how genomic information generates trait variation that forms the basis of both natural and managed communities. As microbiome research advances, however, it has become clear that many of these traits are symbiotic extended phenotypes, being governed by genetic variation encoded not only by the host's own genome, but also by the genomes of myriad cryptic symbionts. Building a reliable G→P map therefore requires accounting for the multitude of interacting genes and even genomes involved in symbiosis. Here we use naturally-occurring genetic variation in 191 strains of the model microbial symbiont <em>Sinorhizobium meliloti</em> paired with two genotypes of the host <em>Medicago truncatula</em> in four genome-wide association studies (GWAS) to determine the genomic architecture of a key symbiotic extended phenotype – partner quality, or the fitness benefit conferred to a host by a particular symbiont genotype, within and across environmental contexts and host genotypes. We define three novel categories of loci in rhizobium genomes that must be accounted for if we want to build a reliable G→P map of partner quality; namely, 1) loci whose identities depend on the environment, 2) those that depend on the host genotype with which rhizobia interact, and 3) universal loci that are likely important in all or most environments.</p> <p><span>IMPORTANCE:</span><strong> </strong>Given the rapid rise of research on how microbiomes can be harnessed to improve host health, understanding the contribution of microbial genetic variation to host phenotypic variation is pressing, and will better enable us to predict the evolution of (and select more precisely for) symbiotic extended phenotypes that impact host health. We uncover extensive context-dependency in both the identity and functions of symbiont loci that control host growth, which makes predicting the genes and pathways important for determining symbiotic outcomes under different conditions more challenging. Despite this context-dependency, we also resolve a core set of universal loci that are likely important in all or most environments, and thus, serve as excellent targets both for genetic engineering and future coevolutionary studies of symbiosis.</p>
Nationwide genomic biobank in Mexico unravels demographic history and complex trait architecture from 6,057 individuals: GWAS summary statistics
<p>Latin America continues to be severely underrepresented in genomics research, and fine-scale genetic histories as well as complex trait architectures remain hidden due to the lack of Big Data. To fill this gap, the Mexican Biobank project genotyped 1.8 million markers in 6,057 individuals from 32 states and 898 sampling localities across Mexico with linked complex trait and disease information creating a valuable nationwide genotype-phenotype database. Through a suite of state-of-the-art methods for ancestry deconvolution and inference of identity-by-descent (IBD) segments, we inferred detailed ancestral histories for the last 200 generations in different Mesoamerican regions, unravelling native and colonial/post-colonial demographic dynamics. We observed large variations in runs of homozygosity (ROH) among genomic regions with different ancestral origins reflecting their demographic histories, which also affect the distribution of rare deleterious variants across Mexico. We analysed a range of biomedical complex traits and identified significant genetic and environmental factors explaining their variation, such as ROH found to be significant predictors for trait variation in BMI and triglycerides.<br> ======================================</p> <p>This dataset contains GWAS summary statistics for the Mexico Biobank Project. Summary statistics for 22 binary and quantitative traits are provided from the full cohort of 5721 individuals from across Mexico, and a subset of 1061 individuals inferred to have more than 90% Native American ancestry.</p>
Data for: Dissecting the genetic architecture of leaf morphology traits in mungbean (Vigna radiata (L.) Wizcek) using genome‐wide association study
<p><span>Mungbean (<em>Vigna radiata</em> (L) Wizcek) is an important pulse crop, increasingly used as a source of protein, fiber, low fat, carbohydrates, minerals, and bioactive compounds in human diets. Mungbean is a dicot plant with trifoliate leaves. Leaves are central to various plant processes like photosynthesis, light interception, and overall canopy structure. The objectives were to study leaf morphological traits, use image analysis to extract leaf traits from images from the Iowa Mungbean Diversity (IMD) panel, develop a regression model for the prediction of leaflet area, and conduct association mapping for leaf morphological traits. We collected more than 5000 leaf images of the IMD panel consisting of 484 accessions over two years (2020 and 2021) with two replications per experiment. Leaf traits were extracted using image analysis, analyzed, and used for association mapping. Morphological diversity included leaflet type (oval or lobed), leaflet size (small, medium, large), lobed angle (shallow, deep), and vein coloration (green, purple). A regression model was developed to predict each ovate leaflet's area (adjusted R<sup>2</sup> = 0.97; residual standard errors of <= 1.10). The candidate genes <em>Vradi01g07560</em>, <em>Vradi05g01240</em>, <em>Vradi02g05730</em>, and <em>Vradi03g00440</em>, are associated with multiple traits (length, width, perimeter, and area) across the leaflets (left, terminal, and right). These are suitable candidate genes for further investigation in their role in leaf development, growth, and function. Future studies will be needed to correlate the observed traits discussed here with yield or important agronomic traits for use as phenotypic or genotypic markers in marker-aided selection methods for mungbean crop improvement.</span></p>
A complex genomic architecture underlies reproductive isolation in a North American Oriole hybrid zone
<p>Natural hybrid zones provide powerful opportunities for identifying the mechanisms that facilitate and inhibit speciation. Documenting the extent of genomic admixture allows us to discern the architecture of reproductive isolation through the identification of isolating barriers. This approach is particularly powerful for characterizing the accumulation of isolating barriers in systems exhibiting varying levels of genomic divergence. Here, we use a hybrid zone between two species--the Baltimore (<em>Icterus galbula</em>) and Bullock's (<em>I. bullockii</em>) orioles--to investigate this architecture of reproductive isolation. We combine whole genome re-sequencing with data from an additional 313 individuals amplityped at ancestry-informative markers to characterize fine-scale patterns of admixture, and to quantify links between genes and the plumage traits. On a genome-wide scale, we document several putative barriers to reproduction, including elevated peaks of divergence above a generally high genomic baseline, a large putative inversion on the Z chromosome, and complex interactions between melanogenesis-pathway candidate genes. Concordant and coincident clines for these different genomic regions further suggest the coupling of pre- and post-mating barriers. Our findings of complex and coupled interactions between pre- and post-mating barriers suggest a relatively rapid accumulation of barriers between these species, and they demonstrate the complexities of the speciation process.</p>
Data from: Genomic and transcriptomic analyses reveal polygenic architecture for ecologically-important functional traits in aspen (Populus tremuloides Michx.)
<p>Intraspecific genetic variation in foundation species such as aspen (<em>Populus</em> <em>tremuloides</em> Michx.) shapes their impact on forest structure and function. Identifying genes underlying ecologically important traits is key to understanding that impact. Previous studies, using single-locus genome-wide association (GWA) analyses to identify candidate genes, have identified fewer genes than anticipated for highly heritable quantitative traits. Mounting evidence suggests that polygenic control of quantitative traits is largely responsible for this "missing heritability" phenomenon. Our research characterized the genetic architecture of 30 ecologically important traits using a common garden of aspen through genomic and transcriptomic analyses. A multilocus association model revealed that most traits displayed a highly polygenic architecture, with most variation explained by loci with small effects (likely below the detection levels of single-locus GWA methods). Consistent with a polygenic architecture, our single-locus GWA analyses found only 38 significant SNPs in 22 genes across 15 traits. Next, we used differential expression analysis on a subset of aspen genets with divergent concentrations of salicinoid phenolic glycosides (key defense traits). This complementary method to traditional GWA discovered 1,243 differentially expressed genes for a polygenic trait. Soft clustering analysis revealed three gene clusters (241 candidate genes) involved in secondary metabolite biosynthesis and regulation. Our work reveals that ecologically important traits governing higher-order community- and ecosystem-level attributes of a foundation forest tree species have complex underlying genetic structures and will require methods beyond traditional GWA analyses to unravel.</p>
Data from: Genomic architecture drives population structuring in Amazonian birds
<p>Geographic barriers are frequently invoked to explain genetic structuring across the landscape. However, inferences on the spatial and temporal origins of population variation have been largely limited to evolutionary neutral models, ignoring the potential role of natural selection and intrinsic genomic processes known as genomic architecture in producing heterogeneity in differentiation across the genome. To test how genomic architecture impacts our ability to reconstruct general patterns of diversification in species that co-occur across geographic barriers, we sequenced the whole genomes of multiple bird populations that are distributed across rivers in southeastern Amazonia. We found that phylogenetic relationships within species and demographic parameters varied across the genome in predictable ways. Genetic diversity was positively associated with recombination rate and negatively associated with species tree support. Gene flow was less pervasive in regions of low recombination, making these windows more likely to retain patterns of population structuring that matched the species tree. We further found that approximately a third of the genome showed evidence of selective sweeps and linked selection, skewing genome-wide estimates of effective population sizes and gene flow between populations towards lower values. In sum, we showed that the effects of intrinsic genomic characteristics and selection can be disentangled from neutral processes to elucidate how inferring spatial patterns of diversification are sensitive to genomic architecture.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.