Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
45
datasets available to search
ShareScore release 0.9.0
Dataset results
45 results for “Genomics statistics”
Nationwide genomic biobank in Mexico unravels demographic history and complex trait architecture from 6,057 individuals: GWAS summary statistics
<p>Latin America continues to be severely underrepresented in genomics research, and fine-scale genetic histories as well as complex trait architectures remain hidden due to the lack of Big Data. To fill this gap, the Mexican Biobank project genotyped 1.8 million markers in 6,057 individuals from 32 states and 898 sampling localities across Mexico with linked complex trait and disease information creating a valuable nationwide genotype-phenotype database. Through a suite of state-of-the-art methods for ancestry deconvolution and inference of identity-by-descent (IBD) segments, we inferred detailed ancestral histories for the last 200 generations in different Mesoamerican regions, unravelling native and colonial/post-colonial demographic dynamics. We observed large variations in runs of homozygosity (ROH) among genomic regions with different ancestral origins reflecting their demographic histories, which also affect the distribution of rare deleterious variants across Mexico. We analysed a range of biomedical complex traits and identified significant genetic and environmental factors explaining their variation, such as ROH found to be significant predictors for trait variation in BMI and triglycerides.<br> ======================================</p> <p>This dataset contains GWAS summary statistics for the Mexico Biobank Project. Summary statistics for 22 binary and quantitative traits are provided from the full cohort of 5721 individuals from across Mexico, and a subset of 1061 individuals inferred to have more than 90% Native American ancestry.</p>
Integrative genomics of the mammalian alveolar macrophage response to intracellular mycobacteria: RNA-seq statistics and results
Open the record for dataset details and reuse information.
Linking genomic offset statistics to the shape of selection gradients
Open the record for dataset details and reuse information.
Data from: Sampling strategy optimization to increase statistical power in landscape genomics: a simulation-based approach
An increasing number of studies are using landscape genomics to investigate local adaptation in wild and domestic populations. The implementation of this approach requires the sampling phase to consider the complexity of environmental settings and the burden of logistic constraints. These important aspects are often underestimated in the literature dedicated to sampling strategies. In this study, we computed simulated genomic datasets to run against actual environmental data in order to trial landscape genomics experiments under distinct sampling strategies. These strategies differed by design approach (to enhance environmental and/or geographic representativeness at study sites), number of sampling locations and sample sizes. We then evaluated how these elements affected statistical performances (power and false discoveries) under two antithetical demographic scenarios. Our results highlight the importance of selecting an appropriate sample size, which should be modified based on the demographic characteristics of the studied population. For species with limited dispersal, sample sizes above 200 units are generally sufficient to detect most adaptive signals, while in random mating populations this threshold should be increased to 400 units. Furthermore, we describe a design approach that maximizes both environmental and geographical representativeness of sampling sites and show how it systematically outperforms random or regular sampling schemes. Finally, we show that although having more sampling locations (between 40 and 50 sites) increase statistical power and reduce false discovery rate, similar results can be achieved with a moderate number of sites (20 sites). Overall, this study provides valuable guidelines for optimizing sampling strategies for landscape genomics experiments.
The search for sexually antagonistic genes: practical insights from studies of local adaptation and statistical genomics
<p>Sexually antagonistic (SA) genetic variation—in which alleles favored in one sex are disfavored in the other—is predicted to be common and has been documented in several animal and plant populations, yet we currently know little about its pervasiveness among species or its population genetic basis. Recent applications of genomics in studies of SA genetic variation have highlighted considerable methodological challenges to the identification and characterization of SA genes, raising questions about the feasibility of genomic approaches for inferring SA selection. The related fields of local adaptation and statistical genomics have previously dealt with similar challenges, and lessons from these disciplines can therefore help overcome current difficulties in applying genomics to study SA genetic variation. Here, we integrate theoretical and analytical concepts from local adaptation and statistical genomics research—including <em>F</em><sub>ST</sub> and <em>F</em><sub>IS</sub> statistics, genome‐wide association studies, pedigree analyses, reciprocal transplant studies, and evolve‐and‐resequence experiments—to evaluate methods for identifying SA genes and genome‐wide signals of SA genetic variation. We begin by developing theoretical models for between‐sex <em>F</em><sub>ST</sub> and <em>F</em><sub>IS</sub>, including explicit null distributions for each statistic, and using them to critically evaluate putative multilocus signals of sex‐specific selection in previously published datasets. We then highlight new statistics that address some of the limitations of <em>F</em><sub>ST</sub>and <em>F</em><sub>IS</sub>, along with applications of more direct approaches for characterizing SA genetic variation, which incorporate explicit fitness measurements. We finish by presenting practical guidelines for the validation and evolutionary analysis of candidate SA genes and discussing promising empirical systems for future work.</p>
Data from: Accurate genomic prediction of Coffea canephora in multiple environments using whole-genome statistical models
Genomic selection have been proposed as the standard method to predict breeding values in animal and plant breeding. Although some crops have benefited from this methodology, studies in Coffea are still emerging. To date, there have been no studies of how well genomic prediction models work across populations and environments for different complex traits in coffee. Considering that predictive models are based on biological and statistical assumptions, it is expected that their performance vary depending on how well these assumptions align with the true genetic architecture of the phenotype. To investigate this, we used data from two recurrent selection populations of Coffea canephora, evaluated in two locations, and single nucleotide polymorphisms identified by Genotyping-by-Sequencing. In particular, we evaluated the performance of 13 statistical approaches to predict three important traits in the coffee — production of coffee beans, leaf rust incidence and yield of green beans. Analyses were performed for predictions within-environment, across locations and across populations to assess the reliability of genomic selection. Overall, differences in the prediction accuracy of the competing models were small, although the Bayesian methods showed a modest improvement over other methods, at the cost of more computation time. As expected, predictive accuracy for within-environment analysis, on average, were higher than predictions across locations and across populations. Our results support the potential of genomic selection to reshape traditional plant breeding schemes. In practice, we expect to increase the genetic gain per unit of time by reducing the length cycle of recurrent selection in coffee.
Source data of the MultiSTAAR manuscript "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies".
<p>This dataset serves as the source data for Figures 2-3 and Extended Data Figures 1-2 of the MultiSTAAR manuscript titled "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies". MultiSTAAR is a statistical framework and computationally-scalable analytical pipeline for functionally-informed multi-trait rare variant analysis in large-scale WGS studies.<br><br><strong>Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and ncRNA multi-trait analysis of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerids (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Figure 3.</strong> TOPMed genetic region (2-kb sliding window) unconditional multi-trait analysis results of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerides (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Extended Data Figure 1.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of fasting glucose (FG) and fasting insulin (FI) using TOPMed data (<em>n</em> = 21,731).<br><br><strong>Extended Data Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of C-reactive protein (CRP), interleukin 6 (IL-6), lipoprotein-associated phospholipase A2 (Lp-PLA2) activity, and lipoprotein-associated phospholipase A2 (Lp-PLA2) mass using TOPMed data (<em>n</em> = 9,380).</p>
Minimizer collision statistics (BLEND: A Fast, Memory-Efficient, and Accurate Mechanism to Find Fuzzy Seed Matches in Genome Analysis)
<p>This dataset includes the statistics for the minimizers that generate the same hash value (i.e., collisions). The hash values are generated using a low-collision hash function and the SimHash technique in BLEND.</p> <p> </p> <p>*collision_stats.txt files include the overall collision statistics for a tool and configuration of the tool (i.e., the number of colliding minimizer pairs with a certain edit distance and their ratio to all number of collisions). For example blend_n3_collision_stats.txt shows the statistics for BLEND where the number of neighbors is set to 3 when running BLEND.</p> <p>_sim.csv files include all minimizer pairs with the same hash value and the edit distance between them.</p> <p> </p>
K-mer collision statistics (BLEND: A Fast, Memory-Efficient, and Accurate Mechanism to Find Fuzzy Seed Matches in Genome Analysis)
<p>This dataset contains 1,077 FASTA files and CSV files. Each FASTA file includes 25-character long sequences similar to each other.</p> <p>We have a CSV file for each tool (i.e., minimap2 and BLEND) and configuration (i.e., different number of neighbors in BLEND). CSV files include the non-identical k-mer pairs (16-mers) that generate the same hash value (i.e., collisions). These k-mers are extracted from sequences that are similar to each other. In each line, we show the hash value of the k-mers, the actual sequene pairs that the k-mers are extracted from, k-mer pairs that generate the same hash value, and the edit distance between these k-mers.</p> <p> </p>
Genome-wide association study of polygenic risk score-defined phenotype suffers from inflated test-statistics
<p>Simulation results from running the following script 100 times: https://github.com/euffelmann/paper-ad_prs_extremes/blob/main/scripts/ad_prs_extremes_simulation.R.</p> <p>These files can be used to reproduce tables and figures in: https://github.com/euffelmann/paper-ad_prs_extremes</p>
Summary statistics from a genome-wide association study of narcolepsy
<p>Type 1 narcolepsy (T1N) is a neurological condition, in which the death of hypocretin-producing neurons in the lateral hypothalamus leads to excessive daytime sleepiness and symptoms of abnormal Rapid Eye Movement (REM) sleep. Known triggers for narcolepsy are influenza-A infection and associated immunization during the 2009 H1N1 influenza pandemic. Here, we genotyped all remaining consented narcolepsy cases worldwide and assembled this with the existing genotyped individuals. We used this multi-ethnic sample in genome wide association study (GWAS) to dissect disease mechanisms and interactions with environmental triggers (5,339 cases and 20,518 controls). Overall, we found significant associations with HLA (2 GWA significant subloci) and 11 other loci. Six of these other loci have been previously reported (<em>TRA</em>, <em>TRB</em>, <em>CTSH</em>, <em>IFNAR1</em>, <em>ZNF365</em> and <em>P2RY11</em>) and five are new (<em>PRF1</em>, <em>CD207</em>, <em>SIRPG</em>, <em>IL27</em> and <em>ZFAND2A</em>). Strikingly, in vaccination-related cases, GWA significant effects were found in <em>HLA</em>, <em>TRA</em>, and in a novel variant near <em>SIRPB1</em>. Furthermore, <em>IFNAR1</em>-associated polymorphisms regulated dendritic cell response to influenza-A infection in vitro (p-value =1.92*10<sup>-25</sup>). A partitioned heritability analysis indicated specific enrichment of functional elements active in cytotoxic and helper T cells. Furthermore, functional analysis showed the genetic variants in <em>TRA</em> and <em>TRB</em> loci act as remarkably strong chain usage QTLs for <em>TRAJ*24</em> (p-value = 0.0017), <em>TRAJ*28</em> (p-value = 1.36*10<sup>-10</sup>) and <em>TRBV*4-2</em> (p-value = 3.71*10-<sup>117</sup>). This was further validated in TCR sequencing of 60 narcolepsy cases and 60 DQB1*06:02 positive controls, where chain usage effects were further accentuated. Together these findings show that the autoimmune component in narcolepsy is defined by antigen presentation, mediated through specific T cell receptor chains, and modulated by influenza-A as a critical trigger.</p>
Data from: Accurate genomic prediction of Coffea canephora in multiple environments using whole-genome statistical models
Open the record for dataset details and reuse information.
The search for sexually antagonistic genes: practical insights from studies of local adaptation and statistical genomics
Open the record for dataset details and reuse information.
Summary statistics from a genome-wide association study of narcolepsy
Open the record for dataset details and reuse information.
Data from: Sampling strategy optimization to increase statistical power in landscape genomics: a simulation-based approach
Open the record for dataset details and reuse information.
Data from: The relative power of genome scans to detect local adaptation depends on sampling design and statistical method
Although genome scans have become a popular approach towards understanding the genetic basis of local adaptation, the field still does not have a firm grasp on how sampling design and demographic history affect the performance of genome scans on complex landscapes. To explore these issues, we compared 20 different sampling designs in equilibrium (i.e. island model and isolation by distance) and nonequilibrium (i.e. range expansion from one or two refugia) demographic histories in spatially heterogeneous environments. We simulated spatially complex landscapes, which allowed us to exploit local maxima and minima in the environment in 'pair' and 'transect' sampling strategies. We compared FST outlier and genetic–environment association (GEA) methods for each of two approaches that control for population structure: with a covariance matrix or with latent factors. We show that while the relative power of two methods in the same category (FST or GEA) depended largely on the number of individuals sampled, overall GEA tests had higher power in the island model and FST had higher power under isolation by distance. In the refugia models, however, these methods varied in their power to detect local adaptation at weakly selected loci. At weakly selected loci, paired sampling designs had equal or higher power than transect or random designs to detect local adaptation. Our results can inform sampling designs for studies of local adaptation and have important implications for the interpretation of genome scans based on landscape data.
Data from: Genome-wide prediction of bacterial effector candidates across six secretion system types using a feature-based statistical framework
Gram-negative bacteria are responsible for hundreds of millions infections worldwide, including the emerging hospital-acquired infections and neglected tropical diseases in the third-world countries. Finding a fast and cheap way to understand the molecular mechanisms behind the bacterial infections is critical for efficient diagnostics and treatment. An important step towards understanding these mechanisms is the discovery of bacterial effectors, the proteins secreted into the host through one of the six common secretion system types. Unfortunately, current prediction methods are designed to specifically target one of three secretion systems, and no accurate "secretion system-agnostic" method is available. Here, we present PREFFECTOR, a computational feature-based approach to discover effector candidates in Gram-negative bacteria, without prior knowledge on bacterial secretion system(s) or cryptic secretion signals. Our approach was first evaluated using several assessment protocols on a manually curated, balanced dataset of experimentally determined effectors across all six secretion systems, as well as non-effector proteins. The evaluation revealed high accuracy of the top performing classifiers in PREFFECTOR, with the small false positive discovery rate across all six secretion systems. Our method was also applied to six bacteria that had limited knowledge on virulence factors or secreted effectors. PREFFECTOR web-server is freely available at: http://korkinlab.org/preffector.
The effect of different statistical methods on the accuracy of predicting genomic selection in beef cattle
Open the record for dataset details and reuse information.
Data from: Genomic selection and association mapping in rice (Oryza sativa): effect of trait genetic architecture, training population composition, marker number and statistical model on accuracy of rice genomic selection in elite, tropical rice breeding lines
Open the record for dataset details and reuse information.
Data from: The relative power of genome scans to detect local adaptation depends on sampling design and statistical method
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.