Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
886
datasets available to search
ShareScore release 0.9.0
Dataset results
886 results for “genetic analysis”
GWAS Summary Statistics from "Sex and statin-related genetic associations at the PCSK9 gene locus – results of genome-wide association meta-analysis"
<p>GWAMA summary statistics of PCSK9 levels stratified by sex and statin useage in Europeans.</p> <p>When using this data, please cite:</p> <p>Pott, J., Kheirkhah, A., Gadin, J.R. <em>et al.</em> Sex and statin-related genetic associations at the <em>PCSK9</em> gene locus: results of genome-wide association meta-analysis. <em>Biol Sex Differ</em> <strong>15</strong>, 26 (2024). https://doi.org/10.1186/s13293-024-00602-6</p> <p>All txt files contain the following columns:</p> <ul> <li>markername (unique SNP ID)</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>EA (effect allele)</li> <li>OA (other allele)</li> <li>EAF (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (in case of double-stratified data: number of studies; in case of single-stratified data: 2, as it is a meta-analysis of the two double-stratified data sets)</li> <li>beta (effect estimate)</li> <li>SE (standard error)</li> <li>pval (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>invalidAssoc (TRUE/FALSE flag if this variant was excluded in our analysis)</li> <li>reason4exclusion (reason why this SNP was excluded)</li> <li>phenotype (phenotyp setting)</li> </ul>
Genetic association analysis of anti-VEGF treatment response in neovascular age-related macular degeneration
<p>Summary statisics of an association study of 6,908,005 genetic variants with anti-VEGF nAMD treatment response in 179 treatment-naïve nAMD probands. This dataset supplements the publication "Genetic Association Analysis of Anti-VEGF Treatment Response in Neovascular Age-Related Macular Degeneration" (DOI: 10.3390/ijms23116094). Details regarding the methods and version numbers can be found in the corresponding manuscript.</p>
Data to reproduce analysis in "Systematic analysis of transcriptional and epigenetic effects of genetic variation in Kupffer cells enables discrimination of cell intrinsic and environment-dependent mechanisms"
<p>Here you can find the datasets necessary to reproduce all analyses described in the Glass lab paper by <a href="https://www.biorxiv.org/content/10.1101/2022.09.22.509046v1">Bennett et al</a>. The python and R code for reproducing analysis and figures can be found on our linked <a href="https://github.com/HunterBennett/KupfferCell_NaturalGeneticVariation">github repository.</a></p> <p>Briefly, this paper explores the effect of natural genetic variation <em>in vivo</em>, using Kupffer cells as a model cell type. We collect and analyze transcriptional and epigenetic data (ATAC-seq, H3K27Ac ChIP-seq) to identify putative <em>trans</em> regulators driving differential gene expression across inbred strains of mice. Additionally, we provide evidence that <em>trans</em> effects control a majority of strain differential genes at homeostasis while <em>cis</em> effects dominate the transcriptional response to an external signal (lipopolysaccharide).</p> <p>References:</p> <p>Hunter Bennett, Ty D. Troutman, Enchen Zhou, Nathanael J. Spann, Verena M. Link, Jason S. Seidman, Christian K. Nickl, Yohei Abe, Mashito Sakai, Martina P. Pasillas, Justin M. Marlman, Carlos Guzman, Mojgan Hosseini, Bernd Schnabl, Christopher K. Glass bioRxiv 2022.09.22.509046; doi: <a href="https://doi.org/10.1101/2022.09.22.509046">https://doi.org/10.1101/2022.09.22.509046</a></p> <p> </p>
Functional genomics analysis to disentangle the role of genetic variants in major depression - Supplementary Tables
<p>This entry contains the data generated by the study "Functional genomics analysis to disentangle the role of genetic variants in major depression" that are part of the Supplementary information of the article describing the study.</p> <p>The entry contains the following data:</p> <p><strong>Supplementary Tables S1-S7</strong></p> <p>Supplementary Table S1. Summary of resources.</p> <p>Supplementary Table S2. Causal GVs for MD.</p> <p>Supplementary Table S3. pGenes functional and disease enrichment analysis.</p> <p>Supplementary Table S4. Fine-mapped MD causal GVs disease enrichment analysis.</p> <p>Supplementary Table S5. Colocalizing GWAS-eQTLs association to disease.</p> <p>Supplementary Table S6. TFBS analysis.</p> <p>Supplementary Table S7. GVs state annotation. </p>
Human pancreatic islet microRNAs implicated in diabetes and related traits by large-scale genetic analysis
<p>Genetic studies have identified ≥240 loci associated with risk of type 2 diabetes (T2D), yet most of these loci lie in non-coding regions, masking the underlying molecular mechanisms. Recent studies investigating mRNA expression in human pancreatic islets have yielded important insights into the molecular drivers of normal islet function and T2D pathophysiology. However, similar studies investigating microRNA (miRNA) expression remain limited. Here, we present data from 63 individuals, the largest sequencing-based analysis of miRNA expression in human islets to date. We characterize the genetic regulation of miRNA expression by decomposing the expression of highly heritable miRNAs into <em>cis</em>- and <em>trans</em>-acting genetic components and mapping <em>cis</em>-acting loci associated with miRNA expression (miRNA-eQTLs). We find (i) 84 heritable miRNAs, primarily regulated by <em>trans</em>-acting genetic effects, and (ii) 5 miRNA-eQTLs. We also use several different strategies to identify T2D-associated miRNAs. First, we colocalize miRNA-eQTLs with genetic loci associated with T2D and multiple glycemic traits, identifying one miRNA, miR-1908, that shares genetic signals for blood glucose and glycated hemoglobin (HbA1c). Next, we intersect miRNA seed regions and predicted target sites with credible set SNPs associated with T2D and glycemic traits and find 32 miRNAs that may have altered binding and function due to disrupted seed regions. Finally, we perform differential expression analysis and identify 14 miRNAs associated with T2D status—including miR-187-3p, miR-21-5p, miR-668, and miR-199b-5p—and 4 miRNAs associated with a polygenic score for HbA1c levels—miR-216a, miR-25, miR-30a-3p, and miR-30a-5p.</p>
A Linked Application of Discrete Differential Evolution Algorithm Coupled with Simulation- Optimization Model and Comparative Analysis by Genetic Algorithm for Discrete Groundwater Management Problems
<p>Complete dataset of publication name as "The complete publication dataset is "A Discrete Differential Evolution- Linear Programming Algorithm for Groundwater Management Problems." You can find all the written codes in the zip file.</p>
A scalable, accurate, and universal analysis framework using individual-level allele frequency for large-scale genetic association studies in an admixed population
<p>Inclusion of individuals with diverse or admixed genetic ancestries is crucial to discover novel findings that may be missed by genomics analyses rooted solely in Caucasian population. Here, we present an analysis framework, SPAmix, which is scalable to a large-scale biobank data analysis including hundreds of thousands of admixed individuals and is universally applicable to various types of complex traits including binary trait, quantitative trait, time-to-event trait, longitudinal traits, etc. For each genetic variant, SPAmix uses genotype data and genetic principal components (PCs) to estimate individual-level allele frequency, which is subsequently used to calibrate p values via a retrospective analysis. A hybrid strategy including saddlepoint approximation (SPA) can greatly increase the accuracy to analyze rare genetic variants, especially if the phenotypic distribution is unbalanced or extremely unbalanced. Compared to Tractor, SPAmix does not require local ancestry information and can be straightforwardly applicable to a multi-way admixed population. Meanwhile, SPAmix can also be extended to SPAmix<sub>local</sub> in which the local ancestry can be incorporated if available. In addition, we propose SPAmix<sub>CCT</sub> to combine the p values of SPAmix and SPAmix<sub>local</sub> via Cauchy combination (CCT). SPAmix<sub>local</sub> performs close to Tractor when analyzing quantitative traits and is more accurate when analyzing binary traits with an unbalanced case-control ratio. And SPAmix<sub>CCT </sub>is an optimal unified approach for various cross-ancestry genetic architectures. Extensive simulation studies and real data analyses of 369,314 UK Biobank individuals from multiple ancestries demonstrated that SPAmix is scalable and can discover novel hits while controlling type I error rates well.</p>
Large-scale longitudinal gradients of genetic diversity: a meta-analysis across six phyla in the Mediterranean basins
Predicting patterns of variation in biodiversity across the globe is a fundamental issue in ecology and evolution. Diversity within species, that is, genetic diversity, is of prime importance for understanding past and present evolutionary patterns, and highlighting areas where conservation might be a priority. However, most studies on spatial patterns of genetic diversity have not considered longitude as a potentially important ecological driver of these patterns. Therefore, we carried out a meta-analysis to examine the longitudinal patterns of genetic diversity in the Mediterranean Basin. Using published literature and a systematic review/meta-analysis framework, we collected data on the genetic diversity of species whose populations occur in the Mediterranean basin. We then calculated a coefficient of correlation between within‐population genetic diversity indices and longitude, and estimated the role of biological, ecological, biogeographic, and marker type factors on the strength and magnitude of this correlation in six phylla. The results of this study were published in the paper titled Large‐scale longitudinal gradients of genetic diversity: a meta‐analysis across six phyla in the Mediterranean basin (Conord et al. 2012).
Genetic data and underlying taxa and GenBank sources of diatoms used in phylogenetic analysis for the diatom genus Nupela
<p>Supplementary material for the manuscript: Kulikovskiy M., Maltsev Y., Glushchenko A., Gusev E., Kapustin D., Kuznetsova I., Kociolek J.P. Preliminary molecular phylogeny of the diatom genus <em>Nupela</em> with the description of a new species and consideration of the interrelationships of taxa in the suborder Neidiineae D.G. Mann sensu E.J. Cox. Fottea</p> <p>Molecular investigation of diatom genera <em>Nupela</em> and <em>Brachysira</em> is conducted using strains from Indonesia and Vietnam. New species from the genus <em>Nupela indonesica</em> sp. nov. is described using combined approach. <em>Nupela lesothensis</em> (Schoeman) Lange-Bertalot is investigated using molecular data too. Phylogenetic analysis shows that <em>Nupela</em> and <em>Brachysira</em> are not closest genera. Morphology of <em>Nupela</em> and it differences from <em>Brachysira</em> is discussed. The genus <em>Nupela</em> is differs from all other diatom taxa by having coalescent hymenes ouside of areolae but not inside. Facultative development of raphe between different <em>Nupela</em> species is discussed.<br> SUPPLEMENT S1. Taxa and DNA sequence data used in phylogenetic analysis.<br> SUPPLEMENT S2. Final alignment of 2-gene DNA sequence data used for phylogenetic analysis in FASTA format.<br> SUPPLEMENT S3. Maximum Likelihood tree of <em>Nupela</em> species (indicated in bold) constructed from a concatenated alignment of 163 partial rbcL and partial 18S rDNA sequences of 1806 characters. Values near the horizontal lines (slash) are bootstrap support from RAxML analyses (<50 are not shown). Species from the centric diatoms were used as an outgroup. Families indicated according COX (2015).</p>
Tara Pacific 18S-based coral host genetic analysis data release version 1
<p>This dataset contains 4 tables and 3 sets of figures related to the primary analysis of the 18S metabarcoding sequencing output. This dataset is only concerned with the identity of the coral host (i.e. not additional protist diversity). The samples included in this dataset have a 'sample-material_label' value of 'CORAL' and 'sampling-protocol_label' value of 'SEQ-CS4L'. They represent the coral samples collected at all 32 of the islands visited in the Tara Pacific expedition.</p>
Time trees and Clock genes: a Systematic Review and Comparative Analysis of Contemporary Avian Migration Genetics (Dataset)
<p>Complete dataset of <em>Clock</em> and <em>Adcyap1</em> alleles, distance matrices, and migration data used in the review and meta-analysis "<strong>Time trees and Clock genes: a Systematic Review and Comparative Analysis of Contemporary Avian Migration Genetics".</strong> </p>
Supplementary materials for Phylogenomics and genetic analysis of solvent-producing Clostridium species
<p>The genus <em>Clostridium</em> is a large and diverse group within the Bacillota (formerly Firmicutes), whose members can encode useful complex traits such as solvent production, gas-fermentation, and lignocellulose breakdown. We describe 270 genome sequences of solventogenic clostridia from a comprehensive industrial strain collection assembled by Professor David Jones that includes 194 <em>C. beijerenckiI, </em>57 <em>C. saccharobutylicum</em>, 4 <em>C. saccharoperbutylacetonicum</em>, 5 <em>C. butyricum</em>, 7 <em>C. acetobutylicum</em>, and 3 <em>C. tetanomorphum </em>genomes. We report methods, analyses and characterization for phylogeny, key attributes, core biosynthetic genes, secondary metabolites, plasmids, prophage/CRISPR diversity, cellulosomes and quorum sensing for the 6 species. The expanded genomic data described here will facilitate engineering of solvent-producing clostridia as well as non-model microorganisms with innately desirable traits. Sequences could be applied in conventional platform biocatalysts such as yeast or <em>Escherichia coli </em>for enhanced chemical production. Recently, gene sequences from this collection were used to engineer <em>Clostridium autoethanogenum</em>, a gas-fermenting autotrophic acetogen, for continuous acetone or isopropanol production<em>, </em>as well as butanol, butanoic acid, hexanol and hexanoic acid production. </p>
[Data from:] Genetic Analysis Reveals Three Novel QTLs Underpinning a Butterfly Egg-Induced Hypersensitive Response-Like Cell Death in Brassica Rapa
<p><strong>Background</strong></p> <p>Cabbage white butterflies (<em>Pieris</em> spp.) can be severe pests of <em>Brassica</em> crops such as Chinese cabbage, Pak choi (<em>Brassica rapa</em>) or cabbages (<em>B. oleracea</em>). Eggs of <em>Pieris</em> spp. can induce a hypersensitive response-like (HR-like) cell death which reduces egg survival in the wild black mustard (<em>B. nigra</em>). Unravelling the genetic basis of this egg-killing trait in <em>Brassica</em> crops could improve crop resistance to herbivory, reducing major crop losses and pesticides use. Here we investigated the genetic architecture of a HR-like cell death induced by <em>P. brassicae</em> eggs in <em>B. rapa.</em></p> <p><strong>Results</strong></p> <p>A germplasm screening of <em>B. rapa</em> 56 accessions, representing the genetic and geographical diversity of a <em>B. rapa</em> core collection, showed phenotypic variation for cell death. An image-based phenotyping protocol was developed to accurately measure size of HR-like cell death and was then used to identify two accessions that consistently showed weak (R-o-18) or strong cell death response (L58). Screening of 160 RILs derived from these two accessions resulted in three novel QTLs for P<em>ieris</em> b<em>rassicae-</em>induced cell death on chromosomes A02 (<em>Pbc1</em>), A03 (<em>Pbc2</em>), and A06 (<em>Pbc3</em>). The three QTLs <em>Pbc1-3</em> contain cell surface receptors, intracellular receptors and other genes involved in plant immunity processes, such as ROS accumulation and cell death formation. Synteny analysis with <em>A. thaliana</em> suggested that <em>Pbc1</em> and <em>Pbc2</em> are novel QTLs associated with this trait, while <em>Pbc3</em> contains also LecRK-I.1, a gene of <em>A. thaliana</em> previously associated with cell death induced by a <em>P. brassicae</em> egg extract.</p> <p><strong>Conclusions</strong></p> <p>This study provides the first genomic regions associated with the <em>Pieris</em> egg-induced HR-like cell death in a <em>Brassica</em> crop species. It is a step closer towards unravelling the genetic basis of an egg-killing crop resistance trait, paving the way for breeders to further fine-map and validate candidate genes.</p>
Local chromatin context dictates the genetic determinants of the heterochromatin spreading reaction. Analysis Code, Numerical and Primary data.
<p>Uploaded under this Zenodo DOI is the following:</p> <p>1. the Analysis Code used for Flow Cytometry analysis in the paper, GO complex analysis (Figure 3) and Hit visualization (Figure 1, 2 S1, S4 Figs).</p> <p>2. The primary Flow Cytometry data from both the initial screen (ScreenFlowFCS) and validation experiments (ValidationFlowFCS) are included as .zip files.</p> <p>3. a .zip folder is uploaded that contains all the analysis code for the ChIP-Seq experiments. </p> <p>4. Excel worksheets that contain the numerical source data for all qPCR bar plots.</p>
Deciphering colorectal cancer genetics through multi-omic analysis of 100,204 cases and 154,587 controls of European and East Asian ancestries
<p><strong>Colorectal cancer (CRC) is a leading cause of mortality worldwide. We conducted a genome-wide association study meta-analysis of 100,204 CRC cases and 154,587 controls of European and Asian ancestry, identifying 205 independent risk associations, of which 50 were unreported. We performed integrative genomic, transcriptomic and methylomic analyses across large bowel mucosa and other tissues. Transcriptome- and methylome-wide association studies revealed an additional 53 risk associations. We identified 155 high confidence effector genes functionally linked to CRC risk, many of which had no previously established role in CRC. These have multiple different functions, and specifically indicate that variation in normal colorectal homeostasis, proliferation, cell adhesion, migration, immunity and microbial interactions determines CRC risk. Cross-tissue analyses indicated that over a third of effector genes most likely act outside the colonic mucosa. Our findings provide insights into colorectal oncogenesis, and highlight potential targets across tissues for new CRC treatment and chemoprevention strategies.</strong></p> <p><strong>The data submitted here are expression and methylation models with LD reference data for the transcriptome-wide (TWAS), methylome-wide (MWAS) and transcript isoform-wide association study (TIsWAS) as described in the manuscript "Deciphering colorectal cancer genetics through multi-omic analysis of 100,204 cases and 154,587 controls of European and East Asian ancestries". Details of the methods are presented in the method section and supplementary information file. </strong></p> <p><strong>TWAS analysis </strong></p> <p>Gene expression models for the six in-house expression datasets were generated using the PredictDB v7 pipeline for a total of 1,077 participants. Elastic net model building with 10-fold cross-validation was performed independently for each dataset. The elastic net models for GTEx v8 Colon Transverse were obtained from the PredictDB data repository (<a href="http://predictdb.org/">http://predictdb.org/</a>) and had been generated using the same pipeline. Models were computed using HapMap2 SNPs ±1Mb from each gene, together with covariate factors estimated using PEER32, clinical covariates when appropriate (age, sex and, where appropriate, case-control status, type of polyp and anatomic location in the colorectum), and three PCs from the individual dataset’s SNP genotype data.</p> <p>Transcript-based TWAS analyses (TIsWAS) were likewise performed by using transcript-level data from the SOCCS, BarcUVa-Seq and GTEx Colon Transverse datasets.</p> <p><strong>MWAS analysis </strong></p> <p>Methylation beta values were calculated based on the manufacturer’s standard, ranging from 0 to 1. Quality control and data normalization were performed in R using the ChAMP software pipeline for the EPIC and 450K arrays. Briefly, we filtered out failed probes with detection P > 0.02 in >5% of samples, probes with <3 reads in >5% of samples per probe and all non-CpG probes. Samples with failed probes >0.1 were also excluded from downstream analyses. We discarded all probes with SNPs within 10bp of the interrogated CpG (from 1,000 Genomes Project, CEU population)34, and probes that ambiguously mapped to multiple locations in the human genome with up to two mismatches33. We only considered probes mapping to autosomes and those overlapping between the EPIC and the 450K arrays. Normalization was achieved using the Beta MIxture Quantile (BMIQ) method. Per probe methylation models were created using the PredictDB pipeline on the normalized methylation matrix and the genotypes as per TWAS eQTL analysis. To optimize power, we restricted our analysis to 263,341-238,443 (for the 450K array) and 377,678 (for the EPIC array) probes annotated to Islands, Shores and Shelves, and discarded “Open Sea” regions. </p>
Construction of a SNP fingerprinting database and population genetic analysis of 329 cauliflower cultivars
<p>The VCF file contains the information of 1662 SNP sites of 820 cauliflower inbred lines that were filtered according to a series of stringent conditions.</p>
Fig. 13 in On Roth's "human fossil" from Baradero, Buenos Aires Province, Argentina: morphological and genetic analysis
Fig. 13 Virtual reconstruction of the skull presented from different views: A frontal; B frontal with template; C posterior; D posterior with template; E left lateral; F right lateral with template; G right lateral; H right lateral with template; I superior; J superior with template; K ventral; L ventral with template. The template is displayed in transparency as a reference of the skull shape without modifications. The scales indicate 5 cm in every case
Fig. 12 3D in On Roth's "human fossil" from Baradero, Buenos Aires Province, Argentina: morphological and genetic analysis
Fig. 12 3D models produced from the CT-scans of the 19 skull bone fragments. See Table 2 for references on which bones are contained in each fragment. Since fragments 17 and 18 did not present any diagnostic features, it was not possible to identify them and were not employed in the reconstruction
Fig. 11 in On Roth's "human fossil" from Baradero, Buenos Aires Province, Argentina: morphological and genetic analysis
Fig. 11 Sex determination based on the SNP coverage for the X and Y chromosomes for the merged sequenced data. Value is plotted ± standard error (SE) for accuracy on the sex determination. Values ± SE above 0.20 on the y axis are considered as genetically male
Fig. 10 Damage plots for the 3 in On Roth's "human fossil" from Baradero, Buenos Aires Province, Argentina: morphological and genetic analysis
Fig. 10 Damage plots for the 3' (A) and 5ʹ ends (B) for the four libraries [each line corresponds to a different library and treatment (shotgun or capture)], obtained with the built-in version of DamageProfiler (v0.4.9) for nf-core/eager (v. 2.4.3). The number of reads would be sufficient to clearly distinguish the substitution patterns characteristic for aDNA
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.