Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

208

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

208 results for “gene enrichment”

Learn how ShareScore rates datasets ↗
zenodo44/100

Raw and processed GO term data to support running GCEA analyses using ensemble-based nulls, as described in the manuscript, 'Overcoming bias in gene category enrichment analyses of brain-wide transcriptomic data'.

<p>Data to support a toolbox for performing gene category enrichment analyses, including against ensembles of null phenotypes.</p> <p>Descriptions of how these data files can be used for this purpose are in the documentation for the toolbox, at https://github.com/benfulcher/GCEA_FalsePositives</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Processed proteomic and phosphoproteomic timeseries from Ostreococcus tauri, with Gene Ontology enrichment, from "A phospho-dawn of protein modification anticipates light onset in the picoeukaryote O. tauri"

<p>Diel regulation of protein levels and protein modification had been less studied than transcript rhythms. These data tables in .XLSX format report partial proteome (Table_S1)&nbsp;and phosphoproteome data (Table_S2), assayed using shotgun mass-spectrometry, from cultures of the alga <em>Ostreococcus tauri&nbsp;</em>under light-dark cycles, sampled at Zeitgeber times (ZT, hours) 0, 4, 8, 12, 16 and 20.&nbsp;10% of quantified proteins but two-thirds of phosphoproteins were rhythmic. Gene Ontology enrichment analysis was applied to infer the functional enrichment of the proteins or phosphoproteins, grouped by their loadings in PCA analysis (Table_S3), by hierarchical clustering (Table_S4) or&nbsp;by the peak time of their rhythmic profile (Table_S5).Prompted by night-peaking and apparently dark-stable proteins, we also tested the proteome of cultures transferred to prolonged darkness for 24, 48, 72 or 96h (Table_S6), where the proteome changed less than under the diel cycle. The raw data are available from ProteomeXchange, with identifiers PXD001734, PXD001735 and PXD002909.</p>

opencc-by-4.0Mar 2023View details →
dryad40/100

R code for differential gene expression and enrichment analyses

<p>The information about the magnitude of differences in thermal plasticity both between and within populations, as well as identification of the underlying molecular mechanisms are key to understanding the evolution of thermal plasticity. In particular, genes underlying variation in the physiological response to temperature can provide raw material for selection acting on plastic traits. Using RNAseq, we investigate the transcriptional response to temperature in males and females from bulb mite populations selected for the increased frequency of one of two discrete male morphs (fighter- and scrambler-selected populations) that differ in relative fitness depending on temperature. We show that different mechanisms underlie the divergence in thermal response between fighter- and scrambler-selected populations at decreased vs. increased temperatures. Temperature decrease to 18°C was associated with higher transcriptomic plasticity of males with more elaborate armaments, as indicated by a significant selection-by-temperature interaction effect on the expression of 40 genes, 38 of which were upregulated in fighter-selected populations in response to temperature decrease. In response to 28°C, no selection-by-temperature interaction in gene expression was detected. Hence, differences in phenotypic response to temperature increase likely depended on genes associated with their distinct morph-specific thermal tolerance. Selection on males also drove gene expression patterns in females. These patterns could be associated with temperature-dependent fitness differences between females from fighter- vs. scrambler-selected populations reported in previous studies. Our study shows that selection for divergent male sexually selected morphologies and behaviors has the potential to drive divergence in metabolic pathways underlying plastic response to temperature in both sexes.</p>

opencc-zeroMay 2024View details →
zenodo40/100

Differential abundance and gene set enrichment in plasma of cancer patients versus controls

<p>Version update: Human Protein Atlas v23 used for gene set enrichment analysis (instead of Human Protein Atlas v18)</p> <p><br>DESeq2 differential abundance output for genes with q &lt; 0.05 and |log<sub>2</sub>&nbsp;fold change| &gt; 1&nbsp;in cancer vs control plasma samples:</p> <ul> <li><em>differentialabundance_pancancer.txt</em>: tables with differentially abundant genes (|log2(fold change)|&gt;1 and adjusted p&gt;0.05) per cancer-control comparison (cancertype) in a pan-cancer plasma sample cohort (25 locally advanced to metastatic cancer types - 7 or&nbsp;8 patients&nbsp;per type - vs 8 cancer-free control donors)</li> <li><em>differentialabundance_threecancer.txt</em>: tables with differentially abundant genes (|log2(fold change)|&gt;1 and adjusted p&gt;0.05) per cancer-control comparison (cancertype) in the three-cancer plasma cohort&nbsp;(ovarian,&nbsp;prostate and uterine cancer - 11 or&nbsp;12&nbsp;patients per type - vs 20&nbsp;cancer-free controls) <ul> <li>Gene_id: Ensembl gene id (GChr38 v91); baseMean: mean of normalized counts for all samples; log2FoldChange: log2 fold change for cancer vs control; lfcSE: standard error for cancer vs control; stat: Wald statistic for cancer vs control; pvalue: Wald test p-value for cancer vs control; padj: Benjamini-Hochberg corrected p-value; cancertype: respective cancer type abbreviation of cancer patient&nbsp;plasma samples that were compared to plasma samples of controls.</li> </ul> </li> </ul> <p>Gene set enrichment analyses based on fold change ranked gene lists (cancer versus control) -&nbsp;results obtained with fgea (v1.22.0):</p> <ul> <li><em>customgenesets.txt</em>: custom gene set lists based on RNA Atlas (&amp;Human Protein Atlas), Tabula Sapiens, GTEX, TCGA data. <ul> <li>Reference: reference to create gene sets (including&nbsp;RNA Atlas,&nbsp;Human Protein Atlas, Tabula Sapiens, GTEX, and&nbsp;TCGA); set: set name; genes: gene list for set</li> </ul> </li> <li><em>GSEA_pancancer.txt</em> &amp; <em>GSEA_threecancer.txt</em>: gene set enrichment results based on fold change ranked gene list (specific cancer type versus controls) in pan-cancer cohort and three-cancer cohort, respectively <ul> <li>Sets:&nbsp;gene set category&nbsp;(HALLMARK&nbsp;and KEGG:&nbsp;Hallmark and Canonical Pathways gene sets&nbsp;obtained from&nbsp;MSigDB (v2022.1); CUSTOM: custom&nbsp;tissue and cell type specific gene sets&nbsp;as defined in&nbsp;<em>customgenesets.txt)</em>; pathway: pathway/set&nbsp;name; pval: enrichment p-value; padj: Benjamini-Hochberg adjusted p-value; log2err: expected error for the standard deviation of the P-value logarithm; ES: enrichment score, same as in Broad GSEA implementation; NES: enrichment score normalized to mean enrichment of random samples of the same size;&nbsp;size: size of the pathway after removing genes without&nbsp;statistic values; leadingEdge: leading edge genes that drive the enrichment;&nbsp;Disease: respective cancer type abbreviation of cancer patient&nbsp;plasma samples that were compared to plasma samples of controls</li> </ul> </li> </ul>

opencc-by-4.0May 2023View details →
zenodo40/100

Summary statistics - Imputed gene associations identify replicable trans-acting genes enriched in transcription pathways and complex traits

<p>Summary statistics for all trans-acting/target gene pairs tested in our manuscript.</p> <p>Preprint available:&nbsp;<a href="https://doi.org/10.1101/471748">https://doi.org/10.1101/471748</a></p>

opencc-by-4.0Jan 2019View details →
zenodo40/100

Enrichment of gene sets altered in dram1 mutant zebrafish larvae

<p>Enrichment of gene sets altered in dram1 mutant zebrafish larvae.</p> <p>A. Gene Ontology categories significantly over and underrepresented in the significant genes differentially regulated between <em>dram1</em><sup>∆19n/∆19n </sup>PBS-injected mutants compared to <em>dram1</em><sup>+/+</sup> larvae.</p> <p>B. Gene sets from the MSigDB C2 database significantly positively correlated to the <em>dram1</em><sup>∆19n/∆19n </sup>mutants transcriptome compared to <em>dram1</em><sup>+/+</sup> larvae.</p> <p>C. Gene sets from the MSigDB C2 database significantly negatively correlated to the <em>dram1</em><sup>∆19n/∆19n </sup>mutants transcriptome compared to <em>dram1</em><sup>+/+</sup> larvae.</p>

opencc-by-4.0Mar 2019View details →
zenodo40/100

Fig. 4 in Phylogenomics of the tropical plant family Ochnaceae using targeted enrichment of nuclear genes and 250+ taxa

Fig. 4. Distribution of specimen ages and the number of loci recovered in the phylogenomic study of Ochnaceae. A, Histogram of the collection years of all Ochnaceae specimens; B &amp; C, Relationship between the year of collection of the specimens and the number of loci recovered for tissue obtained from herbarium material (excluding specimens with silica-dried leaf material), analysed for Ochneae and all the remaining Ochnaceae separately, either using the consensus-alignment (B) or the sample-specific (C) reference-based assembly approach. Pearson correlation coefficients and confidence intervals are given for each group.

opencc-by-4.0Feb 2021View details →
zenodo40/100

Fig. 2 in Phylogenomics of the tropical plant family Ochnaceae using targeted enrichment of nuclear genes and 250+ taxa

Fig. 2. RAxML trees based on the concatenated nuclear loci of Ochnaceae. A, Early-diverging branches of Ochnaceae and relationships within Quiinoideae based on the FAM dataset; B, Phylogenetic relationships of Sauvagesieae, Luxemburgieae and Testuleeae based on the SLT dataset. Numbers on the branches are bootstrap values&gt;50%. Numbers in parentheses after species names correspond to the specimen IDs (only for species with multiple accessions). The indicated classification of subfamilies and tribes follows Schneider &amp; al. (2014).

opencc-by-4.0Feb 2021View details →
zenodo40/100

Fig. 1 in Phylogenomics of the tropical plant family Ochnaceae using targeted enrichment of nuclear genes and 250+ taxa

Fig. 1. Overview of the phylogenetic relationships of the major clades of Ochnaceae based on the FAM dataset together with images of representative species. The classification follows Schneider &amp; al. (2014). Ochninae is by far the most species-rich clade comprising about two-thirds of the family's species and six genera (Brack. = Brackenridgea; Cmp. = Campylospermum, clades A and B; I. = Idertia; Ochna; Ouratea; Rh. = Rhabdophyllum). Letters around the tree refer to the photos (mostly flowers except where indicated) and the relative position of the displayed taxa on the tree. A, Medusagyne oppositifolia (Medusagynoideae); B, Froesia venezuelensis (Quiinoideae); C, Luxemburgia schwackeana (Luxemburgieae); D, Rhytidanthera sulcata; E, Cespedesia spathulata; F, Poecilandra retusa; G, Godoya antioquiensis; H, Wallacea insignis; I, Sauvagesia semicylindrifolia; J, Sauvagesia erecta (Sauvagesieae); K, Infructescence of Lophira lanceolata with accrescent sepals (Lophirinae); L, Flower of Elvasia kollmannii (Elvasiinae); M, Perissocarpa umbellifera; N, Fruiting Rhabdophyllum arnoldianum; O, Brackenridgea nitida; P, Campylospermum glaberrimum; Q, Ochna serrulata; R, Fruit of Ochna integerrima with drupelets sitting on enlarged receptacle; S, Fruit of Ouratea sp. with enlarged red receptable; T, Ouratea sp. — Photos: A, K &amp; N from www.africanplants.senckenberg.de (Dressler &amp; al., 2014–); B by Julio Schneider; C by William Milliken/ Royal Botanic Gardens, Kew; D by Sandra Reinales; E by Reinaldo Aguilar; F, H &amp; M by Francisco Farroñay; G by John Clark; I, J, S &amp; T by Domingos Cardoso; L by Claudio Nicoletti de Fraga; O by John Elliott; P by Warran McCleland; Q by Marja Broersma; R by Pierre Grard.

opencc-by-4.0Feb 2021View details →
dryad40/100

R code for differential gene expression and enrichment analyses

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad36/100

Data from: Targeted enrichment of large gene families for phylogenetic inference: phylogeny and molecular evolution of photosynthesis genes in the Portullugo clade (Caryophyllales)

Hybrid enrichment is an increasingly popular approach for obtaining hundreds of loci for phylogenetic analysis across many taxa quickly and cheaply. The genes targeted for sequencing are typically single-copy loci, which facilitate a more straightforward sequence assembly and homology assignment process. However, this approach limits the inclusion of most genes of functional interest, which often belong to multi-gene families. Here we demonstrate the feasibility of including large gene families in hybrid enrichment protocols for phylogeny reconstruction and subsequent analyses of molecular evolution, using a new set of bait sequences designed for the "portullugo" (Caryophyllales), a moderately sized lineage of flowering plants (∼2200 species) that includes the cacti and harbors many evolutionary transitions to C4 and CAM photosynthesis. Including multi-gene families allowed us to simultaneously infer a robust phylogeny and construct a dense sampling of sequences for a major enzyme of C4 and CAM photosynthesis, which revealed the accumulation of adaptive amino acid substitutions associated with C4 and CAM origins in particular paralogs. Our final set of matrices for phylogenetic analyses included 75–218 loci across 74 taxa, with ∼50% matrix completeness across datasets. Phylogenetic resolution was greatly improved across the tree, at both shallow and deep levels. Concatenation and coalescent-based approaches both resolve the sister lineage of the cacti with strong support: Anacampserotaceae + Portulacaceae, two lineages of mostly diminutive succulent herbs of warm, arid regions. In spite of this congruence, BUCKy concordance analyses demonstrated strong and conflicting signals across gene trees. Our results add to the growing number of examples illustrating the complexity of phylogenetic signals in genomic-scale data.

opencc-zeroDec 2016View details →
dryad36/100

Dataset: T-DNA characterization of genetically modified 3-R-gene late blight resistant potato events with a novel procedure utilizing the Samplix Xdrop® Enrichment Technology

<p>Before commercialization of genetically modified crops, the events carrying the novel DNA must be thoroughly evaluated for agronomic, nutritional, and molecular characteristics. Over the years, Polymerase Chain Reaction-based methods, Southern blot, and short-read sequencing techniques have been utilized for collecting molecular characterization data. Multiple genomic applications are necessary to determine the insert location, flanking sequence analysis, characterization of the inserted DNA, and determination of any interruption of native genes. These techniques are time-consuming and labor-intensive, making it difficult to characterize multiple events. Current advances in sequencing technologies are enabling whole genomic sequencing of modified crops to obtain full molecular characterization. However, in polyploids, such as the tetraploid potato, it is a challenge to obtain whole genomic sequencing coverage that meets regulatory approval of the genetic modification. Here we describe an alternative to labor-intensive applications with a novel procedure using Samplix Xdrop® enrichment technology and next-generation Nanopore sequencing technology to more efficiently characterize the T-DNA insertions of four genetically modified potato events developed by the Feed the Future Global Biotech Potato Partnership: DIA_MSU_UB015, DIA_MSU_UB255, GRA_MSU_UG234 and GRA_MSU_UG265 (derived from regionally important varieties Diamant and Granola). Using the Xdrop® /Nanopore technique, we obtained a very high sequence read coverage within the T-DNA and junction regions. In three of the four events, we were able to use the data to confirm single T-DNA insertions, identify insert locations, identify flanking sequences, and characterize the inserted T-DNA. We further used the characterization data to identify native gene interruption and confirm the stability of the T-DNA across clonal cycles. These results demonstrate the functionality of using the Xdrop® /Nanopore technique for T-DNA characterization. This research will contribute to meeting regulatory safety and regulatory approval requirements for commercialization with small shareholder farmers in target countries within our partnership.</p>

opencc-zeroFeb 2024View details →
zenodo36/100

Pleiotropic Expression Quantitative Trait Loci Are Enriched in Enhancers and Transcription Factor Binding Sites and Impact More Genes

<p>This dataset comprises two files that accompany the article (link to be added upon publication).</p> <h2>1. gwas2eqtl_colocalization_full.tar.gz</h2> <p>This file contains the complete colocalization dataset generated using the code from the following GitHub repository: gwas2eqtl. This dataset is used as input for the pleiotropic eQTL analysis available at gwas2eqtl_pleiotropy, which produces the figures in the article.</p> <p><strong>File structure:</strong></p> <blockquote> <p>.<br>└── gwas417<br>&nbsp; &nbsp; └── coloc<br>&nbsp; &nbsp; &nbsp; &nbsp; ├── ebi-a-GCST000998<br>&nbsp; &nbsp; &nbsp; &nbsp; │ &nbsp; └── pval_5e-08<br>&nbsp; &nbsp; &nbsp; &nbsp; │ &nbsp; &nbsp; &nbsp; └── r2_0.1<br>&nbsp; &nbsp; &nbsp; &nbsp; │ &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; └── kb_1000<br>&nbsp; &nbsp; &nbsp; &nbsp; │ &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; └── window_1000000<br>&nbsp; &nbsp; &nbsp; &nbsp; │ &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ├── Alasoo_2018_ge_macrophage_IFNg+Salmonella.tsv<br>&nbsp; &nbsp; &nbsp; &nbsp; │ &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ├── Alasoo_2018_ge_macrophage_IFNg.tsv<br>...</p> </blockquote> <p>Each TSV file contains the following columns:</p> <blockquote> <p>chrom &nbsp; &nbsp;pos &nbsp; &nbsp;rsid &nbsp; &nbsp;ref &nbsp; &nbsp;alt &nbsp; &nbsp;eqtl_gene_id &nbsp; &nbsp;gwas_beta &nbsp; &nbsp;gwas_pval &nbsp; &nbsp;gwas_id &nbsp; &nbsp;eqtl_beta &nbsp; &nbsp;eqtl_pval &nbsp; &nbsp;eqtl_id &nbsp; &nbsp;PP.H4.abf &nbsp; &nbsp;SNP.PP.H4 &nbsp; &nbsp;nsnps &nbsp; &nbsp;PP.H3.abf &nbsp; &nbsp;PP.H2.abf &nbsp; &nbsp;PP.H1.abf &nbsp; &nbsp;PP.H0.abf &nbsp; &nbsp;coloc_variant_id &nbsp; &nbsp;coloc_region<br>1 &nbsp; &nbsp;109272258 &nbsp; &nbsp;rs4970834 &nbsp; &nbsp;C &nbsp; &nbsp;T &nbsp; &nbsp;ENSG00000168765 &nbsp; &nbsp;-0.12874.25001047052626e-09 &nbsp; &nbsp;ebi-a-GCST000998 &nbsp; &nbsp;-0.250697 &nbsp; &nbsp;0.0893351 &nbsp; &nbsp;Alasoo_2018_ge_macrophage_IFNg+Salmonella &nbsp; &nbsp;0.108283426725895 &nbsp; &nbsp;0.0520205502224409 &nbsp; &nbsp;6 &nbsp; &nbsp;0.000552728832012655 &nbsp; &nbsp;2.09397277427793e-07 &nbsp; &nbsp;0.890881419183881 &nbsp; &nbsp;0.000282215860934432 &nbsp; &nbsp;1_109279544_G_A &nbsp; &nbsp;1:108779544-109779543<br>1 &nbsp; &nbsp;109274968 &nbsp; &nbsp;rs12740374 &nbsp; &nbsp;G &nbsp; &nbsp;T &nbsp; &nbsp;ENSG00000168765 &nbsp; &nbsp;-0.103341 &nbsp; &nbsp;1.63998546891446e-09 &nbsp; &nbsp;ebi-a-GCST000998 &nbsp; &nbsp;-0.197397 &nbsp; &nbsp;0.172673Alasoo_2018_ge_macrophage_IFNg+Salmonella &nbsp; &nbsp;0.108283426725895 &nbsp; &nbsp;0.0857585178966856 &nbsp; &nbsp;6 &nbsp; &nbsp;0.000552728832012655 &nbsp; &nbsp;2.09397277427793e-07 &nbsp; &nbsp;0.890881419183881 &nbsp; &nbsp;0.000282215860934432 &nbsp; &nbsp;1_109279544_G_A &nbsp; &nbsp;1:108779544-109779543<br>1 &nbsp; &nbsp;109275216 &nbsp; &nbsp;rs660240 &nbsp; &nbsp;T &nbsp; &nbsp;C &nbsp; &nbsp;ENSG00000168765 &nbsp; &nbsp;0.1044492.78997299740827e-09 &nbsp; &nbsp;ebi-a-GCST000998 &nbsp; &nbsp;0.214165 &nbsp; &nbsp;0.139318 &nbsp; &nbsp;Alasoo_2018_ge_macrophage_IFNg+Salmonella &nbsp; &nbsp;0.108283426725895 &nbsp; &nbsp;0.0557749486050279 &nbsp; &nbsp;6 &nbsp; &nbsp;0.000552728832012655 &nbsp; &nbsp;2.09397277427793e-07 &nbsp; &nbsp;0.890881419183881 &nbsp; &nbsp;0.000282215860934432 &nbsp; &nbsp;1_109279544_G_A &nbsp; &nbsp;1:108779544-109779543<br>1 &nbsp; &nbsp;109275684 &nbsp; &nbsp;rs629301 &nbsp; &nbsp;G &nbsp; &nbsp;T &nbsp; &nbsp;ENSG00000168765 &nbsp; &nbsp;0.1054716.129993302249e-10 &nbsp; &nbsp;ebi-a-GCST000998 &nbsp; &nbsp;0.197397 &nbsp; &nbsp;0.172673 &nbsp; &nbsp;Alasoo_2018_ge_macrophage_IFNg+Salmonella &nbsp; &nbsp;0.108283426725895 &nbsp; &nbsp;0.22229240584331 &nbsp; &nbsp;6 &nbsp; &nbsp;0.000552728832012655 &nbsp; &nbsp;2.09397277427793e-07 &nbsp; &nbsp;0.890881419183881 &nbsp; &nbsp;0.000282215860934432 &nbsp; &nbsp;1_109279544_G_A &nbsp; &nbsp;1:108779544-109779543<br>1 &nbsp; &nbsp;109278889 &nbsp; &nbsp;rs602633 &nbsp; &nbsp;T &nbsp; &nbsp;G &nbsp; &nbsp;ENSG00000168765 &nbsp; &nbsp;0.1034352.15998134341707e-09 &nbsp; &nbsp;ebi-a-GCST000998 &nbsp; &nbsp;0.226329 &nbsp; &nbsp;0.102673 &nbsp; &nbsp;Alasoo_2018_ge_macrophage_IFNg+Salmonella &nbsp; &nbsp;0.108283426725895 &nbsp; &nbsp;0.0782482718431504 &nbsp; &nbsp;6 &nbsp; &nbsp;0.000552728832012655 &nbsp; &nbsp;2.09397277427793e-07 &nbsp; &nbsp;0.890881419183881 &nbsp; &nbsp;0.000282215860934432 &nbsp; &nbsp;1_109279544_G_A &nbsp; &nbsp;1:108779544-109779543<br>...</p> </blockquote> <p>&nbsp;</p> <p>The dataset provides colocalization statistics for GWAS-eQTL pairs, including posterior probabilities and variant annotations.</p> <h2>2. gwas2eqtl0.1.3.tsv.gz</h2> <p>This file is a filtered version of the colocalization dataset, refined based on cutoffs of PP.H4.abf &ge; 0.75 and SNP.PP.H4 &ge; 0. This subset is utilized in the gwas2eqtl web application for data visualization.</p> <p>Sample Columns:</p> <blockquote> <p>chrom &nbsp; pos19 &nbsp; pos38 &nbsp; cytoband &nbsp; &nbsp; &nbsp; &nbsp;rsid &nbsp; &nbsp;ref &nbsp; &nbsp; alt &nbsp; &nbsp; gwas_trait &nbsp; &nbsp; &nbsp;gwas_class &nbsp; &nbsp; &nbsp;gwas_beta &nbsp; &nbsp; &nbsp; eqtl_gene_symbol &nbsp; &nbsp; &nbsp; &nbsp;eqtl_beta &nbsp; &nbsp; &nbsp; eqtl_id eqtl_gene_id &nbsp; &nbsp;gwas_id gwas_pval &nbsp; &nbsp; &nbsp; eqtl_pval &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;pp_h4_abf &nbsp; &nbsp; &nbsp; snp_pp_h4 &nbsp; &nbsp; &nbsp; tophits_variant_id &nbsp; &nbsp; &nbsp;nsnps<br>1 &nbsp; &nbsp; &nbsp; 1163804 1228424 1p36.33 rs7515488 &nbsp; &nbsp; &nbsp; C &nbsp; &nbsp; &nbsp; T &nbsp; &nbsp; &nbsp; Inflammatory bowel disease &nbsp; &nbsp; &nbsp;Autoimmune dis. 0.0874308 &nbsp; &nbsp; &nbsp; ANKRD65 -0.175816 &nbsp; &nbsp; &nbsp; BrainSeq_ge_brain &nbsp; &nbsp; &nbsp; ENSG00000235098 ebi-a-GCST003043 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.85292266979231e-10 &nbsp; &nbsp;0.000841046 &nbsp; &nbsp; 0.978425254116226 &nbsp; &nbsp; &nbsp; 6.18454060493069e-12 &nbsp; &nbsp;1_1312114_T_C &nbsp; 3<br>1 &nbsp; &nbsp; &nbsp; 1163804 1228424 1p36.33 rs7515488 &nbsp; &nbsp; &nbsp; C &nbsp; &nbsp; &nbsp; T &nbsp; &nbsp; &nbsp; Inflammatory bowel disease &nbsp; &nbsp; &nbsp;Autoimmune dis. 0.0874308 &nbsp; &nbsp; &nbsp; ANKRD65 -0.175816 &nbsp; &nbsp; &nbsp; BrainSeq_ge_brain &nbsp; &nbsp; &nbsp; ENSG00000235098 ieu-a-294 &nbsp; &nbsp; &nbsp; 2.85292266979231e-10 &nbsp; &nbsp; &nbsp; 0.000841046 &nbsp; &nbsp; 0.974019788384412 &nbsp; &nbsp; &nbsp; 7.52286530905187e-12 &nbsp; &nbsp;1_1312114_T_C &nbsp; 4<br>1 &nbsp; &nbsp; &nbsp; 1163804 1228424 1p36.33 rs7515488 &nbsp; &nbsp; &nbsp; C &nbsp; &nbsp; &nbsp; T &nbsp; &nbsp; &nbsp; Inflammatory bowel disease &nbsp; &nbsp; &nbsp;Autoimmune dis. 0.0874308 &nbsp; &nbsp; &nbsp; ANKRD65 -0.293529 &nbsp; &nbsp; &nbsp; CommonMind_ge_DLPFC_naive &nbsp; &nbsp; &nbsp; ENSG00000235098 ebi-a-GCST003043 &nbsp; 2.85292266979231e-10 &nbsp; &nbsp;2.30452e-06 &nbsp; &nbsp; 0.953333690803618 &nbsp; &nbsp; &nbsp; 6.9758581004380506e-15 &nbsp;1_1312114_T_C &nbsp; 6<br>1 &nbsp; &nbsp; &nbsp; 1163804 1228424 1p36.33 rs7515488 &nbsp; &nbsp; &nbsp; C &nbsp; &nbsp; &nbsp; T &nbsp; &nbsp; &nbsp; Inflammatory bowel disease &nbsp; &nbsp; &nbsp;Autoimmune dis. 0.0874308 &nbsp; &nbsp; &nbsp; ANKRD65 -0.293529 &nbsp; &nbsp; &nbsp; CommonMind_ge_DLPFC_naive &nbsp; &nbsp; &nbsp; ENSG00000235098 ieu-a-294 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;2.85292266979231e-10 &nbsp; &nbsp;2.30452e-06 &nbsp; &nbsp; 0.951092499048109 &nbsp; &nbsp; &nbsp; 7.0876793540547e-15 &nbsp; &nbsp; 1_1312114_T_C &nbsp; 7<br>1 &nbsp; &nbsp; &nbsp; 1163804 1228424 1p36.33 rs7515488 &nbsp; &nbsp; &nbsp; C &nbsp; &nbsp; &nbsp; T &nbsp; &nbsp; &nbsp; Inflammatory bowel disease &nbsp; &nbsp; &nbsp;Autoimmune dis. 0.0874308 &nbsp; &nbsp; &nbsp; ANKRD65 -0.510549 &nbsp; &nbsp; &nbsp; FUSION_ge_adipose_naive ENSG00000235098 ebi-a-GCST003043 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.85292266979231e-10 &nbsp; &nbsp;1.2212e-06 &nbsp; &nbsp; &nbsp;0.974412793352836 &nbsp; &nbsp; &nbsp; 1.70246427912903e-11 &nbsp; &nbsp;1_1312114_T_C &nbsp; 6</p> </blockquote> <p>&nbsp;</p> <p>This filtered dataset focuses on high-confidence colocalization events for functional exploration of genetic associations and regulatory mechanisms.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

SaintGSE: Transformer-based efficient and explainable gene set enrichment analysis

<h1>SaintGSE: Transformer-based efficient and explainable gene set enrichment analysis</h1> <div>&nbsp;</div> <div>SaintGSE is an artificial intelligence model designed to predict human gene-pathway relationships using large-scale differentially expressed gene (DEG) datasets. By leveraging an autoencoder and the SAINT transformer model, SaintGSE overcomes challenges in gene expression analysis, such as data scarcity, model compatibility, and interpretability. This project fine-tuned codes from the SAINT project (https://github.com/somepago/saint), licensed under the Apache License 2.0.&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <h2>Key Features</h2> <div>&nbsp;</div> <div>&nbsp; * AI-Driven Pathway Prediction: Uses autoencoders and the SAINT model to analyze gene expression data and predict related signaling pathways.</div> <div>&nbsp;</div> <div>&nbsp; * Osteoarthritis Study: Applied to osteoarthritis (OA) to identify key pathways and potential therapeutic targets.</div> <div>&nbsp;</div> <div>&nbsp; * Explainability: Utilizes Shapley additive explanations (SHAP) to interpret model predictions and identify influential genes.</div> <div>&nbsp;</div> <div>&nbsp;</div> <h2>Installation</h2> <div>&nbsp;</div> <div>Before installation, we recommend to build a conda environment from the attached yml file and activate it.</div> <div>Our code has been tested with python=3.8 on linux.</div> <div>&nbsp;</div> <div>```</div> <div>$ cd /path/to/SaintGSE</div> <div>$ conda env create -f saintgse_env.yml</div> <div>$ conda activate saintgse_env</div> <div>```</div> <div>&nbsp;</div> <div>After downloading all the files, please extract the contents of all compressed directories by running the following command in your terminal:</div> <div>&nbsp;</div> <div>```</div> <div>$ find . -name "*.tar.gz" -exec tar -xzvf {} \;</div> <div>$ rm *.tar.gz</div> <div>```</div> <div>&nbsp;</div> <div>Once the file structure is formed as follows, the preparation for using SaintGSE is complete.</div> <div>&nbsp;</div> <div>```</div> <div>.</div> <div>├── datasets</div> <div>│&nbsp; &nbsp;├── AE_100cycle_model.pth</div> <div>│&nbsp; &nbsp;├── AE_enrichment.tsv</div> <div>│&nbsp; &nbsp;├── AEshap</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp;├── shap_values_latent_dim_1.csv</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp;├── shap_values_latent_dim_2.csv</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp;├── shap_values_latent_dim_3.csv</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp;├── ...</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp;├── ...</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp;└── shap_values_latent_dim_256.csv</div> <div>│&nbsp; &nbsp;├── bestmodels</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp;└── binary</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp; &nbsp; &nbsp;├── Chronic_Myeloid_Leukemia</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp; &nbsp; &nbsp;│&nbsp; &nbsp;└── testrun</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp; &nbsp; &nbsp;│&nbsp; &nbsp; &nbsp; &nbsp;└── saint_gse_model.pth</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp; &nbsp; &nbsp;│── ...</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp; &nbsp; &nbsp;└── Selencompound_Biosynthesis</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;└── testrun</div> <div>│&nbsp; &nbsp;│&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;└── saint_gse_model.pth</div> <div>│&nbsp; &nbsp;├── gene_list.pkl</div> <div>│&nbsp; &nbsp;├── MGI_Gene_Model_Coord.tsv</div> <div>│&nbsp; &nbsp;└── pathway_list_in_DEG.txt</div> <div>```</div> <div>&nbsp;</div> <div> <p>The code in this dataset is also accessible via GitHub. You can find the GitHub repository at the following link:</p> <p><a href="https://github.com/MSjeon27/SaintGSE" target="_new" rel="noopener">https://github.com/MSjeon27/SaintGSE</a></p> </div> <div>&nbsp;</div> <h2>DEG dataset preparation</h2> <div>Prior to SaintGSE analysis, prepare DEG data to be used as input in .tsv format as follows. In the column, the official gene symbol of DEGs is located, and the row adds the log2 fold change value in each DEG group. An example is as follows.</div> <div>&nbsp;</div> <div>```</div> <div>LAP3 CD99 HS3ST1 MAD1L1 LASP1 SNX11</div> <div>'mock-6' vs 'LPS-6' -1.3 0 2.4 0 0.7 0</div> <div>'mock-6' vs 'EBOV-6' -1.3 0 2.3 0 0.6 0</div> <div>```</div> <div>&nbsp;</div> <h2>Usage</h2> <div>&nbsp;</div> <h3>Step 0. Preprocessing the input DEG (from pyDESeq2 result)</h3> <div>&nbsp;</div> <div>Currently, SaintGSE has the function of converting mouse genes into human genes. The preprocessing code serves to change the human or mouse DEG data into the format used for SaintGSE.</div> <div>&nbsp;</div> <div>* human DEGs</div> <div>```</div> <div>$ preprocessing.py --query_fc /path/to/your/DEGs.tsv --out Preprocessed_fc.tsv</div> <div>```</div> <div>&nbsp;</div> <div>* mouse DEGs</div> <div>```</div> <div>$ preprocessing.py --query_fc /path/to/your/DEGs.tsv --org mouse --out Preprocessed_fc.tsv</div> <div>```</div> <div>&nbsp;</div> <div>&nbsp;</div> <h3>Step 1. Training SaintGSE for a target pathway</h3> <div>&nbsp;</div> <div>SaintGSE can be used to analyze new gene expression datasets for pathway prediction:</div> <div>&nbsp;</div> <div>```</div> <div>$ SaintGSE.py --pathway 'Proteins Involved in Osteoarthritis' --pretrain</div> <div>```</div> <div>&nbsp;</div> <div>&nbsp;</div> <h3>Step 2. Prediction through SaintGSE</h3> <div>&nbsp;</div> <div>```</div> <div>$ SaintGSE.py --predict Preprocessed_fc.tsv --pathway 'Proteins Involved in Osteoarthritis'</div> <div>```</div> <div>&nbsp;</div> <div>The results of the predictions are as follows.</div> <div>&nbsp;</div> <div>```</div> <div>tensor([[1.]], device='cuda:0')</div> <div>```</div> <div>&nbsp;</div> <div>This indicates that your DEG data is related to the target signaling path.</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <h3>Step 3. Interpretation the result of SaintGSE (Get Relative SHAP contribution for each DEGs)</h3> <div>```</div> <div>$ Interpret.py -d Preprocessed_fc.tsv -p 'Proteins Involved in Osteoarthritis'</div> <div>```</div> <div>&nbsp;</div> <div>The result of interpretation produces the following files for each sample.</div> <div>&nbsp;</div> <div>```</div> <div>&lt;Sample_name&gt;_&lt;target_pathway&gt;_significant_gene_shap.csv</div> <div>```</div> <div>&nbsp;</div> <div>This represents the relative SHAP contribution for each gene in the DEG data. In the subsequent analysis, it is recommended to focus on the genes with the relative SHAP contributions in the top 35% to 50% as we suggested in the paper, depending on the number of DEGs.</div> <div>&nbsp;</div> <div>&nbsp;</div> <h2>How to Cite</h2> <div>&nbsp;</div> <div>If you use this model or repository in your research, please cite it as follows:</div> <div>&nbsp;</div> <div>```</div> <div>Jeon, MS &amp; Nam, JH et al., "SaintGSE: Transformer-based efficient and explainable gene set enrichment analysis," 2024. GitHub repository. Available at: https://github.com/MSjeon27/SaintGSE</div> <div>```</div> <div>&nbsp;</div> <div>For more information or any questions regarding citation, feel free to contact us (msjeon27@cau.ac.kr).</div> <p>&nbsp;</p>

openapache2.0Nov 2024View details →
zenodo36/100

Gene Enrichment Map Data from gProfiler Analysis - Selected MPK Interactions of Arabidopsis thaliana

<p>Gene enrichment analysis results for the selected predicted MPK interactions are included in the supplementary materials.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Pathways enriched in downstream target genes of miRNAs associated with NT-proBNP and OPN

<p>The full set of&nbsp;30 Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways that were significantly enriched (FDR-adjusted p&lt;0.01) in the predicted mRNA targets of either the OPN- or NT-proBNP-associated miRNAs; 21 of these pathways were significantly enriched in both sets of targets.</p> <p>KEGG: Kyoto Encyclopedia of Genes and Genomes; miRNAs: microRNAs; NT-proBNP: N-terminal pro <a href="http://www.mayomedicallaboratories.com/test-catalog/Clinical+and+Interpretive/83873">B-type natriuretic peptide</a>; OPN: osteopontin</p>

opencc-by-4.0Apr 2019View details →
zenodo36/100

Fig. 3 in Phylogenomics of the tropical plant family Ochnaceae using targeted enrichment of nuclear genes and 250+ taxa

Fig. 3. Continues. For caption, see next part.

opencc-by-4.0Feb 2021View details →
dryad36/100

Data from: Targeted enrichment of large gene families for phylogenetic inference: phylogeny and molecular evolution of photosynthesis genes in the Portullugo clade (Caryophyllales)

Open the record for dataset details and reuse information.

publicSep 2017View details →
dryad36/100

Dataset: T-DNA characterization of genetically modified 3-R-gene late blight resistant potato events with a novel procedure utilizing the Samplix Xdrop® Enrichment Technology

Open the record for dataset details and reuse information.

publicFeb 2024View details →
dryad32/100

Data from: Targeted gene enrichment and high-throughput sequencing for environmental biomonitoring: a case study using freshwater macroinvertebrates

Recent studies have advocated biomonitoring using DNA techniques. In this study, two high-throughput sequencing (HTS)-based methods were evaluated: amplicon metabarcoding of the cytochrome C oxidase subunit I (COI) mitochondrial gene and gene enrichment using MYbaits (targeting nine different genes including COI). The gene-enrichment method does not require PCR amplification and thus avoids biases associated with universal primers. Macroinvertebrate samples were collected from 12 New Zealand rivers. Macroinvertebrates were morphologically identified and enumerated, and their biomass determined. DNA was extracted from all macroinvertebrate samples and HTS undertaken using the illumina miseq platform. Macroinvertebrate communities were characterized from sequence data using either six genes (three of the original nine were not used) or just the COI gene in isolation. The gene-enrichment method (all genes) detected the highest number of taxa and obtained the strongest Spearman rank correlations between the number of sequence reads, abundance and biomass in 67% of the samples. Median detection rates across rare (&lt;1% of the total abundance or biomass), moderately abundant (1–5%) and highly abundant (&gt;5%) taxa were highest using the gene-enrichment method (all genes). Our data indicated primer biases occurred during amplicon metabarcoding with greater than 80% of sequence reads originating from one taxon in several samples. The accuracy and sensitivity of both HTS methods would be improved with more comprehensive reference sequence databases. The data from this study illustrate the challenges of using PCR amplification-based methods for biomonitoring and highlight the potential benefits of using approaches, such as gene enrichment, which circumvent the need for an initial PCR step.

opencc-zeroDec 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record