Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
29,880
datasets available to search
ShareScore release 0.7.1
Dataset results
29,880 results for “gene expression”
Differential Gene Expression Datasets for "Identification of candidate repurposable drugs to combat COVID‑19 using a signature‑based approach"
<p>This dataset has the unfiltered transcriptome differential expression results used in the paper "Identification of candidate repurposable drugs to combat COVID‑19 using a signature‑based approach". </p>
Dataset underlying the study "Enhanced Susceptibility to Tomato Chlorosis Virus (ToCV) in Hsp90- and Sgt1-Silenced Plants: Insights from Gene Expression Dynamics"
<p>This dataset is underlying the scientific publication titled "Enhanced Susceptibility to Tomato Chlorosis Virus (ToCV) in Hsp90- and Sgt1-Silenced Plants: Insights from Gene Expression Dynamics", published in the <a href="https://www.mdpi.com/1999-4915/15/12/2370">Viruses</a> journal. </p><p>The dataset includes a time-course transcriptome analysis using RNA-seq of naïve (no whitefly and no virus), mock (non-viruliferous whiteflies) and ToCV (ToCV_viruliferous whiteflies)-treated tomato samples at 2, 7, and 14 days post-infection (dpi) and viral small RNAs derived from Tomato plants infected with ToCV at 14 dpi. The dataset provided here has been deposited in full by the authors in the European Nucleotide Archive (ENA) at EMBL-EBI under accession number PRJEB67704 (<a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB67704"><strong>https://www.ebi.ac.uk/ena/browser/view/PRJEB67704</strong></a><br><br>The provided information in the dataset are further discussed and interpreted in detail, as well as their subsequent results, in the scientific publication.</p><p>This research was conducted within the VIRTIGATION project, which is part of the EU Open Research Data pilot. This project has received funding from the European Union's Horizon 2020 research and innovation program under grant agreement No. 101000570.</p>
Dataset from the analysis of biological activity of endophytic strain Serratia quinivorans KP32, the expression of biocontrol-related genes and the activity of antioxidant enzymes in bacterial cells treated with pathogenic fungi filtrates
<p>This dataset contains the data from the analyses published in the article entitled "Genetic Determinants of Antagonistic Interactions and the Response of New Endophytic Strain <i>Serratia quinivorans</i> KP32 to Fungal Phytopathogens" in the International Journal of Molecular Sciences (https://doi.org/10.3390/ijms232415561). The data consist of results collected for studies on the antifungal activity of KP32 strain towards four fungal phytopathogens, results of primer efficiency determination and studies on the expression of genes potentially involved in biocontrol after treatment of KP32 strain with the fungal phytopathogens filtrates. Additionally, absorbances from activity tests for catalase (CAT) and superoxide dismutase (SOD) in the strain treated with fungal pathogens are included.</p>
Processed data to accompany "Clonally heritable gene expression imparts a layer of diversity within cell types"
<p>This is the processed data underlying the paper "Clonally heritable gene expression imparts a layer of diversity within cell types" by Mold, Weissman, et al. Data has been gone through preprocessing steps, using the Python Notebooks found at <a href="https://github.com/MartyWeissman/ClonalOmics/tree/main/Data">https://github.com/MartyWeissman/ClonalOmics/tree/main/Data</a>. </p> <p>Smaller files are provided in .csv (comma-separated-value) format and larger files such as expression matrices are provided in .loom format (<a href="https://anndata.readthedocs.io/en/latest/">using the AnnData package</a>).</p> <p> </p> <p> </p>
Raw differential gene expression data, data S1, from: Molecular cascades and cell type-specific signatures in ASD revealed by single cell genomics
<p>Genomic profiling in post-mortem brain from autistic individuals has consistently revealed convergent molecular changes. What drives these changes and how they relate to genetic susceptibility in this complex condition is not understood. We performed deep single nuclear RNA sequencing (snRNAseq) to examine cell composition and transcriptomics, identifying dysregulation of cell type-specific gene regulatory networks (GRNs) in autism, which we corroborated using snATAC-seq and spatial transcriptomics. Transcriptomic changes were primarily cell type-specific, involving multiple cell types, most prominently interhemispheric and callosal-projecting neurons, interneurons within superficial laminae, and distinct glial reactive states involving oligodendrocytes, microglia, and astrocytes. Autism-associated GRN drivers and their targets were enriched in rare and common genetic risk variants, connecting autism genetic susceptibility and cellular and circuit alterations in the human brain. This data is the raw differential gene expression comparing ASD versus CTL subjects for each cell cluster. </p>
Effect of shade and nitrogen content on Arabidopsis Col-0 and cytokinin mutants abcg14 and cypDM mRNA-seq gene expression processed tables.
<p>This dataset is an add-on for Gautrat et al., containing processed files for the mRNAseq data in tab delimited txt format.</p> <p>Here, you can obtain the raw counts file, the normalized CPM values, and the normalized logCPM values</p> <p>The RNA-seq raw data supporting the conclusions of this article have been deposited in ArrayExpress (Kolesnikov et al., 2015) at EMBL-EBI (www.ebi.ac.uk/arrayexpress), under accession numbers E-MTAB-13638.</p> <p>All relevant custom r scripts are available at https://github.com/aromanowski/shade_N_ck</p>
Data from: Gene expression differences between western redcedar seedlings resistant and susceptible to cedar leaf blight
<p>Western redcedar (<em>T. plicata</em>) is an important Cupressaceae both at economic and cultural levels in the Pacific Northwest of North America. In adult trees, the species produces one of the most weathering-resistant heartwoods among conifers, making it one of the preferred species for outdoor applications. However, young <em>T. plicata</em> plants are susceptible to infection with cedar leaf blight (<em>D. thujina</em>), an important foliar pathogen that can be devastating in nurseries and small-spaced plantations. Despite that, variability in the resistance against <em>D. thujina</em> in <em>T. plicata</em> has been documented, and such a variability can be used to breed <em>T. plicata</em> for resistance against the pathogen. This investigation aimed to discern the phenotypic and gene expression differences between resistant and susceptible <em>T. plicata</em> seedlings to shed light on the potential constitutive resistance mechanisms against cedar leaf blight in western redcedar. The study consisted of two parts. First, the histological differences between four resistant and four susceptible families that were never infected with the pathogen were investigated. And second, the differences between one resistant and one susceptible family that were infected and not infected with the pathogen were analyzed at the chemical (C, N, mineral nutrients, lignin, fiber, starch, and terpenes) and gene expression (RNA-Seq) levels. The histological part showed that <em>T. plicata</em> seedlings resistant to <em>D. thujina</em> had constitutively thicker cuticles and lower stomata densities than susceptible plants. The chemical analyses revealed that, regardless of their infection status, resistant plants had higher foliar concentrations of sabinene and α-thujene, and higher levels of expression of transcripts that code for leucine-rich repeat receptor-like protein kinases and for bark storage proteins. In conclusion, the data collected in this study shows that constitutive differences at the phenotypic (histological and chemical) and gene expression level exist between <em>T. plicata</em> seedlings susceptible and resistant to <em>D. thujina</em>. Such differences have potential use for marker-assisted selection and breeding for resistance against cedar leaf blight in western redcedar in the future.</p>
FIGURE 5 in Differential expression of HPG-axis genes in autotetraploids derived from red crucian carp Carassius auratus red var., × blunt snout bream Megalobrama amblycephala,
FIGURE 5 Mean (+SD) relative expression of gnrh2, fshb, lhb, fshr and lhr messenger (m)RNA in (a) the breeding season () 2n, and () 4n and (b) the non-breeding season in Carassius auratus red var. () 2n, and () 4n. (RCC,) and autotetraploid C. auratus red var. ♀ × Megalobrama amblycephala ♂ (4nRR,). T, gene detected in the testis; O, gene detected in the ovary. *, significant difference between RCC and 4nRR (P <0.05)
FIGURE 4 Deduced amino-acid sequences for the Gnrh2 in Differential expression of HPG-axis genes in autotetraploids derived from red crucian carp Carassius auratus red var., × blunt snout bream Megalobrama amblycephala,
FIGURE 4 Deduced amino-acid sequences for the Gnrh2 () and Lhr () genes in Carassius auratus red var. (RCC) and autotetraploid C. auratus red var. ♀ × Megalobrama amblycephala ♂ (4nRR)
FIGURE 2 in Differential expression of HPG-axis genes in autotetraploids derived from red crucian carp Carassius auratus red var., × blunt snout bream Megalobrama amblycephala,
FIGURE 2 (a) The mature eggs (scale bar = 100 μm) and (b) mature sperm (scale bar = 10 μm) of autotetraploid Carrasius auratus red var. ♀ x Megalobrama amblycephala ♂ (4nRR)
FIGURE 3 in Differential expression of HPG-axis genes in autotetraploids derived from red crucian carp Carassius auratus red var., × blunt snout bream Megalobrama amblycephala,
FIGURE 3 Reverse-transcription (RT)-PCR analysis of the expression of (a) gnrh2, (b) fshb, (c) lhb, (d) fshr and (e) lhr messenger (m)RNA in various tissues of autotetraploid Carrasius auratus red var. ♀ x Megalobrama amblycephala ♂ (4nRR). The upper strip of each panel (a)–(e) shows the positive control of actin gene while the lower strip of each panel shows the RT-PCR amplification of the target gene
FIGURE 1 in Differential expression of HPG-axis genes in autotetraploids derived from red crucian carp Carassius auratus red var., × blunt snout bream Megalobrama amblycephala,
FIGURE 1 The gonadal structure of Carassius auratus red var. [RCC; (a)–(c)] and autotetraploids C. auratus red var. ♀ × Megalobrama amblycephala ♂ [4nRR; (d)–(f)]: (a) ovary of 7 month-old RCC containing many phase II and a few phase III oocytes; (b) ovary of 12 month-old RCC showing many mature phase IV ova; (c) testis of 12 month-old RCC with numerous mature sperms () and a small amount of spermatocytes () in the lobules of testes; (d) ovary of 7 month-old 4nRR containing phase II and a few phase III oocytes; (e) ovary of 12 month-old 4nRR with numerous mature phase IV ova; (f) testis of 12 month-old 4nRR with numerous mature sperms () and a small amount of spermatocytes () in the lobules of testes, the scale bars: (a), (b), (d), and (e) = 100 μm; (c) and (f) = 10 μm
Supplementary data of article "Domestication has altered gene expression and secondary metabolites in pea seed coat".
<p><strong>Table S1.</strong> Excel- GO_terms_MF_selected_WGCNA_modules.</p> <p><strong>Table S2.</strong> Excel- GO_terms_MF_DEGs_UP_and_DOWN.</p> <p><strong>Table S3.</strong> Excel- GO_terms_MF_DEGs_summary.</p> <p><strong>Table S4.</strong> Excel- List of DEGs involved in flavonoid pathway found in WILD gene set.</p> <p><strong>Table S5.</strong> Protein recoveries calculated for individual pea protein samples. Numbers 1, 2, 3 denote treatment groups corresponding to seed developmental stages (D1, D2 and mature seeds, respectively). Letters a–d denote biological replicates within the treatment groups.</p> <p><strong>Table S6.</strong> Excel- Annotation of proteins differentially expressed in wild and domesticated pea seed coat samples.</p> <p><strong>Table S7.</strong> Primary metabolites identified by spectral similarity library search and/or co-elution with authentic standards in pea seed coats aqua methanolic extracts. Metabolite analysis relied on GC-EI-Q-MS analysis after derivatization of the lyophilized extracts with methoxamine hydrochloride (MOA) and <em>N</em>-methyl-<em>N</em>-(trimethylsilyl)trifluoroacetamide (MSTFA).</p> <p><strong>Table S8.</strong> Primary metabolites detected in the aq. methanolic extracts of mature Cameor seed coats demonstrating statistically significant up- and down-regulation in comparison to those of wild JI261.</p> <p><strong>Table S9.</strong> Primary metabolites of mature JI92 seed coats demonstrating statistically significant up- and down-regulation in comparison with those of wild JI261.</p> <p><strong>Table S10.</strong> Primary metabolites of mature JI1794 seed coats demonstrating statistically significant up- and down-regulation in comparison with those of JI261.</p> <p><strong>Table S11.</strong> Primary metabolites of mature JI64 seed coats demonstrating statistically significant up- and down-regulation in comparison with those of JI261.</p> <p><strong>Table S12.</strong> Mass analyzer settings applied for QqTOF-MS experiments in analysis of seed coat (cell wall) hydrolyzates and reference authentic standards.</p> <p><strong>Table S13.</strong> Cell wall-bound metabolites extracted from the seed coats of wild (JI64, JI1794, JI261) and domesticated (Cameor, JI92) peas upon alkali hydrolysis of corresponding isolated and purified cell wall material.</p> <p><strong>Table S14.</strong> Excel- Coordinates of markers in S-plot obtained from OPLS-DA analysis (FIA-ESI-HRTMS, negative ionization, lock mass uncorrected).</p> <p><strong>Table S15.</strong> List of identified significantly differential metabolites rising during seed coat development.</p> <p><strong>Table S16.</strong> List of identified significantly differential metabolites decreasing during seed coat development (positive ionization mode).</p> <p><strong>Table S17.</strong> List of identified metabolites with significantly higher content in wild compared cultivated genotypes in older developmental stages (D5-6).</p> <p><strong>Table S18.</strong> Excel- Expression of genes encoding enzymes of monolignol pathway in seed coats (SC) and embryos (E) of domesticated (Cameor, JI92 and <em>Pisum abyssinicum</em> PI358617) and wild (JI64, JI1794, JI261) peas over five seed developmental stages (13, 17, 20, 23, 28 DAP, labelled as 1-5). PAL: phenylalanine ammonia-lyase, C4H: cinnamate-4-hydroxylase, 4CL: 4-coumaroyl: CoA ligase, HCT: hydroxycinnamoyl CoA:shikimate hydroxycinnamoyltransferase, COMT: caffeic acid O-methyltransferase, CSE: caffeoyl shikimate esterase, CAD: cinnamyl alcohol dehydrogenase, CCR: cinnamoyl CoA reductase, CCoAMT: caffeoyl CoA-3-methyltransferase, F5H: ferulate-5-hydroxylase</p> <p><strong>Table S19.</strong> Studied metabolites of phenylpropanoid pathway.</p> <p><strong>Table S20.</strong> Instrument settings used in the proteomics LIT-Orbitrap-MS and -MS/MS experiments.</p> <p><strong>Table S21.</strong> Procedures and specific settings for data processing and post-processing of the proteomics data.</p> <p><strong>Table S22.</strong> Gas chromatographic (GC) separation conditions and electron ionization-quadrupole-mass spectrometry (EI-Q-MS) settings for GC-EI-Q-MS analysis of the primary metabolites in pea seed coats.</p> <p><strong>Table S23.</strong> Chromatographic conditions used for UHPLC separation of seed coat (cell wall) hydrolyzates and reference authentic standards.</p> <p><strong>Table S24.</strong> Variable parameters of MS/cIMS/MS measurements.</p> <p><strong>Figure S1.</strong> The dynamics of gene expression between studied developmental stages within all genotypes (a) or among genotypes in particular developmental stages (b).</p> <p><strong>Figure S2.</strong> Twelve representative groups of transcription factors described within 20 gene modules of pea SC. Visualized by Cytoscape 3.9.0.</p> <p><strong>Figure S3.</strong> SDS-PAGE electropherograms of the total protein fractions isolated from the seed coats of JI92 (a, c, e) and JI64 (b, d, f) seeds before and after tryptic hydrolysis. Numbers 1, 2, 3 denote seed developmental stages: DS1, DS2 and mature seeds, respectively. Letters a-d denote biological replicates. The aliquots (10 μg) of samples before hydrolysis (a, b), the incompletely digested aliquots left on filter unit after peptide elution (c, d) and aliquots of tryptic hydrolysates (corresponding to 5 μg of protein), (e, f) were loaded on gels. Inter-gel normalization relied on the total density of the Protein Ladder (PageRuler™ Prestained Protein Ladder #26616, 10–180 kDa) lane (St); the ND (non-digested) sample represents a reference protein not subjected to hydrolysis.</p> <p><strong>Figure S4.</strong> The numbers of tryptic peptides (a), possible proteins (b), and non-redundant proteins (protein groups) (c) identified in domesticated JI92 seed coats at developmental stages D1, D2 and D6. The tryptic digests (<em>n</em> =&thinsp;3) obtained from seed coats were analyzed by nano-high performance liquid chromatography-electrospray ionization linear ion trap-orbital trap mass spectrometry (nanoHPLC-ESI-LIT-Orbitrap-MS) operated in positive DDA mode.</p> <p><strong>Figure S5.</strong> The numbers of tryptic peptides (a), possible proteins (b), and non-redundant proteins (protein groups, c) identified in wild pea JI64 seed coats at D1, D2 and D6 stages. The tryptic digests (<em>n</em> =&thinsp;3), obtained from pea seedlings, were analyzed by nano-high performance liquid chromatography-electrospray ionization linear ion trap-orbital trap mass spectrometry (nanoHPLC-ESI-LIT-Orbitrap-MS) operated in positive DDA mode.</p> <p><strong>Figure S6.</strong> Principal component analysis (PCA) with score plot representation (a) accomplished for seed coat proteins differentially expressed at developmental stages D1 and D2 and in the mature state (D6) and hierarchical clustering with a heatmap representation (b).</p> <p><strong>Figure S7.</strong> Functional annotation (accomplished with the Mercator MapMan v3.6 tool) of the pea seed coat proteins isolated in stage D1. White and black columns denote the functional groups of the proteins, which were more expressed in the developing seeds of domesticated JI92 and wild JI64, respectively.</p> <p><strong>Figure S8.</strong> Functional annotation (accomplished with the Mercator MapMan v3.6 tool) of the pea seed coat proteins isolated in stage D2. White and black columns denote the functional groups of the proteins, which were more expressed in the developing seeds of the domesticated JI92 and wild JI64, respectively.</p> <p><strong>Figure S9.</strong> Prediction of sub-cellular localization of the proteins more expressed in the developing seeds of JI92 and JI64 with the BUSCA prediction tool.</p> <p><strong>Figure S10.</strong> Evaluation of the differences in the metabolic profiles of the mature seeds obtained from the wild JI261 and domesticated Cameor by principal component analysis (PCA).</p> <p><strong>Figure S11.</strong> Representation of the differences in the metabolic profiles of the mature seed coats obtained from the wild JI261 and domesticated Cameor by the t-test with Volcano plot representation (a) and the top 30 differentially abundant metabolites demonstrating the most pronounced differences of corresponding GC-MS signals associated with seed dormancy (b).</p> <p><strong>Figure S12.</strong> Evaluation of the differences in the metabolic profiles of the mature seeds obtained from the wild JI261 and domesticated JI92 by principal component analysis (PCA) with score plot representation (a) and hierarchical clustering with heatmap representation (b).</p> <p><strong>Figure S13.</strong> Principal component analysis (PCA) illustrating distribution of metabolic profiles of mature seed coats of two wild pea genotypes, JI1794 and JI261.</p> <p><strong>Figure S14.</strong> Principal component analysis (PCA) demonstrates the distribution of mature seed coat metabolic profiles of two wild pea genotypes, JI64 and JI261(control).</p> <p><strong>Figure S15.</strong> Evaluation of the differences in the patterns of the cell wall-bound metabolites obtained from mature seed coats of wild JI261 and domesticated Cameor: principal component analysis (PCA) with score plot representation (a), hierarchical clustering with heatmap representation (b) and <em>t</em>-test analysis with the Volcano-plot representation (c).</p> <p><strong>Figure S16.</strong> Statistical analysis (<em>t</em>-test with Volcano plot representation) characterizing the differences between the levels of mature seed coat cell wall-bound metabolites of Cameor compared with those of wild JI261.</p> <p><strong>Figure S17.</strong> Principal component analysis (PCA) illustrates the distribution of mature seed coat metabolic profiles of domesticated JI92 and wild JI261.</p> <p><strong>Figure S18.</strong> Principal component analysis (PCA) shows the distribution of metabolic profiles of mature seed coats of two wild pea genotypes, JI1794 and JI261, control.</p> <p><strong>Figure S19.</strong> Principal component analysis (PCA) demonstrates the distribution of mature seed coat metabolic profiles of two wild genotypes, JI64 and control JI261.</p> <p><strong>Figure S20.</strong> Annotated cell wall-bound metabolites extracted from the seed coats of the dormant wild pea genotype JI261 and the seed coats from several pea genotypes varying in their dormancy (Cameor, JI92, JI64, and JI1794) upon alkali hydrolysis of corresponding isolated and purified cell wall material. </p> <p><strong>Figure S21.</strong> Ion mobility separation of <em>m/z</em> 299.0841.</p> <p><strong>Figure S22.</strong> Ion mobility separation of <em>m/z</em> 701.1907. </p> <p><strong>Figure S23.</strong> Ion mobility separation of <em>m/z</em> 619.1041.</p> <p><strong>Figure S24.</strong> Ion mobility separation of <em>m/z</em> 631.1017.</p> <p><strong>Figure S25.</strong> Ion mobility separation of <em>m/z</em> 641.1139.</p> <p><strong>Figure S26.</strong> Ion mobility separation of <em>m/z</em> 771.1346. </p> <p><strong>Figure S27.</strong> Reconstructed chromatograms of p-hydroxybenzoic and salicylic acids in DS5 of dormant JI64 and domesticated landraces JI92 (LC/HRTMS, negative ionization mode).</p> <p><strong>Figure S28.</strong> Module-trait relationship depiction showing the correlation between expression of the gene modules and the abundance of identified metabolites of the monolignol pathway.</p>
Distinct tissue-dependent composition and gene expression of human fetal innate lymphoid cells
<p>Countmatrix and metadata for gene expression in bulk NK cells and CD304+ ILC3s isolated from human fetal liver, lung, intestine, and skin. </p>
Consensus molecular environment of schizophrenia risk genes in co-expression networks shifting across age and brain regions
<p>This is the online data repository accompanying the following manuscript:<br><strong>Consensus molecular environment of schizophrenia risk genes in coexpression networks shifting across age and brain regions</strong></p> <p><em>Giulio Pergola<sup>1,2,3,*</sup>, Madhur Parihar<sup>1</sup>, Leonardo Sportelli<sup>1,2</sup>, Rahul Bharadwaj<sup>1</sup>, Christopher Borcuk<sup>2</sup>, Eugenia Radulescu<sup>1</sup>, Loredana Bellantuono<sup>2,5</sup>, Giuseppe Blasi<sup>2,4</sup>, Qiang Chen<sup>1</sup>, Joel E. Kleinman<sup>1,3</sup>, Yanhong Wang<sup>1</sup>, Srinidhi Rao Sripathy<sup>1</sup>, Brady J. Maher<sup>1,3,7</sup>, Alfonso Monaco<sup>5,9</sup>, Fabiana Rossi<sup>1,2</sup>, Joo Heon Shin<sup>1</sup>, Thomas M. Hyde<sup>1,3,6</sup>, Alessandro Bertolino<sup>2,4,*</sup>, Daniel R. Weinberger<sup>1,7,8,*</sup></em></p> <p> </p> <p><strong>Affiliations:</strong></p> <p><em>1)Lieber Institute for Brain Development, Johns Hopkins Medical Campus, Baltimore, MD (USA)<br>2)Group of Psychiatric Neuroscience, Department of Translational Biomedicine and Neuroscience, University of Bari Aldo Moro, Bari, Italy<br>3)Department of Psychiatry and Behavioral Sciences, Johns Hopkins University School of Medicine, Baltimore, Maryland<br>4)Azienda Ospedaliero-Universitaria Consorziale Policlinico, Bari, Italy<br>5)Istituto Nazionale di Fisica Nucleare (INFN), Bari, Italy<br>6)Department of Neurology, Johns Hopkins University School of Medicine, Baltimore, Maryland<br>7)Department of Neuroscience, Johns Hopkins University School of Medicine, Baltimore, Maryland<br>8)Department of Genetic Medicine, Johns Hopkins University School of Medicine, Baltimore, Maryland<br>9)Dipartimento Interateneo di fisica, Università degli Studi di Bari Aldo Moro, Bari, Italy</em></p> <p> </p> <p><strong>Abstract:</strong></p> <p><em>Schizophrenia is a neurodevelopmental brain disorder whose genetic risk is associated with shifting clinical phenomena across the life span. We investigated the convergence of putative schizophrenia risk genes in brain coexpression networks in postmortem human prefrontal cortex (DLPFC), hippocampus, caudate nucleus, and dentate gyrus granule cells, parsed by specific age periods (total N = 833). The results support an early prefrontal involvement in the biology underlying schizophrenia and reveal a dynamic interplay of regions in which age parsing explains more variance in schizophrenia risk compared to lumping all age periods together. Across multiple data sources and publications, we identify 28 genes that are the most consistently found partners in modules enriched for schizophrenia risk genes in DLPFC; twenty-three are previously unidentified associations with schizophrenia. In iPSC-derived neurons, the relationship of these genes with schizophrenia risk genes is maintained. The genetic architecture of schizophrenia is embedded in shifting coexpression patterns across brain regions and time, potentially underwriting its shifting clinical presentation.</em></p> <p> </p> <p><strong>Citation:</strong> <em>Giulio Pergola et al. ,Consensus molecular environment of schizophrenia risk genes in coexpression networks shifting across age and brain regions.Sci. Adv.9, eade2812(2023).DOI:10.1126/sciadv.ade2812</em></p> <p> </p> <p><strong>Data Files:<br>DLPFC hit.genes_kb_200__online.version.zip: </strong><br>Interactive Sankey plot for age-parsed DLPFC networks with SCZ genes (200 kbp list) only. For Sankey plots, hover mouse over the links to see the list of genes. Also supports zoom, drag and selection.<br><strong>DLPFC hit.genes_kb_200__paper.version.zip:</strong><br>Interactive Sankey plot for age-parsed DLPFC networks with SCZ genes (200 kbp list) only. For paper version of the figure, smaller modules are merged into a macro-module (lightgrey color)<br><strong>DLPFC all.genes_kb_200__online.version.zip:</strong><br>Interactive Sankey plot for age-parsed DLPFC networks with all genes<br><strong>DLPFC all.genes_kb_200__paper.version.zip:</strong><br>Interactive Sankey plot for age-parsed DLPFC networks with all genes. For paper version of the figure, smaller modules are merged into a macro-module (lightgrey color)<br><strong>HP hit.genes_kb_200__online.version.zip:</strong><br>Interactive Sankey plot for age-parsed Hippocampus networks with SCZ genes (200 kbp list) only<br><strong>HP hit.genes_kb_200__paper.version.zip:</strong><br>Interactive Sankey plot for age-parsed Hippocampus networks with SCZ genes (200 kbp list) only. For paper version of the figure, smaller modules are merged into a macro-module (lightgrey color)<br><strong>HP all.genes_kb_200__online.version.zip:</strong><br>Interactive Sankey plot for age-parsed Hippocampus networks with all genes<br><strong>HP all.genes_kb_200__paper.version.zip:</strong><br>Interactive Sankey plot for age-parsed Hippocampus networks with all genes. For paper version of the figure, smaller modules are merged into a macro-module (lightgrey color)<br><strong>Modulewise SCZ enrichment(1.0).xlsx:</strong><br>Excel file contains module level SCZ enrichment results for all networks<br><strong>wide_form_test_slidingwindow_NC_SchizoNew(v1.4)_final.xlsx:</strong><br>Excel file contains WGCNA output for sliding window networks<br><strong>wide_form_WGCNA(v3.7.1)_final.xlsx:</strong><br>Excel file contains WGCNA output for our generated networks and from previously published networks<br><strong>libdnetworks(NC).preprocessed.exp.RData: </strong><br>Preprocessed ranknormalised expression assay for age-parsed/nonparsed NC networks (DLPFC, HP, CAUDATE, DENTATE). For fixed window and sliding window study.<br><strong>libdnetworks(SCZ).preprocessed.exp.RData: </strong><br>Preprocessed ranknormalised expression assay for nonparsed SCZ networks (DLPFC, HP, CAUDATE, DENTATE). For the sliding window study.<br><strong>sample_matched_HP_DG_qsva(NC).preprocessed.exp.RData:</strong><br>Preprocessed ranknormalised expression assay for the sample-matched HP-DG. QSVA removed pipeline. For Cell population enrichment study.<br><strong>sample_matched_HP_DG_noqsva(NC).preprocessed.exp.RData:</strong><br>Preprocessed ranknormalised expression assay for the sample-matched HP-DG. No QSVA removed pipeline. For Cell population enrichment study.<br><strong>stemcell.preprocessed.exp.RData:</strong><br>Preprocessed ranknormalised expression assay for the iPSC network. For replication in human iPSC data study. Neuronal samples averaged for each “RealGenome”.<br><strong>SCZ.ref.list.sciadv.ade2812.rds</strong>: List of All Biotypes/ Protein Coding Schizophrenia reference genelist for following bins: PGC3, 0 kbp, 20 kbp, 50 kbp, 100 kbp, 150 kbp, 200 kbp, 250 kbp, 500 kbp.</p> <p> </p> <p>Accompanying code can be found at: <a href="https://github.com/LieberInstitute/Brain_WGCNA">https://github.com/LieberInstitute/Brain_WGCNA</a><br>Data from this repository is also available at: <a href="https://nets.libd.org/age_wgcna/">https://nets.libd.org/age_wgcna/</a></p> <p> </p> <p>For any data inquiries please contact:<br><strong>Giulio Pergola: </strong><a href="mailto:Giulio.Pergola@libd.org"><strong>Giulio.Pergola@libd.org</strong></a></p> <p> </p>
Gene expression and splicing counts from 49 tissues from GTEx v6p genome build hg19 - non-strand specific
<p><strong>Dataset description:</strong></p> <p>49 folders, each corresponding to one tissue from GTEx v6p and containing the following files:</p> <ol> <li> <p>geneCounts: gene-level counts </p> </li> <li> <p>k_j: split counts spanning from one exon to another.</p> </li> <li> <p>k_theta: non-split counts covering a splice site</p> </li> <li> <p>n_psi3: total split counts from a given acceptor site</p> </li> <li> <p>n_psi5: total split counts from a given donor site</p> </li> <li> <p>n_theta: total split and non-split counts for a given splice site</p> </li> <li> <p>Sample annotation describing each sample from the dataset</p> </li> <li> <p>Description file with global information from the dataset</p> </li> </ol> <p>The gene counts were originated using the GTF file from <a href="http://www.gencodegenes.org/human/release_29lift37.html">release 29 of GENCODE</a>, and the split and non-split counts contain only the annotated junctions from the same release. Statistics are reported only for GENCODE-annotated introns and splice sites, in compliance with the regulations of the GTEx consortium. For a description of the samples, methods, and protocols, see the GTEx publication specified below.</p> <p><strong>Use: </strong>The count matrices are intended to help researchers that are interested in using RNA-Seq data with the purpose of diagnostics. Researchers can merge their own dataset with the downloaded ones, provided the tissue, genome build, strand, and paired-end specifications match. Afterwards, the <a href="https://github.com/gagneurlab/drop">Detection of RNA outliers Pipeline (DROP) </a> can be used to compute gene expression and splicing outliers.<br> <strong>Organism:</strong> Homo sapiens<br> <strong>Genome assembly:</strong> hg19<br> <strong>Gene annotation:</strong> gencode29<br> <strong>Strand specific: </strong>FALSE<br> <strong>Paired end: </strong>TRUE<br> <strong>Protocol: </strong>poly(A) enrichment</p> <p><strong>Contact:</strong> Vicente A. Yepez, yepez at in.tum.de; Christian Mertes, mertes at in.tum.de; Julien Gagneur, gagneur at in.tum.de</p> <p><strong>Citation:</strong> Write the following in the "Data availability" section of the manuscript or similar replacing the three citations by the ones from the References section below:</p> <blockquote> <p><strong>The count matrices for the GTEx samples <cite GTEx publication, see below> were downloaded from Zenodo (doi: 10.5281/zenodo.5596755) and were generated through DROP <cite DROP, see below> using the release 29 of the GENCODE annotation <cite GENCODE, see below>. </strong></p> </blockquote> <p>Also, write the following in the Acknowledgements section:<br> </p> <blockquote> <p><strong>The Genotype-Tissue Expression (GTEx) Project was supported by the Common Fund of the Office of the Director of the National Institutes of Health, and by NCI, NHGRI, NHLBI, NIDA, NIMH, and NINDS. The raw data used for the analyses described in this manuscript were obtained from the GTEx Portal on June 12, 2017, under accession number dbGaP phs00424.v6.p1.</strong></p> </blockquote> <p><br> </p>
Genome-wide gene expression noise in Escherichia coli is condition-dependent and determined by propagation of noise through the regulatory network
<p>In this repository we provide raw and processed datasets for the article: “Genome-wide gene expression noise in <em>Escherichia coli </em>is condition-dependent and determined by propagation of noise through the regulatory network<strong>” </strong>by Arantxa Urchueguía, Luca Galbusera, Dany Chauvin, Gwendoline Bellement, Thomas Julou and Erik van Nimwegen.</p> <p>A preprint is available under the following DOI: <a href="https://doi.org/10.1101/795369">https://doi.org/10.1101/795369</a>. </p> <p>The repository consists of the following datasets: </p> <p><strong>1. preprocessed_datasets.zip(~22GB)</strong></p> <ul> <li>This dataset contains raw data from the flow cytometry experiments (FACS Canto II, BD Bioscience) in all measured conditions in RData format. Raw fcs files were processed with the tools described in the publication ''Using fluorescence flow cytometry data for single-cell gene expression analysis in bacteria" published here: <a href="https://doi.org/10.1371/journal.pone.0240233">https://doi.org/10.1371/journal.pone.0240233</a>. The tools themselves are available here: <a href="https://github.com/vanNimwegenLab/E-Flow">https://github.com/vanNimwegenLab/E-Flow</a>. Included in the files are the outputs of these processing tools together with all raw values that came directly from the flow cytometer. The file <em>directory_structure_in_preprocessed </em>contains information about how the files are organized.</li> </ul> <p><strong>2. info_files: </strong>This is a set of csv files containing detailed information about the experiments done to acquire the preprocessed_datasets as well as annotation files that we used to retrieve promoter information. </p> <p><strong>3. processed_datasets:</strong> These files correspond to the processed datasets from the raw Rdata files under 1 above. The processed data provide mean and variance estimates in fluorescence of E.coli promoters across the different growth conditions. Note that we discarded flow cytometry measurements from promoter/growth-condition combinations that contained abnormal fluorescence distributions (due to contamination) as well as measurements from reporters with annotation mismatches. The folder contains the following clean dataset files that were used in the paper:</p> <ul> <li><strong>FULL_dataset_mean_var_wreplicates:</strong> In this dataset we include the processed means and variances (in both logarithmic and linear scale) of all promoters in each condition. Included as well are replicate measurements for some conditions.. We also include the name and Blattner number of the gene immediately downstream of each promoter, the DNA sequence of each promoter, and regulatory information (number of unique inputs for transcription factors sites and their names) which we obtained from RegulonDB v 10.5 (<a href="https://doi.org/10.1093/nar/gky1077">https://doi.org/10.1093/nar/gky1077</a>). </li> <li><strong>dataset_with_noise_estimates: </strong>In this dataset we provide noise estimates for all promoters expressed above an expression threshold (mean GFP fluorescence at least as large as autofluorescence). Note that the noise estimate correspond to the difference between the promoter’s variance in log-expression and the minimal variance as a function of its mean expression (i.e. the so called noise floor was subtracted). Apart from the mean, variance, noise and promoter features (sequence, name of gene downstream, number of unique regulatory inputs and name of the TFs binding), we also include the parameters used for fitting the minimal noise, i.e. noise floor, in each of the conditions. </li> <li><strong>time_course_data_SI</strong>: This dataset contains mean and variance measurements of one of the plates of the library measured at different time points during growth in Minimal media 0.4M NaCl: 0h (just after dilution), 1h, 2h, 3h, 5h, 6.5h, 8.5h, 10h and 11h. </li> <li><strong>growth_curves_SI</strong>: Growth data (OD<sub>600</sub> as a function of time) for a subset of the promoters from the library across different growth conditions.</li> <li><strong>singlecell_areas_SI: </strong>Single-cell areas estimated using agar patches of cells growing in each condition. Each row of the table contains data for a single-cell. </li> <li><strong>synthetic_promoters_dataset: </strong>This dataset contains mean, variance and noise measurements of a set of constitutive promoters from <a href="https://doi.org/10.7554/eLife.05856.001">https://doi.org/10.7554/eLife.05856.001</a> across different conditions.</li> <li><strong>MARA_results:</strong> All transcription factor activities results explaining measured noise levels in each condition. This data has been obtained after performing Motif Activity Response Analysis on the noise levels of all measured promoters in each condition.</li> </ul>
Immune disease variants modulate gene expression in regulatory CD4+ T cells
<p>We mapped genetic regulation (QTL) of gene expression and chromatin activity in Tregs and we identified 133 colocalizing loci with immune disease variants.<br> For the time being, the preprint DOI: <a href="https://doi.org/10.1101/654632">10.1101/654632</a></p>
Predicted gene expression in ancestrally diverse populations leads to discovery of susceptibility loci for lifestyle and cardiometabolic traits
<p>Full summary statistics for the publication "Predicted gene expression in ancestrally diverse populations leads to discovery of susceptibility loci for lifestyle and cardiometabolic traits". </p> <p>The files, bmi.UKBBsummary.txt and height.UKBBsummary.txt, contain tissue specific associations with body mass index (BMI) and height respectively. The suffix UKBB450k indicates results from all ~450,000 European ancestry individuals in UK Biobank. The suffix UKBB50k corresponds to results from a subset of 50,000 Europeans in the UK Biobank. The suffix PAGE corresponds to results from ~50,000 individuals in the Population Architecture using Genomics and Epidemiology (PAGE) study. </p> <p>The file PAGE_PrediXcan_associations.txt includes trait~tissue specific GReX associations for 25 traits. The first field specifies the tissue.trait.gene of the association results.</p>
Summary Statistics from "Genetically regulated gene expression and proteins revealed discordant effects" (LWAS of biomarker)
<p>Summary statistics of 92 blood protein levels. The corresponding publication is currently under revision.</p> <p> The zipped txt file is tab-delimited and contains the following columns:</p> <ul> <li>protein: protein name abbreviation</li> <li>cytoband: genomic region</li> <li>gene: gene name abbreviation</li> <li>setting: either "combined" (adj. for sex & age) or sex-stratified ("males", "females"; adj. for age)</li> <li>variant_id_hg19: SNP ID according to hg19</li> <li>variant_id_hg38: SNP ID according to hg19</li> <li>chr: chromosome</li> <li>pos_hg19: base position according to hg19</li> <li>pos_hg38: base position according to hg19</li> <li>effect_allele: also known as counted allele in additive model</li> <li>other_allele: not-counted allele</li> <li>eaf: effect allele frequency</li> <li>maf: minor allele frequency</li> <li>info: imputation info score</li> <li>n_samples: number of samples</li> <li>beta: effect estimate</li> <li>se: standard error</li> <li>zscore: Z-statistic</li> <li>pvalue: p-value</li> <li>FDR: FDR by gene and setting</li> <li>BBFDR: hierarchical FDR by setting</li> <li>hierFDR: TRUE if SNP is significant after hierarchical FDR</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.