Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
650
datasets available to search
ShareScore release 0.7.1
Dataset results
650 results for “Differential expression analysis”
Datasets, reproducible codes, and results for evaluating differential expression analysis methods on population-level RNA-seq data
<p>This upload contains the necessary R codes and data to reproduce the FDR and Power results described in our correspondence "Neglecting normalization impact in semi-synthetic RNA-seq data simulation generates artificial false positives" to Li Y, Ge X, Peng F, Li W, Li JJ, Exaggerated false positives by popular differential expression methods when analyzing human population samples, <em>Genome Biology</em> 23, 79, 2022, DOI: 10.1186/s13059-022-02648-4.</p>
Data from "Corset: enabling differential gene expression analysis for de novo assembled transcriptomes"
<p>This dataset contains de novo transcriptome assemblies for three publicly available RNA-seq dataset (SRA055442, SRR453566-SRR453571 and GSE37704 ). For each assembly we also provide a table with the read counts per contig, the output from corset (clusters and counts), and the results from a genome-based analysis. This dataset was used to assess the performance of the corset software. More detail is provided in the paper: Nadia M Davidson and Alicia Oshlack,<strong> </strong>Corset: enabling differential gene expression analysis for de novo assembled transcriptomes, <em>Genome Biology</em> 2014, <strong>15</strong>:410. http://genomebiology.com/2014/15/7/410/abstract</p>
Integrating differential expression and weighted correlation network analysis for identifying genes controlling shoot development in Sorghum bicolor
<p>Supplementery materials of journal article "Integrating differential expression and weighted correlation network analysis for identifying genes controlling shoot development in <em>Sorghum bicolor</em>"</p>
Pelagomonas calceolata gene expression levels in different nitrogen conditions and differential expression analysis.
<p>These files contains the expression levels and DESeq2 results of <em>Pelagomonas calceolata</em> genes cultivated with different nitrate conditions. Two strains of <em>P. calceolata </em>(RCC100 and RCC697) were cultivated and their RNAs reads were aligned on the predicted genes of <em>P. calceolata</em> RCC100 genome: <a href="https://www.ncbi.nlm.nih.gov/Traces/wgs/CAKKNE01?display=download" rel="nofollow">https://www.ncbi.nlm.nih.gov/Traces/wgs/CAKKNE01?display=download</a></p> <p>The following culture conditions were analysed :</p> <p>882 µM of Nitrate (RCC100 and RCC697) </p> <p>441 µM of Nitrate (RCC100)</p> <p>220 µM of Nitrate (RCC100 and RCC697)</p> <p>50 µM of Nitrate (RCC697)</p> <p>882 µM Cyanate (RCC100)</p> <p>882 µM Ammonia (RCC100)</p> <p>441 µM Urea (RCC100)</p> <p><a href="../api/records/12582059/draft/files/20230427_RCC100-Nitrate_transcriptomes_rawcounts.tsv/content" target="_blank" rel="noopener noreferrer">20230427_RCC100-Nitrate_transcriptomes_rawcounts.tsv</a> : the file contains the raw read counts of RCC100 in 6 culture conditions in triplicate + the gene names = 19 columns.</p> <p><a href="../api/records/12582059/draft/files/20230427_RCC100-Nitrate_transcriptomes_TPM.tsv/content" target="_blank" rel="noopener noreferrer">20230427_RCC100-Nitrate_transcriptomes_TPM.tsv</a> : same data normalized in transcript per kb per million mapped reads (TPM).</p> <p><a href="../api/records/12582059/draft/files/20230427_RCC100-Nitrate_transcriptomes_rawcounts.tsv/content" target="_blank" rel="noopener noreferrer">20230427_RCC697-Nitrate_transcriptomes_rawcounts.tsv</a> : the file contains the raw read counts of RCC697 of 3 culture conditions in triplicate + the gene names = 10 columns.</p> <p><a href="../api/records/12582059/draft/files/20230427_RCC100-Nitrate_transcriptomes_TPM.tsv/content" target="_blank" rel="noopener noreferrer">20230427_RCC697-Nitrate_transcriptomes_TPM.tsv</a> : same data normalized in transcript per kb per million mapped reads (TPM).</p> <p><span>Differential expression analysis (DESeq2) was performed by pairwise comparisons between the standard condition (882 µM nitrate) and low-nitrate conditions (50, 220 or 441 µM nitrate) or changing nitrogen sources (882 µM ammonium, 882 µM cyanate and 441 µM urea). Each DESeq-results_RCCxxx_xxx.tsv file contains 6 columns : <em>P.calceolata </em>gene name, base Mean, log2 Fold Change, standard error value (lfcSE), pvalue and adjusted pvalue (padj).<br></span></p>
Extended data tables to Haering and Habermann, F1000Res, RNfuzzyApp: an R shiny RNA-seq data analysis app for visualisation, differential expression analysis, time-series clustering and enrichment analysis
<p><b>Background</b> </p> <p>RNA-seq is a widely adopted affordable method for large scale gene expression profiling. However, user-friendly and versatile tools for wet-lab biologists to analyse RNA-seq data beyond standard analyses such as differential expression, are rare. Especially, the analysis of time-series data is difficult for wet-lab biologists lacking advanced computational training. Furthermore, most meta-analysis tools are tailored for model organisms and not easily adaptable to other species.</p> <p><b>Results</b></p> <p>With RNfuzzyApp, we provide a user-friendly, web-based R-shiny app for differential expression analysis, as well as time-series analysis of RNA-seq data. RNfuzzyApp offers several methods for normalization and differential expression analysis of RNA-seq data, providing easy-to-use toolboxes, interactive plots and downloadable results. For time-series analysis, RNfuzzyApp presents the first web-based, automated pipeline for soft clustering with the Mfuzz R package, including methods to aid in cluster number selection, Mfuzz loop computations, cluster overlap analysis, as well as cluster enrichments.</p> <p><b>Conclusion</b></p> <p>RNfuzzyApp is an intuitive, easy to use and interactive R shiny app for RNA-seq differential expression and time-series analysis, offering a rich selection of interactive plots, providing a quick overview of raw data and generating rapid analysis results. Furthermore, its orthology assignment, enrichment analysis, as well as ID conversion functions are accessible to non-model organisms.</p>
RNAseq data: Analysis of circRNA expression in human neuronal differentiation
<p>This dataset contains sequencing read count data related to samples from differentiating human neuroepithelial stem cells (NES) collected at days zero (NES), five (D5) and 28 (D28) of differentiation. Details on how samples were collected and how data was generated and processed are described below.</p> <p> </p> <p><em>Sample preparation</em></p> <p>NES were seeded on tissue culture flasks coated with 20 μg/ml poly-L-ornithine (Sigma-Aldrich P3655), and 1 μg/ml laminin (Sigma-Aldrich L2020). Cells were grown in DMEM/F12+GlutaMAX medium (ThermoFisher 31331093) supplemented with 0.05X B27 (ThermoFisher 17504044), 1X N2 (ThermoFisher 17502001), 10 ng/ml bFGF (fisher scientific CTP0261), 10 ng/ml EGF (PeproTech AF-100-15) and 10 U/ml penicillin/streptomycin (ThermoFisher 15140122). Medium was exchanged 50% daily and cells maintained in 5% CO2 at 37ºC, passaging once 100% confluent and seeding at a density of 5x104 cells/cm2. Neural differentiation was induced by growth factor withdrawal the day after plating with media B27 concentration increased to 0.5X. Media was exchanged 50% every second day up until D15, after which media was supplemented with 0.4 ug/ml laminin and exchanged 50% every three days. </p> <p> </p> <p><em>RNA extraction and sequencing</em></p> <p>Cells were lysed in TRIzol reagent (ThermoFischer 15596026) before separating with chloroform and mixing the aqueous phase with isopropanol as per manufacturer directions. RNA was then isolated from the isopropanol/chloroform solution using the ReliaPrep RNA Cell Miniprep kit (Promega Z6010). Libraries were prepared with Illumina Truseq Stranded total RNA RiboZero GOLD kit and sequenced on the NovaSeq6000 platform with a 2x151 setup using NovaSeqXp workflow in S4 mode flowcell.</p> <p> </p> <p><em>Data generation</em></p> <p>Raw reads were processed using cutadapt v3.2 to trim adaptor sequences and low-quality base pairs and discard short reads (options: -m 20 -e 0.1 -q 20 -O 1). The GRCh37 genome assembly was used for all alignment, annotation, and downstream analysis steps. Trimmed read weres alignment to the GRCh37 genome assembly using TopHat v2.0.9 tophat_fusion (with Bowtie v1.1.2 and Samtools v0.1.19) with –fusion-min-dist 200. BAM files have been anonymised by removal of potentially identifiable genetic variant information using BAMboozle v0.5.0 (Ziegenhain & Sandberg, 2021) with default settings. This BAM files and corresponding index (.bai) files are provided here with naming convention "<em>label.</em>bam" Information on sample labels and corresponding conditions is provided in the file 'metadata.txt'.</p> <p><br> </p>
Results of the differential gene expression analysis in SIV infection in Chlorocebus sabaeus and Macaca mulatta
<p>Results of differential gene expression analysis in SIV infection in Chlorocebus sabaeus and Macaca mulatta.</p> <p>From the transcriptome data repository MACE (http://mace.ihes.fr)</p>
The raw microarray data and the differential expression analysis results from "Manipulating the growth environment through co-culture to enhance stress tolerance and viability of probiotic strains in the gastrointestinal tract".
<p>The signal data for each spot were subsequently quantified by using Feature Extraction software (Agilent Technologies).M1.txt to M5.txt: monoculture; C1.txt to C5.txt: co-culture; P1.txt to P5.txt: pH-controlled monoculture. The differential expression analysis results were obtained by using limma.</p>
Extended data tables to Haering and Habermann, F1000Res, RNfuzzyApp: an R shiny RNA-seq data analysis app for visualisation, differential expression analysis, time-series clustering and enrichment analysis
Open the record for dataset details and reuse information.
Replicated differential expression analysis in a green-brown polymorphic grasshopper reveals role of beta-carotene-binding protein in body coloration
Open the record for dataset details and reuse information.
Processed datasets and codes for differential expression analysis on polulation-level RNA-seq data
<p>This version includes codes and data necessary to reproduce all results in our response to the correspondences ("Response to 'Neglecting normalization impact in semi‑synthetic RNA‑seq data simulation generates artificial false positives' and 'Winsorization greatly reduces false positives by popular differential expression methods when analyzing human population samples'") (<a href="https://doi.org/10.1186/s13059-024-03232-8">https://doi.org/10.1186/s13059-024-03232-8</a>).</p> <p>It also includes a README file to guide the reproduction of the results in our original publication and resources for the goodness of fit test in the original publication, "Exaggerated False Positives by Popular Differential Expression Methods When Analyzing Human Population Samples" (<a href="https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02648-4">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02648-4</a>).</p>
Differential expression analysis to aluminum toxicity in Citrus x limonia Osbeck
<p>Here we deliver the Differential gene expression on genes response to aluminum toxicity in <em>Citrus</em> x<em> limonia. </em>Root apices of ‘Mandarin’ lime plants grown for 60 days in nutrient solutions either with 1480 mM Al<sup>3+</sup> or 0 mM Al<sup>3+</sup> were analyzed by RNA-seq.</p> <p>Clean reads were mapped to the sweet orange (<em>Citrus sinensis</em>) genome (Xu et al. 2013). Gene expression levels were calculated by CPM (Counts per million) reads. We used HTSeq ver. 0.6.1 (Anders et al. 2015) CPM estimation. The differentially expressed genes (DEGs) here reported by NOIseq ver. 2.16.0 (Tarazona et al. 2016). </p> <p> </p> <p><strong>Results:</strong></p> <p><br> Number of differentially expressed (DE) features (Probability > 0.7): 3,351</p> <p>Up-regulated (M > 0): 1,664<br> Down-regulated (M < 0): 1,687</p> <p>All software were run on OmicsBox interface.</p> <p><strong>References:</strong></p> <p>Anders S., Pyl PT. and Huber W. (2015). HTSeq--a Python framework to work with high-throughput sequencing data. Bioinformatics (Oxford, England), 31(2), 166-9.</p> <p>OmicsBox - Bioinformatics made easy. BioBam Bioinformatics (Version 2.0.36). March 3, 2019. www.biobam.com/omicsbox.</p> <p>Tarazona S., Furio-Tari P., Turra D., Pietro AD., Nueda MJ., Ferrer A. and Conesa A. (2015). Data quality aware analysis of differential expression in RNA-seq with NOISeq R/Bioc package. Nucleic acids research, 43(21), e140.</p> <p>Xu Q, Chen L-L, Ruan X, et al (2013) The draft genome of sweet orange (Citrus sinensis). Nat Genet 45:59–66.</p> <p> </p> <p>Legend: </p> <p>Regulation - UP or DOWN = differentially expressed genes, UPregulated or DOWNregulated</p> <p>Citrus_40_Al_2 - Normalized CPM for root apexes under 1480 mM Al<sup>3+</sup> </p> <p>Citrus_0_Al_1 - Normalized CPM for root apexes under 0 mM Al<sup>3+</sup></p>
Single cell multiomic analysis identifies key genes differentially expressed in innate lymphoid cells from COVID-19 patients
<p>Innate lymphoid cells (ILCs) are enriched at mucosal surfaces where they respond rapidly to environmental stimuli and contribute to both tissue inflammation and healing. To gain insight into the role of ILCs in the pathology and recovery from COVID-19 infection, we employed a multi-omic approach consisting of Abseq and targeted mRNA sequencing to respectively probe the surface marker expression, transcriptional profile and heterogeneity of ILCs in peripheral blood of patients with COVID-19 compared with healthy controls. We found that the frequency of ILC1 and ILC2 cells was significantly increased in COVID-19 patients. Moreover, all ILC subsets displayed a significantly higher frequency of CD69-expressing cells, indicating a heightened state of activation. ILC2s from COVID-19 patients had the highest number of significantly differentially expressed (DE) genes. The most notable genes DE in COVID-19 vs healthy participants included a) genes associated with responses to virus infections and b) genes that support ILC self-proliferation, activation and homeostasis. In addition, differential gene regulatory network analysis revealed ILC-specific regulons and their interactions driving the differential gene expression in each ILC. Overall, this study provides mechanistic insights into the characteristics of ILC subsets activated during COVID-19 infection.</p>
Table with results of differential gene expression analysis in the controlled environment for fin tissues.
<p>Differences between transcriptomes of three Cottus fish lineages were assessed under controlled, laboratory conditions. Two tissues were investigated: fins and livers. Present table shows results of the differential gene expression analysis performed on fin tissues of Cottus fish. Base-mean, Log-2-fold change. standard error, statistics and associated p-values and FDR-corrected p-values are given for every contrast possible in our experimental design.</p>
Table with results of differential gene expression analysis in the controlled environment for liver tissues.
<p>Differences between transcriptomes of three Cottus fish lineages were assessed under controlled, laboratory conditions. Two tissues were investigated: fins and livers. Present table shows results of the differential gene expression analysis performed on liver tissues of Cottus fish. Base-mean, Log-2-fold change. standard error, statistics and associated p-values and FDR-corrected p-values are given for every contrast possible in our experimental design.</p>
RNA-Seq analysis to identify differentially expressed genes in top and bottom leaves under Alternaria brassicicola infection
<p>The broccoli plants were infected with Alternaria brassicicola and RNA samples were extracted for control and inoculated plants at 10 days post inoculation. </p>
curatedPCaData supplementary data table for differential gene expression analysis
<p>This tab-separated plaintext file contains differential gene expression analyses reported for the curatedPCaData data resource publication.</p>
Single cell multiomic analysis identifies key genes differentially expressed in innate lymphoid cells from COVID-19 patients
Open the record for dataset details and reuse information.
Data from: Differential gene expression analysis of symbiotic and aposymbiotic Exaiptasia anemones under immune challenge with Vibrio coralliilyticus
Anthozoans are a class of Cnidarians that includes scleractinian corals, anemones and their relatives. Despite a global rise in disease epizootics impacting scleractinian corals, little is known about the immune response of this key group of invertebrates. To better characterize the anthozoan immune response, we used the model anemone Exaiptasia pallida to explore the genetic links between the anthozoan-algal symbioses and immunity in a two-factor RNA-Seq experiment using both symbiotic and aposymbiotic(menthol-bleached) Exaiptasia pallida exposed to the bacterial pathogen Vibrio coralliilyticus. Multivariate and univariate analyses of Exaiptasia gene expression demonstrated that exposure to live Vibrio coralliilyticus had strong and significant impacts on transcriptome-wide gene expression for both symbiotic and aposymbiotic anemones, but we did not observe strong interactions between symbiotic state and Vibrio exposure. There were 4,164 significantly differentially expressed (DE) genes for Vibrio exposure, 1,114 DE genes for aposymbiosis, and 472 DE genes for the additive combinations of Vibrio and aposymbiosis. KEGG enrichment analyses identified 11 pathways - involved in immunity (5), transport and catabolism (4) and cell growth and death (2) - that were enriched due to both Vibrio and/or aposymbiosis. Immune pathways showing strongest differential expression included complement, coagulation, nucleotide-binding and oligomerization domain (NOD), and Toll for Vibrio exposure and coagulation and apoptosis for aposymbiosis.
Comparative analysis of statistical methods used for detecting differential expression in label-free mass spectrometry proteomics - Data Supplement
<p>This the is Data Supplement for the article "Comparative analysis of statistical methods used for detecting differential expression in label-free mass spectrometry proteomics" submitted to the Journal of Proteomics 2015.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.