Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,614
datasets available to search
ShareScore release 0.9.0
Dataset results
2,614 results for “RNA-seq analysis”
Mammary single-cell RNA-seq analysis and prostate cancer survival as a function of H2AFJ expression for the paper entitled: The histone variant H2A.J is enriched in luminal epithelial cells
<p>H2A.J is a poorly studied mammalian-specific variant of histone H2A. We used immunohistochemistry to study its localization in various human and mouse tissues. H2A.J showed cell-type specific expression with a striking enrichment in luminal epithelial cells of multiple glands including those of breast, prostate, pancreas, thyroid, stomach, and salivary glands. H2A.J was also highly expressed in many carcinoma cell lines and in particular, those derived from luminal breast and prostate cancer. H2A.J thus appears to be a novel marker for luminal epithelial cancers. Knocking-out the H2AFJ gene in T47D luminal breast cancer cells reduced the expression of several estrogen-responsive genes which may explain its putative tumorigenic role in luminal-B breast cancer.</p>
Datasets, reproducible codes, and results for evaluating differential expression analysis methods on population-level RNA-seq data
<p>This upload contains the necessary R codes and data to reproduce the FDR and Power results described in our correspondence "Neglecting normalization impact in semi-synthetic RNA-seq data simulation generates artificial false positives" to Li Y, Ge X, Peng F, Li W, Li JJ, Exaggerated false positives by popular differential expression methods when analyzing human population samples, <em>Genome Biology</em> 23, 79, 2022, DOI: 10.1186/s13059-022-02648-4.</p>
Robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data
<p>Data used to test the robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data, described in <a href="https://doi.org/10.1186/s13059-020-1949-z">Holland et al. 2020</a>.</p> <p>The folder <em>data </em>contains<em> </em>raw data and the folder <em>output</em> contains intermediate and final results of all analyses. </p> <p>The associated analyses code and more information are available on <a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq">GitHub</a>.</p> <p> </p> <p><strong>Abstract</strong></p> <p><strong>Background</strong></p> <p>Many functional analysis tools have been developed to extract functional and mechanistic insight from bulk transcriptome data. With the advent of single-cell RNA sequencing (scRNA-seq), it is in principle possible to do such an analysis for single cells. However, scRNA-seq data has characteristics such as drop-out events and low library sizes. It is thus not clear if functional TF and pathway analysis tools established for bulk sequencing can be applied to scRNA-seq in a meaningful way.</p> <p><strong>Results</strong></p> <p>To address this question, we perform benchmark studies on simulated and real scRNA-seq data. We include the bulk-RNA tools PROGENy, GO enrichment, and DoRothEA that estimate pathway and transcription factor (TF) activities, respectively, and compare them against the tools SCENIC/AUCell and metaVIPER, designed for scRNA-seq. For the in silico study, we simulate single cells from TF/pathway perturbation bulk RNA-seq experiments. We complement the simulated data with real scRNA-seq data upon CRISPR-mediated knock-out. Our benchmarks on simulated and real data reveal comparable performance to the original bulk data. Additionally, we show that the TF and pathway activities preserve cell type-specific variability by analyzing a mixture sample sequenced with 13 scRNA-seq protocols. We also provide the benchmark data for further use by the community.</p> <p><strong>Conclusions</strong></p> <p>Our analyses suggest that bulk-based functional analysis tools that use manually curated footprint gene sets can be applied to scRNA-seq data, partially outperforming dedicated single-cell tools. Furthermore, we find that the performance of functional analysis tools is more sensitive to the gene sets than to the statistic used.</p> <p> </p> <p>For questions related to the data please write an email to christian.holland@bioquant.uni-heidelberg.de or use the <a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq/issues">GitHub issue system</a>.</p>
Data for "Tuning parameters of dimensionality reduction methods for single-cell RNA-seq analysis"
<p>The files named <code>df_scran.csv</code>, <code>df_seurat.csv</code>, <code>df_zinbwave.csv</code>, <code>df_dca.csv</code>, and <code>df_scvi.csv</code> contain one row per configuration that we ran successfully.</p> <p>The files named <code>DATASET.METHOD.h5ad</code> are encoded with anndata <code>v0.7.0</code> (be careful as they are not readable with previous versions) and contain 100 embeddings each. The embeddings are in the <code>obsm</code> attribute of the object. All the embeddings can be listed with the <code>obsm_keys()</code> method. The name of the embedding contains the parameters used to generate that embedding and are written like that <code>method=zinbwave.dims=10.epsilon=1000.features=300.gene_covariate=0</code>.</p> <p> </p> <p>For questions on this dataset please contact fraimundo@google.com</p>
Training material for small RNA-seq data analysis (Galaxy Training Network tutorial)
<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes small RNA-seq (sRNA-seq) data from a study published by Harrington et al. (DOI:10.1186/s12864-017-3692-8) to detect differential abundance of various classes of endogenous short interfering RNAs (esiRNAs). The goal of this study was to investigate "connections between differential retroTn and hp-derived esiRNA processing and cellular location, and to investigate the potential link between mRNA 3’ end cleavage and esiRNA biogenesis." To this end, sRNA-seq libraries were constructed from triplicate <em>Drosophila</em> tissue culture samples under conditions of either control RNAi or RNAi knockdown of a factor involved in mRNA 3’ end processing, <em>Symplekin</em>. This dataset (GEO Accession: GSE82128) consists of single-end, size-selected, non-rRNA-depleted sRNA-seq libraries. Because of the long processing time for the large original files, we have downsampled the original raw data files to include only reads that align to a subset of interesting transcript features including: (1) transposable elements, (2) <em>Drosophila</em> piRNA clusters, (3) <em>Symplekin</em>, and (4) genes encoding mass spectrometry-defined protein binding partners of <em>Symplekin</em> from Additional File 2 in the indicated paper by Harrington et al. More details on features 1 and 2 can be found here: https://github.com/bowhan/piPipes/blob/master/common/dm3/genomic_features (piRNA_Cluster, Trn). All features are from the <em>Drosophila</em> genome Apr. 2006 (BDGP R5/<em>dm3</em>) release.</p>
Data for manuscript "rMATS-turbo: an efficient and flexible computational tool for alternative splicing analysis of large-scale RNA-seq data"
<p>Output files generated by rMATS-turbo for the two example datasets described in the manuscript titled "rMATS-turbo: an efficient and flexible computational tool for alternative splicing analysis of large-scale RNA-seq data".</p> <table> <tbody> <tr> <td>File</td> <td>Description</td> <td>Cell lines</td> <td>BioProject</td> </tr> <tr> <td>PC3E-GS689.tar.gz</td> <td>Compressed folder containing all 36 rMATS-turbo output files for Example 1 described in the manuscript</td> <td>PC3E and GS689 cell lines</td> <td>PRJNA438990</td> </tr> <tr> <td>CCLE.tar.gz</td> <td>Compressed folder containing all 36 rMATS-turbo output files for Example 2 described in the manuscript</td> <td>1,019 CCLE human cancer cell lines</td> <td>PRJNA523380</td> </tr> </tbody> </table> <p>A detailed description of the output files is available in the manuscript and the rMATS-turbo software GitHub repository (https://github.com/Xinglab/rmats-turbo).</p>
Extended data tables to Haering and Habermann, F1000Res, RNfuzzyApp: an R shiny RNA-seq data analysis app for visualisation, differential expression analysis, time-series clustering and enrichment analysis
<p><b>Background</b> </p> <p>RNA-seq is a widely adopted affordable method for large scale gene expression profiling. However, user-friendly and versatile tools for wet-lab biologists to analyse RNA-seq data beyond standard analyses such as differential expression, are rare. Especially, the analysis of time-series data is difficult for wet-lab biologists lacking advanced computational training. Furthermore, most meta-analysis tools are tailored for model organisms and not easily adaptable to other species.</p> <p><b>Results</b></p> <p>With RNfuzzyApp, we provide a user-friendly, web-based R-shiny app for differential expression analysis, as well as time-series analysis of RNA-seq data. RNfuzzyApp offers several methods for normalization and differential expression analysis of RNA-seq data, providing easy-to-use toolboxes, interactive plots and downloadable results. For time-series analysis, RNfuzzyApp presents the first web-based, automated pipeline for soft clustering with the Mfuzz R package, including methods to aid in cluster number selection, Mfuzz loop computations, cluster overlap analysis, as well as cluster enrichments.</p> <p><b>Conclusion</b></p> <p>RNfuzzyApp is an intuitive, easy to use and interactive R shiny app for RNA-seq differential expression and time-series analysis, offering a rich selection of interactive plots, providing a quick overview of raw data and generating rapid analysis results. Furthermore, its orthology assignment, enrichment analysis, as well as ID conversion functions are accessible to non-model organisms.</p>
RNA-seq based analysis of gene expression in thyroids of wild-type and Keap1 knockdown mice after exposure to excess iodide.
<p>C57BL/6J Keap1flox/flox mice were developed by Prof. Masayuki Yamamoto (DOI: 10.1016/j.bbrc.2005.10.185). These mice express lower levels of Keap1 because of the loxP site insertions (DOI: 10.1128/MCB.01591-09) and are designated as Keap1KD. 3-4 months old WT and Keap1KD mice fed a standard diet (KLIBA NAFAG 3242, Switzerland) containing 1.5 mg/kg sodium iodine were given normal tap water with or without 0.05% sodium iodide (Sigma, St Louis, MO, USA) for 7 days. Hence, the following groups of mice were included:</p> <p>WT mice on control water (n=8)-designated as wt in the file</p> <p>WT mice exposed to excess iodide (WT-IOD, n=6)- designated as wti in the file</p> <p>Keap1KD mice on control water (n=8)- designated as kp in the file</p> <p> and Keap1KD mice exposed to excess iodide (Keap1KD-IOD, n=6)- designated as kpI in the file. </p> <p> Mice were maintained in the animal facility of the Department of Physiology at the University of Lausanne in temperature-, light-, and humidity-controlled rooms with a 12-hour light/dark cycle. All animal procedures were in accordance with Swiss legislature and the study was approved by the Canton of Vaud SCAV.</p> <p>RNA from individual mouse thyroids was prepared as previously described (DOI: 10.1007/978-1-4939-3756-1_25) and was submitted to Alithea Genomics (Switzerland). The bulk RNA barcoding and sequencing (BRB-seq) libraries were generated and sequenced as described previously (DOI: 10.1186/s13059-019-1671-x) to a depth of approximately 1.2 million raw reads per sample.</p> <p>This research was funded by the Swiss National Science Foundation Research Grants 310030_212558, IZCOZ0_205415 and IZCOZ0_177070</p> <p></p> <p></p>
Extended data tables to Haering and Habermann, F1000Res, RNfuzzyApp: an R shiny RNA-seq data analysis app for visualisation, differential expression analysis, time-series clustering and enrichment analysis
Open the record for dataset details and reuse information.
Data from: Single cell RNA-seq analysis reveals that prenatal arsenic exposure results in long-term, adverse effects on immune gene expression in response to Influenza A infection
<p>Arsenic exposure via drinking water is a serious environmental health concern. Epidemiological studies suggest a strong association between prenatal<i> </i>arsenic exposure and subsequent childhood respiratory infections, as well as morbidity from respiratory diseases in adulthood, long after systemic clearance of arsenic.<i> </i>We investigated the impact of exclusive prenatal arsenic exposure on the inflammatory immune response and respiratory health after an adult influenza A (IAV) lung infection. C57BL/6J mice were exposed to 100 ppb sodium arsenite<i> in utero,</i> and subsequently infected with IAV (H1N1) after maturation to adulthood. Assessment of lung tissue and bronchoalveolar lavage fluid (BALF) at various time points post IAV infection reveals greater lung damage and inflammation in arsenic exposed mice versus control mice. Single-cell RNA sequencing analysis of immune cells harvested from IAV infected lungs suggests that the enhanced inflammatory response is mediated by dysregulation of innate immune function of monocyte derived macrophages, neutrophils, NK cells, and alveolar macrophages. Our results suggest that prenatal arsenic exposure results in lasting effects on the adult host innate immune response to IAV infection, long after exposure to arsenic, leading to greater immunopathology. This study provides the first direct evidence that exclusive prenatal exposure to arsenic in drinking water causes predisposition to a hyperinflammatory response to IAV infection in adult mice, which is associated with significant lung damage.</p>
Sample dataset from RNA-Seq analysis of Ostreococcus tauri
<p>Sample dataset from:</p> <p>Lelandais et al. "<em>Ostreococcus tauri</em> is a new model green alga for studying iron metabolism in eukaryotic phytoplankton", <em>BMC Genomics</em> (2016).<br> DOI <a href="https://bmcgenomics.biomedcentral.com/articles/10.1186/s12864-016-2666-6">10.1186/s12864-016-2666-6</a></p>
Processed datasets and codes for differential expression analysis on polulation-level RNA-seq data
<p>This version includes codes and data necessary to reproduce all results in our response to the correspondences ("Response to 'Neglecting normalization impact in semi‑synthetic RNA‑seq data simulation generates artificial false positives' and 'Winsorization greatly reduces false positives by popular differential expression methods when analyzing human population samples'") (<a href="https://doi.org/10.1186/s13059-024-03232-8">https://doi.org/10.1186/s13059-024-03232-8</a>).</p> <p>It also includes a README file to guide the reproduction of the results in our original publication and resources for the goodness of fit test in the original publication, "Exaggerated False Positives by Popular Differential Expression Methods When Analyzing Human Population Samples" (<a href="https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02648-4">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02648-4</a>).</p>
Diamond hits: Another lesson from unmapped reads – in depth analysis of RNA-Seq reads from various horse tissues
<p>Diamond hits - predicted peptides were compared against NCBI RefSeq protein (NR) databases using Diamond software</p> <p>Data: loin adipose - AD, hoof lamina - LM, liver - LI, longissimus muscle - LO, left lung - LU, heart left ventricle - LV, ovary- OV and parietal cortex – PC<br> 1 – horse ECA_UCD_AH1<br> 2 – horse ECA_UCD_AH2</p>
Diamond hits_1: Another lesson from unmapped reads – in depth analysis of RNA-Seq reads from various horse tissues
<p>Diamond hits - de novo assembled transcripts were compared against NCBI RefSeq Nucleotide (NT) using Diamond software</p> <p> </p> <p>Data: loin adipose - AD, hoof lamina - LM, liver - LI, longissimus muscle - LO, left lung - LU, heart left ventricle - LV, ovary- OV and parietal cortex – PC<br> 1 – horse ECA_UCD_AH1<br> 2 – horse ECA_UCD_AH2</p> <p> </p>
Unmapped reds: Another lesson from unmapped reads – in depth analysis of RNA-Seq reads from various horse tissues
<p>Unmapped reds – all RNA-Seq reads that did not map to the reference EquCab3.0 genome using Tophat2 software</p> <p>Data: loin adipose - AD, hoof lamina - LM, liver - LI, longissimus muscle - LO, left lung - LU, heart left ventricle - LV, ovary- OV and parietal cortex – PC<br> 1 – horse ECA_UCD_AH1<br> 2 – horse ECA_UCD_AH2</p>
Predicted peptides: Another lesson from unmapped reads – in depth analysis of RNA-Seq reads from various horse tissues
<p>Predicted peptides – peptides predicted based on the assembled transcripts using TransDecoder software</p> <p> </p> <p>Data: loin adipose - AD, hoof lamina - LM, liver - LI, longissimus muscle - LO, left lung - LU, heart left ventricle - LV, ovary- OV and parietal cortex – PC<br> 1 – horse ECA_UCD_AH1<br> 2 – horse ECA_UCD_AH2</p> <p> </p>
Assembled transcripts: Another lesson from unmapped reads – in depth analysis of RNA-Seq reads from various horse tissues
<p>Assembled transcripts – putative transcripts de novo assembled with Trinity software</p> <p> </p> <p>Data: loin adipose - AD, hoof lamina - LM, liver - LI, longissimus muscle - LO, left lung - LU, heart left ventricle - LV, ovary- OV and parietal cortex – PC<br> 1 – horse ECA_UCD_AH1<br> 2 – horse ECA_UCD_AH2</p> <p> </p>
Simulated RNA-seq data for differential splicing analysis with covariates
<p>The repository includes alignments of simulated RNA-seq data for evaluating differential splicing detection with covariates. Starting from an empirical transcript expression matrix trained on an RNA-seq data set from lung fibroblasts (GenBank A# SRR493366) and using GENCODE v.41 as reference, 11.5 million 100 bp long paired-end reads were generated per sample, from 2,000 genes with two or more expressed isoforms. RNA-seq data was simulated for one ‘condition’, with values ‘control’, ‘disease’ and ‘stage2’, with one covariate, ‘biological sex’, with values ‘M’ and ‘F’. 5 samples each were simulated for each (condition x sex) category. Changes were simulated in the expression (DE) and/or the splicing ratio (DS) of genes as follows. Changes in expression (DE) were simulated by either halving or doubling the expression level of the gene. Changes in splicing ratios (DS) were simulated by swapping the expression levels of the gene’s top two transcript isoforms. All RNA-seq data was mapped to the hg38 genome with the spliced alignment tool STAR v2.7.10a.</p> <p> <em><u>Pairwise comparison alignment set</u></em>: Differences due to ‘condition’ between two states, ‘control’ and ‘disease’, were simulated at 600 genes, including 200 DE, 200 DS and 200 DE+DS genes. Differences in ‘biological sex’ (covariate) were represented as changes in 300 genes, including 100 from each of the DS, DE and DE+DS categories. Hence, the target gene set for differential splicing ratio (DSR)<em> pairwise comparisons </em>consists of the pooled 200 DS and 200 DS+DE genes differentially spliced between the ‘control’ and ‘disease’ states, while for differential splicing abundance (DSA)<em> pairwise comparisons </em>the target gene set is the set of 600 modified genes, 200 in each of the DS, DE and DS+DE categories.</p> <p> <em><u>Multiway (3-way) comparison alignment set:</u></em> Additional changes between ‘disease’ and ‘stage2’ were made to 100 of the previously modified genes, as well as to a set of 200 additional genes not encountered previously, for each of the categories DE, DS and DE+DS. Therefore, for <em>DSR three-way comparisons</em>, the target gene set represents the 800 genes simulated as being DS or DE+DS between any of the ‘control’, ‘disease’ and ’stage2’ categories, while for the <em>multi-way DSA comparisons</em> the target is the full set of 1,200 genes (400 DE, 400 DS and 400 DE+DS) simulated to have changed between any of the 'control', 'disease’ and ‘stage2’ states.</p> <p> <em><u>Further details:</u></em> See the ‘key’ directories in each package for the gene lists.</p>
Recovery and analysis of transcriptome subsets from pooled single-cell RNA-seq libraries
<p>Processed data files for manuscript: "Recovery and analysis of transcriptome subsets from pooled single-cell RNA-seq libraries" <a href="https://doi.org/10.1093/nar/gky1204">https://doi.org/10.1093/nar/gky1204</a> . Scripts for generating figures are found here: https://github.com/rnabioco/scrna-subsets</p>
Training material for analysis small RNA-seq data (Galaxy Training Network tutorial)
<p>The data provided here is part of the Galaxy Training Network tutorial for analysis of small RNA-seq (sRNA-seq) data using mirdeep2 and miranda. This dataset is provided by INRA (Le Rheu, France).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.