Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

21,236

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

21,236 results for “RNA-seq”

Learn how ShareScore rates datasets ↗
zenodo48/100

bollito: a flexible pipeline for comprehensive single-cell RNA-seq analyses - Melanoma tutorial

<p>Downsampled version of the melanoma dataset originally published by&nbsp;<em><a href="https://genome.cshlp.org/content/28/9/1353">Ho et al </a>(1)</em>. The&nbsp;dataset is composed by cells from the 451Lu cell line. There&nbsp;are two samples available:</p> <table> <tbody> <tr> <td><strong>Name</strong></td> <td><strong>Description</strong></td> <td><strong>R1/R2</strong></td> </tr> <tr> <td>451LU</td> <td>Parental cell line</td> <td>2500K_451LU_L003_R*_001.fastq.gz</td> </tr> <tr> <td>451LUBR3</td> <td>Vemurafenib-resistant sample treated with targeted BRAF inhibitors</td> <td>500K_451LUBR3_L004_R*_001.fastq.gz</td> </tr> </tbody> </table> <p><br> (1)&nbsp;Ho YJ, Anaparthy N, Molik D, et al. Single-cell RNA-seq analysis identifies markers of resistance to targeted BRAF inhibitors in melanoma cell populations.&nbsp;<em>Genome Res</em>. 2018;28(9):1353-1363. doi:10.1101/gr.234062.117</p>

opencc-by-4.0Feb 2022View details →
zenodo48/100

RNA-seq Dataset from Crowley et. al. 2015

<p>This dataset was uploaded to this&nbsp;repository&nbsp;by Joshua P. Zitovsky and Michael I. Love&nbsp;with permission from the original authors, due to the fact that the original dataset is not currently hosted in a stable repository.&nbsp;If this dataset is used in original research leading to published work, please cite the&nbsp;Crowley et. al. 2015 paper.</p> <p>This repository&nbsp;contains an RNA-seq dataset based on the allelic expression study by Crowley et. al. (2015). The study took mice from three divergent inbred strains (CAST/EiJ, PWK/PhJ and WSB/EiJ) and performed a diallel cross. The data set contains allele-specific expression (ASE) counts for 72 mice and 23,297 genes in the resulting cross, with 12 mice of each possible parent combination, and an equal number of males and females within each parent combination. Sequencing was performed with the Illumina HiSeq 2000 platform to generate 100-bp paired-end reads and following the TruSeq RNA Sample Preparation v2 protocol.</p>

opencc-by-4.0Sep 2019View details →
zenodo44/100

Simulated RNA-seq data

<p>Simulated RNA-seq data shows that histograms from p value sets with around one hundred&nbsp;&nbsp;true effects out of 20,000 features can be classified as &#39;uniform&#39;.&nbsp;RNA-seq data was simulated with polyester R package <a href="https://doi.org/10.1093/bioinformatics/btv272">(Frazee, 2015)</a> on 20,000 transcripts from human transcriptome&nbsp;using grid of 3, 6, and 10 replicates and 100, 200, 400, and 800 effects for two groups.&nbsp;Fold changes were set to 0.5 and 2.&nbsp;Differential expression was assessed using DESeq2 R package <a href="https://doi.org/10.1186/s13059-014-0550-8">(Love, 2014)</a> using default settings&nbsp;and group 1 versus group 2 contrast.&nbsp;Effects denotes in facet labels the number of true effects and N denotes number of replicates.&nbsp;Red line denotes QC threshold used for dividing p histograms into discrete classes.&nbsp;Workflow and code used to run this simulation is available on <a href="https://github.com/rstats-tartu/simulate-rnaseq">rstats-tartu/simulate-rnaseq</a>.</p> <p>&nbsp;</p> <p>Files</p> <ul> <li>de_simulation_results.csv -- merged and processed DE analysis results of simulated data.</li> <li>simulate-reads-2021-01-25.tar.gz -- raw DE analysis results&nbsp;on 20,000 transcripts from human transcriptome&nbsp;using grid of 3, 6, and 10 replicates and 100, 200, 400, and 800 effects for two groups.&nbsp;Fold changes were set to 0.5, 1, and 2.&nbsp;Differential expression was assessed using DESeq2 with default settings.</li> <li>simulate-rnaseq.tar.gz -- snakemake workflow and input fasta file&nbsp;to simulate RNA-seq data with polyester and analyse results with DESeq2. Adjust settings in config.yaml to customise simulation. Includes software to run workflow on Linux, given that <a href="https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh">Conda</a> and <a href="https://snakemake.readthedocs.io/en/stable/index.html">snakemake</a> are installed.</li> </ul> <p>The simulate-rnaseq.tar.gz&nbsp;archive can be re-executed on a vanilla machine that only has Conda and Snakemake installed via:</p> <pre><code class="language-bash">tar -xf simulate-rnaseq.tar.gz snakemake --use-conda -n</code></pre> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Training material for de novo transcriptome reconstruction from RNA-seq data

<p>The data provided here are part of a Galaxy tutorial that analyzes RNA-seq data from a study published by Wu et al., 2014 (DOI:10.1101/gr.164830.113). The goal of this study was to investigate "the dynamics of occupancy and the role in gene regulation of the transcription factor Tal1, a critical regulator of hematopoiesis, at multiple stages of hematopoietic differentiation." To this end, RNA-seq libraries were constructed from multiple mouse cell types including G1E - a GATA-null immortalized cell line derived from targeted disruption of GATA-1 in mouse embryonic stem cells - and megakaryocytes. This RNA-seq data was used to determine differential gene expression between G1E and megakaryocytes and later correlated with Tal1 occupancy. This dataset (GEO Accession: GSE51338) consists of biological replicate, paired-end, polyA selected RNA-seq libraries. Because of the long processing time for the large original files, we have downsampled the original raw data files to include only reads that align to chromosome 19 and a subset of interesting genomic loci identified by Wu et al.</p>

opencc-by-4.0Jan 2017View details →
zenodo44/100

Expanding and improving analyses of nucleotide recoding RNA-seq experiments with the EZbakR suite

<p>Data necessary to reproduce figures in manuscript titled "Expanding and improving analyses of nucleotide recoding RNA-seq experiments with the EZbakR suite". Includes:</p> <ol> <li>Compressed arrow dataset used to produce Figure S6 and S7 (subtlseq_data.tar.gz)</li> <li>Compressed arrow dataset used to produce Figure 7 (subtlseq_perturbations_data.tar.gz)</li> <li>eCLIP DDX3X peak calls from ENCODE used in Figure 7 (ENCFF901BYH_DDX3X_eCLIP.bed)</li> <li>Annotations used to process data (Hs_ensembl_lvl1_and_2.gtf and Hs_ensembl.gtf)</li> <li>Processed data to produce Figures 3C and S2 (cB_ensembl_totRNAsubtlseq.csv.gz and cB_ensemblLvl1and2_totRNAsubtlseq.csv.gz, respectively).</li> <li>Simulated data originally used in bakR publication (Vock and Simon, 2023) used to make Figure 6 of EZbakR suite paper.</li> <li>Processed data from the nanodynamo paper (Tarrerro et al. 2024) used to make Figure S9</li> <li>Table from Ietswaart et al. 2024 of kinetic parameter estimates and PUND calls from that study (mmc2.xlsx)</li> </ol> <p>Also includes supplemental tables of:</p> <ol> <li>List of genes producing transcripts predicting to undergo nuclear decay (PUNDs; Supplemental_Table_PUNDs.csv)</li> <li>Estimates for mature RNA synthesis, nuclear degradation, nuclear export, and cytoplasmic degradation rate constants from Ietswaart et al., 2024 total-cytoplasmic-nuclear TimeLapse-seq dataset (Supplemental_Table_NucCytoEsts.csv)</li> <li>Estimates for premature RNA synthesis, premature RNA processing, and mature RNA degreadation obtained from EZbakR analysis of Ietswaart et al., 2024 total RNA TimeLapse-seq dataset (Supplemental_Table_PtoMests.csv).</li> </ol> <p>Scripts to reproduce figures can be found at: https://github.com/isaacvock/EZbakRsuite_paper_code</p> <p>Updated to include data necessary to reproduce new figures/panels in revisions.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Data archive: CICT for single cell RNA-seq network inference

<p>This archive contains benchmarking input data and results for using single cell gene expression data to infer gene regulatory networks (GRN) by the Causal Inference with Composition of Transactions (CICT) method and a selected set of published methods. This accompanies the manuscript "Robust discovery of gene regulatory networks from single-cell gene expression data by Causal Inference Using Composition of Transactions" (Shojaee and Huang, Brief in Bioinform 2023. DOI: 10.1093/bib/bbad370). The CICT code is available at the GitHub repo (https://github.com/hlab1/scRNAseqWithCICT/).</p><p>The original CICT algorithm was described in Shojaee et al. (arXiv:1608.02658, 2016). The benchmarked methods were included in the BEELINE benchmarking pipeline (Pratapa et al., Nat Methods 2020), to which we added DEEPDRIM (Chen et al., Brief Bioinform 2021), SCENIC (Aibar et al., Nat Methods 2017), Inferelator 3.0 (Gibbs et al., Bioinformatics 2022), and CellOracle (Kamimoto et al., Nature 2023). The output directory names are (subdirectories within each dataset):</p><p>* CICT_ewMIshrink_RFmaxdepth10_RFntrees20/: CICT for simulated data<br>* CICT_v2/: CICT for experimental data<br>* CELLORACLEDB/: CellOracle for experimental data<br>* DEEPDRIM72_ewMIshrink_RFmaxdepth10_RFntrees20/: DEEPDRIM for simulated data<br>* DEEPDRIM72_v2/: DEEPDRIM for experimental data<br>* INFERELATOR38_ewMIshrink_RFmaxdepth10_RFntrees20/: Inferelator-Prior for simulated data<br>* INFERELATOR38_v2/: Inferelator-Prior for experimental data<br>* INFERELATOR34_ewMIshrink_RFmaxdepth10_RFntrees20/: Inferelator-NoPrior for experimental data<br>* INFERELATOR34_v2/: Inferelator-NoPrior for experimental data<br>* GENIE3/: GENIE3<br>* GRNBOOST2/: GRNBOST2<br>* LEAP/: LEAP<br>* PIDC/: PIDC<br>* PPCOR/: PPCOR<br>* SCENICDB/: SCENIC for experimental data<br>* SCNS/: SCNS<br>* SCODE/: SCODE<br>* SCRIBE/: SCRIBE<br>* SINCERITIES/: SINCERITIES<br>* SINGE/: SINGE<br>* RANDOM/: RANDOM</p><p>The methods were benchmarked against two kinds of scRNA-seq datasets:<br>* Simulated datasets produced by the SERGIO simulator from a synthetic network (Dibaeinia et al., Cell Systems 2020), including complete datasets and datasets with dropouts with shape parameter k=6.5 and rate parameter q=10, 30, 50, 70, 80.&nbsp;<br>* Experimental datasets compiled by the BEELINE pipeline, evaluated at three different levels L0, L1 and L2, with three types of ground truth networks.<br>&nbsp; &nbsp; * Evaluation levels:<br>&nbsp;&nbsp; &nbsp; &nbsp; &nbsp;* L0: 500 highly varying genes plus TFs<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* L1: 1000 highly varying genes plus TFs<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* L2: 500 highly varying genes, TFs and 500 genes randomly selected that excluded the 1000 highly varying genes from L1.<br>&nbsp; &nbsp; * Types of ground truths:<br>&nbsp;&nbsp; &nbsp; &nbsp; &nbsp;* Cell-type-specific ChIP-seq ground truth (L0, L1, L2)<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* Non-specific ChIP-seq ground truth (L0_ns, L1_ns, L2_ns)<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* Loss-of-function/gain-of-function ground truth (L0_lofgof, L1_lofgof, L2_lofgof)</p><p>The directory structure is organized in accordance with the BEELINE benchmarking pipeline. For complete details please please see the BEELINE documentation (https://murali-group.github.io/Beeline/) and Github repo (https://github.com/Murali-group/Beeline).</p><p>&nbsp;</p>

opencc-by-nc-sa-4.0Jun 2023View details →
zenodo44/100

Predicting Phenotypic Traits Using a Massive RNA-seq Dataset

<h2><strong>Abstract</strong></h2><p>The included datasets are a conglomerate of all available <i>Arabidopsis thaliana</i> RNA-seq data available from NCBI as of November 2022 processed to count data. In addition, the associated annotation files from NCBI BioProject database and processed versions of this data is included. Data has been processed according to the "Data Description Methods" in the manuscript titled "Predicting Phenotypic Traits Using a Massive RNA-seq Dataset" (in publication). The associated Methods can be found at this repository:<a href="https://gitlab.com/ficklinlab-public/modeling-with-transcriptomics"> https://gitlab.com/ficklinlab-public/modeling-with-transcriptomics</a>. These datasets can be used for exploring machine learning methods for predicting both continuous (Age) and categorical (Tissue) phenotypic traits using gene expression. Additionally, the gene expression data can be used on its own for the investigation of gene expression in <i>Arabidopsis thaliana.</i></p><h3><strong>Note to Researchers</strong></h3><p>This repository contains all of the datasets and information necessary to recreate the experiments in our paper. However, if may be that you are interested in our dataset for testing your own hypotheses/programs. If this is the case, we predict that you are looking for one or more of the the following 5 datasets<br>&nbsp;</p><h3><strong>Note on File Compression</strong></h3><p>All files in this repository are compressed using bzip2 to conserve space and allow for easier file transfer. The unzip command on linux systems is `bzip2 -d FILE_NAME`. For other computer systems (Windows and Apple) please consult your user manual.</p><h3><strong>Description All Datasets:</strong></h3><p><strong>Title:</strong> Gene Expression Count Data of all <i>Arabidopsis thaliana</i> data available from NCBI SRA as of November 2022<br><strong>Abstract: </strong>Gene Expression Count data was created using the workflow GEMmaker. The resulting Gene Expression Matrix (GEM) was then normalized and thresholded. The following 4 files are normalizations of the same data for Trimmed Mean of M values (TMM), Median Ratios Normalization (MRN), Transcripts Per kilobase Million (TPM), and No Normalization (NoNo) respectively. Additionally, Each file is included as a tsv and a python pickle. The tsv file is human readable, whereas the pickle file can be read into memory substantially faster. Format for tsv is each row represents a sample and each column represents a gene. <strong>NCBI_Nov2022_SRR_runinfo.csv</strong> is the starting file from NCBI which reports SRR information for each sample. <strong>Note 1 to Researchers: </strong>MRN normalization performed the best in our experiments and is likely what you want to use if you are doing additional expermentation with this dataset. Otherwise start with NoNo and perform your own normalizations. <strong>Note 2 to Researchers:</strong> the 54547 dataset will need to be thresholded prior to use. We include it in addition to the 32432 datasets in case you wish to try a different thresholding to the one outlined in our manuscript.&nbsp;<br><strong>Author:</strong> John Anthony Hadish<br><strong>Data Type: </strong>Gene Expression Count Data<br><strong>Organism:</strong> <i>Arabidopsis thaliana</i><br><strong>Files:</strong></p><p><strong>NCBI_Nov2022_SRR_runinfo.csv - </strong>Arabidopsis RNA-seq SRA RunInfo Retrieved from NCBI November 2022. This is the unprocessed data.<br><strong>Dataset_54547_NoFilter_raw.pkl </strong>- Raw File Before thresholding (".pkl" format). Same as NoNo normalization without thresholding.<br><strong>Dataset_54547_NoFilter_raw.tsv </strong>- Raw File Before thresholding (".tsv" format). Same as NoNo normalization without thresholding.<br><strong>Dataset_32432_MRN.pkl</strong> - MRN normalized (".pkl" format)<br><strong>Dataset_32432_MRN.tsv - </strong>MRN normalized (".tsv" format)<br><strong>Dataset_32432_NoNo.pkl - </strong>NoNo normalized (".pkl" format)<br><strong>Dataset_32432_NoNo.tsv - </strong>NoNo normalized (".tsv" format)<br><strong>Dataset_32432_TMM.pkl - </strong>TMM normalized (".pkl" format)<br><strong>Dataset_32432_TMM.tsv - </strong>TMM normalized (".tsv" format)<br><strong>Dataset_32432_TPM.pkl - </strong>TPM normalized (".pkl" format)<br><strong>Dataset_32432_TPM.tsv - </strong>TPM normalized (".tsv" format)<br><br><br><strong>Title: </strong>Meta Data Arabidopsis Age and Tissue<br><strong>Abstract: </strong>Meta Data for Age and Tissue after processing. In our experiment this was used as response variable to gene expression. Shared columns are "bio_sample", "bioproject_name", "experiment". In addition to these processed datasets, <strong>NCBI_Nov2022_BioSample_data.tsv </strong>is the unprocessed starting material for these two data frames.<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>Metadata on phenotypes. ".tsv" format&nbsp;<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:&nbsp;</strong><br><strong>NCBI_Nov2022_BioSample_data.tsv - </strong>Arabidopsis BioSample data retrieved from NCBI November 2022. This is the unprocessed data.<br><strong>df_metadata_tissue.tsv</strong> - Tissue Annotations for 24876 samples<br><strong>df_metadata_age.tsv</strong><i><strong> - </strong></i>Age Annotations for 16078 samples. In addition to shared columns includes<i> "</i>days<i>_</i>age"(how many days old the sample is converted to days) and "annotation_age" (how the annotation was reported for this sample in the raw data file-- i.e. "days", "weeks" etc.)<br><br><br><strong>Title: </strong>Machine Learning Dataset for <i>Arabidopsis thaliana</i> <strong>Age</strong><br><strong>Abstract: </strong>The dataset used for Machine learning on the phenotype Age that is a combination of the Gene Expression Matrix and the Annotation Matrix. Consists of a list of 4 for the train and test splits.<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>Gene Expression Matrix and Annotations Combined, split into train and test&nbsp;<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:</strong><br><strong>Dataset_Age_TrainTestSplits_mrn.pkl</strong> - MRN normalized<br><strong>Dataset_Age_TrainTestSplits_NoNo.pkl </strong>- NoNo normalized<br><strong>Dataset_Age_TrainTestSplits_tmm.pkl </strong>- TMM normalized<br><strong>Dataset_Age_TrainTestSplits_tpm.pkl - </strong>TPM normalized<br><br><br><strong>Title: </strong>Machine Learning Dataset for <i>Arabidopsis thaliana</i> <strong>Tissue</strong><br><strong>Abstract: </strong>The dataset used for Machine learning on the phenotype Tissue that is a combination of the Gene Expression Matrix and the Annotation Matrix. Consists of a list of 4 for the train and test splits. Saved as python ".pkl" files.<br><strong>Author:</strong> John Anthony Hadish<br><strong>Data Type: </strong>Gene Expression Matrix and Annotations Combined, split into train and test. Saved as python ".pkl" files.<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:</strong><br><strong>Dataset_Tissue_TrainTestSplits_mrn.pkl </strong>- MRN normalized<br><strong>Dataset_Tissue_TrainTestSplits_NoNo.pkl </strong>- NoNo normalized<br><strong>Dataset_Tissue_TrainTestSplits_tmm.pkl </strong>- TMM normalized<br><strong>Dataset_Tissue_TrainTestSplits_tpm.pkl </strong>- TPM normalized<br><strong>Dataset_Tissue_TrainTestSplits_mrn_4category.pkl </strong>- MRN for the tissue-4 dataset<br><br><br><strong>Title: </strong>BioProject Names<br><strong>Abstract:</strong> Three Column File With BioProject Name, BioSample Name, and Experiment Name<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>".tsv"<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:</strong><br><strong>BioProject_Names_All.tsv</strong><br><br><br><strong>Title:</strong> Manuscript Supplemental Material<br><strong>Abstract:</strong> Supplemental tables and figures described in the manuscript (included with manuscript and here for convenience). Please see manuscript for additional information.<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>".tsv", ".png".pdf"<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:</strong><br><strong>Supplemental_Figures.zip </strong>- Supplemental figures from the manuscript. Includes description of each figure.<br><strong>Supplemental_Tables.zip </strong>- Supplemental tables from the manuscript. Includes description of each table.<br><br>&nbsp;</p><p><strong>Title:</strong> Splits of data for 3 experiments<br><strong>Abstract:</strong> 2 column tsv files. The first column is the experiment (sample) name, and the second column is if it is included in the train or test data. <strong>Included here to make sure pkl files are reproducible in case the pkl package breaks in the future.</strong> Not used by scripts, included to prevent future potential loss of data.<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>".tsv"<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:</strong><br><strong>Dataset_Tissue_TrainTestSplits_4category_namesOnly.tsv</strong><br><strong>Dataset_Tissue_TrainTestSplits_namesOnly.tsv</strong><br><strong>Dataset_Age_TrainTestSplits_namesOnly.tsv</strong></p><p>&nbsp;</p><p><strong>Title:</strong> Git Code Repository<br><strong>Abstract:</strong> A tar bz2 compression of the git repository containing all of the code created for this manuscript. The same code found in this file is also avalible on GitLab at the link: https://gitlab.com/ficklinlab-public/modeling-with-transcriptomics<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>Git Repository, python code<br><strong>Files:</strong><br><strong>modeling-with-transcriptomics-main.tar.bz2</strong> - Compressed Git repository of all code used in paper.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Single-cell RNA-seq profiles of tumor-bearing mice treated with PAGln with or without anti-PD-1

<p>single-cell RNA sequencing (scRNA-seq) profiles of&nbsp; tumor-bearing mice treated using Phenylacetylglutamine (PAGln) with or without anti-PD-1 were performed to compare the alterations of immune microenvironment affected by PAGln under the condition of anti-PD-1 treatment.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Formatted TCGA clinical and RNA-Seq data for colon adenocarcinoma (COAD) and rectum adenocarcinoma (READ)

<p>COAD/READ/COADREAD_rnaseq_fpkm.txt files contain TCGA RNA-Seq data in FPKM normalisation for colorectal adenocarcinoma (COAD), rectum adenocarcinoma (READ) or combined (COADREAD).</p> <p>COAD/READ/COADREAD_rnaseq_tpm.txt files contain TCGA RNA-Seq data in TPM normalisation for colorectal adenocarcinoma (COAD), rectum adenocarcinoma (READ) or combined (COADREAD).</p> <p>COAD/READ/COADREAD_clinical_raw.xlsx&nbsp;files contain TCGA clinical data for patients with&nbsp;colorectal adenocarcinoma (COAD), rectum adenocarcinoma (READ) or combined (COADREAD).</p> <p>COAD/READ/COADREAD_rnaseq_clinical_raw.xlsx&nbsp;files contain corresponding information of TCGA clinical data and RNA-Seq data for patients with&nbsp;colorectal adenocarcinoma (COAD), rectum adenocarcinoma (READ) or combined (COADREAD).</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Vignettes: Removing unwanted variation from TCGA RNA-Seq data.

<p>This repository contains all datasets that are required for the vignettes of&nbsp;&nbsp;R.Molania et.al bioRxiv paper (https://www.biorxiv.org/content/10.1101/2021.11.01.466731v1).</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Evaluating the influence of structural properties on proximity metric performance in single cell RNA-seq data - Datasets

<p>Includes raw and processed copies of the scRNA-seq datasets used for the paper: &#39;<strong>How does data structure impact cell-cell similarity? Evaluating the influence of structural properties on proximity metric performance in single cell RNA-seq data.&#39;</strong></p> <p><strong>Real scRNA-seq.zip </strong>contains the Abundant (subset1) and Rare (subset 2) subsets generated to represent discretely structured datasets (sourced from<strong> </strong> Wegmann et al. 2019) and the continuously structured data (sourced from Popescu et al. 2019).</p> <p><strong>Simulated scRNA-seq.zip</strong> contains the Abundant, Moderately-Rare and Ultra-Rare subsets for discretely and continuously structured datasets. All data was simulated using the PROSSTT package in Python 3.8, as well as the dataset containing the labels to re-produce Figure 3 of the manuscript.</p> <p><strong>Results.zip </strong>contains the results for all datasets from the full analysis, in a pickled python dictionary. Code to read in and visualise results is available on the projects github</p> <p>The scripts for the dataset generation, processing and visualisation of results are available at <a href="https://github.com/Ebony-Watson/scProximitE">our github for the scProcimitE package</a>, and documentation is available <a href="https://ebony-watson.github.io/scProximitE/">here</a>.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

RNA-seq dataset for Integrative functional genomic analyses implicate specific molecular pathways and circuits in autism

<p>Data to be used along with <a href="https://github.com/neelroop/asd-development-coexpression-2013">code</a> from 2013 paper that was originally on a site hosted at UCLA, but may no longer be accessible.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Dataset of "Single-Cell RNA-Seq Reveals Transcriptional Heterogeneity in Latent and Reactivated HIV-infected Cells"

<p><strong>Detailed quantitative analysis of GFP expression in SAHA and TCR-treated cells &amp; Computational analysis of&nbsp;bulk and single-cell RNA-Seq data.</strong></p> <p>&nbsp;</p> <p><em><strong>Detailed quantitative analysis of GFP expression in SAHA and TCR-treated cells.</strong></em></p> <p>Cells were prepared for single cell analysis at the Genome Technology Facility (GTF) of the University of Lausanne. Cells were loaded on Fluidigm C1 IFC plates (5-10 &mu;m), with run ID smart33, smart34 and smart35, corresponding to untreated, SAHA- and TCR-treated conditions respectively. After single cell capture on the Fluidigm C1 IFC plate, each chamber was inspected visually by microscopy and pictures were captured with a Zeiss Axiovert 200 M fluorescence microscope equipped with a Roper Scientific CoolSnap HQ camera using a Plan-Neofluar 10X lens (smart34 run) or 20X lens (for smart35 run). For each capture chamber, pictures in bright field and FITC channel were taken with the MetaMorph 6.3 software. Picture analysis was then performed using ImageJ 1.50b software (open access software: website). Brightness and contrast were adjusted for qualitative assessment of the pictures.</p> <p><em><strong>Computational analysis of&nbsp;bulk and single-cell RNA-Seq data.</strong></em></p> <p>Upon bulk or single cell isolation, RNA extraction and library preparation was performed according to Illumina protocols. Bulk and single-cell RNA-Seq data analysis are detailed here.</p> <p>&nbsp;</p> <p>Linked to the paper published in Cell Reports (doi:10.1016/j.celrep.2018.03.102):&nbsp;</p> <p><strong>Single-Cell RNA-Seq Reveals Transcriptional Heterogeneity&nbsp;in Latent and Reactivated HIV-infected Cells</strong></p> <p>Despite effective treatment, HIV can persist in latent reservoirs, which represent a major obstacle towards HIV eradication. Targeting and reactivating latent cells is challenging due to the heterogeneous nature of HIV infected cells. Here, we used a primary model of HIV latency and single-cell RNA sequencing to characterize transcriptional heterogeneity during HIV latency and reactivation. Our analysis identified transcriptional programs leading to successful reactivation of HIV expression.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2018View details →
zenodo44/100

Sample datasets for Galaxy RNA-seq tutorial

<p>This is downsampled dataset from&nbsp;http://dx.doi.org/10.1038/nprot.2016.095. It was prepared as follows:</p> <ol> <li>Mapping data from&nbsp;&nbsp;http://dx.doi.org/10.1038/nprot.2016.095 against hg38 using HISAT2</li> <li>Restricting resulting BAM datasets to chrX:70,000,000-80,000,000</li> <li>Extracting reads using picard SamToFastq tool</li> </ol> <p>File rnaseq_sex.tab contains mapping between accession numbers and sex of the sequences individuals.&nbsp;</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

Single-cell RNA-seq of breast cancer infiltrating T cells (case 1)

<p>Single cell suspensions were generated from two individual TNBC primary&nbsp;tumor samples (this&nbsp;entry contains case 2) and the viable cells were FACS sorted for CD3<sup>+</sup> T cells.&nbsp;Sorted cells were then counted and assessed for viability. Single cell library preparation was carried out as per the 10X Genomics Chromium Single cell protocol for the v2 reagent kit (10X Genomics, Pleasanton, CA, USA). Cell suspensions were loaded onto a Chromium Single Cell Chip along with the reverse transcription (RT) mastermix and single cell 3&rsquo; gel beads. Following generation of single cell gel bead-in-emulsions (GEMs), reverse transcription was performed using a C1000 Touch Thermal Cycler with a Deep Well Reaction Module (Bio-Rad Laboratories, Hercules, CA, USA). Amplified cDNA was purified using SPRIselect beads (Beckman Coulter, Lane Cove, NSW, Australia) and sheared to approximately 200bp with a Covaris S2 instrument (Covaris, Woburn, MA, USA) using the manufacturer&rsquo;s recommended parameters. Sequencing libraries were generated with unique sample indices (SI) for each sample. Libraries were sequenced on an Illumina HiSeq 2500 High Output Mode using V4 clustering and sequencing chemistry.</p> <p>This dataset&nbsp;contains&nbsp;the raw .bcl files.</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

Single-cell RNA-seq of breast cancer infiltrating T cells (case 2)

<p>Single cell suspensions were generated from two individual TNBC primary&nbsp;tumor samples (this&nbsp;entry contains case 1) and the viable cells were FACS sorted for CD3<sup>+</sup> T cells.&nbsp;Sorted cells were then counted and assessed for viability. Single cell library preparation was carried out as per the 10X Genomics Chromium Single cell protocol for the v2 reagent kit (10X Genomics, Pleasanton, CA, USA). Cell suspensions were loaded onto a Chromium Single Cell Chip along with the reverse transcription (RT) mastermix and single cell 3&rsquo; gel beads (this sample was divided into two channels). Following generation of single cell gel bead-in-emulsions (GEMs), reverse transcription was performed using a C1000 Touch Thermal Cycler with a Deep Well Reaction Module (Bio-Rad Laboratories, Hercules, CA, USA). Amplified cDNA was purified using SPRIselect beads (Beckman Coulter, Lane Cove, NSW, Australia) and sheared to approximately 200bp with a Covaris S2 instrument (Covaris, Woburn, MA, USA) using the manufacturer&rsquo;s recommended parameters. Sequencing libraries were generated with unique sample indices (SI) for each sample. Libraries were sequenced on an Illumina HiSeq 2500 High Output Mode using V4 clustering and sequencing chemistry.</p> <p>This dataset&nbsp;contains&nbsp;the raw .bcl files.</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

CellSIUS provides sensitive and specific detection of rare cell populations from complex single cell RNA-seq data: Codes and processed data

<p>Codes and processed data to reproduce the analysis discussed in:&nbsp;</p> <p>Wegmann <em>et Al.</em>,<strong> CellSIUS provides sensitive and specific detection of rare cell<br> populations from complex single cell RNA-seq data</strong>, Genome Biology 2019 (Accepted)<br> &nbsp;</p>

openapache2.0Jun 2019View details →
zenodo44/100

BaRTv1.0: an improved barley reference transcript dataset to determine accurate changes in the barley transcriptome using RNA-seq

<p>Background<br> Time consuming computational assembly and quantification of gene expression and splicing analysis from RNA-seq data vary considerably. Recent fast non-alignment tools such as Kallisto and Salmon overcome these problems, but these tools require a high quality, comprehensive reference transcripts dataset (RTD), which are rarely available in plants.</p> <p>Results<br> A high-quality, non-redundant barley gene RTD and database (Barley Reference Transcripts &ndash; BaRTv1.0) has been generated. BaRTv1.0, was constructed from a range of tissues, cultivars and abiotic treatments and transcripts assembled and aligned to the barley cv. Morex reference genome (Mascher et al., 2017). Full-length cDNAs from the barley variety Haruna nijo (Matsumoto et al., 2011) determined transcript coverage, and high-resolution RT-PCR validated alternatively spliced (AS) transcripts of 86 genes in five different organs and tissue. These methods were used as benchmarks to select an optimal barley RTD. BaRTv1.0-Quantification of Alternatively Spliced Isoforms (QUASI) was also made to overcome inaccurate quantification due to variation in 5&rsquo; and 3&rsquo; UTR ends of transcripts. BaRTv1.0-QUASI was used for accurate transcript quantification of RNA-seq data of five barley organs/tissues. This analysis identified 20,972 significant differentially expressed genes, 2,791 differentially alternatively spliced genes and 2,768 transcripts with differential transcript usage.</p> <p>Conclusion<br> A high confidence barley reference transcript dataset consisting of 60,444 genes with 177,240 transcripts has been generated. Compared to current barley transcripts, BaRTv1.0 transcripts are generally longer, have less fragmentation and improved gene models that are well supported by splice junction reads. Precise transcript quantification using BaRTv1.0 allows routine analysis of gene expression and AS.</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

An example RNA-seq count table using nf-core/rnaseq

<p>An example RNA-seq count table generated from data deposited under GSE40419 (lung adenocarcinoma study) and nf-core/rnaseq for testing downstream RNA-seq analyses.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

ZIRFs: zero-inflated random forests for estimating gene regulatory networks from single cell RNA-seq data (assessment of predictive accuracy and VIM stability)

<p>We developed a zero-inflated random forests (ZIRFs) algorithm to produce a metric of connection strength&nbsp;between regulator genes and target genes. This file contains SCENIC results for the aorta and diaphragm tissue data sets from the Tabula Muris Consortium results. SCENIC is a genetic regulatory network analysis published by Aibar et al. (2017). The purpose of the data sets and R source code are described by README files in each directory.</p>

opencc-by-3.0-usJul 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record