Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,780
datasets available to search
ShareScore release 0.7.1
Dataset results
5,780 results for “mRNA”
Pan-cancer analysis of mRNA stability for decoding tumour post-transcriptional programs
<p>Supplemental data and analysis files for Perron et al.: "Pan-cancer analysis of mRNA stability for decoding tumour post-transcriptional programs" (<a href="https://www.nature.com/articles/s42003-022-03796-w">https://www.nature.com/articles/s42003-022-03796-w</a>). The .tar.gz files contain read counts associated with various RNA-seq analyses. The .rds files are single R object files that contain various analysis results tables. The .csv files also contain analysis results or sample metadata tables. See <a href="http://csg.lab.mcgill.ca/sup/pancancer_stability/">http://csg.lab.mcgill.ca/sup/pancancer_stability/</a> for a full description of the files.</p>
Enhanced Biosafety of the Sleeping Beauty Transposon System by Using mRNA as Source of Transposase to Efficiently and Stably Transfect Retinal Pigment Epithelial Cells
<p>Raw data of the publication "Enhanced Biosafety of the Sleeping Beauty Transposon System by Using mRNA as Source of Transposase to Efficiently and Stably Transfect Retinal Pigment Epithelial Cells".</p> <p>Abstract: Neovascular age-related macular degeneration (nvAMD) is characterized by choroidal<br> neovascularization (CNV), which leads to retinal pigment epithelial (RPE) cell and photoreceptor<br> degeneration and blindness if untreated. Since blood vessel growth is mediated by endothelial cell<br> growth factors, including vascular endothelial growth factor (VEGF), treatment consists of repeated,<br> often monthly, intravitreal injections of anti-angiogenic biopharmaceuticals. Frequent injections are<br> costly and present logistic difficulties; therefore, our laboratories are developing a cell-based gene<br> therapy based on autologous RPE cells transfected ex vivo with the pigment epithelium derived factor<br> (PEDF), which is the most potent natural antagonist of VEGF. Gene delivery and long-term expression<br> of the transgene are enabled by the use of the non-viral Sleeping Beauty (SB100X) transposon system<br> that is introduced into the cells by electroporation. The transposase may have a cytotoxic effect and a<br> low risk of remobilization of the transposon if supplied in the form of DNA. Here, we investigated<br> the use of the SB100X transposase delivered as mRNA and showed that ARPE-19 cells as well as<br> primary human RPE cells were successfully transfected with the Venus or the PEDF gene, followed<br> by stable transgene expression. In human RPE cells, secretion of recombinant PEDF could be detected<br> in cell culture up to one year. Non-viral ex vivo transfection using SB100X-mRNA in combination<br> with electroporation increases the biosafety of our gene therapeutic approach to treat nvAMD while<br> ensuring high transfection efficiency and long-term transgene expression in RPE cells.</p>
Associated data from: An end-to-end workflow to study newly synthesized mRNA following rapid protein depletion in Saccharomyces cerevisiae
<p>This dataset includes two custom BED files described in "An end-to-end workflow to study newly synthesized mRNA following rapid protein depletion in <em>Saccharomyces cerevisiae</em>" (Ridenour and Donczew, submitted), which were used to define counting windows for processing SLAM-seq data in SLAM-DUNK (version 0.4.3) [1]. The BED files contain all annotated open reading frames (ORFs) in the<em> Saccharomyces cerevisiae</em> genome or the <em>Schizosaccharomyces</em><em> pombe</em> genome and were created using BEDOPS (version 2.4.3) [2]. All ORFs were then extended 250 bp beyond their stop position to capture 3′ untranslated regions (UTRs) using SAMtools (version 1.14) [3] and BEDTools (version 2.30.0) [4]. The reference genome annotations for <em>S. cerevisiae</em> strain S288C (version R64-3-1, RefSeq Assembly GCF_000146045.2) and <em>S. pombe</em> strain 972h- (version ASM294v2, RefSeq Assembly GCF_000002945.1) were retrieved from the NCBI Datasets repository. The <em>S. cerevisiae </em>chromosome names were modified to reflect standard nomenclature (https://www.yeastgenome.org/).</p>
Arabidopsis thaliana circadian mRNA-seq gene expression processed tables from Romanowski et al., TPJ 2020.
<p>This dataset is an add-on for Romanowski et al., TPJ 2020 (https://doi.org/10.1111/tpj.14776) containing processed files for the circadian RNAseq data in tab delimited txt format.</p> <p><br> Here, you can the raw counts file, the normalized CPM values, and the full JTK result (without recalculated circadian phases, just the original ones). All genes with a read density > 0.05 in at least one timepoint were considered expressed. The read density is calculated as the amount of reads divided by the effective length of a gene (total reads / length). Genes rd file is also included.</p> <p>Some useful notes:<br> 1) Counts were assigned using ASpli and the AtRTDv2 annotation (34,212 genes).<br> 2) After filtering by rd we had a total of 18,503 expressed genes.<br> 3) 13,256 genes passed the QL F-tests.<br> 4) 9,127 genes were rhythmic according to JTK_cycle. </p> <p>For detailed protocols, please see Romanowski et al., TPJ 2020 (https://doi.org/10.1111/tpj.14776)</p> <p>The RNA-seq raw data supporting the conclusions of this article have been deposited in ArrayExpress (Kolesnikov et al., 2015) at EMBL-EBI (www.ebi.ac.uk/arrayexpress), under accession numbers E-MTAB-7933.</p> <p>All relevant custom r scripts are available at https://github.com/aromanowski/Circadian_rhythms_and_alternative_splicing</p>
Multi-Dimensional Data Viewer (MDV) user manual for data exploration: "Systematic analysis of YFP traps reveals common discordance between mRNA and protein across the nervous system"
<table> <tbody> <tr> <td> <p> Please also see the latest version of the repository:<br> <a href="https://doi.org/10.5281/zenodo.6374011">https://doi.org/10.5281/zenodo.6374011</a> and<br> our website: <a href="https://ilandavis.com/jcb2023-yfp">https://ilandavis.com/jcb2023-yfp</a></p> </td> </tr> </tbody> </table> <p> </p> <p>The explosion in the volume of biological imaging data challenges the available technologies for data interrogation and its intersection with related published bioinformatics data sets. Moreover, intersection of highly rich and complex datasets from different sources provided as flat csv files requires advanced informatics skills, which is time consuming and not accessible to all. Here, we provide a “user manual” to our new paradigm for systematically filtering and analysing a dataset with more than 1300 microscopy data figures using Multi-Dimensional Viewer (MDV) -<a href="https://mdv.molbiol.ox.ac.uk/projects/mdv_project/7012?view=RNA+%2F+Protein+Distribution">link</a>, a solution for interactive multimodal data visualisation and exploration. The primary data we use are derived from our published systematic analysis of 200 YFP traps reveals common discordance between mRNA and protein across the nervous system (<a href="https://doi.org/10.1083/jcb.202205129">eprint link</a>). This manual provides the raw image data together with the expert annotations of the mRNA and protein distribution as well as associated bioinformatics data. We provide an explanation, with specific examples, of how to use MDV to make the multiple data types interoperable and explore them together. We also provide the open-source python code <a href="https://github.com/ilandavislab/Annotate.OMERO.Fig">(github link)</a> used to annotate the figures, which could be adapted to any other kind of data annotation task.</p>
CGGA mRNA-array_301 Expression Data
<p><strong>Abstract:</strong></p> <p>The Chinese Glioma Datasets (CGGA) are comprehensive and valuable collections of data related to glioma, a type of brain tumor, originating from Chinese patients. The CGGA is a data portal for the storage and interactive exploration of cross-omics data, including nearly 2000 primary and recurrent glioma samples. This dataset's gene expression profile was measured experimentally using Agilent-014850 Whole Human Genome Microarray. The Sample IDs serve as unique identifiers for each sample. </p> <p><strong>Inspiration:</strong></p> <p>This dataset was uploaded to UBRITE for GTKB project. </p> <p><strong>Instruction:</strong></p> <p>z-normalization was performed on all the samples </p> <p><strong>Acknowledgments:</strong></p> <p>Zhao, Z., Zhang, KN., Wang, QW., et al. Chinese Glioma Genome Atlas (CGGA): A Comprehensive Resource with Functional Genomic Data from Chinese Glioma Patients (2021). Genomics, Proteomics & Bioinformatics 19(1):1-12.</p> <p>Fang, S., Liang, J., Qian, T., et al. (2017). Anatomic Location of Tumor Predicts the Accuracy of Motor Function Localization in Diffuse Lower-Grade Gliomas Involving the Hand Knob Area. AMERICAN JOURNAL OF NEURORADIOLOGY. 38(10): 1990-1997</p> <p>Wang, Y., Wang, Y., Fan, X., et al. (2017). Putamen involvement and survival outcomes in patients with insular low-grade gliomas. JOURNAL OF NEUROSURGERY. 126(6): 1788-1794.</p> <p> </p> <p><strong>U-BRITE last update: </strong>07/27/2023</p>
CGGA mRNA seq 325 Clinical Data
<p><strong>Abstract:</strong></p> <p>The Chinese Glioma Datasets (CGGA) are comprehensive and valuable collections of data related to glioma, a type of brain tumor, originating from Chinese patients. The CGGA is a data portal for the storage and interactive exploration of cross-omics data, including nearly 2000 primary and recurrent glioma samples. This dataset contains phenotype information. The Sample IDs serve as unique identifiers for each sample. </p> <p><strong>Inspiration:</strong></p> <p>This dataset was uploaded to UBRITE for GTKB project. </p> <p><strong>Acknowledgments:</strong></p> <p>Zhao, Z., Zhang, KN., Wang, QW., et al. Chinese Glioma Genome Atlas (CGGA): A Comprehensive Resource with Functional Genomic Data from Chinese Glioma Patients (2021). Genomics, Proteomics & Bioinformatics 19(1):1-12.</p> <p>Fang, S., Liang, J., Qian, T., et al. (2017). Anatomic Location of Tumor Predicts the Accuracy of Motor Function Localization in Diffuse Lower-Grade Gliomas Involving the Hand Knob Area. AMERICAN JOURNAL OF NEURORADIOLOGY. 38(10): 1990-1997.</p> <p>Wang, Y., Wang, Y., Fan, X., et al. (2017). Putamen involvement and survival outcomes in patients with insular low-grade gliomas. JOURNAL OF NEUROSURGERY. 126(6): 1788-1794.</p>
Test data for running snakePipes : mRNA-seq workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the mRNA-seq workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for mouse (<strong>GRCm38</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>
Widespread Polycistronic Transcripts in Fungi Revealed by Single-Molecule mRNA Sequencing
<p>Genes in prokaryotic genomes are often arranged into clusters and co-transcribed into poly- cistronic RNAs. Isolated examples of polycistronic RNAs were also reported in some higher eukaryotes but their presence was generally considered rare. Here we developed a long- read sequencing strategy to identify polycistronic transcripts in several mushroom forming fungal species including Plicaturopsis crispa, Phanerochaete chrysosporium, Trametes ver- sicolor, and Gloeophyllum trabeum. We found genome-wide prevalence of polycistronic transcription in these Agaricomycetes, involving up to 8% of the transcribed genes. Unlike polycistronic mRNAs in prokaryotes, these co-transcribed genes are also independently transcribed. We show that polycistronic transcription may interfere with expression of the downstream tandem gene. Further comparative genomic analysis indicates that polycis- tronic transcription is conserved among a wide range of mushroom forming fungi. In sum- mary, our study revealed, for the first time, the genome prevalence of polycistronic transcription in a phylogenetic range of higher fungi. Furthermore, we systematically show that our long-read sequencing approach and combined bioinformatics pipeline is a generic powerful tool for precise characterization of complex transcriptomes that enables identifica- tion of mRNA isoforms not recovered via short-read assembly.</p>
Additional Data: Poised PABP-RNA hubs implement signal-dependent mRNA decay in development
<p>This repository contains processed data resulting from iCLIP experiments that were analysed in the following paper:"<strong>Poised PABP-RNA hubs implement signal-dependent mRNA decay in development</strong>"<br>The paper is published at Nature Structural and Molecular BIology.</p> <h2><br>Archived data</h2> <p>Data archived in this repository include:</p> <ol> <li>Data derived from iCLIP experiments targeting LIN28A, PABPC1, and PABPC4, that were analysed in the manuscript (see iCLIP.zip). Raw data is available from ENA, with the accession code PRJEB60519. <ol> <li>Sample descriptions are given in iCLIP-SampleAnnotation.csv</li> <li>Crosslink files in BED6 format (individual replicates and merged replicates)</li> <li>Peak files generated with the Clippy peak caller in BED6 format</li> <li>K-mer enrichment around high-confidence crosslink sites in the 3'-UTRs, calculated by the PEKA software</li> </ol> </li> <li>Expression values (salmon quantfiles) for 3'-seq experiments, specified in "QuantseqExperimentsAnnotation.tsv", are available in "SalmonQuantfiles.zip". Raw data is available from ENA, with the accession code PRJEB60519.</li> <li>Source code of the nextflow pipeline, which was used on the iMaps webserver to analyse iCLIP data and produce the files archived here (see imaps-nf-0.30.zip).</li> <li>A list of naive genes, that were analysed in the manuscript (see NaiveGeneIds.csv).</li> </ol> <h2>Details on iCLIP data generation</h2> <p>iCLIP data for LIN28A-WT (in 2iL and FGF2 treated cells), LIN28A-S200A (in FGF2 treated cells) as well as for PABPC1 and PABPC4 (in LIN28A KO cells with and without LIN28A overexpression), were analysed on iMaps Goodwright server (<a href="https://imaps.goodwright.com/">https://imaps.goodwright.com/</a>). The LIN28A iCLIPs were analysed on 18th of July, 2022; the PABPC iCLIPs were analysed on 26th of December, 2022. The code and settings used in the pipeline (release v0.30) can be viewed at <a href="https://github.com/goodwright/imaps-nf">https://github.com/goodwright/imaps-nf </a>, and is also archived here - (imaps-nf-0.30.zip)<br> </p> <ul> <li>First, reads were demultiplexed using Ultraplex and barcodes were trimmed from the reads. The default Ultraplex settings were applied, as denoted below:</li> </ul> <blockquote> <p>adapter='AGATCGGAAGAGCGGTTCAG'<br>adapter2='AGATCGGAAGAGCGTCGTG'<br>barcodes='barcode.csv',<br>final_min_length=20<br>fiveprimemismatches=1<br>ignore_no_match=False<br>ignore_space_warning=False<br>inputfastq='MOD4878A1-merged.fastq.gz',<br>keep_barcode=False,<br>min_trim=3,<br>outputprefix='demux',<br>phredquality=30,<br>phredquality_5_prime=0,<br>sbatchcompression=False,<br>threads=10,<br>threeprimemismatches=0,<br>ultra=False</p> </blockquote> <p> </p> <ul> <li>TrimGalore was used to run FASTQC and quality trim the reads and remove reads with length less than 10 nt:</li> </ul> <blockquote> <p>trim_galore --fastqc --length 10 -q 20 --cores 8 --gzip file.fastq.gz</p> </blockquote> <p> </p> <ul> <li>Reads were then premapped to rRNA, tRNA sequences referred to as small RNA, smRNA, using mouse genome build (GRCm39 GENCODE M28 annotation) with Bowtie v1.3.0 (Langmead et al., 2009)</li> </ul> <blockquote> <p>bowtie --threads 12 --sam -x $INDEX -q --un file.unmapped.fastq -v 2 -m 100 --norc --best --strata file.fq.gz 2</p> </blockquote> <p> </p> <ul> <li>Reads that did not map with Bowtie were then aligned with STAR v2.7.9a (Dobin et al., 2013) to mouse genome build (GRCm39 GENCODE M28 annotation).</li> </ul> <blockquote> <p>STAR \<br>--genomeDir star \<br>--readFilesIn file.unmapped.fastq.gz \<br>--runThreadN 12 \<br>--outFileNamePrefix 1_R1. \<br>\<br>--sjdbGTFfile Homo_sapiens_filtered.gtf \<br>--outSAMattrRGline 'ID:1_R1' 'SM:1_R1' \<br> --readFilesCommand zcat --outSAMtype BAM SortedByCoordinate --quantMode TranscriptomeSAM --outFilterMultimapNmax 1 --outFilterMultimapScoreRange 1 --outSAMattributes All --alignSJoverhangMin 8 --alignSJDBoverhangMin 1 --outFilterType BySJout --alignIntronMin 20 --alignIntronMax 1000000 --outFilterScoreMin 10 --alignEndsType Extend5pOfRead1 --twopassMode Basic</p> </blockquote> <p> </p> <ul> <li>PCR-duplicates were removed using UMI-tools (Smith, Heger and Sudbery, 2017)</li> </ul> <blockquote> <p>java -jar /UMICollapse/umicollapse.jar \<br> bam \<br> -i file.Aligned.sortedByCoord.out.bam \<br> -o file.dedup.bam \<br> --umi-sep rbc:</p> </blockquote> <p> </p> <ul> <li>The nucleotide preceding each sequencing read was assigned as the crosslink event.</li> </ul> <p> </p> <ul> <li>Peaks of crosslinking signal were identified with Clippy v1.4.1, using the default settings.</li> </ul> <p> </p> <ul> <li>Obtained peaks and crosslink sites were used to run PEKA v1.0.0 (Kuret et al., 2022), using the default settings.</li> </ul> <p> </p> <ul> <li>For Clippy and PEKA, the GENCODE primary assembly annotation M28 was filtered to retain only entries with transcript support level 1 or 2, in genes where such transcripts were available, and used to produce a segmentation file with the <em>get_segments</em> function from the iCount tool (Curk, 2019).</li> </ul> <p> </p> <ul> <li>All files generated during data processing are available from the iMaps Goodwright webserver for analysis of CLIP data (see <a href="https://imaps.goodwright.com/collections/882635250203/">https://imaps.goodwright.com/collections/882635250203/</a> and <a href="https://imaps.goodwright.com/collections/340215254997/">https://imaps.goodwright.com/collections/340215254997/</a> for LIN28A and PABPC1/4 iCLIPs, respectively).</li> </ul> <h2>Source data</h2> <p>Raw sequencing reads, from which the data enclosed here were derived, are accessible at ENA (PRJEB60519).<br>The raw sequencing reads and all data produced by the analysis pipeline is also available at the iMaps webserver (see <a href="https://imaps.goodwright.com/collections/882635250203/">https://imaps.goodwright.com/collections/882635250203/</a> and <a href="https://imaps.goodwright.com/collections/340215254997/">https://imaps.goodwright.com/collections/340215254997/</a> for LIN28A and PABPC1/4 iCLIPs, respectively); and on the updated Flow webserver (see <a href="https://app.flow.bio/projects/882635250203/">https://app.flow.bio/projects/882635250203/</a> and <a href="https://app.flow.bio/projects/340215254997/">https://app.flow.bio/projects/340215254997/ </a>for LIN28A and PABPC1/4 iCLIPs, respectively).</p> <h2>Downstream computational analysis of enclosed data</h2> <p>The code, used to analyse the data enclosed here and train the CNN to predict transcript stability in naive-to-primed transition based on 3'UTR nucleotide sequence, is available at GitHub (<a href="https://github.com/ulelab/LIN28A_RNPreassembly_bioinformatics">https://github.com/ulelab/LIN28A_RNPreassembly_bioinformatics</a>) and archived on Zenodo (<a href="../doi/10.5281/zenodo.10054297">https://zenodo.org/doi/10.5281/zenodo.10054297</a><strong>).</strong></p>
SARS-CoV-2 mRNA vaccines induce persistent human germinal centre responses
<p>These are the<strong> processed</strong> BCR repertoire bulk sequencing data described in <a href="https://doi.org/10.1038/s41586-021-03738-2">Turner & O'Halloran et al., Nature, 2021</a> (Fig 3b-d; Extended Data Fig 3; Extended Data Table 6). The corresponding <strong>raw</strong> sequencing reads are available on SRA under <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA731610">BioProject PRJNA731610</a>.</p> <p><strong>Summary</strong>: Bulk-sorted total plasmablasts from PBMCs and germinal centre B cells at 4 weeks after primary immunization from 3 vaccinees who had no prior history of infection with SARS-CoV-2. </p> <p><strong>Code: </strong>Code along with Docker container for reproducing the NGS data-based figures and analyses in the published paper can be <a href="https://github.com/julianqz/wustl_published/tree/main/nature_2021">found on GitHub</a>.</p> <p><strong>Metadata file</strong>: WU368_turner_et_al_nature_2021_meta.tsv</p> <p>Abbreviations:</p> <ul> <li>LN = lymph node</li> <li>PB = plasmablast</li> <li>GC = germinal centre</li> <li>mAb = monoclonal antibody</li> </ul> <p><strong>BCR data file</strong>: WU368_turner_et_al_nature_2021_bcr.tsv.gz</p> <p>In addition to the processed bulk sequences, also included are the heavy chains of 37 mAbs that had been validated to be spike-binding and that were used together with the bulk sequences for clonal lineage inference. The mAbs are annotated as "mab" in the "seq_type" column.</p> <p><strong>BCR data column descriptions</strong></p> <p>The columns largely follow the <a href="https://changeo.readthedocs.io/en/stable/standard.html">AIRR-C Rearrangement format</a>. The main deviation is that CDR3s are used, as opposed to IMGT-defined "junctions". Non-standard columns are noted below.</p> <ul> <li>v_call_genotyped: V gene annotation reassigned after individualized genotyping by <a href="https://tigger.readthedocs.io/en/stable/">TIgGER</a></li> <li>germline_[vdj]_call: clonal consensus germline sequence reconstructed via <a href="https://changeo.readthedocs.io/en/stable/methods/germlines.html">`CreateGermlines.py --cloned` using Change-O</a></li> <li>isotype: IGH[ADEGM]</li> <li>cdr3: CDR3 nucleotide sequence</li> <li>cdr3_length: CDR3 nucleotide sequence length</li> <li>cdr3_aa: CDR3 amino acid sequence</li> <li>collapse_count: number of duplicate IMGT-aligned V(D)J sequences that were collapsed by <a href="https://alakazam.readthedocs.io/en/stable/topics/collapseDuplicates/">`alakazam::collapseDuplicates`</a></li> <li>donor: vaccinee</li> <li>sample: sample ID (arbitrary)</li> <li>timepoint: time point at which sample was collected</li> <li>tissue: tissue from which sample was collected</li> <li>sorting: FACS sorting</li> <li>seq_type: sequence type (mAb or bulk)</li> <li>nuc_RS_19_312: number of replacement and silent mutations between IMGT-numbered nucleotide positions 19-312 along IGHV sequences, calculated by <a href="https://shazam.readthedocs.io/en/stable/topics/calcObservedMutations/">`shazam::calcObservedMutations`</a></li> <li>nuc_denom_19_312: number of informative nucleotide positions for counting mutations, excluding non-A/T/G/C positions (such as "N", "-", ".")</li> <li>nuc_RS_freq_19_312: nucleotide-level mutation frequency (= nuc_RS_19_312 / nuc_denom_19_312)</li> </ul>
Regulation of mature mRNA levels by RNA processing efficiency
<p>Data from the research paper "Regulation of mature mRNA levels by RNA processing efficiency" by Henfrey, C., Murphy, S., and Tellier, M.:</p> <p>-Highest expressed transcript annotation for protein-coding genes, gencode v38.</p> <p>-15 txt (BED) files for chromatin vs nucleoplasm enrichment gene sets: HeLa full gene sets, canonical protein only sets, chromatin RNA seq subsamples, mNET seq subsamples, Raji gene sets.</p> <p>-Proteomics data table</p> <p>-mRNA half-life table</p> <p>-Splicing efficiency for POINT-seq, ChrRNA-seq, NucRNA-seq table</p> <p>-Ser2-P mNET-seq readthrough index data table</p> <p>-Splicing efficiency for siLuc/siEX3 table (ChrRNA-seq, NucRNA-seq)</p> <p>-RMATs output tables for alternative splicing results (siEX3 vs siLuc)</p> <p>-Tables for TSS:TES quantifications (mNET-seq(CTD) vs log2FoldChange, chr/nuc/mnet siEX3 vs siLuc)</p>
Summary statistics data for "Genetic Analyses Support the Contribution of mRNA N6-methyladenosine (m6A) Modification to Human Diseases Heritability"
<p>We included the summary statistics data associated with our manuscript "<strong>Genetic Analyses Support the Contribution of mRNA <em>N</em><sup>6</sup>-methyladenosine (m<sup>6</sup>A) Modification to Human Diseases Heritability". </strong></p> <p>We also included the newly imputed genotype data for the 60 YRI individuals involved in our study, the joint m<sup>6</sup>A peaks tested (locations of the molecular phetnotype in BED12 format) and normalized log odds ratio (enrichment) of these joint peaks (molecular phenotype data). </p>
mRNA expression data of genes related to mitochondrial quality control in hepatopancreas of the two marine bivalves, Mytilus edulis and Crassostrea gigas, during short-term hypoxia/reoxygenation stress
<p>Coastal environments commonly experience strong oxygen fluctuations. Resulting hypoxia/reoxygenation stress can negatively affect mitochondrial functions, since oxygen deficiency impairs ATP generation, whereas a surge of oxygen causes mitochondrial damage by oxidative stress mechanisms. Marine intertidal bivalves are adapted to fluctuating oxygen conditions, yet the underlying molecular mechanisms that sustain mitochondrial integrity and function during oxygen fluctuations are not yet well understood. We used targeted mRNA expression analysis to determine the potential involvement of the mitochondrial quality control mechanisms in responses to short-term hypoxia (24 h at <0.01% O<sub>2</sub>) and subsequent reoxygenation (1.5 h at 21% O<sub>2</sub>) in two hypoxia-tolerant marine bivalves, the Pacific oysters <em>Crassostrea gigas</em> and the blue mussels <em>Mytilus edulis</em>. To test these hypotheses, We focused on the transcript levels of the following marker genes: for mitochondrial fission and fusion - <em>mfn</em>2 (encoding mitofusin 2), <em>opa</em>1 (mitochondrial dynamin-like 120kDa protein), <em>dnm</em>1<em>l </em>(dynamin-1-like protein), <em>mff</em> (mitochondrial fission factor), <em>fis</em>1 (mitochondrial fission protein 1); for protein and DNA quality control - <em>tsfm</em> (encoding mitochondrial translation elongation factor Ts), <em>lonp</em>1 (mitochondrial Lon protease), <em>spg</em>7 (paraplegin), <em>oma</em>1 (mitochondrial metalloendopeptidase OMA1), <em>clpB</em> (mitochondrial caseinolytic matrix peptidase chaperone subunit B), <em>atp</em>23 (mitochondrial inner membrane protease ATP23), <em>twnk</em> (mitochondrial twinkle mtDNA helicase); and for mitophagy - <em>mieap</em> (encoding mitochondrial eating protein), <em>hyou</em>1 (hypoxia upregulated protein 1), <em>prkn</em> (parkin), <em>pink</em>1 (PTEN- induced kinase 1), and <em>pgam</em>5 (mitochondrial serine/threonine protein phosphatase PGAM5). The revealed species-specific differences in the expression of the mitochondrial quality control pathways shed light on the potentially important mechanisms of mitochondrial protection against H/R-induced damage that might contribute to hypoxia tolerance in marine bivalves. </p>
Effect of shade and nitrogen content on Arabidopsis Col-0 and cytokinin mutants abcg14 and cypDM mRNA-seq gene expression processed tables.
<p>This dataset is an add-on for Gautrat et al., containing processed files for the mRNAseq data in tab delimited txt format.</p> <p>Here, you can obtain the raw counts file, the normalized CPM values, and the normalized logCPM values</p> <p>The RNA-seq raw data supporting the conclusions of this article have been deposited in ArrayExpress (Kolesnikov et al., 2015) at EMBL-EBI (www.ebi.ac.uk/arrayexpress), under accession numbers E-MTAB-13638.</p> <p>All relevant custom r scripts are available at https://github.com/aromanowski/shade_N_ck</p>
Code for generating figures and analyzing amplicon sequencing of human mRNA and reporter mRNA targeted with type III-A CRISPR complex from Streptococcus thermophiles
<p>This dataset contains code for analyzing amplicon sequencing data and generating figures in the manuscript by Anna Nemudraia, Artem Nemudryi, and Blake Wiedenheft (2024), "Repair of CRISPR-guided RNA breaks enables site-specific RNA excision in human cells." </p> <p>Amplicon sequencing data has been deposited to NCBI Sequence Read Archive (SRA) under BioProject PRJNA1099688. The description of read files deposited to SRA can be found in the spreadsheet ./code_for_sequencing_data_analysis/SRA_read_files_description.xlsx</p> <p>The code for analyzing amplicon sequencing data can be found in the archive "code_for_sequencing_data_analysis.tar.gz." Output files from this analysis were used to generate figures. Figures were generated using the ggplot2 package in R and finalized in CorelDRAW.</p> <p>Code for generating figures can be found in the archive "code_for_generating_figures.tar.gz". </p> <p>Any questions or requests regarding the data or the code should be addressed to Dr. Artem Nemudryi at artem.nemudryi@gmail.com.</p>
Pre-vaccination and early B cell signatures of the antibody response to SARS-CoV-2 mRNA vaccine
<p>The data presented in Code repository for Kardava, L., Rachmaninoff, N., Lau, W. W., Buckner, C. M., Trihemasava, K., Blazkova, J., ... & Moir, S. (2022). Early human B cell signatures of the primary antibody response to mRNA vaccination. Proceedings of the National Academy of Sciences, 119(28), e2204607119.<a href="https://www.pnas.org/doi/epdf/10.1073/pnas.2204607119">https://www.pnas.org/doi/epdf/10.1073/pnas.2204607119</a> are made available here.</p> <p>All code to reproduce the figures can be found here: https://github.com/niaid/COVID_Vaccine_Bcells</p> <p><a href="https://zenodo.org/api/files/f93859d0-b062-4def-8b17-c0f21ee36f09/all_subjects_cd19_positive_and_keys.zip">all_subjects_cd19_positive_and_keys.zip</a> contains a CSV file of all CD19+ cells with flowSOM clusters shown. Accompanying files allow for matching of timepoint and subject information.</p> <p><a href="https://zenodo.org/api/files/f93859d0-b062-4def-8b17-c0f21ee36f09/All_subjects_FCS_files_deidentified.zip">All_subjects_FCS_files_deidentified.zip</a> contains the raw fcs files and is organized by timepoint and subject.</p>
UPF3A and UPF3B are redundant and modular activators of nonsense-mediated mRNA decay in human cells
<p>Source data for the publication: UPF3A and UPF3B are redundant and modular activators of nonsense-mediated mRNA decay in human cells.<br> Includes raw image data (e.g. agarose gels, western blots, northern blots), quantifications, qPCR raw Ct values and other supporting material.</p>
SHAPE datasets from manuscript: Modulation of pre-mRNA structure by hnRNP proteins regulates alternative splicing of MALT1
<p>Normalized SHAPE reactivity for the MALT1 M1 minigene RNA constructs (wildtype, variant 1, and variant 2) reported in the manuscript titled 'Modulation of pre-mRNA structure by hnRNP proteins regulates alternative splicing of <em>MALT1'</em> .</p> <p> </p>
Germinal centre-driven maturation of B cell response to SARS-CoV-2 mRNA vaccination
<p>These are the<strong> processed</strong> BCR repertoire and transcriptomics data described in <a href="https://doi.org/10.1038/s41586-022-04527-1">Kim & Zhou et al., <em>Nature</em>, 2022</a>. The <strong>raw</strong> sequencing data new to this study are available on SRA under BioProject <a href="https://www.ncbi.nlm.nih.gov/sra/?term=PRJNA777934">PRJNA777934</a>. This study also used BCR repertoire data from <a href="https://doi.org/10.1038/s41586-021-03738-2">Turner & O'Halloran et al., <em>Nature</em>, 2021</a> (<a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA731610">PRJNA731610</a>) and <a href="https://doi.org/10.1016/j.immuni.2021.08.013">Schmitz, Turner & Liu et al., <em>Immunity</em>, 2021</a> (<a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA741267">PRJNA741267</a>).</p> <p> </p> <p><strong>Code</strong></p> <p>Code along with Docker containers for reproducing the NGS data-based figures and analyses in the published paper can be <a href="https://github.com/julianqz/wustl_published/tree/main/nature_2022">found on GitHub</a>.</p> <p> </p> <p><strong>Metadata</strong></p> <p>File: WU368_kim_et_al_nature_2022_meta.tsv</p> <p>Notes:</p> <ul> <li>Sample breakdown by `sequence_type` (132 total) <ul> <li>73 bulk BCR sequencing samples (`bulk`) <ul> <li>57 new</li> <li>5 from <a href="https://doi.org/10.1038/s41586-021-03738-2">Turner & O'Halloran et al., <em>Nature</em>, 2021</a></li> <li>11 from <a href="https://doi.org/10.1016/j.immuni.2021.08.013">Schmitz, Turner & Liu et al., <em>Immunity</em>, 2021</a>.</li> </ul> </li> <li>56 10x Genomics single-cell VDJ + 5' gene expression samples (`tgx`)</li> <li>3 samples from <a href="https://doi.org/10.1038/s41586-021-03738-2">Turner & O'Halloran et al., <em>Nature</em>, 2021</a> (`mab`) corresponding to a total of 37 S-binding mAbs previously reported. These are not the same as the 2099 recombinant mAbs generated in this study (see below).</li> </ul> </li> <li>Sample collection time was originally recorded in days in the `timepoint` column. Timepoints were referenced in weeks in the manuscript, as shown in the `timepoint_ms` column.</li> <li>`bio_rep` and `tech_rep` = biological replicate and technical replicate respectively.</li> </ul> <p>Abbreviations:</p> <ul> <li>LN = lymph node</li> <li>BM = bone marrow</li> <li>PB = plasmablast</li> <li>GC = germinal centre</li> <li>LLPC = long-lived plasma cell</li> <li>NS = no sorting</li> <li>mAb = monoclonal antibody</li> </ul> <p> </p> <p><strong>Information on the 2099 recombinant mAbs generated in this study</strong></p> <p>File: WU368_kim_et_al_nature_2022_mabs.tsv</p> <p>Notes on columns:</p> <ul> <li>`h_sequence_id` and `l_sequence_id`: Sequence IDs of the heavy and light chains respectively.</li> <li>`elisa`: ELISA results for binding to SARS-CoV-2 S (`TRUE` = positive).</li> </ul> <p> </p> <p><strong>Processed BCR data - heavy chains</strong></p> <p>File: WU368_kim_et_al_nature_2022_bcr_heavy.tsv</p> <p><em>Analysis was based on heavy chain-based clonal inference.</em></p> <p>Notes on columns:</p> <p>The columns largely follow the <a href="https://changeo.readthedocs.io/en/stable/standard.html">AIRR-C Rearrangement format</a>. The main deviation is that CDR3s were used, as opposed to IMGT-defined "junctions". Nonetheless, junction-related columns are included here as some repositories such as <a href="https://gateway.ireceptor.org/login"><em>iReceptor</em></a> use these. Non-standard columns are noted below.</p> <ul> </ul> <ul> <li>`cell_id`: Only sequences from single-cell samples and the 37 mAbs from Turner & O'Halloran et al., <em>Nature</em>, 2021 have cell IDs following the format `[donor]_[sample]@[id]`. `NA` for bulk sequences.</li> <li>`sequence_id`: Sequence IDs follow the format `[donor]_[sample]@[id]`.</li> <li>`v_call_genotyped`: V gene annotation reassigned after individualized genotyping by <a href="https://tigger.readthedocs.io/en/stable/">TIgGER</a>.</li> <li>`germline_[vdj]_call`: Clonal consensus germline calls after corresponding clonal consensus sequence were reconstructed via <a href="https://changeo.readthedocs.io/en/stable/methods/germlines.html">`CreateGermlines.py --cloned` from Change-O</a>.</li> <li>`isotype`: IGH[ADEGM].</li> <li>`cdr3`: CDR3 nucleotide sequence.</li> <li>`cdr3_length`: CDR3 nucleotide sequence length.</li> <li>`cdr3_aa`: CDR3 amino acid sequence.</li> <li>`collapse_count`: Number of duplicate IMGT-aligned V(D)J sequences that were collapsed by <a href="https://alakazam.readthedocs.io/en/stable/topics/collapseDuplicates/">`alakazam::collapseDuplicates`</a>.</li> <li>`donor`, `timepoint`, `tissue`, `sorting`, `seq_type`: Propagated as is from the metadata file. <ul> <li>In `seq_type`, `tgx` corresponds to 10x Genomics data; `mab` corresponds specifically to the 37 S-binding mAbs from Turner & O'Halloran et al., <em>Nature</em>, 2021.</li> </ul> </li> <li>`timepoint_2`: Same as `timepoint`, except that `d28+d35` and `d201+d208` were treated as `d28` (week 4) and `d201` (week 29) respectively as described in Materials & Methods.</li> <li>`gex_anno`: Cell type identity annotation based on transcriptomic profiles. Mapped from `anno_leiden_0.18` from WU368_kim_et_al_nature_2022_gex_b_cells.h5ad.</li> <li>`compartment`: B cell compartment. <ul> <li>ABC = activated B cell. LNPC = lymph node plasma cell. RMB = resting memory B cell.</li> <li>Minor differences in terminology <ul> <li>The manuscript refers to the memory compartment as MBCs, whereas the terminology used in the data is RMB. As described in Materials & Methods, analysis involving the memory compartment used specifically d201 bulk-sequenced memory sorts from blood. To get these sequences, subset `s_pos_clone`, `seq_type`, `compartment`, and `timepoint_2` to, respectively, `TRUE`, `bulk`, `RMB`, and `d201`. </li> <li>The manuscript uses the term BMPC (bone marrow plasma cell), whereas the data uses the term LLPC.</li> </ul> </li> </ul> </li> <li>`clone_id`: B cell clonal lineage IDs follow the format `[donor]@[id]`.</li> <li>`s_pos_clone`: `TRUE` if a sequence belonged to a B cell clone that was designated as S-binding by virtue of containing one of the recombinant mAbs that tested positive via ELISA or one of the S-binding mAbs from Turner & O'Halloran et al., <em>Nature</em>, 2021.</li> <li>`expressed_id`: mAb IDs for the 2099 recombinant mAbs generated in this study (mapped from `mab_id` from WU368_kim_et_al_nature_2022_mabs.tsv) and the 37 mAbs from Turner & O'Halloran et al., <em>Nature</em>, 2021. `NA` for everything else.</li> <li>`elisa`: ELISA results for binding of recombinant mAbs to SARS-CoV-2 S. `TRUE` if positive. `NA` if not tested.</li> <li>`nuc_RS_19_312`: number of replacement and silent mutations between IMGT-numbered nucleotide positions 19-312 along IGHV sequences, calculated by <a href="https://shazam.readthedocs.io/en/stable/topics/calcObservedMutations/">`shazam::calcObservedMutations`</a>.</li> <li>`nuc_denom_19_312`: number of informative nucleotide positions for counting mutations, excluding non-A/T/G/C positions (such as "N", "-", ".").</li> <li>`nuc_RS_freq_19_312`: nucleotide-level mutation frequency (= nuc_RS_19_312 / nuc_denom_19_312).</li> </ul> <p> </p> <p><strong>Processed BCR data - light chains</strong></p> <p>File: WU368_kim_et_al_nature_2022_bcr_light.tsv</p> <p><em>Light chains were not used for heavy chain-based clonal inference or analysis.</em></p> <p> </p> <p><strong>Processed transcriptomics data</strong></p> <p>Files:</p> <ul> <li>WU368_kim_et_al_nature_2022_gex_all_cells.h5ad (clustering all cells)</li> <li>WU368_kim_et_al_nature_2022_gex_b_cells.h5ad (re-clustering only the B cells)</li> </ul> <p>Notes:</p> <ul> <li>The `h5ad` files can be imported into <a href="https://scanpy.readthedocs.io/en/stable/index.html">Scanpy</a> as an <a href="https://scanpy.readthedocs.io/en/stable/usage-principles.html#anndata">AnnData object</a>.</li> <li>Each `AnnData` object has 3 `.layers`, each representing a version of the count matrix. <ul> <li>`raw_counts`: Imported from `<a href="https://support.10xgenomics.com/single-cell-gene-expression/software/pipelines/6.0/using/aggregate">cellranger aggr</a>` output by `scanpy.read_10x_mtx`.</li> <li>`log_norm`: Log-noramlized expression values outputted by `scanpy.pp.normalize_total` followed by `scanpy.pp.log1p`.</li> <li>`scaled`: The `log_norm` layer scaled to unit variance and zero mean by `scanpy.pp.scale`. </li> </ul> </li> <li>The `gene_name` and `biotype` columns in `.var` were extracted from GENCODE v32 GTF.</li> <li>Columns in `.obs` (each row corresponds to a cell) <ul> <li>`n_feature`: The `n_genes_by_counts` column produced by `scanpy.pp.calculate_qc_metrics`, renamed. The number of genes expressed. This is before subsetting the genes.</li> <li>`n_umi`: The `total_counts` column produced by `scanpy.pp.calculate_qc_metrics`, renamed. The total UMI counts in a cell.</li> <li>`pct_mt`: The `pct_counts_mt` column produced by `scanpy.pp.calculate_qc_metrics`, renamed. The percentage of counts in mitochondrial genes.</li> <li>`n_hkg`: The number of housekeeping genes for which expression was detected.</li> <li>`n_gene_expressed`: The total number of genes for which expression was detected. This is after subsetting the genes.</li> <li>`pre_qc_bcr`: `TRUE` if a cell also had paired BCR data available. Produced by cross-referencing the cellular barcodes in `cell_barcodes.json` outputted by `cellranger vdj`. At this point the BCR data had not gone through the QC process in the BCR processing pipeline (hence `pre_qc`). </li> <li>`leiden_[resolution]`: Cluster assignment by `scanpy.tl.leiden`.</li> <li>`anno_leiden_[resolution]`: Cell type identity annotations based on transcriptomic profiles. This was mapped onto the `gex_anno` column in the processed heavy chain BCR data.</li> </ul> </li> <li>UMAP coordinates can be found in `.obsm["X_umap"]`.</li> <li>`.X` has been set to `None` in order to reduce file size.</li> </ul> <p>In addition, the preprocessed count matrix outputted by `<a href="https://support.10xgenomics.com/single-cell-gene-expression/software/pipelines/6.0/using/aggregate">cellranger aggr</a>` is available from <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE195673">GEO under BioProject PRJNA777934</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.