Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
36,856
datasets available to search
ShareScore release 0.9.0
Dataset results
36,856 results for “rna”
Protein-RNA complex simulation
<p>Simulation trajectory of TSEN/pre-tRNAArgTCT generated with GROMACS. </p> <p> </p> <p>A hybrid model of truncated TSEN/pre-tRNAArgTCT with the additional TSEN2 domain from AlphaFold Jumper et al, 2021</p> <p>and in silico modeled intron bases 37 to 43 was subjected to all-atom molecular dynamics simulations</p> <p>for assessment of flexibility. Simulation trajectory can be visualised with PyMOL or VMD.</p> <p> </p> <p>The protein was described by the AMBER14ff (Maier, et al 2015), and the RNA with the OL3 force field (Zgarbova et al, 2011)</p> <p>and the TIP3P water model was employed (Jorgensen et al, 1983).</p>
Structural Interaction Fingerprints and Machine Learning for predicting and explaining binding of small molecule ligands to RNA: a benchmark dataset
<p><b>Structural Interaction Fingerprints and Machine Learning for predicting and explaining binding of small molecule ligands to RNA: a benchmark dataset.</b></p><p>Ribonucleic acids (RNA) play crucial roles in living organisms as they are involved in key processes necessary for proper cell functioning. Some RNA molecules, such as bacterial ribosomes and precursor messenger RNA, are targets of small molecule drugs, while others, e.g., bacterial riboswitches or viral RNA motifs are considered as potential therapeutic targets. Thus, the continuous discovery of new functional RNA increases the demand for developing compounds targeting them and for methods for analyzing RNA—small molecule interactions. We recently developed fingeRNAt - a software for detecting non-covalent bonds formed within complexes of nucleic acids with different types of ligands. The program detects several non-covalent interactions, such as hydrogen and halogen bonds, ionic, Pi, inorganic ion- and water-mediated, lipophilic interactions, and encodes them as computational-friendly Structural Interaction Fingerprint (SIFt). Here we present the application of SIFts accompanied by machine learning methods for binding prediction of small molecules to RNA targets. We show that SIFt-based models outperform the classic, general-purpose scoring functions in virtual screening. We discuss the aid offered by Explainable Artificial Intelligence in the analysis of the binding prediction models, elucidating the decision-making process, and deciphering molecular recognition processes.</p>
Regulation of mature mRNA levels by RNA processing efficiency
<p>Data from the research paper "Regulation of mature mRNA levels by RNA processing efficiency" by Henfrey, C., Murphy, S., and Tellier, M.:</p> <p>-Highest expressed transcript annotation for protein-coding genes, gencode v38.</p> <p>-15 txt (BED) files for chromatin vs nucleoplasm enrichment gene sets: HeLa full gene sets, canonical protein only sets, chromatin RNA seq subsamples, mNET seq subsamples, Raji gene sets.</p> <p>-Proteomics data table</p> <p>-mRNA half-life table</p> <p>-Splicing efficiency for POINT-seq, ChrRNA-seq, NucRNA-seq table</p> <p>-Ser2-P mNET-seq readthrough index data table</p> <p>-Splicing efficiency for siLuc/siEX3 table (ChrRNA-seq, NucRNA-seq)</p> <p>-RMATs output tables for alternative splicing results (siEX3 vs siLuc)</p> <p>-Tables for TSS:TES quantifications (mNET-seq(CTD) vs log2FoldChange, chr/nuc/mnet siEX3 vs siLuc)</p>
Alterations in RNA editing in skeletal muscle following exercise training in individuals with Parkinson's disease
<p>Parkinson’s Disease (PD) is the second most common neurodegenerative disease behind Alzheimer’s Disease, currently affecting more than 10 million people worldwide. The progression of PD results in the loss of function due to neurodegeneration and neuroinflammation. The etiology of PD is multifactorial, including both genetic and environmental origins. We explored changes in RNA editing, specifically editing through the actions of the Adenosine Deaminases Acting on RNA (ADARs), in the progression of PD. Analysis of ADAR editing of skeletal muscle transcriptomes from PD patients and controls, including those that engaged in a rehabilitative exercise training program revealed significant differences in ADAR editing patterns based on age, disease status, and following rehabilitative exercise. Further, deleterious editing events in protein coding regions were identified in multiple genes with known associations to PD pathogenesis. Our findings of differential ADAR editing complement findings of changes in transcriptional network identified by a recent Lavin et al. 2020 (<a href="https://doi.org/10.3389/fphys.2020.00653">https://doi.org/10.3389/fphys.2020.00653)</a> study and offer insights into dynamic ADAR editing changes associated with PD pathogenesis. VCF files were generated using AIDD (Plonski et al., 2020) (<a href="https://doi.org/10.1186/s12859-020-03888-6">https://doi.org/10.1186/s12859-020-03888-6</a>).</p>
Datasets, reproducible codes, and results for evaluating differential expression analysis methods on population-level RNA-seq data
<p>This upload contains the necessary R codes and data to reproduce the FDR and Power results described in our correspondence "Neglecting normalization impact in semi-synthetic RNA-seq data simulation generates artificial false positives" to Li Y, Ge X, Peng F, Li W, Li JJ, Exaggerated false positives by popular differential expression methods when analyzing human population samples, <em>Genome Biology</em> 23, 79, 2022, DOI: 10.1186/s13059-022-02648-4.</p>
Joint embedding of vertebrate brain single-cell RNA-Seq using sequence or structure
<p>Embeddings of single-cell RNA-Seq data from three adult vertebrate brain datasets into Orthogroup feature space or Structural cluster feature space. Orthogroups were generated using OrthoFinder v5.5.0; Structural clusters were assigned by using FoldSeek to cluster AlphaFold-v4 structural predictions.<br> <br> The three datasets used as the basis for these embeddings were:</p> <ul> <li>sample <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM3768152">"Brain8"</a> from the <a href="https://www.frontiersin.org/articles/10.3389/fcell.2021.743421/full">Jiang et al. 2021</a> zebrafish cell atlas (files beginning with GSM3768152)</li> <li>sample <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM2906405">"Brain1"</a> from the <a href="https://www.sciencedirect.com/science/article/pii/S0092867418301168#sec4">Han et al. 2018</a> mouse cell atlas (files beginning with GSM2906405)</li> <li>sample <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM6214268">"Xenopus_brain_COL65"</a> from the <a href="https://www.nature.com/articles/s41467-022-31949-2">Liao et al. 2022</a> Xenopus laevis adult cell atlas (files beginning with GSM6214268)</li> </ul> <p>For each dataset, we also generated a standardized cell type annotation file based on the author's originally provided cell type annotation data. The first column is the cell barcode for that species and the second column is the original study's cell type annotation for that cell.</p> <p>For the Xenopus brain data, we removed around ~18k cells that were not annotated in the original data to simplify data analyses - these are reflected in the files with the "subsampled" suffix. Subsampled versions of the data are also available for the joint embedding space (prefixed with "DrerMmusXlae").</p> <p>For the final datasets used in our analyses, we also provide features x cell matrices as .h5ad files for smaller file sizes and faster loading using Scanpy. </p> <p>For visualizing our UMAP plots of our top200 embedding space, we provide ".tsv" files with a variety of metrics and the x and y positions of each cell in the UMAP. See "DrerMmusXlae_adultbrain_FoldSeek_plotlydata.tsv" and "DrerMmusXlae_adultbrain_OrthoFinder_plotlydata.tsv"</p> <p>These data are part of the Arcadia Science Pub titled <a href="https://doi.org/10.57844/arcadia-vw5e-2670">"Comparing gene expression across species based on protein structure instead of sequence"</a>.</p>
RNA-Seq data from: Hox genes modulate physical forces to differentially shape small and large intestinal epithelia
<p>Hox genes are highly conserved, master regulators of spatial patterning in the embryo, but how these factors trigger regional morphogenesis has largely remained a mystery. In the developing gut, Hox genes help demarcate identities of the small and large intestines early in embryogenesis, which ultimately leads to their specialization in both form and function. While the midgut forms villi, the hindgut develops flat, brain-like sulci that resolve into heterogeneous outgrowths. Combining mechanical measurements and mathematical modeling, we demonstrate that the posterior Hox gene Hoxd13 regulates biophysical phenomena that shape the hindgut lumen. We further show that Hoxd13 acts through the TGFβ pathway to thicken, stiffen, and promote isotropic growth of the subepithelial mesenchyme; together, these features lead to hindgut surface buckling. TGFβ, in turn, promotes collagen deposition to affect mesenchymal geometry and growth. We thus identify a cascade of events downstream of positional genetic identity that direct posterior intestinal morphogenesis. </p> <p>To identify genes and pathways that are directly or indirectly regulated by Hoxd13 to affect posterior gut morphogenesis in the chick, we compared mesodermal transcriptomes of wild-type midgut and hindgut intestinal samples, as well as mesodermal samples from a Hoxd13-overexpressing midgut at E12 and E14. Tissues were dissected and endoderm layers were removed manually before RNA extraction and downstream processing. Unbiased clustering was used to identify genes commonly differentially expressed in the hindgut and Hoxd13-misexpressing midgut. This submission contains bulk RNA-seq raw data (fastq.bz2 files) and processed .txt files with read counts. Experiment information is provided in .xlsx Metadata file used for NCBI GEO submission.</p>
High-quality RNA residues: RNA2023
<p>Introduction<br> --------------------------------------------------------------------------------<br> This is the RNA2023 dataset by the Richardson Lab at Duke University</p> <p>These are high-quality residues from high-quality, low-redundancy RNA chains in the PDB.</p> <p>For a similar set of quality-filtered protein residues, see the top2018 datasets at:<br> <a href="https://doi.org/10.5281/zenodo.4626149">https://doi.org/10.5281/zenodo.4626149</a><br> <a href="https://doi.org/10.5281/zenodo.5115232">https://doi.org/10.5281/zenodo.5115232</a></p> <p> </p> <p>Corresponding authors<br> --------------------------------------------------------------------------------<br> dcrjsr at kinemage.biochem.duke.edu<br> christopher.sci.williams at gmail.com</p> <p><br> Usage recommendations<br> --------------------------------------------------------------------------------<br> RNA residues that fail the filtering criteria described below have been removed from the files. As a result, these files can be considered pre-filtered and will return only results for residues of good model quality with supporting experimental data.</p> <p>Files already contain hydrogens added by Reduce in the context of the original full models.</p> <p>Two datasets are provided. The standard dataset is rna2023_pruned. We recommend this version as the default. The RNA backbone conformational space is highly diverse, and some real conformations fall below the statistical threshold for recognition as suites. Therefore we do not recommend excluding suite outliers from the dataset except in specialty cases. We also provide a rna2023_nosuiteout dataset. In this case, no residues having "!!" outlier suite identifications are permitted. This set may be useful in specialist cases where suite outliers are undesireable or where losing some real conformations is an acceptable sacrifice for maximal filtering.</p> <p>Each dataset also has a mmCIF version.</p> <p>Note: Chains are named based on author chain ids, except for 8b0x, chain a. To avoid conflicts with 8b0x chain A in file systems that do not support case-sensitive file names, 8b0x chain a has been renamed to chain AB, matching its PDB/mmCIF designation.</p> <p><br> Additional files<br> --------------------------------------------------------------------------------<br> rna2023_pdbmetadata.csv contains information on release date, resolution, title, authors, etc for each source pdb.</p> <p>rna2023_chain_list contains a list of all included chains, plus statistics on the number residues from the original chain passed the quality filters.</p> <p>rna2023_suitename_table.csv and rna2023_suitename_table_nosuiteout.csv contain suitename identifications of rotameric RNA backbone conformations (1a, 1c, 2u, 6d, etc) precomputed for convenience.</p> <p><br> Filtering criteria: Chain level<br> --------------------------------------------------------------------------------<br> The chain list was derived from http://rna.bgsu.edu/rna3dhub/nrlist, version 3.150 as of 2020/10/28, with a 1.9Å resolution cutoff.</p> <p>We added 6ugg chain A and two recent EM ribosome structures: 8a3d and 8b0x</p> <p>After residue-level filtering, chains with no complete suites were removed.</p> <p><br> Filtering criteria: Residue level<br> --------------------------------------------------------------------------------<br> Even excellent structures usually contain some poorly-resolved regions. Residue-level filtering helps avoid including these regions in otherwise high-quality data</p> <p>Residues are required to meet the following validation quality contain:<br> No sugar pucker outliers<br> No steric overlaps or "clashes", as per Probe >= 0.5Å<br> No covalent bond or angle geometry outliers<br> Optionally, no !! suite outliers</p> <p>Residues from xray structures are required for meet the following fit-to-map criteria:<br> Average of worst 2 atoms' 2Fo-Fc map values >= 1.2<br> Average of worst 2 atoms' RSCC scores >= 0.7<br> No atoms modeled at partial occupancy</p> <p>Residues from em structures are required for meet the following fit-to-map criteria:<br> RSCC >= 0.7<br> Residue inclusion fraction = 1.0 or >= 0.95, depending on structure<br> No atoms modeled at partial occupancy</p> <p>Filtering is documented in each pruned file. See USER DOC lines in .pdb and data_rna2023_dataset loops in .cif</p> <p><br> Version history<br> --------------------------------------------------------------------------------<br> Version 1.0 Jun 30, 2023<br> Initial version</p>
Supplements for "Understanding the interaction between a human transferrin receptor aptamer-short double stranded RNA conjugate and its cell membrane target by in silico methods"
<p>Supplements for "Understanding the interaction between a human transferrin receptor aptamer-short double stranded RNA conjugate and its cell membrane target by in silico methods". </p> <p>This supplement includes the following files:</p> <p> </p> <p>1. Structures of the most stable Protein-Aptamer complexes predicted from HADDOCK</p> <p>Haddock_Cluster1.pdb <br> Haddock_Cluster2.pdb <br> Haddock_Cluster3.pdb </p> <p>2. Structure of the most stable conformation aligned with Protein-transferring complex PDB</p> <p>cluster1_aligned.pdb <br> transferrin_aligned.pdb </p> <p>3. MM-GBSA decomposition analysis of the three replicas for Protein-Aptamer</p> <p>aptamer_new_rep01_Decomp.dat <br> aptamer_new_rep02_Decomp.dat <br> aptamer_new_rep03_Decomp.dat </p> <p>4. MM-GBSA decomposition analysis of the three replicas for Protein-Aptamer-Conjugate<br> conjugate_new_rep01_Decomp.dat <br> conjugate_new_rep02_Decomp.dat <br> conjugate_new_rep03_Decomp.dat <br> </p> <p><br> <br> </p>
Reads-per-UMI tables across single-cell RNA sequencing protocols
<p>Data analyzed in <a href="https://www.biorxiv.org/content/10.1101/2023.08.02.551637v1">Lause, Ziegenhain et al. (2023)</a>.</p> <p>Code to obtain these tables from public data sources is available on <a href="https://github.com/berenslab/read-normalization">github</a>.</p> <p> </p> <p>Each row in the table is a UMI-tag detected in a certain cell (column RG) attached to a molecule from a specific gene (column GE) with a certain barcode (column UB). Column N gives the number of times the UMI was detected for that gene and cell.</p> <p>Data sources and protocols are given with the respective file names below.</p> <p><strong>Johnsson2022_Smartseq3_PE.hd1.txt.gz</strong>: Mouse fibroblasts profiled with <strong>Smart-seq3</strong> paired-end; accession E-MTAB-10148, sample plate2,<br> <a href="https://doi.org/10.1038/s41588-022-01014-1">Paper</a><br> <br> <strong>Hagemann-Jensen2020_Smartseq3_SE.hd1.txt.gz: </strong>Mouse fibroblasts profiled with <strong>Smart-seq3</strong> single-end; accession E-MTAB-8735, sample Smartseq3.Fibroblasts.smFISH<br> <a href="https://doi.org/10.1038/s41587-020-0497-0">Paper</a><br> <br> <strong>Hagemann-Jensen2022_Smartseq3xpress.hd1.txt.gz: </strong>HEK293 cells profiled with <strong>Smart-seq3Xpress</strong>; accession E-MTAB-11467.<br> <a href="https://www.biorxiv.org/content/10.1101/2021.07.10.451889v1">Paper</a><br> <br> <strong>Ziegenhain2017.hd1.txt.gz: </strong>Mouse embryonic stem cells profiled by <strong>CEL-seq2, Drop-seq, MARS-seq, </strong>and<strong> SCRB-seq</strong>; GEO accession GSE75790<br> <a href="https://doi.org/10.1016/j.molcel.2017.01.023">Paper</a></p>
CRISPR-based engineering of RNA viruses
<p>CRISPR RNA-guided endonucleases have enabled precise editing of DNA. However, options for editing RNA remain limited. Here, we combine sequence-specific RNA cleavage by CRISPR ribonucleases with programmable RNA repair to make precise deletions and insertions in RNA. This work establishes a new recombinant RNA technology with immediate applications for the facile engineering of RNA viruses.</p> <p> </p> <p>This dataset contains code for analyzing sequencing data and generating figures in the manuscript.</p>
Simultaneous estimation of gene regulatory network structure and RNA kinetics from single cell gene expression
<p>Supplemental Data 1 is single-cell response to rapamycin count data first sequenced in this work and deposited in GEO with accession GSE242556. It is a 173348 rows × 5847 columns TSV.GZ file where the first row is a header, the first 5843 columns are integer gene counts, and the final 4 columns ('Gene', 'Replicate', 'Pool', and 'Experiment') are cell-specific metadata.</p> <p>Supplemental Data 2 is bulk response to rapamycin count data first sequenced in this work. It is a 33 rows × 5847 columns TSV.GZ file where the first row is a header, the first 5843 columns are integer gene counts, and the final 4 columns ('Oligo', 'Time', 'Replicate', and 'Sample_barcode') are sample-specific metadata.</p> <p>Supplemental Data 3 is single-cell count data published as GSE125162 and re-analyzed with the pipeline used for single-cell quantification in this work. It is a 65068 rows × 5850 columns TSV.GZ file where the first row is a header, the first 5843 columns are integer gene counts, and the final 7 columns ('Condition', 'Sample', 'Genotype_Group', 'Genotype_Individual', 'Genotype', 'Replicate', 'Cell_Barcode') are cell-specific metadata.</p> <p>Supplemental Data 4 is the four deep learning models trained in this work. It is a TAR.GZ file containing the final biophysical transcription/decay model, the pre-trained decay model, the velocity prediction model, and the count prediction model. Each model file is an h5 file containing a pytorch model that can be loaded with supirfactor\_dynamical.read().</p> <p>Supplemental Data 5 is the prior knowledge network used to constrain the models for TF interpretability. It is a 1574 rows × 204 columns [Genes x TFs] TSV.GZ file where the first row is a header with TF names, the first column is an index of gene names, and TF-gene interactions are indicated by non-zero values in the matrix. There are 2799 TF-gene interactions.</p> <p><br> Supplemental Table 6 is the oligonucleotide sequences used in this work. It is a TSV file with a header row.</p> <p>Supplemental Table 7 is the yeast strains used in this work. It is a TSV file with a header row.</p> <p>Supplemental Table 8 is gene metadata used in this work (e.g. Ribosomal Protein gene labels, etc). It is a TSV file with a header row.</p> <p>Supplemental Table 9 is FY4/5 growth curve data generated in this work. It is a 20 rows × 7 columns TSV file where the first row is a header with replicate IDs, the first column is an index of times in minutes, and values are cell densities in YPD culture, in units of 10$^6$ cells / mL.</p> <p>Supplemental Data 10 is a TAR.GZ file containing the yeast SacCer3 genome, modified to add UTR sequences, that was used to generate transcripts for kallisto pseudoalignment in this work.</p>
Robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data
<p>Data used to test the robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data, described in <a href="https://doi.org/10.1186/s13059-020-1949-z">Holland et al. 2020</a>.</p> <p>The folder <em>data </em>contains<em> </em>raw data and the folder <em>output</em> contains intermediate and final results of all analyses. </p> <p>The associated analyses code and more information are available on <a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq">GitHub</a>.</p> <p> </p> <p><strong>Abstract</strong></p> <p><strong>Background</strong></p> <p>Many functional analysis tools have been developed to extract functional and mechanistic insight from bulk transcriptome data. With the advent of single-cell RNA sequencing (scRNA-seq), it is in principle possible to do such an analysis for single cells. However, scRNA-seq data has characteristics such as drop-out events and low library sizes. It is thus not clear if functional TF and pathway analysis tools established for bulk sequencing can be applied to scRNA-seq in a meaningful way.</p> <p><strong>Results</strong></p> <p>To address this question, we perform benchmark studies on simulated and real scRNA-seq data. We include the bulk-RNA tools PROGENy, GO enrichment, and DoRothEA that estimate pathway and transcription factor (TF) activities, respectively, and compare them against the tools SCENIC/AUCell and metaVIPER, designed for scRNA-seq. For the in silico study, we simulate single cells from TF/pathway perturbation bulk RNA-seq experiments. We complement the simulated data with real scRNA-seq data upon CRISPR-mediated knock-out. Our benchmarks on simulated and real data reveal comparable performance to the original bulk data. Additionally, we show that the TF and pathway activities preserve cell type-specific variability by analyzing a mixture sample sequenced with 13 scRNA-seq protocols. We also provide the benchmark data for further use by the community.</p> <p><strong>Conclusions</strong></p> <p>Our analyses suggest that bulk-based functional analysis tools that use manually curated footprint gene sets can be applied to scRNA-seq data, partially outperforming dedicated single-cell tools. Furthermore, we find that the performance of functional analysis tools is more sensitive to the gene sets than to the statistic used.</p> <p> </p> <p>For questions related to the data please write an email to christian.holland@bioquant.uni-heidelberg.de or use the <a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq/issues">GitHub issue system</a>.</p>
Bulk RNA-Seq PBMC data of SLE patients and healthy volunteers/ profiling of 29 individual immune cell types as well as PBMCs of healthy donors
<p>This Zenodo project contains processed gene expression data from two publicly available data sets. It includes the gene expression data of peripheral blood mononuclear cells (PBMCs) of systemic lupus erythematosus (SLE) patients as well as healthy volunteers (GSE122459). The project also comprises the bulk RNA-Seq profiling of 29 immune cell types as well as PBMCs of healthy individuals (GSE107011). In both cases, the raw RNA-Seq data was downloaded, aligned and processed. The gene expression data is available in form of a count matrix (GSE107011) or count matrix and transcript-per-million (TPM) values (GSE122459). For the latter, an annotation file is attached. Further details are provided in the information file. </p>
Chemical-genetic interrogation of RNA polymerase mutants reveals structure-function relationships and physiological tradeoffs
<p>The multi-subunit bacterial RNA polymerase (RNAP) and its associated regulators carry out transcription and integrate myriad regulatory signals. Numerous studies have interrogated the inner workings of RNAP, and mutations in genes encoding RNAP drive adaptation of <i>Escherichia coli</i> to many health- and industry-relevant environments, yet a paucity of systematic analyses has hampered our understanding of the fitness benefits and trade-offs from altering RNAP function. Here, we conduct a chemical-genetic analysis of a library of RNAP mutants. We discover phenotypes for non-essential insertions, show that clustering mutant phenotypes increases their predictive power for drawing functional inferences, and demonstrate that some RNA polymerase mutants both decrease average cell length and confer insensitivity to killing by cell-wall targeting antibiotics. Our findings demonstrate that RNAP chemical-genetic interactions provide a general platform for interrogating structure-function relationships <i>in vivo</i> and for identifying physiological trade-offs of mutations, including those relevant for disease and biotechnology. This strategy should have broad utility for illuminating the role of other important protein complexes.</p>
Data for "Tuning parameters of dimensionality reduction methods for single-cell RNA-seq analysis"
<p>The files named <code>df_scran.csv</code>, <code>df_seurat.csv</code>, <code>df_zinbwave.csv</code>, <code>df_dca.csv</code>, and <code>df_scvi.csv</code> contain one row per configuration that we ran successfully.</p> <p>The files named <code>DATASET.METHOD.h5ad</code> are encoded with anndata <code>v0.7.0</code> (be careful as they are not readable with previous versions) and contain 100 embeddings each. The embeddings are in the <code>obsm</code> attribute of the object. All the embeddings can be listed with the <code>obsm_keys()</code> method. The name of the embedding contains the parameters used to generate that embedding and are written like that <code>method=zinbwave.dims=10.epsilon=1000.features=300.gene_covariate=0</code>.</p> <p> </p> <p>For questions on this dataset please contact fraimundo@google.com</p>
Single cell RNA-seq Data - Dissecting the functional reprogramming of the microenvironment in bone marrow fibrosis at the single cell level
<p>We provide results regarding the bioinformatic analysis of scRNA-seq from distinct bone marrow fibrosis mouse models and human samples.</p> <p> </p> <p>These include:</p> <p>Robjects&Markdown - R markdown and R objects with QC statistics, UMAP and final scRNA-seq data sets.</p> <p>Markers - Excel tables with cluster specific marker genes.</p> <p>DE Genes - Excel table with DE genes when comparing cells in control vs. disease condition per cluster.</p> <p>GO Analysis - Gene enrichment analysis of either DE genes. These are divided by either UP or down regulated genes.</p>
EI Single-Cell RNA-Seq Workshop 2020
<p>Datasets to be used for the "Single-Cell RNA-Seq Workshop 2020" at the Earlham Institute, Norwich, UK.</p>
Supplementary materials for "Relative Information Gain: Shannon entropy-based measure of the relative structural conservation in RNA alignments"
<p>Supplementary materials for "Relative Information Gain: Shannon entropy-based measure of the relative structural conservation in RNA alignments". These include precalculated RNA Blocks, MBRs (Matrix of Bear encoded RNA), sPSSMs (structural Position Specific Scoring Matrix), RIG (Relative Information Gain) scores, and plots calculated for 3016 Rfam 14.1 families. In particular:</p> <ul> <li><strong>alignments.zip:</strong> zipped file containing the structural alignments for each Rfam family.</li> <li><strong>RNA_Blocks.zip</strong>: zipped file containing the RNA blocks used to derive different substitution matrices.</li> <li><strong>MBRs.zip</strong>: zipped file containing the substitution matrices.</li> <li><strong>sPSSMs.zip</strong>: zipped file containing the structural Position Specific Scoring Matrices.</li> <li><strong>RIGs.zip</strong>: zipped file containing the RIG scores.</li> <li><strong>entropy.zip</strong>: zipped file containing the (rescaled) entropy.</li> <li><strong>plots.zip</strong>: zipped file containing the plots. </li> </ul> <p>All the scripts to build all these files are available at <a href="https://github.com/helmercitterich-lab/RIG">https://github.com/helmercitterich-lab/RIG</a>.</p>
Beyondcell: targeting cancer therapeutic heterogeneity in single-cell RNA-seq
<p><strong><a href="https://gitlab.com/bu_cnio/Beyondcell">Beyondcell</a> </strong>is a methodology for the identification of drug vulnerabilities in single cell RNA-seq data. To this end, <strong>Beyondcell</strong> focuses on the analysis of drug-related commonalities between cells by classifying them into distinct therapeutic clusters. We have validated the tool in a population of MCF7-AA cells exposed to 500nM of bortezomib and collected at different time points: t0 (before treatment), t12, t48 and t96 (72h treatment followed by drug wash and 24h of recovery) obtained from <a href="https://www.nature.com/articles/s41586-018-0409-3"><strong><em>Ben-David U, et al., Nature, 2018</em></strong></a>. Here, you can find the integrated Seurat object obtained from this analysis. This object is meant to help users follow <strong>Beyondcell's</strong> <a href="https://gitlab.com/bu_cnio/Beyondcell/-/tree/master/tutorial/analysis_workflow">analysis workflow</a>.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.