Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

36,843

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

36,843 results for “RNA”

Learn how ShareScore rates datasets ↗
zenodo44/100

Characteristics of human and viral RNA binding sites and site clusters recognized by SRSF1 and RNPS1

<p>This dataset was developed for the following article:</p> <p>&nbsp;Rogan PK, Mucaki EJ and Shirley BC. A proposed molecular mechanism for pathogenesis of severe RNA-viral pulmonary infections [version 1; peer review: awaiting peer review].&nbsp;<em>F1000Research</em>&nbsp;2020,&nbsp;<strong>9</strong>:943 (<a href="https://doi.org/10.12688/f1000research.25390.1">https://doi.org/10.12688/f1000research.25390.1</a>)</p> <p><strong>Section 1. Extended Data Tables</strong></p> <p>This archive contains the extended data tables for the research article &quot;A proposed mechanism for molecular pathogenesis of severe RNA-viral pulmonary infections&quot;. These tables provide&nbsp;SRSF1, RNPS1 and hnRNP A1 binding site and information-dense cluster counts across various RNA viral genomes [including multiple SARS-CoV-2 and influenza strains] and the human transcriptome, the estimated SARS-CoV-2 doubling time necessary for viral genome SRSF1 binding site availability to exceed sites within the host transcriptome, and an analysis of influenza, dengue, and aplastic anemia patients misdiagnosed as irradiated by established radiation gene signatures.These tables are:</p> <p><strong>Section 1 - Table 1.</strong> RNPS1 and hnRNPA1 binding sites and Information-Dense Clusters for RNPS1 and<br> hnRNPA1 in RNA Virus Genomes<br> <strong>Section 1 - Table 2A.</strong> Detailed Analysis of Information-Dense Clusters for SRSF1 (Replicate 1) in RNA Virus<br> Genomes<br> <strong>Section 1 - Table 2B.</strong> Detailed Analysis of Information-Dense Clusters for SRSF1 (Replicate 2) in RNA Virus<br> Genomes<br> <strong>Section 1 - Table 2C.</strong> Detailed Analysis of Information-Dense Clusters for RNPS1 in RNA Virus Genomes<br> <strong>Section 1 - Table 2D.</strong> Detailed Analysis of Information-Dense Clusters for hnRNP A1 in RNA Virus<br> Genomes<br> <strong>Section 1 - Table 3.</strong>&nbsp;Binding Site Analysis of Multiple Coronavirus Strains (Both Strands)<br> <strong>Section 1 - Table 4A.</strong>&nbsp;Binding Site Analysis of Multiple Influenza A (H3N2) Strains (Negative Strand Only)<br> <strong>Section 1 - Table 4B.</strong>&nbsp;Binding Site Analysis of Multiple Influenza A (H3N2) Strains (Both Strands)<br> <strong>Section 1 - Table 5.</strong>&nbsp;SRSF1, RNPS1 and hnRNPA1 Binding Sites and Information-Dense Clusters by Gene<br> <strong>Section 1 - Table 6A.</strong> Transcriptome-Wide Information Dense Clusters Intersecting DRIP- and DRIPc-seq<br> Intervals<br> <strong>Section 1 - Table 6B.&nbsp;</strong>Exome-Wide Information Dense Clusters within DRIP- and DRIPc-seq Intervals<br> <strong>Section 1 - Table 6C.</strong>&nbsp;Transcriptome-Wide Scan of Strong Binding Sites Intersecting DRIP- and DRIPc-seq<br> Intervals<br> <strong>Section 1 - Table 6D.&nbsp;</strong>Exome-Wide Scan of Strong Binding Sites within DRIP- and DRIPc-seq Intervals<br> <strong>Section 1 - Table 7.</strong> Rate of False Positives for Influenza, Dengue Virus and Aplastic Anemia Using<br> Radiation Signatures<br> <strong>Section 1 - Table 8.</strong> Radiation Model Genes Contributing to False Positives for Patients with Influenza A,<br> Dengue Virus, and Aplastic Anemia<br> <strong>Section 1 - Table 9A.</strong>&nbsp;Doubling Time of SARS-CoV-2 Needed to Exceed Host Transcriptome SRSF1 Binding<br> Sites (Positive-Strand Sites Only)<br> <strong>Section 1 - Table 9B.</strong>&nbsp;Doubling Time of SARS-CoV-2 Needed to Exceed Host Transcriptome SRSF1 Binding<br> Sites (Both Strands Considered)</p> <p><strong>Section 2.&nbsp; All SRSF1, hnRNPA1 and RNPS1 binding site tracks for human and viral genomes</strong></p> <p>We provide bedgraph tracks which provide the location and strength of binding sites (and binding site clusters) for SRSF1, RNPS1 and hnRNPA1 across the human transcriptome (GRCh37), the human exome (including +/-300nt surrounding the exon; non-intergenic only), and for all viral genome investigated in this study (Coronavirus, Dengue, HIV-1 [two strains] and Influenza [two strains]). Note that if no clusters were found for a particular viral genome, a file for said genome will not be present in the Zenodo archive.</p> <p>Folder &ldquo;Cluster-to-DRIPseq-Intersection-Tracks&rdquo; contain tracks which indicate where binding site clusters have been identified, intersected with DRIP-seq and DRIPc-seq intervals which indicate where there is evidence of R-Loop formation in the human genome. The DRIP-seq dataset (GSE68845) is not strand specific. DRIPc-seq (GSE70189) is strand specific, and has been taken into account in the intersection (e.g. tracks only list positive strand clusters found in positive-strand DRIPc-seq intervals).</p> <p>Due to sheer size, the human transcriptome and exome tracks which indicate the location of individual binding sites are split into two separate files (separated by strand). While the custom tracks containing human binding site information are designed to be uploaded to the UCSC Genome Browser, files containing transcriptome-wide binding site information may be too large to be uploaded and may require further filtering (i.e. by chromosome).</p> <p>To be classified as a cluster, binding sites on the same strand must have <em>Ri</em> values which sum to &gt;50 bits, each binding site must have a neighboring site within 25nt, and all binding sites in the cluster must have <em>R<sub>i</sub></em> greater than a minimum bit threshold. For human transcriptomes and exomes, this bit minimum was set to <em>R<sub>sequence</sub></em>. The bit minimum for viral binding sites was set to 0.1 * <em>R<sub>sequence</sub></em>. The information density-based clustering algorithm utilized in this work is described in&nbsp; Lu and Rogan 2018 (<a href="https://f1000research.com/articles/7-1933/v2">https://f1000research.com/articles/7-1933/v2</a>) and archived source code is available through Zenodo (<a href="https://dx.doi.org/10.5281/zenodo.1892051">https://dx.doi.org/10.5281/zenodo.1892051</a>).</p> <p><strong>Section 3. Binding site clusters - lollipop plots</strong></p> <p>Lollipop plots present the genomic coordinates and information densities of clusters across the human transcriptome, human exome, and viral genomes (Coronavirus, Dengue, HIV-1 [two strains] and Influenza [one strain]). The height of the &quot;lollipop&quot; corresponds to the information density of a cluster. Labels above &quot;lollipops&quot; present the start and end genomic coordinate (GRCh37) of the cluster followed by the number of sites in the cluster enclosed in brackets. Lollipop plots associated with human transcriptomes/exomes each contain a single gene. Influenza has 8 segments and each segment requires its own plot, other viral genomes examined are presented in a single plot.</p> <p>File naming convention for human plots:</p> <ul> <li>RBP_Gene.png</li> <li>e.g. RNPS1_ADK.png</li> </ul> <p>File naming convention for viral plots (elements in square brackets do not always appear):</p> <ul> <li>Virus[.InfluenzaSegment].RiThreshold.Strand.RBP.png</li> <li>e.g. Wuhan-Hu-1.complete-genome.4.2-bits.PosStrand.hnRNPA1.png</li> </ul> <p>The specified Ri threshold indicates all binding sites which comprise a cluster have <em>R<sub>i</sub></em> greater-than or equal to the threshold.</p> <p><strong>Section 4. Ri(b,l) matrices for all binding sites scanned</strong></p> <p>The information theory-based position weight matrices for the following RNA binding proteins (RBP) used in this study: SRSF1, hnRNPA1 and RNPS1. We investigated binding using two different RNPS1 binding models. While similar, these two models contained binding site information on opposing sides of the binding site motif which is why we found it prudent to scan with both models.</p> <p>Structure of each file:</p> <p>Line #1: Start position, End position and<em> R<sub>sequence</sub></em> [average strength of sequences used to generate the model]</p> <p>Subsequent lines describe the information on each position of the binding site:</p> <ul> <li>First four columns: <em>R<sub>i</sub></em> contribution of nucleotide at this position of the matrix [A, C, G, T]</li> <li>Row #5: Position of the matrix</li> <li>Last four columns: Number of binding sites used to generate model with a particular nucleotide at this position of the matrix [A, C, G, T]</li> </ul> <p>Example:</p> <p>-2.965775&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 1.282153&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0.034225&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; -4.906891&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 19&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 8&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0</p> <p>At zero position of the matrix (first nucleotide), a &lsquo;C&rsquo; would have a positive contribution to binding site strength, a &lsquo;G&rsquo; would be relatively neutral, and an &lsquo;A&rsquo; or &lsquo;T&rsquo; would negatively contribute to binding site strength.</p> <p>Generation of R<sub>i</sub>(b,l) matrices and computation of <em>R<sub>i</sub></em> values and can be accomplished by utilizing the Delila package (<a href="https://alum.mit.edu/www/toms/delila/delilaprograms.html">https://alum.mit.edu/www/toms/delila/delilaprograms.html</a>).</p> <p><strong>Section 5. Ri and intersite distance - histograms</strong></p> <p>Two sets of histograms present <em>R<sub>i</sub></em> distribution and intersite distance distribution across the human transcriptome, human exome, and viral genomes (Coronavirus, Dengue, HIV-1 [two strains] and Influenza [one strain]).&nbsp;</p> <p>File naming convention for human plots (elements in square brackets do not always appear):</p> <ul> <li>[IntersiteDistancesThreshold-]Human-[DRIPc]-AllChrs-RBP[-RiThreshold].png</li> <li>e.g. IntersiteDistances500-Human-AllChrs-hnRNPA1-4.6-bits.png</li> </ul> <p>File naming convention for viral plots (elements in square brackets do not always appear):</p> <ul> <li>[IntersiteDistancesThreshold-]Strand-RBP-Virus[.InfluenzaSegment][-RiThreshold].png</li> <li>e.g. IntersideDistances1000-PosStrandOnly-SRSF1-top50000sitesReplicate1-HIV-1-Strain-B.png</li> </ul> <p>Intersite distance thresholds of 500 or 1000 were assigned for all intersite distance histograms. Any distances above the corresponding threshold were excluded from the plot. Plots presenting <em>R<sub>i</sub></em> distributions contain a dashed line indicating <em>R<sub>sequence</sub></em> if it is visible within the scope of the plot.</p> <p><strong>Section 6. Perl Scripts and Descriptions</strong></p> <p>This archive contains all Perl scripts discussed in this archive&#39;s associated manuscript&nbsp;and a document file which describes them (&quot;Perl-Script-Descriptions-Page.docx&quot;). The programs and their general functions are as follows:</p> <p>&ldquo;ClusterToDRIPseqAnalysisProgram.pl&rdquo; &ndash; reports which information-dense clusters are located within DRIPc- and/or DRIP-seq intervals (individually and by gene)</p> <p>&ldquo;ClusterToDRIPseqAnalysisProgram.GeneDensityFinder.pl&rdquo; &ndash; uses the output from script &ldquo;ClusterToDRIPseqAnalysisProgram.pl&rdquo; to determine the number and the density of information-dense clusters within a gene (total clusters within the gene and those within DRIPc-seq intervals)</p> <p>&ldquo;calculateIntersiteDistance.pl&rdquo; &ndash; determines the distance between all binding sites in the same gene from a list of genomic coordinates</p> <p>&ldquo;removeOutliersHigherThanN.pl&rdquo; &ndash; discards intersite distances computed by script &ldquo;calculateIntersiteDistance.pl&rdquo; that are greater than a specified threshold</p> <p>&ldquo;getStatisticsOnCol.pl&rdquo; &ndash;&nbsp;calculates the count, geometric mean, median, arithmetic mean, and standard deviation of values from the output of script &ldquo;removeOutliersHigherThanN.pl&rdquo;</p> <p>&ldquo;ScanDataSummaryProgram.pl&rdquo; &ndash;&nbsp;determines the number of binding sites (above a specified <em>R<sub>i</sub></em> threshold) found within known genes (the program also reports the total expression of those genes using external A549 and pneumocyte expression datasets) from binding site coordinate data</p> <p>&ldquo;TotalBindingSitePerCellCalculator.pl&rdquo; &ndash;&nbsp;estimates the number of binding sites expressed in a single A549 or pneumocyte cell at any given time.</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

SARS-CoV-2 Nidoviral RNA Uridylate‐Specific Endoribonuclease (NSP15); A Target Enabling Package

<p>The non-structural protein 15 (NSP15, NendoU<sup>SARS-CoV-2</sup>) from severe acute respiratory syndrome 2 virus (SARS-CoV-2) is an uridylate-specific endoribonuclease, likely responsible in the viral immune evasion mechanism. This TEP provides a set of reagents for further interrogation of the molecular function of NSP15. We have established a purification protocol for the active protein for biochemical and structural studies. Moreover, we have crystallised the protein and performed a crystallographic fragment screen which yielded several hits. Data generated here will be used for the development of enzyme inhibitors that would illuminate the biological role of the gene product, and eventually point the way to new antiviral therapies.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Simulated RNA-seq data

<p>Simulated RNA-seq data shows that histograms from p value sets with around one hundred&nbsp;&nbsp;true effects out of 20,000 features can be classified as &#39;uniform&#39;.&nbsp;RNA-seq data was simulated with polyester R package <a href="https://doi.org/10.1093/bioinformatics/btv272">(Frazee, 2015)</a> on 20,000 transcripts from human transcriptome&nbsp;using grid of 3, 6, and 10 replicates and 100, 200, 400, and 800 effects for two groups.&nbsp;Fold changes were set to 0.5 and 2.&nbsp;Differential expression was assessed using DESeq2 R package <a href="https://doi.org/10.1186/s13059-014-0550-8">(Love, 2014)</a> using default settings&nbsp;and group 1 versus group 2 contrast.&nbsp;Effects denotes in facet labels the number of true effects and N denotes number of replicates.&nbsp;Red line denotes QC threshold used for dividing p histograms into discrete classes.&nbsp;Workflow and code used to run this simulation is available on <a href="https://github.com/rstats-tartu/simulate-rnaseq">rstats-tartu/simulate-rnaseq</a>.</p> <p>&nbsp;</p> <p>Files</p> <ul> <li>de_simulation_results.csv -- merged and processed DE analysis results of simulated data.</li> <li>simulate-reads-2021-01-25.tar.gz -- raw DE analysis results&nbsp;on 20,000 transcripts from human transcriptome&nbsp;using grid of 3, 6, and 10 replicates and 100, 200, 400, and 800 effects for two groups.&nbsp;Fold changes were set to 0.5, 1, and 2.&nbsp;Differential expression was assessed using DESeq2 with default settings.</li> <li>simulate-rnaseq.tar.gz -- snakemake workflow and input fasta file&nbsp;to simulate RNA-seq data with polyester and analyse results with DESeq2. Adjust settings in config.yaml to customise simulation. Includes software to run workflow on Linux, given that <a href="https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh">Conda</a> and <a href="https://snakemake.readthedocs.io/en/stable/index.html">snakemake</a> are installed.</li> </ul> <p>The simulate-rnaseq.tar.gz&nbsp;archive can be re-executed on a vanilla machine that only has Conda and Snakemake installed via:</p> <pre><code class="language-bash">tar -xf simulate-rnaseq.tar.gz snakemake --use-conda -n</code></pre> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

G-quadruplex in the gene of the large subunit of plant RNA polymerase II: billion years old story

<p><strong>Supplementary material to the journal article</strong></p> <p>Consist of:</p> <p>Supplementary material S1: Analyzed <em>RPB1 </em>sequences in 40 plant species together with detailed characteristics and G-quadruplex prediction using four different computational approaches.</p> <p>Supplementary material S2: G4 locus is the most conserved within the&nbsp;<em>RPB1</em> gene (40 bp long potential G4 locus is the most conserved site in the whole ~ 6000 bp long <em>RPB1</em> gene. See the histogram below the alignment:&nbsp;the position of the G4 locus is depicted, together with the horizontal red dashed line indicating relative nucleotide&nbsp;conservation&nbsp;among aligned sequences of<em>&nbsp;RPB1</em>)</p> <p>Supplementary material S3: Multiple alignment of G4 locus of&nbsp;<em>RPB1</em> paralogs in <em>Arabidopsis thaliana&nbsp;</em>centered to G4 locus of <em>RPB1&nbsp;</em></p> <p>Supplementary material S4: Modelled 3D structure of G4 from <em>Bathycoccus prasinos</em> in PDB format</p> <p>Supplementary material S5: Gel electrophoresis and ThT staining of the selected G4-forming sequences</p> <p>Supplementary material S6: All analyzed <em>RPB1</em> sequences in FASTA format</p> <p>Supplementary material S7: Aligned <em>RPB1</em> sequences in FASTA format</p> <p>Supplementary material S8: <em>RPB1</em> paralogs&nbsp;in <em>Arabidopsis thaliana</em></p> <p>Supplementary material S9: Spectral composition of light used in the UV experiment. Analysis of emitted light was performed by Ocean Optics (HR4000CG-UV-NIR, USA) device.</p> <p>Supplementary material S10: Difference CD spectra - comparison without and with previous UV treatment</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Training material for de novo transcriptome reconstruction from RNA-seq data

<p>The data provided here are part of a Galaxy tutorial that analyzes RNA-seq data from a study published by Wu et al., 2014 (DOI:10.1101/gr.164830.113). The goal of this study was to investigate "the dynamics of occupancy and the role in gene regulation of the transcription factor Tal1, a critical regulator of hematopoiesis, at multiple stages of hematopoietic differentiation." To this end, RNA-seq libraries were constructed from multiple mouse cell types including G1E - a GATA-null immortalized cell line derived from targeted disruption of GATA-1 in mouse embryonic stem cells - and megakaryocytes. This RNA-seq data was used to determine differential gene expression between G1E and megakaryocytes and later correlated with Tal1 occupancy. This dataset (GEO Accession: GSE51338) consists of biological replicate, paired-end, polyA selected RNA-seq libraries. Because of the long processing time for the large original files, we have downsampled the original raw data files to include only reads that align to chromosome 19 and a subset of interesting genomic loci identified by Wu et al.</p>

opencc-by-4.0Jan 2017View details →
zenodo44/100

Expanding and improving analyses of nucleotide recoding RNA-seq experiments with the EZbakR suite

<p>Data necessary to reproduce figures in manuscript titled "Expanding and improving analyses of nucleotide recoding RNA-seq experiments with the EZbakR suite". Includes:</p> <ol> <li>Compressed arrow dataset used to produce Figure S6 and S7 (subtlseq_data.tar.gz)</li> <li>Compressed arrow dataset used to produce Figure 7 (subtlseq_perturbations_data.tar.gz)</li> <li>eCLIP DDX3X peak calls from ENCODE used in Figure 7 (ENCFF901BYH_DDX3X_eCLIP.bed)</li> <li>Annotations used to process data (Hs_ensembl_lvl1_and_2.gtf and Hs_ensembl.gtf)</li> <li>Processed data to produce Figures 3C and S2 (cB_ensembl_totRNAsubtlseq.csv.gz and cB_ensemblLvl1and2_totRNAsubtlseq.csv.gz, respectively).</li> <li>Simulated data originally used in bakR publication (Vock and Simon, 2023) used to make Figure 6 of EZbakR suite paper.</li> <li>Processed data from the nanodynamo paper (Tarrerro et al. 2024) used to make Figure S9</li> <li>Table from Ietswaart et al. 2024 of kinetic parameter estimates and PUND calls from that study (mmc2.xlsx)</li> </ol> <p>Also includes supplemental tables of:</p> <ol> <li>List of genes producing transcripts predicting to undergo nuclear decay (PUNDs; Supplemental_Table_PUNDs.csv)</li> <li>Estimates for mature RNA synthesis, nuclear degradation, nuclear export, and cytoplasmic degradation rate constants from Ietswaart et al., 2024 total-cytoplasmic-nuclear TimeLapse-seq dataset (Supplemental_Table_NucCytoEsts.csv)</li> <li>Estimates for premature RNA synthesis, premature RNA processing, and mature RNA degreadation obtained from EZbakR analysis of Ietswaart et al., 2024 total RNA TimeLapse-seq dataset (Supplemental_Table_PtoMests.csv).</li> </ol> <p>Scripts to reproduce figures can be found at: https://github.com/isaacvock/EZbakRsuite_paper_code</p> <p>Updated to include data necessary to reproduce new figures/panels in revisions.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Data archive: CICT for single cell RNA-seq network inference

<p>This archive contains benchmarking input data and results for using single cell gene expression data to infer gene regulatory networks (GRN) by the Causal Inference with Composition of Transactions (CICT) method and a selected set of published methods. This accompanies the manuscript "Robust discovery of gene regulatory networks from single-cell gene expression data by Causal Inference Using Composition of Transactions" (Shojaee and Huang, Brief in Bioinform 2023. DOI: 10.1093/bib/bbad370). The CICT code is available at the GitHub repo (https://github.com/hlab1/scRNAseqWithCICT/).</p><p>The original CICT algorithm was described in Shojaee et al. (arXiv:1608.02658, 2016). The benchmarked methods were included in the BEELINE benchmarking pipeline (Pratapa et al., Nat Methods 2020), to which we added DEEPDRIM (Chen et al., Brief Bioinform 2021), SCENIC (Aibar et al., Nat Methods 2017), Inferelator 3.0 (Gibbs et al., Bioinformatics 2022), and CellOracle (Kamimoto et al., Nature 2023). The output directory names are (subdirectories within each dataset):</p><p>* CICT_ewMIshrink_RFmaxdepth10_RFntrees20/: CICT for simulated data<br>* CICT_v2/: CICT for experimental data<br>* CELLORACLEDB/: CellOracle for experimental data<br>* DEEPDRIM72_ewMIshrink_RFmaxdepth10_RFntrees20/: DEEPDRIM for simulated data<br>* DEEPDRIM72_v2/: DEEPDRIM for experimental data<br>* INFERELATOR38_ewMIshrink_RFmaxdepth10_RFntrees20/: Inferelator-Prior for simulated data<br>* INFERELATOR38_v2/: Inferelator-Prior for experimental data<br>* INFERELATOR34_ewMIshrink_RFmaxdepth10_RFntrees20/: Inferelator-NoPrior for experimental data<br>* INFERELATOR34_v2/: Inferelator-NoPrior for experimental data<br>* GENIE3/: GENIE3<br>* GRNBOOST2/: GRNBOST2<br>* LEAP/: LEAP<br>* PIDC/: PIDC<br>* PPCOR/: PPCOR<br>* SCENICDB/: SCENIC for experimental data<br>* SCNS/: SCNS<br>* SCODE/: SCODE<br>* SCRIBE/: SCRIBE<br>* SINCERITIES/: SINCERITIES<br>* SINGE/: SINGE<br>* RANDOM/: RANDOM</p><p>The methods were benchmarked against two kinds of scRNA-seq datasets:<br>* Simulated datasets produced by the SERGIO simulator from a synthetic network (Dibaeinia et al., Cell Systems 2020), including complete datasets and datasets with dropouts with shape parameter k=6.5 and rate parameter q=10, 30, 50, 70, 80.&nbsp;<br>* Experimental datasets compiled by the BEELINE pipeline, evaluated at three different levels L0, L1 and L2, with three types of ground truth networks.<br>&nbsp; &nbsp; * Evaluation levels:<br>&nbsp;&nbsp; &nbsp; &nbsp; &nbsp;* L0: 500 highly varying genes plus TFs<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* L1: 1000 highly varying genes plus TFs<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* L2: 500 highly varying genes, TFs and 500 genes randomly selected that excluded the 1000 highly varying genes from L1.<br>&nbsp; &nbsp; * Types of ground truths:<br>&nbsp;&nbsp; &nbsp; &nbsp; &nbsp;* Cell-type-specific ChIP-seq ground truth (L0, L1, L2)<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* Non-specific ChIP-seq ground truth (L0_ns, L1_ns, L2_ns)<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* Loss-of-function/gain-of-function ground truth (L0_lofgof, L1_lofgof, L2_lofgof)</p><p>The directory structure is organized in accordance with the BEELINE benchmarking pipeline. For complete details please please see the BEELINE documentation (https://murali-group.github.io/Beeline/) and Github repo (https://github.com/Murali-group/Beeline).</p><p>&nbsp;</p>

opencc-by-nc-sa-4.0Jun 2023View details →
zenodo44/100

Predicting Phenotypic Traits Using a Massive RNA-seq Dataset

<h2><strong>Abstract</strong></h2><p>The included datasets are a conglomerate of all available <i>Arabidopsis thaliana</i> RNA-seq data available from NCBI as of November 2022 processed to count data. In addition, the associated annotation files from NCBI BioProject database and processed versions of this data is included. Data has been processed according to the "Data Description Methods" in the manuscript titled "Predicting Phenotypic Traits Using a Massive RNA-seq Dataset" (in publication). The associated Methods can be found at this repository:<a href="https://gitlab.com/ficklinlab-public/modeling-with-transcriptomics"> https://gitlab.com/ficklinlab-public/modeling-with-transcriptomics</a>. These datasets can be used for exploring machine learning methods for predicting both continuous (Age) and categorical (Tissue) phenotypic traits using gene expression. Additionally, the gene expression data can be used on its own for the investigation of gene expression in <i>Arabidopsis thaliana.</i></p><h3><strong>Note to Researchers</strong></h3><p>This repository contains all of the datasets and information necessary to recreate the experiments in our paper. However, if may be that you are interested in our dataset for testing your own hypotheses/programs. If this is the case, we predict that you are looking for one or more of the the following 5 datasets<br>&nbsp;</p><h3><strong>Note on File Compression</strong></h3><p>All files in this repository are compressed using bzip2 to conserve space and allow for easier file transfer. The unzip command on linux systems is `bzip2 -d FILE_NAME`. For other computer systems (Windows and Apple) please consult your user manual.</p><h3><strong>Description All Datasets:</strong></h3><p><strong>Title:</strong> Gene Expression Count Data of all <i>Arabidopsis thaliana</i> data available from NCBI SRA as of November 2022<br><strong>Abstract: </strong>Gene Expression Count data was created using the workflow GEMmaker. The resulting Gene Expression Matrix (GEM) was then normalized and thresholded. The following 4 files are normalizations of the same data for Trimmed Mean of M values (TMM), Median Ratios Normalization (MRN), Transcripts Per kilobase Million (TPM), and No Normalization (NoNo) respectively. Additionally, Each file is included as a tsv and a python pickle. The tsv file is human readable, whereas the pickle file can be read into memory substantially faster. Format for tsv is each row represents a sample and each column represents a gene. <strong>NCBI_Nov2022_SRR_runinfo.csv</strong> is the starting file from NCBI which reports SRR information for each sample. <strong>Note 1 to Researchers: </strong>MRN normalization performed the best in our experiments and is likely what you want to use if you are doing additional expermentation with this dataset. Otherwise start with NoNo and perform your own normalizations. <strong>Note 2 to Researchers:</strong> the 54547 dataset will need to be thresholded prior to use. We include it in addition to the 32432 datasets in case you wish to try a different thresholding to the one outlined in our manuscript.&nbsp;<br><strong>Author:</strong> John Anthony Hadish<br><strong>Data Type: </strong>Gene Expression Count Data<br><strong>Organism:</strong> <i>Arabidopsis thaliana</i><br><strong>Files:</strong></p><p><strong>NCBI_Nov2022_SRR_runinfo.csv - </strong>Arabidopsis RNA-seq SRA RunInfo Retrieved from NCBI November 2022. This is the unprocessed data.<br><strong>Dataset_54547_NoFilter_raw.pkl </strong>- Raw File Before thresholding (".pkl" format). Same as NoNo normalization without thresholding.<br><strong>Dataset_54547_NoFilter_raw.tsv </strong>- Raw File Before thresholding (".tsv" format). Same as NoNo normalization without thresholding.<br><strong>Dataset_32432_MRN.pkl</strong> - MRN normalized (".pkl" format)<br><strong>Dataset_32432_MRN.tsv - </strong>MRN normalized (".tsv" format)<br><strong>Dataset_32432_NoNo.pkl - </strong>NoNo normalized (".pkl" format)<br><strong>Dataset_32432_NoNo.tsv - </strong>NoNo normalized (".tsv" format)<br><strong>Dataset_32432_TMM.pkl - </strong>TMM normalized (".pkl" format)<br><strong>Dataset_32432_TMM.tsv - </strong>TMM normalized (".tsv" format)<br><strong>Dataset_32432_TPM.pkl - </strong>TPM normalized (".pkl" format)<br><strong>Dataset_32432_TPM.tsv - </strong>TPM normalized (".tsv" format)<br><br><br><strong>Title: </strong>Meta Data Arabidopsis Age and Tissue<br><strong>Abstract: </strong>Meta Data for Age and Tissue after processing. In our experiment this was used as response variable to gene expression. Shared columns are "bio_sample", "bioproject_name", "experiment". In addition to these processed datasets, <strong>NCBI_Nov2022_BioSample_data.tsv </strong>is the unprocessed starting material for these two data frames.<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>Metadata on phenotypes. ".tsv" format&nbsp;<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:&nbsp;</strong><br><strong>NCBI_Nov2022_BioSample_data.tsv - </strong>Arabidopsis BioSample data retrieved from NCBI November 2022. This is the unprocessed data.<br><strong>df_metadata_tissue.tsv</strong> - Tissue Annotations for 24876 samples<br><strong>df_metadata_age.tsv</strong><i><strong> - </strong></i>Age Annotations for 16078 samples. In addition to shared columns includes<i> "</i>days<i>_</i>age"(how many days old the sample is converted to days) and "annotation_age" (how the annotation was reported for this sample in the raw data file-- i.e. "days", "weeks" etc.)<br><br><br><strong>Title: </strong>Machine Learning Dataset for <i>Arabidopsis thaliana</i> <strong>Age</strong><br><strong>Abstract: </strong>The dataset used for Machine learning on the phenotype Age that is a combination of the Gene Expression Matrix and the Annotation Matrix. Consists of a list of 4 for the train and test splits.<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>Gene Expression Matrix and Annotations Combined, split into train and test&nbsp;<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:</strong><br><strong>Dataset_Age_TrainTestSplits_mrn.pkl</strong> - MRN normalized<br><strong>Dataset_Age_TrainTestSplits_NoNo.pkl </strong>- NoNo normalized<br><strong>Dataset_Age_TrainTestSplits_tmm.pkl </strong>- TMM normalized<br><strong>Dataset_Age_TrainTestSplits_tpm.pkl - </strong>TPM normalized<br><br><br><strong>Title: </strong>Machine Learning Dataset for <i>Arabidopsis thaliana</i> <strong>Tissue</strong><br><strong>Abstract: </strong>The dataset used for Machine learning on the phenotype Tissue that is a combination of the Gene Expression Matrix and the Annotation Matrix. Consists of a list of 4 for the train and test splits. Saved as python ".pkl" files.<br><strong>Author:</strong> John Anthony Hadish<br><strong>Data Type: </strong>Gene Expression Matrix and Annotations Combined, split into train and test. Saved as python ".pkl" files.<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:</strong><br><strong>Dataset_Tissue_TrainTestSplits_mrn.pkl </strong>- MRN normalized<br><strong>Dataset_Tissue_TrainTestSplits_NoNo.pkl </strong>- NoNo normalized<br><strong>Dataset_Tissue_TrainTestSplits_tmm.pkl </strong>- TMM normalized<br><strong>Dataset_Tissue_TrainTestSplits_tpm.pkl </strong>- TPM normalized<br><strong>Dataset_Tissue_TrainTestSplits_mrn_4category.pkl </strong>- MRN for the tissue-4 dataset<br><br><br><strong>Title: </strong>BioProject Names<br><strong>Abstract:</strong> Three Column File With BioProject Name, BioSample Name, and Experiment Name<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>".tsv"<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:</strong><br><strong>BioProject_Names_All.tsv</strong><br><br><br><strong>Title:</strong> Manuscript Supplemental Material<br><strong>Abstract:</strong> Supplemental tables and figures described in the manuscript (included with manuscript and here for convenience). Please see manuscript for additional information.<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>".tsv", ".png".pdf"<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:</strong><br><strong>Supplemental_Figures.zip </strong>- Supplemental figures from the manuscript. Includes description of each figure.<br><strong>Supplemental_Tables.zip </strong>- Supplemental tables from the manuscript. Includes description of each table.<br><br>&nbsp;</p><p><strong>Title:</strong> Splits of data for 3 experiments<br><strong>Abstract:</strong> 2 column tsv files. The first column is the experiment (sample) name, and the second column is if it is included in the train or test data. <strong>Included here to make sure pkl files are reproducible in case the pkl package breaks in the future.</strong> Not used by scripts, included to prevent future potential loss of data.<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>".tsv"<br><strong>Organism: </strong><i>Arabidopsis thaliana</i><br><strong>Files:</strong><br><strong>Dataset_Tissue_TrainTestSplits_4category_namesOnly.tsv</strong><br><strong>Dataset_Tissue_TrainTestSplits_namesOnly.tsv</strong><br><strong>Dataset_Age_TrainTestSplits_namesOnly.tsv</strong></p><p>&nbsp;</p><p><strong>Title:</strong> Git Code Repository<br><strong>Abstract:</strong> A tar bz2 compression of the git repository containing all of the code created for this manuscript. The same code found in this file is also avalible on GitLab at the link: https://gitlab.com/ficklinlab-public/modeling-with-transcriptomics<br><strong>Author: </strong>John Anthony Hadish<br><strong>Data Type: </strong>Git Repository, python code<br><strong>Files:</strong><br><strong>modeling-with-transcriptomics-main.tar.bz2</strong> - Compressed Git repository of all code used in paper.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Additional Data: Poised PABP-RNA hubs implement signal-dependent mRNA decay in development

<p>This repository contains processed data resulting from iCLIP experiments that were analysed in the following paper:"<strong>Poised PABP-RNA hubs implement signal-dependent mRNA decay in development</strong>"<br>The paper is published at Nature Structural and Molecular BIology.</p> <h2><br>Archived data</h2> <p>Data archived in this repository include:</p> <ol> <li>Data derived from iCLIP experiments targeting LIN28A, PABPC1, and PABPC4, that were analysed in the manuscript (see iCLIP.zip). Raw data is available&nbsp;from ENA, with the accession code PRJEB60519. <ol> <li>Sample descriptions are given in iCLIP-SampleAnnotation.csv</li> <li>Crosslink files in BED6 format (individual replicates and merged replicates)</li> <li>Peak files generated with the Clippy peak caller in BED6 format</li> <li>K-mer enrichment around high-confidence crosslink sites in the 3'-UTRs, calculated by the PEKA software</li> </ol> </li> <li>Expression values (salmon quantfiles)&nbsp; for 3'-seq experiments, specified in "QuantseqExperimentsAnnotation.tsv", are available in "SalmonQuantfiles.zip".&nbsp;Raw data is available from ENA, with the accession code PRJEB60519.</li> <li>Source code of the nextflow pipeline, which was used on the iMaps webserver to analyse iCLIP data and produce the files archived here (see imaps-nf-0.30.zip).</li> <li>A list of naive genes, that were analysed in the manuscript (see NaiveGeneIds.csv).</li> </ol> <h2>Details on iCLIP data generation</h2> <p>iCLIP data for LIN28A-WT (in 2iL and FGF2 treated cells), LIN28A-S200A (in FGF2 treated cells) as well as for PABPC1 and PABPC4 (in LIN28A KO cells with and without LIN28A overexpression), were analysed on iMaps Goodwright server (<a href="https://imaps.goodwright.com/">https://imaps.goodwright.com/</a>). The LIN28A iCLIPs were analysed on 18th of July, 2022; the PABPC iCLIPs were analysed on 26th of December, 2022. The code and settings used in the pipeline (release v0.30) can be viewed at <a href="https://github.com/goodwright/imaps-nf">https://github.com/goodwright/imaps-nf </a>, and is also archived here - (imaps-nf-0.30.zip)<br>&nbsp;</p> <ul> <li>First, reads were demultiplexed using Ultraplex and barcodes were trimmed from the reads. The default Ultraplex settings were applied, as denoted below:</li> </ul> <blockquote> <p>adapter='AGATCGGAAGAGCGGTTCAG'<br>adapter2='AGATCGGAAGAGCGTCGTG'<br>barcodes='barcode.csv',<br>final_min_length=20<br>fiveprimemismatches=1<br>ignore_no_match=False<br>ignore_space_warning=False<br>inputfastq='MOD4878A1-merged.fastq.gz',<br>keep_barcode=False,<br>min_trim=3,<br>outputprefix='demux',<br>phredquality=30,<br>phredquality_5_prime=0,<br>sbatchcompression=False,<br>threads=10,<br>threeprimemismatches=0,<br>ultra=False</p> </blockquote> <p>&nbsp;</p> <ul> <li>TrimGalore was used to run FASTQC and quality trim the reads and remove reads with length less than 10 nt:</li> </ul> <blockquote> <p>trim_galore --fastqc --length 10 -q 20 --cores 8 --gzip file.fastq.gz</p> </blockquote> <p>&nbsp;</p> <ul> <li>Reads were then premapped to rRNA, tRNA sequences referred to as small RNA, smRNA, using mouse genome build (GRCm39 GENCODE M28 annotation) with Bowtie v1.3.0 (Langmead et al., 2009)</li> </ul> <blockquote> <p>bowtie --threads 12 --sam -x $INDEX -q --un file.unmapped.fastq -v 2 -m 100 --norc --best --strata file.fq.gz 2</p> </blockquote> <p>&nbsp;</p> <ul> <li>Reads that did not map with Bowtie were then aligned with STAR v2.7.9a (Dobin et al., 2013) to mouse genome build (GRCm39 GENCODE M28 annotation).</li> </ul> <blockquote> <p>STAR \<br>--genomeDir star \<br>--readFilesIn file.unmapped.fastq.gz \<br>--runThreadN 12 \<br>--outFileNamePrefix 1_R1. \<br>\<br>--sjdbGTFfile Homo_sapiens_filtered.gtf \<br>--outSAMattrRGline 'ID:1_R1' 'SM:1_R1' \<br>&nbsp;--readFilesCommand zcat --outSAMtype BAM SortedByCoordinate --quantMode TranscriptomeSAM --outFilterMultimapNmax 1 --outFilterMultimapScoreRange 1 --outSAMattributes All --alignSJoverhangMin 8 --alignSJDBoverhangMin 1 --outFilterType BySJout --alignIntronMin 20 --alignIntronMax 1000000 --outFilterScoreMin 10 --alignEndsType Extend5pOfRead1 --twopassMode Basic</p> </blockquote> <p>&nbsp;</p> <ul> <li>PCR-duplicates were removed using UMI-tools (Smith, Heger and Sudbery, 2017)</li> </ul> <blockquote> <p>java -jar /UMICollapse/umicollapse.jar \<br>&nbsp;&nbsp;bam \<br>&nbsp;&nbsp;-i file.Aligned.sortedByCoord.out.bam \<br>&nbsp;&nbsp;-o file.dedup.bam \<br>&nbsp;&nbsp;--umi-sep rbc:</p> </blockquote> <p>&nbsp;</p> <ul> <li>The nucleotide preceding each sequencing read was assigned as the crosslink event.</li> </ul> <p>&nbsp;</p> <ul> <li>Peaks of crosslinking signal were identified with Clippy v1.4.1, using the default settings.</li> </ul> <p>&nbsp;</p> <ul> <li>Obtained peaks and crosslink sites were used to run PEKA v1.0.0 (Kuret et al., 2022), using the default settings.</li> </ul> <p>&nbsp;</p> <ul> <li>For Clippy and PEKA, the GENCODE primary assembly annotation M28 was filtered to retain only entries with transcript support level 1 or 2, in genes where such transcripts were available, and used to produce a segmentation file with the <em>get_segments</em> function from the iCount tool (Curk, 2019).</li> </ul> <p>&nbsp;</p> <ul> <li>All files generated during data processing are available from the iMaps Goodwright webserver for analysis of CLIP data (see <a href="https://imaps.goodwright.com/collections/882635250203/">https://imaps.goodwright.com/collections/882635250203/</a> and <a href="https://imaps.goodwright.com/collections/340215254997/">https://imaps.goodwright.com/collections/340215254997/</a> for LIN28A and PABPC1/4 iCLIPs, respectively).</li> </ul> <h2>Source data</h2> <p>Raw sequencing reads, from which the data enclosed here were derived, are accessible at ENA (PRJEB60519).<br>The raw sequencing reads and all data produced by the analysis pipeline is also available at the iMaps webserver (see <a href="https://imaps.goodwright.com/collections/882635250203/">https://imaps.goodwright.com/collections/882635250203/</a> and <a href="https://imaps.goodwright.com/collections/340215254997/">https://imaps.goodwright.com/collections/340215254997/</a> for LIN28A and PABPC1/4 iCLIPs, respectively); and on the updated Flow webserver (see <a href="https://app.flow.bio/projects/882635250203/">https://app.flow.bio/projects/882635250203/</a> and <a href="https://app.flow.bio/projects/340215254997/">https://app.flow.bio/projects/340215254997/ </a>for LIN28A and PABPC1/4 iCLIPs, respectively).</p> <h2>Downstream computational analysis of enclosed data</h2> <p>The code, used to analyse the data enclosed here and train the CNN to predict transcript stability in naive-to-primed transition based on 3'UTR nucleotide sequence, is available at GitHub (<a href="https://github.com/ulelab/LIN28A_RNPreassembly_bioinformatics">https://github.com/ulelab/LIN28A_RNPreassembly_bioinformatics</a>) and archived on Zenodo (<a href="../doi/10.5281/zenodo.10054297">https://zenodo.org/doi/10.5281/zenodo.10054297</a><strong>).</strong></p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Single-cell RNA-seq profiles of tumor-bearing mice treated with PAGln with or without anti-PD-1

<p>single-cell RNA sequencing (scRNA-seq) profiles of&nbsp; tumor-bearing mice treated using Phenylacetylglutamine (PAGln) with or without anti-PD-1 were performed to compare the alterations of immune microenvironment affected by PAGln under the condition of anti-PD-1 treatment.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

RNA-Protein Interaction Prediction Using Network-Guided Deep Learning

<p>RNA-protein interactions are critical to various life processes, including fundamental translation and gene regulation. Identifying these interactions is vital for understanding the mechanisms underlying life processes. Then, ZHMolGraph is an advanced pipeline that integrates graph neural network sampling strategy and unsupervised large language models to enhance binding predictions for novel RNAs and proteins.</p> <div>&nbsp;</div>

openmit-licenseJul 2024View details →
zenodo44/100

The Yin and Yang of RNA in neurodegeneration | Talk - I PhasAGE International Conference

<p>The&nbsp;<strong>I PhasAGE international conference</strong>&nbsp;brought together members of the PhasAGE consortium as well as outstanding international speakers showcasing high impact achievements in the field of liquid-liquid phase separation in aging and late-onset diseases.</p> <p>For details on conference program please see:&nbsp;https://phasage.eu/phasage-conference-1/&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Formatted TCGA clinical and RNA-Seq data for colon adenocarcinoma (COAD) and rectum adenocarcinoma (READ)

<p>COAD/READ/COADREAD_rnaseq_fpkm.txt files contain TCGA RNA-Seq data in FPKM normalisation for colorectal adenocarcinoma (COAD), rectum adenocarcinoma (READ) or combined (COADREAD).</p> <p>COAD/READ/COADREAD_rnaseq_tpm.txt files contain TCGA RNA-Seq data in TPM normalisation for colorectal adenocarcinoma (COAD), rectum adenocarcinoma (READ) or combined (COADREAD).</p> <p>COAD/READ/COADREAD_clinical_raw.xlsx&nbsp;files contain TCGA clinical data for patients with&nbsp;colorectal adenocarcinoma (COAD), rectum adenocarcinoma (READ) or combined (COADREAD).</p> <p>COAD/READ/COADREAD_rnaseq_clinical_raw.xlsx&nbsp;files contain corresponding information of TCGA clinical data and RNA-Seq data for patients with&nbsp;colorectal adenocarcinoma (COAD), rectum adenocarcinoma (READ) or combined (COADREAD).</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Automatic learning of hydrogen-bond fixes in an AMBER RNA force field - dataset

<p>Supporting data related to manuscript &quot;Automatic learning of hydrogen-bond fixes in an AMBER RNA force field&quot;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

VirHunter: a deep learning-based method for detection of novel RNA viruses in plant sequencing data

<p>This storage contains 2&nbsp;archives: toy datasets to test the training of the VirHunter and weights of the&nbsp; fully trained VirHunter models for 3 host species &nbsp;(peach, grapevine, sugar beet) and&nbsp;for fragment sizes 500 and 1000.&nbsp; .</p> <p>The toy dataset consists of 3 archived files: &#39;viruses.fasta&#39;, &#39;host.fasta&#39;, &#39;bacteria.fasta&#39;.</p> <p>&#39;viruses.fasta&#39; contains 10000 randomly selected plant viruses from the virus dataset described in the paper.</p> <p>&#39;host.fasta&#39; consists of peach chromosome 2.</p> <p>&#39;bacteria.fasta&#39; consists of 10 bacterial genomes selected randomly:&nbsp;GCF_000284415, GCF_000590555, GCF_001548055, GCF_002795265, GCF_003330825,&nbsp;GCF_003957805, GCF_005845345,&nbsp;GCF_009176625,&nbsp;GCF_010748935, GCF_014681765</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Paramecium Polycomb Repressive Complex 2 physically interacts with the small RNA binding PIWI protein to repress transposable elements

<p>Polycomb Repressive Complex 2 (PRC2) maintains transcriptionally silent genes in a repressed state via deposition of histone H3 K27 trimethyl (me3) marks. PRC2 has also been implicated in silencing transposable elements (TEs), yet how PRC2 is targeted to TEs remains unclear. To address this question, we identified proteins that physically interact with the <em>Paramecium</em> Enhancer-of-zeste Ezl1 enzyme, which catalyzes H3K9me3 and H3K27me3 deposition at TEs. We show that the <em>Paramecium</em> PRC2 core complex comprises four subunits, each required <em>in vivo</em> for catalytic activity. We also identify PRC2 cofactors, including the RNA interference (RNAi) effector Ptiwi09, which are necessary to target H3K9me3 and H3K27me3 to TEs. We find that the physical interaction between PRC2 and the RNAi pathway is mediated by a RING finger protein and that small RNA recruitment of PRC2 to TEs is analogous to the small RNA recruitment of H3K9 methylation SU(VAR)3-9 enzymes.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

SARS-CoV-2 RNA levels in Scotland's wastewater

<p>Nationwide, wastewater-based monitoring was newly established in Scotland to track the levels of SARS-CoV-2 viral RNA shed into the sewage network, during the COVID-19 pandemic. We present a curated, reference data set produced by this national programme, from May 2020 to February 2022.</p> <p>Viral levels were analysed by RT-qPCR assays of the N1 gene, on RNA extracted from wastewater sampled at 122 locations. Locations were sampled up to four times per week, typically once or twice per week, and in response to local needs.</p> <p>These wastewater data are contributing to estimates of disease prevalence and the viral reproduction number (R) in Scotland and in the UK.</p> <p>We report sampling site locations with geographical coordinates, the total population in the catchment for each site, and the information necessary for data normalisation, such as the incoming wastewater flow values and ammonia concentration, when these were available. The methodology for viral quantification and data analysis is briefly described, with links to detailed protocols online. Check the README for details and the project <a href="https://biordm.github.io/COVID-Wastewater-Scotland/">COVID-WW Website</a></p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Phase separation of hnRNP A1 upon specific RNA-binding observed by magnetic resonance

<p>Experimental data, <a href="https://mmmx.info">MMMx</a> restraint and ensemble analysis files (.mcx), restraint data, raw ensembles, and ensemble lists with populations (.ens) pertaining to the manuscript &quot;Phase separation of hnRNP A1 upon specific RNA-binding observed by magnetic resonance&quot; <a href="https://www.biorxiv.org/content/10.1101/2022.03.21.485092v1">available at bioRxiv</a> and submitted to a peer-reviewed journal.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Vignettes: Removing unwanted variation from TCGA RNA-Seq data.

<p>This repository contains all datasets that are required for the vignettes of&nbsp;&nbsp;R.Molania et.al bioRxiv paper (https://www.biorxiv.org/content/10.1101/2021.11.01.466731v1).</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Evaluating the influence of structural properties on proximity metric performance in single cell RNA-seq data - Datasets

<p>Includes raw and processed copies of the scRNA-seq datasets used for the paper: &#39;<strong>How does data structure impact cell-cell similarity? Evaluating the influence of structural properties on proximity metric performance in single cell RNA-seq data.&#39;</strong></p> <p><strong>Real scRNA-seq.zip </strong>contains the Abundant (subset1) and Rare (subset 2) subsets generated to represent discretely structured datasets (sourced from<strong> </strong> Wegmann et al. 2019) and the continuously structured data (sourced from Popescu et al. 2019).</p> <p><strong>Simulated scRNA-seq.zip</strong> contains the Abundant, Moderately-Rare and Ultra-Rare subsets for discretely and continuously structured datasets. All data was simulated using the PROSSTT package in Python 3.8, as well as the dataset containing the labels to re-produce Figure 3 of the manuscript.</p> <p><strong>Results.zip </strong>contains the results for all datasets from the full analysis, in a pickled python dictionary. Code to read in and visualise results is available on the projects github</p> <p>The scripts for the dataset generation, processing and visualisation of results are available at <a href="https://github.com/Ebony-Watson/scProximitE">our github for the scProcimitE package</a>, and documentation is available <a href="https://ebony-watson.github.io/scProximitE/">here</a>.</p>

opencc-by-4.0Apr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record