Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
36,843
datasets available to search
ShareScore release 0.7.1
Dataset results
36,843 results for “RNA”
Conserved regulation of RNA processing in somatic cell reprogramming
<p><strong>Data set 1. Transcript expression across human RNA-Seq samples: estimated read counts. </strong>The file contains estimated read counts, generated by kallisto (<a href="https://pachterlab.github.io/kallisto/">https://pachterlab.github.io/kallisto/</a>), for human transcripts and RNA-Seq samples used in this study (see Additional file 2 of the accompanying publication). The format is a compressed (GZIP) tab-separated transcript-by-sample matrix. Ensembl transcript identifiers and a combined Sequence Read Archive study/sample name identifier serve as row and column names, respectively.</p> <p><strong>Data set 2. Transcript expression across murine RNA-Seq samples: estimated read counts. </strong>As in Data set 1, but for mouse transcripts.</p> <p><strong>Data set 3. Transcript expression across simian RNA-Seq samples: estimated read counts. </strong>As in Data set 1, but for chimpanzee transcripts.</p> <p><strong>Data set 4. Transcript expression across across human RNA-Seq samples: estimated transcript abundances. </strong>As in Data set 1, but instead of read counts, transcript abundances in transcripts per million (TPM), as estimated by kallisto (<a href="https://pachterlab.github.io/kallisto/">https://pachterlab.github.io/kallisto/</a>), are listed. Format, column and row names as in Data set 1.</p> <p><strong>Data set 5. Transcript expression across murine RNA-Seq samples: estimated transcript abundances. </strong>As in Data set 4, but for mouse transcripts.</p> <p><strong>Data set 6. Transcript expression across simian RNA-Seq samples: estimated transcript abundances. </strong>As in Data set 4, but for chimpanzee transcripts.</p> <p><strong>Data set 7. Differential expression analyses across human RNA-Seq sample groups: log fold changes. </strong>The file contains log fold changes, inferred by edgeR (<a href="http://bioconductor.org/packages/release/bioc/html/edgeR.html">http://bioconductor.org/packages/release/bioc/html/edgeR.html</a>), for human genes and the RNA-Seq sample group contrasts listed in Additional file 3 of the accompanying publication in a compressed (GZIP) TSV gene-by-comparison matrix. Ensembl gene identifiers and a descriptive contrast identifier serve as row and column names, respectively.</p> <p><strong>Data set 8. Differential expression analyses across murine RNA-Seq sample groups: log fold changes. </strong>As in Data set 7, but for mouse genes.</p> <p><strong>Data set 9. Differential expression analyses across simian RNA-Seq sample groups: log fold changes. </strong>As in Data set 7, but for chimpanzee genes.</p> <p><strong>Data set 10. Differential expression analyses across human RNA-Seq sample groups: false discovery rates. </strong>The file contains false discovery rates (FDR) for the differential expression analyses summarized in Data set 7. Format, column and row names as in Data set 7.</p> <p><strong>Data set 11. Differential expression analyses across murine RNA-Seq sample groups: false discovery rates. </strong>As in Data set 10, but for mouse genes.</p> <p><strong>Data set 12. Differential expression analyses across simian RNA-Seq sample groups: false discovery rates. </strong>As in Data set 10, but for chimpanzee genes.</p> <p><strong>Data set 13. Quantification of alternative splicing events across human RNA-Seq samples. </strong>The file contains ‘percent spliced in’ (PSI) values computed by SUPPA (<a href="https://github.com/comprna/SUPPA">https://github.com/comprna/SUPPA</a>) for annotated alternative splicing events (inferred from the transcript annotation of the human genome, Ensembl release 84; <a href="http://www.ensembl.org/">http://www.ensembl.org/</a>). The format is a compressed (GZIP) tab-separated transcript-by-sample matrix. SUPPA-provided event identifiers and a combined Sequence Read Archive study/sample name identifier serve as row and column names, respectively.</p> <p><strong>Data set 14. Quantification of alternative splicing events across murine RNA-Seq samples. </strong>As in Data set 13, but for mouse alternative splicing events.</p> <p><strong>Data set 15. Differential splicing analyses across human RNA-Seq sample groups: differences in ‘percent spliced in’ (ΔPSI). </strong>The file contains ΔPSI values for human alternative splicing events (as in Data set 13). The RNA-Seq sample group contrasts are listed in Additional file 3 of the accompanying publication. Values were inferred by SUPPA’s diffSplice functionality (<a href="https://github.com/comprna/SUPPA">https://github.com/comprna/SUPPA</a>). The format is a compressed (GZIP) tab-separated gene-by-comparison matrix. SUPPA event identifiers and a descriptive contrast identifier serve as row and column names, respectively.</p> <p><strong>Data set 16. Differential splicing analyses across murine RNA-Seq sample groups: differences in ‘percent spliced in’ (ΔPSI). </strong>As in Data set 15, but for mouse alternative splicing events.</p> <p><strong>Data set 17. Differential splicing analyses across human RNA-Seq sample groups: P values. </strong>The file contains P values for the differential splicing analysis of human alternative splicing events summarized in Data set 15. Format, column and row names as in Data set 15.</p> <p><strong>Data set 18. Differential splicing analyses across murine RNA-Seq sample groups: P values. </strong>The file contains P values for the differential splicing analysis of mouse alternative splicing events summarized in Data set 16. Format, column and row names as in Data set 15.</p> <p><strong>Data set 19. Transcript expression across murine RNA-Seq time course data: estimated read counts. </strong>As in Data set 2, but for the time course data generated for the accompanying publication.</p> <p><strong>Data set 20. Transcript expression across murine RNA-Seq time course data: estimated transcript abundances. </strong>As in Data set 5, but for the time course data generated for the accompanying publication.</p> <p><strong>Data set 21. Quantification of alternative splicing events across murine RNA-Seq time course data. </strong>As in Data set 14, but for the time course data generated for the accompanying publication.</p>
Tree-ring measurements from permanent study plot in old-growth hemlock-hardwood forest, Dukes RNA, Hiawatha NF, Marquette Co., MI
This package includes tree growth-ring widths for increment cores collected from a long-term 'macroplot' established in old-growth hemlock-northern hardwoods forest at the Dukes Research Natural Area/Dukes Experimental Forest in the Hiawatha National Forest in Marquette Co., MI. Tree demographic monitoring data for the entire ca. 3.0 ha macroplot, from 1992 to 2019, are available in the EDI package edi.1526.1. In 1993, 1994 and 1995, increment cores were taken for all 'core-able' trees greater than ~ 10 cm diameter for a subsection of the macroplot about 1 ha in area. Trees that were obviously badly rotten and hollow or steeply leaning were not cored. Cores are not cross-dated. See Methods for more details. This data-package may be cross-referenced to the demographic data in edi.1526.1 using stem numbers.
Stream cross-section profiles in the Andrews Experimental Forest and Hagan Block RNA 1978 to 2011
This database consists of periodically resurveyed cross-section profiles for 5 sites (stream reaches) located in the H.J. Andrews Experimental Forest on 3rd to 5th-order stream channels over the period 1978-present, and also includes historical data from Shorter Creek within the Andrews and two sites (North and South) on the Hagan Creek in the Hagan Block RNA. Each of these sites includes 10-20 cross sections spaced about 10-50 m apart. Observations include cross-sectional geometry based on surveys relative to a horizontal line between permanent stations on opposite sides of the stream, substrate type based on pebble counts and notes on extent of bedrock, and photo-documentation of the sites. Sampling was primarily done annually until 2000 and opportunistically since then depending on funding and noted geomorphic changes in response to floods. Sampling at least every 5 years is desired. Data tables cover 1. elevation and distance measurements between fixed end points, 2. pebble counts for particle size, 3. measures of extent of exposed bedrock, 4. vegetation (1984 only) in terms of % cover by spp., % cover of substrate types, and plant measurements for making biomass estimates with allometric equations, but such estimates are not presented. The original raw cross-section data for these sites has been adjusted to align profiles for consecutive year comparisons at each cross-section location to correct for imprecision in the survey methodology. John Faustini, in 1998, further corrected this cross-section data from 1978-1998 for the five main sites in database GS019. Arianna Goodman, in 2021, performed additional corrections on the data, including more recent collections, 1978-2011, also in database GS019. The Shorter Creek cross-section surveys took place in 1978-1982 only. The Shorter Creek cross-section surveys took place in 1978-1982 only.
Supplementary Data for MOCCASIN: A method for correcting known and unknown confounders in RNA-Seq-based splicing analysis
<p>Contents</p> <ol> <li><strong>moccasin_paper_env.yaml</strong>: conda environment file with R and Python packages and modules needed to reproduce analyses.</li> <li><strong>FigureReproduction.zip</strong>: data and code to reproduce main and supplemental figures.</li> <li><strong>MOCCASIN_ExampleDataset.zip</strong>: A small subset of the simulated data with example code to run MOCCASIN.</li> <li><strong>encode_corrected.zip</strong>: Folder with batch-corrected ENCODE differential splicing quantifications (dPSI).</li> </ol> <p> </p> <p> </p> <p>(1) <strong>moccasin_paper_env.yaml</strong></p> <p>Use the moccasin_paper_env.yaml file to create a conda environment from which all analyses for the paper can be reproduced.</p> <pre><code class="language-bash"># need to first install conda. See here: # https://docs.conda.io/en/latest/miniconda.html # Next, create a conda environment: conda env create --name moccasin_paper_env --file moccasin_paper_env.yaml --force # Activate the environment: conda activate moccasin_paper_env</code></pre> <p><br> The only Python packages not included in this environment are MAJIQ & VOILA. Please see majiq.biocipers.org for installation instructions.</p> <p> </p> <p> </p> <p>(2) <strong>FigureReproduction.zip</strong></p> <p>Within FigureReproduction are folders with code and data to reproduce the main and supplemental figures of the publication. Each folder contains data, script(s) and a README.txt with instructions on how to reproduce figures.</p> <p> </p> <p> </p> <p>(3) <strong>MOCCASIN_ExampleDataset.zip</strong></p> <p>Within this folder is an example dataset to test MOCCASIN. The README.txt file contains detailed line-by-line instructions for how to run MOCCASIN and do post-MOCCASIN analyses. In this example, we show how to run MOCCASIN on a group of .majiq samples with one known confounding effect. Also demonstrated is how to run an "explore unknown residuals" analysis as described in the detailed methods in the supplemental of the paper. </p> <p> </p> <p> </p> <p>(4) <strong>encode_corrected.zip</strong></p> <p>Includes a file called ENCODE_BeforeAndAfterMOCCASIN.voila.tsv.zip which includes LSV quantifications before and after MOCCASIN. Each row in the file represents a junction from an LSV. Each column header starts with the prefix "BeforeMOCCASIN" or "AfterMOCCASIN" and headers ending in dPSI corresponds to the dPSI of an ENCODE knockdown vs control experiment. </p>
RDF Representation of RNA Metabolism Evolution data - version 3 (diagrammed in https://zenodo.org/deposit/47641/)
<p>Version 3 (replaces http://doi.org/10.5281/zenodo.50496)<br> <br> Protein complexes involved in RNA Metabolism; individual proteins and their orthologues through a wide range of fungal species spanning much of the kingdom (using yeast as the primary seed for orthology search, and using the EMBL-EBI orthologue database to identify orthologues). For each family of orthologues, the protein domain structure is determined, and then the presence/absence of that domain is evaluated in each of the species. The data is presented in RDF, and is visualized in the form of Heat Maps in http://doi.org/10.5281/zenodo.47641</p>
Cherri - Accurate detection of functional RNA-RNA interactions sites
<p><strong>CheRRI</strong> - Pipeline for the Identification of putative RNA-RNA interaction sites.</p> <p> </p> <p>This repository contains all CheRRI's models computed and mentioned in the content.txt, listing data and their descriptions. All models can be used to classify interaction sites in CheRRI's eval mode.</p> <p> </p> <p>The source code for CheRRI is avalbile on <a href="https://github.com/BackofenLab/Cherri#install-cherri-conda-package">GitHub</a> and can be cited using this Software Heritage citation:</p> <ul> <li><span>Müller T, Mautner S, Videm P, Eggenhofer F, Raden M, Backofen R (2024) CheRRI - Accurate classification of the biological relevance of putative RNA-RNA interaction sites (Version 0.8). [Computer software]. Software Heritage, <a href="https://archive.softwareheritage.org/swh:1:snp:ebac091117f9c46fb5f0fedd3ef23ec2905ced6c;origin=https://github.com/BackofenLab/Cherri">https://archive.softwareheritage.org/swh:1:snp:ebac091117f9c46fb5f0fedd3ef23ec2905ced6c;origin=https://github.com/BackofenLab/Cherri</a></span></li> </ul> <div> <div> <div> <p>The pipeline contains Machine Learning segments which were annotated using DOME:</p> </div> </div> </div> <ul> <li><span><a href="https://dome.ds-wizard.org/projects/74d0e01c-6374-41e9-93b8-2889d6a8fe25">https://dome.ds-wizard.org/projects/74d0e01c-6374-41e9-93b8-2889d6a8fe25</a></span></li> </ul>
xPore: Identification of differential RNA modifications from nanopore direct RNA sequencing
<p>xPore is a Python package for identification and quantification of differential RNA modifications from direct RNA sequencing.</p> <p>The detailed usage is documented at <a href="https://xpore.readthedocs.io/en/latest/">https://xpore.readthedocs.io/en/latest</a>, while all scripts and source code are available at <a href="https://github.com/GoekeLab/xpore">https://github.com/GoekeLab/xpore</a>.</p> <p>All the preprocessed datasets used in the paper are provided here. </p> <p>Please cite our paper below when using these data.<br> Ploy N. Pratanwanich et al. "Detection of differential RNA modifications from direct RNA sequencing of human cell lines." bioRxiv (2020).</p>
bollito: a flexible pipeline for comprehensive single-cell RNA-seq analyses - Melanoma tutorial
<p>Downsampled version of the melanoma dataset originally published by <em><a href="https://genome.cshlp.org/content/28/9/1353">Ho et al </a>(1)</em>. The dataset is composed by cells from the 451Lu cell line. There are two samples available:</p> <table> <tbody> <tr> <td><strong>Name</strong></td> <td><strong>Description</strong></td> <td><strong>R1/R2</strong></td> </tr> <tr> <td>451LU</td> <td>Parental cell line</td> <td>2500K_451LU_L003_R*_001.fastq.gz</td> </tr> <tr> <td>451LUBR3</td> <td>Vemurafenib-resistant sample treated with targeted BRAF inhibitors</td> <td>500K_451LUBR3_L004_R*_001.fastq.gz</td> </tr> </tbody> </table> <p><br> (1) Ho YJ, Anaparthy N, Molik D, et al. Single-cell RNA-seq analysis identifies markers of resistance to targeted BRAF inhibitors in melanoma cell populations. <em>Genome Res</em>. 2018;28(9):1353-1363. doi:10.1101/gr.234062.117</p>
RNA sequencing data for bleomycin exposed THP-1 macrophages
<p>This dataset contains normalized counts matrices, from dds_deseq objects, from DeSeq2 analysis of RNA sequencing data, from THP-1 macrophages exposed to multiple doses of bleomycin in the range of 0-100µg/ml for 24H, 48H or 72H.</p>
RNA datasets to derive predictors for immune checkpoint inhibitor therapy of non-small cell lung cancer
<p>Nanostring nCounter datasets and corresponding clinical data of tumor samples of patients with advanced NSCLC who received anti-PD-1 immuntherapy. Prospectively divided into a discovery and a validation cohort.</p> <p>Please cite the corresponding publication in Annals of Oncology (10.1093/annonc/mdz049)</p>
RNA-seq Dataset from Crowley et. al. 2015
<p>This dataset was uploaded to this repository by Joshua P. Zitovsky and Michael I. Love with permission from the original authors, due to the fact that the original dataset is not currently hosted in a stable repository. If this dataset is used in original research leading to published work, please cite the Crowley et. al. 2015 paper.</p> <p>This repository contains an RNA-seq dataset based on the allelic expression study by Crowley et. al. (2015). The study took mice from three divergent inbred strains (CAST/EiJ, PWK/PhJ and WSB/EiJ) and performed a diallel cross. The data set contains allele-specific expression (ASE) counts for 72 mice and 23,297 genes in the resulting cross, with 12 mice of each possible parent combination, and an equal number of males and females within each parent combination. Sequencing was performed with the Illumina HiSeq 2000 platform to generate 100-bp paired-end reads and following the TruSeq RNA Sample Preparation v2 protocol.</p>
Datasets and Jupyter notebook for the structural analysis of protein-RNA interface evolution
<p>The present repository contains data and code related to our manuscript "Structural comparison of protein-RNA homologous interfaces reveals widespread overall conservation contrasted with versatility in polar contacts". In the manuscript, we analyze the evolution of protein-RNA interfaces by building a dataset of protein-RNA interologs (homologous interfaces) and exploring how interface contacts are conserved between homologous interfaces, as well as possible explanations for non-conserved contacts.</p> <p>This repository contains the following files:</p> <ul> <li>DataAnalysisNotebook.ipynb is a Jupyter notebook to reproduce contact conservation analysis and all figures from our manuscript, and to explore data</li> <li>env.yaml is an environment file in order to build a Conda/Mamba environment to run the Jupyter notebook </li> <li>2022-02-21-PDB.csv contains data from the PDB about 3D structures of complexes containing interacting protein and RNA chains (PDB structure identifier, chain identifiers, experimental technique and resolution)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.tsv contains more detailed information about interacting protein and RNA chains from these complexes (PDB and chain identifiers, protein and RNA size, interface size and number of contacts)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.txt.selectXE_2.50_p30_r10_pi5_ri5_rep_bc-100.out_RNAcl_0.99.tsv contains the same detailed information, restricted to the filtered dataset used as a starting point in our interolog search pipeline</li> <li>PDBinterfaceAlign.csv contains information about the structural alignment of pairs of protein-RNA interactions (structural alignment TM-scores, sequence identity and coverage)</li> <li>DataInterologsParam.tsv contains information about a pre-filtered set of 2587 potential interologs (including interface RMSD, sequence identity and coverage and interface size)</li> <li>DataInterologsContactsFixedSASA.tsv contains detailed information about conserved and non-conserved contacts in the final set of 2022 interologs (atomic contacts, apolar contacts, hydrogen bonds, salt bridges and stacking information for aminoacid-nucleotide pairs, as well as information about whether each belongs to the interface, secondary structures, and the aminoacid surface accessibility and evolutionary conservation metrics) - compared to version 1, the calculation of solvent accessibility was fixed for a number of interolog pairs</li> <li>DataCons.csv contains precomputed contact conservation metrics for each of the 2022 interolog pairs, for fast reproduction of manuscript figures</li> <li>DataInterologsContactsResampledMaintainStructSeqId.tsv, DataInterologsContactsShuffled.tsv and DataInterologsShuffled.tsv relate to baselines computed for contact conservation assessment</li> <li>clan.txt, clan_membership.txt, ecod.latest.domains.uniq.txt, rfam_interfaces_977.txt, DataGroupsECOD.tsv, DataGroupesRFAM.tsv, DataGroupsRFAMClan.tsv, DataInterfaceGroupsECOD.tsv and DataInterfaceGroupsRFAM.tsv relate to the ECOD (respectively Rfam) classification of protein domains (respectively RNA) in protein-RNA interfaces from our dataset</li> <li>ListeIntraHbonds.pkl and ListeIntraSaltBridges.pkl are pickle-format data files containing intra-molecular hydrogen bonds and salt bridges (respectively) that are used to analyse scenarii of compensation for non-conserved polar contacts.</li> </ul>
Expansion of the global RNA virome reveals diverse clades of bacteriophages
<p>This deposit is intended to contain the various data generated as part of the RNA Virus in MetaTranscriptomes project ("RVMT"). This initial version is released ahead of time, near the time of submission, in hopes of providing a long lasting resource for the general scientific community. Note well - The authors listed in this initial version release are a partial list only. The RNA Virus in MetaTranscriptomes consortium is a project with over 90 researches from various institutions (see below).</p> <p>High-throughput RNA sequencing offers broad opportunities to explore the Earth RNA virome. Mining 5,150 diverse metatranscriptomes uncovered >2.5 million RNA virus contigs. Analysis of >330,000 RNA-dependent RNA polymerases (RdRPs) shows that this expansion corresponds to a 5-fold increase of the known RNA virus diversity. Gene content analysis revealed multiple protein domains previously not found in RNA viruses and implicated in virus-host interactions. Extended RdRP phylogeny supports the monophyly of the five established phyla and reveals two putative additional bacteriophage phyla and numerous putative additional classes and orders. The dramatically expanded phylum <em>Lenarviricota</em>, consisting of bacterial and related eukaryotic viruses, now accounts for a third of the RNA virome. Identification of CRISPR spacer matches and bacteriolytic proteins suggests that subsets of picobirnaviruses and partitiviruses, previously associated with eukaryotes, infect prokaryotic hosts.</p> <p>The RNA Virus in metatranscriptomes consortium:<br> Adrienne B. Narrowe, Alexander J. Probst, Alexander Sczyrba, Annegret Kohler, Armand Séguin, Ashley Shade, Barbara J. Campbell, Björn D. Lindahl, Brandi Kiel Reese, Breanna M. Roque, Chris DeRito, Colin Averill, Daniel Cullen, David A. C. Beck, David A. Walsh, David M. Ward, Dongying Wu, Emiley Eloe-Fadrosh, Eoin L. Brodie, Erica B. Young, Erik A. Lilleskov, Federico J. Castillo, Francis M. Martin, Gary R. LeCleir, Graeme T. Attwood, Hinsby Cadillo-Quiroz, Holly M. Simon, Ian Hewson, Igor V. Grigoriev, James M. Tiedje, Janet K. Jansson, Janey Lee, Jean S. VanderGheynst, Jeff Dangl, Jeff S. Bowman, Jeffrey L. Blanchard, Jennifer L. Bowen, Jiangbing Xu, Jillian F. Banfield, Jody W Deming, Joel E. Kostka, John M. Gladden, Josephine Z Rapp, Joshua Sharpe, Katherine D. McMahon, Kathleen K. Treseder, Kay D. Bidle, Kelly C. Wrighton, Kimberlee Thamatrakoln, Klaus Nusslein, Laura K. Meredith, Lucia Ramirez, Marc Buee, Marcel Huntemann, Marina G. Kalyuzhnaya, Mark P Waldrop, Matthew B Sullivan, Matthew O. Schrenk, Matthias Hess, Michael A. Vega, Michelle A. O’Malley, Monica Medina, Naomi E. Gilbert, Nathalie Delherbe, Olivia U. Mason, Paul Dijkstra, Peter F. Chuckran, Petr Baldrian, Philippe Constant, Ramunas Stepanauskas, Rebecca A. Daly, Regina Lamendella, Robert J Gruninger, Robert M. McKay, Samuel Hylander, Sarah L. Lebeis, Sarah P Esser, Silvia G. Acinas, Steven S. Wilhelm, Steven W. Singer, Susannah S. Tringe, Tanja Woyke, TBK Reddy, Terrence H. Bell, Thomas Mock, Tim McAllister, Vera Thiel, Vincent J. Denef, Wen-Tso Liu, Willm Martens-Habbena, Xiao-Jun Allen Liu, Zachary S. Cooper, Zhong Wang. For the full list of authors and related information, please see the spreadsheet tittle "Table S9 - Consortium coauthorship" available in this collection in the folder named "Tables".</p>
Understory plant community data from repeated plot sampling (1978-2019) in old-growth northern hardwood forest, northern Michigan (Dukes RNA, Hiawatha National Forest)
This data-set includes long-term, permanent-plot-based data for understory plant communities in old-growth mixed northern hardwood-hemlock forest and forested peatland in the Upper Great Lakes region. Data for over 900 understory quadrats (all associated with long-term canopy data from larger permanent plots) included multiple (2-5) remeasurements over 23-40 years, with longest periods and most remeasurements for upland forest types. The Dukes Research Natural Area (RNA) (https://www.fs.usda.gov/research/nrs/rnas/locations/dukes) in the Hiawatha National Forest (Marquette Co., MI) includes ca. 100 ha of largely unlogged, original forest. Publications cited below include more detailed information about the site. About half of the RNA supports upland forests intergrading from hemlock (Tsuga candensis) dominance to mixtures of hemlock and northern hardwoods species. Sugar maple (Acer saccharum) is dominant over much of the upland area, with, locally, significant admixtures of beech (Fagus grandifolia), yellow birch (Betula alleghaniensis), and red maple (Acer rubrum). Topographic relief is very slight with total elevational change within the RNA only about 10 m. The stand is within a few km of the western limit of the continuous range of beech. In 1935, 248 continuing forest inventory (CFI) plots (circular, 0.2 acre) were established on a regular grid throughout the RNA, and these have been the subject of repeated sampling through 2018-2019 and support continuing long-term study addressing canopy tree communities (canopy data to be deposited in a separate project). Examples of resulting publications are cited elsewhere in metadata, and can provide more detailed information about the RNA. In 1978-80, U.S. Forest Service researchers, directed by Jan Schultz and Frederick Metzger, initiated studies of understory communities, including herbaceous species and woody seedlings. Data were derived from four sub-quadrats within each of the CFI plots. These quadrats were re-estab
RNAPosers: Machine Learning Classifiers For RNA-Ligand Poses [Data Set]
<ul> <li>This dataset contains the decoys poses used to train and test RNAPosers, a set of RNA-ligand pose classifiers.</li> <li>The folder of each RNA-ligand complex (identified using its PDB ID) contains: <ul> <li>Ligand SMILES: lig.smi</li> <li>Ligand coordinate: lig.sd</li> <li>Receptor coordinate: receptor.mol2</li> <li>Pose coordinates: poses.sd</li> <li>Pose similarity data: rmsd.txt</li> </ul> </li> </ul>
SIRAH-CoV2 initiative: RNA binding domain of nucleocapsid phosphoprotein (PDB id:6VYO)
<p>This dataset contains the trajectory of a 10 microseconds-long coarse-grained molecular dynamics simulation of SARS-CoV2 RNA binding domain of the nucleocapsid phosphoprotein in its APO form with Zn ions bound (PDB id:6VYO, Bioassembly 1). Simulations were performed using the SIRAH force field running with the Amber18 package at the Uruguayan National Center for Supercomputing (ClusterUY) under the conditions reported in <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00006">Machado et al. JCTC 2019</a>, adding 150 mM NaCl according to <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00953">Machado & Pantano JCTC 2020</a>. Zinc ions were parameterized as reported in <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.0c00160">Klein et al. 2020</a>.</p> <p>The file 6VYO_SIRAHcg_rawdata.tar contains all the raw information required to visualize (on VMD), analyze, backmap, and eventually continue the simulations using Amber18 or higher. Step-By-Step tutorials for running, visualizing, and analyzing CG trajectories using <a href="https://academic.oup.com/bioinformatics/article/32/10/1568/1743152">SirahTools</a> can be found at www.sirahff.com. Additionally, the file 6VYO_SIRAHcg_10us_prot.tar contains only the protein coordinates, while 6VYO_SIRAHcg_10us_prot_skip10ns.tar contains one frame every 10ns.</p> <p>To take a quick look at the trajectory:</p> <p>1- Untar the file 6VYO_SIRAHcg_10us_prot_skip10ns.tar</p> <p>2- Open the trajectory on VMD using the command line:</p> <p>vmd 6VYO_SIRAHcg_prot_10us_skip10ns.prmtop 6VYO_SIRAHcg_prot_10us_skip10ns.ncrst 6VYO_SIRAHcg_prot_10us_skip10ns.nc -e sirah_vmdtk.tcl</p> <p>Note that you can use normal VMD drawing methods as vdw, licorice, etc., and coloring by restype, element, name, etc. </p> <p>This dataset is part of the SIRAH-CoV2 initiative.</p> <p>For further details, please contact Florencia Klein (fklein@pasteur.edu.uy) or Sergio Pantano (spantano@pasteur.edu.uy).</p>
SIRAH-CoV2 initiative: NSP9 RNA binding protein (PDBid:6W4B)
<p>This dataset contains the trajectory of a 10 microseconds-long coarse-grained molecular dynamics simulation of SARS-CoV2 NSP9 RNA binding protein (PDB id: 6W4B, Bioassembly 1). Simulations have been performed using the SIRAH force field running with the Amber18 package at the Uruguayan National Center for Supercomputing (ClusterUY) under the conditions reported in <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00006">Machado et al. JCTC 2019</a>, adding 150 mM NaCl according to <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00953">Machado & Pantano JCTC 2020</a>. </p> <p>The file 6W4B_SIRAHcg_rawdata.tar contains all the raw information required to visualize (on VMD), analyze, backmap, and eventually continue the simulations using Amber18 or higher. Step-By-Step tutorials for running, visualizing, and analyzing CG trajectories using <a href="https://academic.oup.com/bioinformatics/article/32/10/1568/1743152">SirahTools</a> can be found at www.sirahff.com.</p> <p>Additionally, the file 6W4B_SIRAHcg_10us_prot.tar contains only the protein coordinates, while 6W4B_SIRAHcg_10us_prot_skip10ns.tar contains one frame every 10ns.</p> <p>To take a quick look at the trajectory:</p> <p>1- Untar the file 6W4B_SIRAHcg_10us_prot_skip10ns.tar</p> <p>2- Open the trajectory on VMD using the command line:</p> <p>vmd 6W4B_SIRAHcg_prot.prmtop 6W4B_SIRAHcg_prot.ncrst 6W4B_SIRAHcg_prot_10us_skip10ns.nc -e sirah_vmdtk.tcl</p> <p>Note that you can use normal VMD drawing methods as vdw, licorice, etc., and coloring by restype, element, name, etc. </p> <p>This dataset is part of the SIRAH-CoV2 initiative.</p> <p>For further details, please contact Sergio Pantano (spantano@pasteur.edu.uy).</p>
SIRAH-CoV2 initiative: Nucleocapsid protein N-terminal RNA binding domain (PDB id:6M3M)
<p>This dataset contains the trajectory of a 10 microseconds-long coarse-grained molecular dynamics simulation of SARS-CoV2 Nucleocapsid protein N-terminal RNA binding domain (PDB id:6M3M). Simulations have been performed using the SIRAH force field running with the Amber18 package at the Uruguayan National Center for Supercomputing (ClusterUY) under the conditions reported in <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00006">Machado et al. JCTC 2019</a>, adding 150 mM NaCl according to <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00953">Machado & Pantano JCTC 2020</a>. </p> <p>The files 6M3M_SIRAHcg_rawdata.tar contains all the raw information required to visualize (on VMD), analyze, backmap, and eventually continue the simulations using Amber18 or higher. Step-By-Step tutorials for running, visualizing, and analyzing CG trajectories using <a href="https://academic.oup.com/bioinformatics/article/32/10/1568/1743152">SirahTools</a> can be found at www.sirahff.com.</p> <p>Additionally, the file 6M3M_SIRAHcg_10us_prot.tar contains only the protein coordinates, while 6M3M_SIRAHcg_10us_prot_skip10ns.tar contains one frame every 10ns.</p> <p>To take a quick look at the trajectory:</p> <p>1- Untar the file 6M3M_SIRAHcg_10us_prot_skip10ns.tar</p> <p>2- Open the trajectory on VMD using the command line:</p> <p>vmd 6W4B_SIRAHcg_prot.prmtop 6W4B_SIRAHcg_prot.ncrst 6W4B_SIRAHcg_prot_10us_skip10ns.nc -e sirah_vmdtk.tcl</p> <p>Note that you can use normal VMD drawing methods as vdw, licorice, etc., and coloring by restype, element, name, etc. </p> <p>This dataset is part of the SIRAH-CoV2 initiative.</p> <p>For further details, please contact Florencia Klein (fklein@pasteur.edu.uy) or Sergio Pantano (spantano@pasteur.edu.uy).</p>
SIRAH-CoV2 initiative: RNA-dependent RNA polymerase in complex with cofactors Nsp7 and Nsp8 (PDB id:7BTF)
<p>This dataset contains the trajectory of a 10 microseconds-long coarse-grained molecular dynamics simulation of SARS-CoV2 RNA-dependent RNA polymerase in complex with cofactors Nsp7 and Nsp8 and Zinc (PDB id: 7BTF). Simulations have been performed using the SIRAH force field running with the Amber18 package at the Uruguayan National Center for Supercomputing (ClusterUY) under the conditions reported in <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00006">Machado et al. JCTC 2019</a>, adding 150 mM NaCl according to <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00953">Machado & Pantano JCTC 2020</a>. Zinc ions were parameterized as reported in <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.0c00160">Klein et al. 2020</a>.</p> <p>The files 7BTF_SIRAHcg_rawdata_0-2us.tar, 7BTF_SIRAHcg_rawdata_2-6us.tar, and 7BTF_SIRAHcg_rawdata_6-10us.tar, contain all the raw information required to visualize (on VMD), analyze, backmap, and eventually continue the simulations using Amber18 or higher. Step-By-Step tutorials for running, visualizing, and analyzing CG trajectories using <a href="https://academic.oup.com/bioinformatics/article/32/10/1568/1743152">SirahTools</a> can be found at www.sirahff.com.</p> <p>Additionally, the file 7BTF_SIRAHcg_10us_prot.tar contains only the protein coordinates, while 7BTF_SIRAHcg_10us_prot_skip10ns.tar contains one frame every 10ns.</p> <p>To take a quick look at the trajectory:</p> <p>1- Untar the file 7BTF_SIRAHcg_10us_prot_skip10ns.tar</p> <p>2- Open the trajectory on VMD 1.9.3 using the command line:</p> <p>vmd 7BTF_SIRAHcg_prot.prmtop 7BTF_SIRAHcg_prot.ncrst 7BTF_SIRAHcg_prot_10us_skip10ns.nc -e sirah_vmdtk.tcl</p> <p>Note that you can use normal VMD drawing methods as vdw, licorice, etc., and coloring by restype, element, name, etc. </p> <p>This dataset is part of the SIRAH-CoV2 initiative.</p> <p>For further details, please contact Martin Soñora (msonora@pasteur.edu.uy) Sergio Pantano (spantano@pasteur.edu.uy).</p>
RNA sequencing dataset for prediction of liver hepatocellular carcinoma using SIMON analysis
<p>The LIHC dataset was used for data mining and for the generation of machine learning model for the detection of liver hepatocellular carcinoma cells (LIHC) using the SIMON platform as described in the "SIMON: open-source knowledge discovery platform" publication (<a href="https://doi.org/10.1101/2020.08.16.252767">https://doi.org/10.1101/2020.08.16.252767</a>). The LIHC dataset was obtained from the <em>GSEABenchmarkeR</em> package ( <a href="https://doi.org/10.1093/bib/bbz158">https://doi.org/10.1093/bib/bbz158</a>) and it contains RNA expression data from 374 liver hepatocellular carcinoma (LIHC) cells and 50 adjacent normal cells.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.