Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
40
datasets available to search
ShareScore release 0.9.0
Dataset results
40 results for “Protein-RNA”
Datasets and Jupyter notebook for the structural analysis of protein-RNA interface evolution
<p>The present repository contains data and code related to our manuscript "Structural comparison of protein-RNA homologous interfaces reveals widespread overall conservation contrasted with versatility in polar contacts". In the manuscript, we analyze the evolution of protein-RNA interfaces by building a dataset of protein-RNA interologs (homologous interfaces) and exploring how interface contacts are conserved between homologous interfaces, as well as possible explanations for non-conserved contacts.</p> <p>This repository contains the following files:</p> <ul> <li>DataAnalysisNotebook.ipynb is a Jupyter notebook to reproduce contact conservation analysis and all figures from our manuscript, and to explore data</li> <li>env.yaml is an environment file in order to build a Conda/Mamba environment to run the Jupyter notebook </li> <li>2022-02-21-PDB.csv contains data from the PDB about 3D structures of complexes containing interacting protein and RNA chains (PDB structure identifier, chain identifiers, experimental technique and resolution)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.tsv contains more detailed information about interacting protein and RNA chains from these complexes (PDB and chain identifiers, protein and RNA size, interface size and number of contacts)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.txt.selectXE_2.50_p30_r10_pi5_ri5_rep_bc-100.out_RNAcl_0.99.tsv contains the same detailed information, restricted to the filtered dataset used as a starting point in our interolog search pipeline</li> <li>PDBinterfaceAlign.csv contains information about the structural alignment of pairs of protein-RNA interactions (structural alignment TM-scores, sequence identity and coverage)</li> <li>DataInterologsParam.tsv contains information about a pre-filtered set of 2587 potential interologs (including interface RMSD, sequence identity and coverage and interface size)</li> <li>DataInterologsContactsFixedSASA.tsv contains detailed information about conserved and non-conserved contacts in the final set of 2022 interologs (atomic contacts, apolar contacts, hydrogen bonds, salt bridges and stacking information for aminoacid-nucleotide pairs, as well as information about whether each belongs to the interface, secondary structures, and the aminoacid surface accessibility and evolutionary conservation metrics) - compared to version 1, the calculation of solvent accessibility was fixed for a number of interolog pairs</li> <li>DataCons.csv contains precomputed contact conservation metrics for each of the 2022 interolog pairs, for fast reproduction of manuscript figures</li> <li>DataInterologsContactsResampledMaintainStructSeqId.tsv, DataInterologsContactsShuffled.tsv and DataInterologsShuffled.tsv relate to baselines computed for contact conservation assessment</li> <li>clan.txt, clan_membership.txt, ecod.latest.domains.uniq.txt, rfam_interfaces_977.txt, DataGroupsECOD.tsv, DataGroupesRFAM.tsv, DataGroupsRFAMClan.tsv, DataInterfaceGroupsECOD.tsv and DataInterfaceGroupsRFAM.tsv relate to the ECOD (respectively Rfam) classification of protein domains (respectively RNA) in protein-RNA interfaces from our dataset</li> <li>ListeIntraHbonds.pkl and ListeIntraSaltBridges.pkl are pickle-format data files containing intra-molecular hydrogen bonds and salt bridges (respectively) that are used to analyse scenarii of compensation for non-conserved polar contacts.</li> </ul>
Protein-RNA complex simulation
<p>Simulation trajectory of TSEN/pre-tRNAArgTCT generated with GROMACS. </p> <p> </p> <p>A hybrid model of truncated TSEN/pre-tRNAArgTCT with the additional TSEN2 domain from AlphaFold Jumper et al, 2021</p> <p>and in silico modeled intron bases 37 to 43 was subjected to all-atom molecular dynamics simulations</p> <p>for assessment of flexibility. Simulation trajectory can be visualised with PyMOL or VMD.</p> <p> </p> <p>The protein was described by the AMBER14ff (Maier, et al 2015), and the RNA with the OL3 force field (Zgarbova et al, 2011)</p> <p>and the TIP3P water model was employed (Jorgensen et al, 1983).</p>
Deep learning model for characterizing protein-RNA interactions from sequence at single-base resolution
<p> </p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip.h5</a> - This file contains the training, validation, and test data for the Reformer model.</p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip_bc.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip_bc.h5</a> - This file contains the training, validation, and test data for the Reformer-BC model.</p> <p><a href="https://zenodo.org/api/records/14027315/draft/files/Reformer-code.zip/content" target="_blank" rel="noopener">Reformer-code.zip</a> - This file contains the training code of Reformer.</p>
FTDMP docking results for protein-protein, protein-DNA, protein-RNA benchmarks
<p><strong>FTDMP docking results for protein-protein, protein-DNA, protein-RNA benchmarks. </strong></p> <p>FTDMP is a software system for running docking experiments and scoring/ranking multimeric models. This dataset contains FTDMP docking results for protein-protein, protein-DNA, protein-RNA benchmarks. The FTDMP framework itself is available at https://github.com/kliment-olechnovic/ftdmp. </p> <p>Every *.tar.gz file in this dataset contains two folders: results for unbound-unbound and bound-bound docking. These folders contain results for the benchmark cases:</p> <p>252 folders with results for the protein-protein docking benchmark cases [1].<br>47 folders with results for the protein-DNA docking benchmark cases [2].<br>42 folders with results for the protein-RNA docking benchmark cases [3-6]. </p> <p>Every folder is named according to the PDB ID of the complex. The folders contain:</p> <p>1. A subfolder named <em>relaxed_top_complexes</em>. This subfolder contains 200 pdb files of relaxed [7] top docking models.<br>2. A text file named <em>scoring_results-ranks.txt</em>. It contains the names of the models (that are in the relaxed_top_complexes folder) in the ranked order. This means that the first model in the file is considered to be the best prediction by the FTDMP framework.<br>3. A text file named <em>cad_scores.txt</em>. It contains interface CAD-score and binding site CAD-score [8] results for every model.<br>4. A text file named <em>rmsd_results.txt</em>, which is available only for protein-DNA and protein-RNA cases. The file contains ligand-RMSD values for the models, where the DNA/RNA is considered as the ligand.<br>5. A text file named <em>DockQ_results.txt</em>, which is available only for the protein-protein docking cases. The file contains DockQ [9] results for every model, as well as model accuracy based on CAPRI criteria (Incorrect, Acceptable, Medium, High)<br>6. A text file named <em>binding_site_CAD-scores.txt</em>, which contains the binding site CAD-score from the <strong>protein</strong> side for RNA and DNA docking. This binding site CAD-score shows how accurately the ligand (DNA/RNA) was docked to the protein without taking the orientation of the ligand into consideration. In the case of protein-protein docking the binding site CAD-score file is available only for antibody-antigen docking targets and contains the binding site (epitope) CAD-score for the antigen. </p> <p>The ligand-RMSD, CAD-scores, and DockQ scores were all calculated by comparing the models to the corresponding targets. The target structures are available at <span>https://zenodo.org/records/10517524</span>. These target structures have the same residue numbering as the models available here. </p> <p>REFERENCES </p> <p>[1] Guest, J. D., et al. (2021). An expanded benchmark for antibody-antigen docking and affinity prediction reveals insights into antibody recognition determinants. Structure, 29(6), 606–621.e5.<br>[2] van Dijk, M., Bonvin, A.M. (2008). A protein-DNA docking benchmark. Nucleic Acids Res, 36, e88. <br>[3] Perez-Cano, L., et. Al. (2012). A protein-RNA docking benchmark (II): extended set from experimental and homology modeling data. Proteins, 80(7): 1872-1882. <br>[4] Huang, S.Y., Zou, X. (2013). A nonredundant structure dataset for benchmarking protein-RNA computational docking. J Comput Chem, 34(4): 311-318. <br>[5] Nithin, C., et. al. (2017). A non-redundant protein-RNA docking benchmark version 2.0. Proteins, 85(2) :256-267. <br>[6] Zheng, J., et al. (2020). P3DOCK: a protein-RNA docking webserver based on template-based and template-free docking. Bioinformatics, 36(1), 96–103. <br>[7] Eastman, P., et al.(2017). OpenMM 7: Rapid development of high performance algorithms for molecular dynamics. PLOS Comp. Biol., 13(7): e1005659. <br>[8] Olechnovic, K., Venclovas, C. (2020). Contact area-based structural analysis of proteins and their complexes using CAD-score. Methods Mol Biol, 2112, 75.<br>[9] Basu, S., Wallner, B. (2016). DockQ: A Quality Measure for Protein-Protein Docking Models. PLoS ONE 11(8): e0161879. </p>
The dynamics of protein-RNA interfaces using all-atom molecular dynamics simulations
<p>We investigated to characterize the dynamics of protein-RNA complexes and their interfaces at molecular level by performing a more systematic analysis. To get insights on the dynamics of protein-RNA complexes, all-atom MD simulations were generated for the manuscript "The dynamics of protein-RNA interfaces using all-atom molecular dynamics simulations". Nine protein-RNA complexes are studied in this work: 1ASY (an aspartyl-tRNA synthase/tRNA), 1JBS (a ribotoxin restrictocin/SRD RNA), 1MMS (a ribosomal protein L11/23S), 1OOA (a nuclear factor NF-kappaB p105 subunit/RNA aptamer), 1RKJ (a nucleolin/pre-rRNA), 2R8S (a FAB/P4-P6 RNA ribozyme domain), 2VPL (a 50S ribosomal protein/mRNA), 2ZM5 (a tRNA delta(2)-isopentenylpyrophosphate transferase/tRNA), 3IEV (a GTP-binding protein era/3' end of 16S rRNA). </p><p>Each folder for a complex is organised as followed:</p><ul><li>in <strong>complex</strong> there are the dry MD simulations for the complex protein-RNA with the starting structure</li><li>in <strong>protein</strong> there are the dry MD simulations for the unbound protein with the starting structure</li><li>in <strong>rna</strong> there are the dry MD simulations for the unbound RNA with the starting structure</li></ul><p>In each folder, all the trajectory files are named : <strong>md_(times of simulations).xtc</strong> and the starting structure called : <strong>start.gro</strong>.</p>
The dataset used in "A max-margin model for predicting residue-base contacts in protein-RNA interactions"
<p>The dataset used in "A max-margin model for predicting residue-base contacts in protein-RNA interactions"</p>
An in situ method for identification of transcriptome-wide protein-RNA interactions in cells
GEO Series GSE240326. Homo sapiens; Mus musculus. 78 samples. Type: Other; Expression profiling by high throughput sequencing.
SpyCLIP: An easy-to-use and high-throughput compatible platform for efficient characterization of protein-RNA interactions
GEO Series GSE114720. Homo sapiens. 13 samples. Type: Other.
A plant tethering system for the functional study of protein-RNA interactions in vivo
GEO Series GSE182403. Arabidopsis thaliana. 4 samples. Type: Non-coding RNA profiling by high throughput sequencing.
Transcriptome-wide high-throughput mapping of protein-RNA occupancy profiles using POP-seq
GEO Series GSE142460. Homo sapiens. 6 samples. Type: Other.
iCLIP and NGS to map protein-RNA interactions for U1-snRNP proteins in Trypanosoma brucei
GEO Series GSE43848. Trypanosoma brucei. 5 samples. Type: Other.
Enhanced CLIP uncovers IMP protein-RNA targets in human pluripotent stem cells important for cell adhesion and survival [RNA-Seq]
GEO Series GSE78508. Homo sapiens. 5 samples. Type: Expression profiling by high throughput sequencing.
Enhanced CLIP uncovers IMP protein-RNA targets in human pluripotent stem cells important for cell adhesion and survival
GEO Series GSE78509. Homo sapiens. 15 samples. Type: Other; Expression profiling by high throughput sequencing.
Enhanced CLIP uncovers IMP protein-RNA targets in human pluripotent stem cells important for cell adhesion and survival [eCLIP-Seq]
GEO Series GSE78507. Homo sapiens. 10 samples. Type: Other.
Chromatin accessibility associates with protein-RNA correlation in human cancer
GEO Series GSE162515. Homo sapiens. 178 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
An in situ method for identification of transcriptome-wide protein-RNA interactions in cells [in_situ_STAMP II]
GEO Series GSE248461. Homo sapiens. 14 samples. Type: Expression profiling by high throughput sequencing.
Genome-wide analysis of HDAC9 and BRG1 protein-RNA interactions in TAA by CLIP-seq
GEO Series GSE93744. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
Discovery of CD80 and CD86 as recent activation markers on regulatory T cells by protein-RNA single-cell analysis
GEO Series GSE150060. Homo sapiens. 6 samples. Type: Other.
An in situ method for identification of transcriptome-wide protein-RNA interactions in cells [in_situ_STAMP - Long-Read]"
GEO Series GSE240325. Homo sapiens. 2 samples. Type: Other.
An in situ method for identification of transcriptome-wide protein-RNA interactions in cells [isstamp_lr]
GEO Series GSE263870. Homo sapiens. 6 samples. Type: Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.