Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

40

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

40 results for “Protein-RNA”

Learn how ShareScore rates datasets ↗
zenodo48/100

Datasets and Jupyter notebook for the structural analysis of protein-RNA interface evolution

<p>The present repository contains data and code related to our manuscript "Structural comparison of protein-RNA homologous interfaces reveals widespread overall conservation contrasted with versatility in polar contacts". In the manuscript, we analyze the evolution of protein-RNA interfaces by building a dataset of protein-RNA interologs (homologous interfaces) and exploring how interface contacts are conserved between homologous interfaces, as well as possible explanations for non-conserved contacts.</p> <p>This repository contains the following files:</p> <ul> <li>DataAnalysisNotebook.ipynb is a Jupyter notebook to reproduce contact conservation analysis and all figures from our manuscript, and to explore data</li> <li>env.yaml is an environment file in order to build a Conda/Mamba environment to run the Jupyter notebook&nbsp;</li> <li>2022-02-21-PDB.csv contains data from the PDB about 3D structures of complexes containing interacting protein and RNA chains (PDB structure identifier, chain identifiers, experimental technique and resolution)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.tsv contains more detailed information about interacting protein and RNA chains from these complexes (PDB and chain identifiers, protein and RNA size, interface size and number of contacts)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.txt.selectXE_2.50_p30_r10_pi5_ri5_rep_bc-100.out_RNAcl_0.99.tsv contains the same detailed information, restricted to the filtered dataset used as a starting point in our interolog search pipeline</li> <li>PDBinterfaceAlign.csv contains information about the structural alignment of pairs of protein-RNA interactions (structural alignment TM-scores, sequence identity and coverage)</li> <li>DataInterologsParam.tsv contains information about a pre-filtered set of 2587 potential interologs (including interface RMSD, sequence identity and coverage and interface size)</li> <li>DataInterologsContactsFixedSASA.tsv contains detailed information about conserved and non-conserved contacts in the final set of 2022 interologs (atomic contacts, apolar contacts, hydrogen bonds, salt bridges and stacking information for aminoacid-nucleotide pairs, as well as information about whether each belongs to the interface, secondary structures, and the aminoacid surface accessibility and evolutionary conservation metrics) - compared to version 1, the calculation of solvent accessibility was fixed for a number of interolog pairs</li> <li>DataCons.csv contains precomputed contact conservation metrics for each of the 2022 interolog pairs, for fast reproduction of manuscript figures</li> <li>DataInterologsContactsResampledMaintainStructSeqId.tsv, DataInterologsContactsShuffled.tsv and DataInterologsShuffled.tsv relate to baselines computed for contact conservation assessment</li> <li>clan.txt, clan_membership.txt, ecod.latest.domains.uniq.txt, rfam_interfaces_977.txt, DataGroupsECOD.tsv, DataGroupesRFAM.tsv, DataGroupsRFAMClan.tsv, DataInterfaceGroupsECOD.tsv and DataInterfaceGroupsRFAM.tsv relate to the ECOD (respectively Rfam) classification of protein domains (respectively RNA) in protein-RNA interfaces from our dataset</li> <li>ListeIntraHbonds.pkl and ListeIntraSaltBridges.pkl are pickle-format data files containing intra-molecular hydrogen bonds and salt bridges (respectively) that are used to analyse scenarii of compensation for non-conserved polar contacts.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo44/100

Protein-RNA complex simulation

<p>Simulation trajectory of TSEN/pre-tRNAArgTCT generated with GROMACS.&nbsp;</p> <p>&nbsp;</p> <p>A hybrid model of truncated TSEN/pre-tRNAArgTCT with the additional TSEN2 domain from AlphaFold Jumper et al, 2021</p> <p>and in silico modeled intron bases 37 to 43 was subjected to all-atom molecular dynamics simulations</p> <p>for assessment of flexibility. Simulation trajectory can be visualised with PyMOL or VMD.</p> <p>&nbsp;</p> <p>The protein was described by the AMBER14ff (Maier, et al 2015), and the RNA with the OL3 force field (Zgarbova et al, 2011)</p> <p>and the TIP3P water model was employed (Jorgensen et al, 1983).</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Deep learning model for characterizing protein-RNA interactions from sequence at single-base resolution

<p>&nbsp;</p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip.h5</a> - This file contains the training, validation, and test data for the Reformer model.</p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip_bc.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip_bc.h5</a> - This file contains the training, validation, and test data for the Reformer-BC model.</p> <p><a href="https://zenodo.org/api/records/14027315/draft/files/Reformer-code.zip/content" target="_blank" rel="noopener">Reformer-code.zip</a> - This file contains the training code of Reformer.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

FTDMP docking results for protein-protein, protein-DNA, protein-RNA benchmarks

<p><strong>FTDMP docking results for protein-protein, protein-DNA, protein-RNA benchmarks.&nbsp;</strong></p> <p>FTDMP is a software system for running docking experiments and scoring/ranking multimeric models. This dataset contains FTDMP docking results for protein-protein, protein-DNA, protein-RNA benchmarks. The FTDMP framework itself is available at https://github.com/kliment-olechnovic/ftdmp.&nbsp;</p> <p>Every *.tar.gz file in this dataset contains two folders: results for unbound-unbound and bound-bound docking. These folders contain results for the benchmark cases:</p> <p>252 folders with results for the protein-protein docking benchmark cases [1].<br>47 folders with results for the protein-DNA docking benchmark cases [2].<br>42 folders with results for the protein-RNA docking benchmark cases [3-6].&nbsp;</p> <p>Every folder is named according to the PDB ID of the complex. The folders contain:</p> <p>1. A subfolder named <em>relaxed_top_complexes</em>. This subfolder contains 200 pdb files of relaxed [7] top docking models.<br>2. A text file named <em>scoring_results-ranks.txt</em>. It contains the names of the models (that are in the relaxed_top_complexes folder) in the ranked order. This means that the first model in the file is considered to be the best prediction by the FTDMP framework.<br>3. A text file named <em>cad_scores.txt</em>. It contains interface CAD-score and binding site CAD-score [8] results for every model.<br>4. A text file named <em>rmsd_results.txt</em>, which is available only for protein-DNA and protein-RNA cases. The file contains ligand-RMSD values for the models, where the DNA/RNA is considered as the ligand.<br>5. A text file named <em>DockQ_results.txt</em>, which is available only for the protein-protein docking cases. The file contains DockQ [9] results for every model, as well as model accuracy based on CAPRI criteria (Incorrect, Acceptable, Medium, High)<br>6. A text file named <em>binding_site_CAD-scores.txt</em>, which contains the binding site CAD-score from the <strong>protein</strong> side for RNA and DNA docking. This binding site CAD-score shows how accurately the ligand (DNA/RNA) was docked to the protein without taking the orientation of the ligand into consideration. In the case of protein-protein docking the binding site CAD-score file is available only for antibody-antigen docking targets and contains the binding site (epitope) CAD-score for the antigen. &nbsp;</p> <p>The ligand-RMSD, CAD-scores, and DockQ scores were all calculated by comparing the models to the corresponding targets. The target structures are available at <span>https://zenodo.org/records/10517524</span>. These target structures have the same residue numbering as the models available here.&nbsp;</p> <p>REFERENCES&nbsp;</p> <p>[1] Guest, J. D., et al. (2021). An expanded benchmark for antibody-antigen docking and affinity prediction reveals insights into antibody recognition determinants. Structure, 29(6), 606&ndash;621.e5.<br>[2] van Dijk, M., Bonvin, A.M. (2008). A protein-DNA docking benchmark. Nucleic Acids Res, 36, e88.&nbsp;<br>[3] Perez-Cano, L., et. Al. (2012). A protein-RNA docking benchmark (II): extended set from experimental and homology modeling data. Proteins, 80(7): 1872-1882.&nbsp;<br>[4] Huang, S.Y., Zou, X. (2013). A nonredundant structure dataset for benchmarking protein-RNA computational docking. J Comput Chem, 34(4): 311-318.&nbsp;<br>[5] Nithin, C., et. al. (2017). A non-redundant protein-RNA docking benchmark version 2.0. Proteins, 85(2) :256-267.&nbsp;<br>[6] Zheng, J., et al. (2020). P3DOCK: a protein-RNA docking webserver based on template-based and template-free docking. Bioinformatics, 36(1), 96&ndash;103.&nbsp;<br>[7] Eastman, P., et al.(2017). OpenMM 7: Rapid development of high performance algorithms for molecular dynamics. PLOS Comp. Biol., 13(7): e1005659. &nbsp;<br>[8] Olechnovic, K., Venclovas, C. (2020). Contact area-based structural analysis of proteins and their complexes using CAD-score. Methods Mol Biol, 2112, 75.<br>[9] Basu, S., Wallner, B. (2016). DockQ: A Quality Measure for Protein-Protein Docking Models. PLoS ONE 11(8): e0161879.&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

The dynamics of protein-RNA interfaces using all-atom molecular dynamics simulations

<p>We investigated to characterize the dynamics of protein-RNA complexes and their interfaces at molecular level by performing a more systematic analysis. To get insights on the dynamics of protein-RNA complexes, all-atom MD simulations were generated for the manuscript "The dynamics of protein-RNA interfaces using all-atom molecular dynamics simulations". Nine protein-RNA complexes are studied in this work: 1ASY (an aspartyl-tRNA synthase/tRNA), 1JBS (a ribotoxin restrictocin/SRD RNA), 1MMS (a ribosomal protein L11/23S), 1OOA (a nuclear factor NF-kappaB p105 subunit/RNA aptamer), 1RKJ (a nucleolin/pre-rRNA), 2R8S (a FAB/P4-P6 RNA ribozyme domain), 2VPL (a 50S ribosomal protein/mRNA), 2ZM5 (a tRNA delta(2)-isopentenylpyrophosphate transferase/tRNA), 3IEV (a GTP-binding protein era/3' end of 16S rRNA).&nbsp;</p><p>Each folder for a complex is organised as followed:</p><ul><li>in <strong>complex</strong> there are the dry MD simulations for the complex protein-RNA with the starting structure</li><li>in <strong>protein</strong> there are the dry MD simulations for the unbound protein with the starting structure</li><li>in <strong>rna</strong> there are the dry MD simulations for the unbound RNA with the starting structure</li></ul><p>In each folder, all the trajectory files are named : <strong>md_(times of simulations).xtc</strong> and the starting structure called : <strong>start.gro</strong>.</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

The dataset used in "A max-margin model for predicting residue-base contacts in protein-RNA interactions"

<p>The dataset used in &quot;A max-margin model for predicting residue-base contacts in protein-RNA interactions&quot;</p>

opencc-by-4.0Oct 2021View details →
geo24/100

An in situ method for identification of transcriptome-wide protein-RNA interactions in cells

GEO Series GSE240326. Homo sapiens; Mus musculus. 78 samples. Type: Other; Expression profiling by high throughput sequencing.

openGEO-OpenJun 2024View details →
geo24/100

SpyCLIP: An easy-to-use and high-throughput compatible platform for efficient characterization of protein-RNA interactions

GEO Series GSE114720. Homo sapiens. 13 samples. Type: Other.

openGEO-OpenJan 2019View details →
geo24/100

A plant tethering system for the functional study of protein-RNA interactions in vivo

GEO Series GSE182403. Arabidopsis thaliana. 4 samples. Type: Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenMay 2022View details →
geo24/100

Transcriptome-wide high-throughput mapping of protein-RNA occupancy profiles using POP-seq

GEO Series GSE142460. Homo sapiens. 6 samples. Type: Other.

openGEO-OpenJun 2020View details →
geo24/100

iCLIP and NGS to map protein-RNA interactions for U1-snRNP proteins in Trypanosoma brucei

GEO Series GSE43848. Trypanosoma brucei. 5 samples. Type: Other.

openGEO-OpenMay 2014View details →
geo24/100

Enhanced CLIP uncovers IMP protein-RNA targets in human pluripotent stem cells important for cell adhesion and survival [RNA-Seq]

GEO Series GSE78508. Homo sapiens. 5 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenApr 2016View details →
geo24/100

Enhanced CLIP uncovers IMP protein-RNA targets in human pluripotent stem cells important for cell adhesion and survival

GEO Series GSE78509. Homo sapiens. 15 samples. Type: Other; Expression profiling by high throughput sequencing.

openGEO-OpenApr 2016View details →
geo24/100

Enhanced CLIP uncovers IMP protein-RNA targets in human pluripotent stem cells important for cell adhesion and survival [eCLIP-Seq]

GEO Series GSE78507. Homo sapiens. 10 samples. Type: Other.

openGEO-OpenApr 2016View details →
geo24/100

Chromatin accessibility associates with protein-RNA correlation in human cancer

GEO Series GSE162515. Homo sapiens. 178 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenDec 2020View details →
geo24/100

An in situ method for identification of transcriptome-wide protein-RNA interactions in cells [in_situ_STAMP II]

GEO Series GSE248461. Homo sapiens. 14 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJun 2024View details →
geo24/100

Genome-wide analysis of HDAC9 and BRG1 protein-RNA interactions in TAA by CLIP-seq

GEO Series GSE93744. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2018View details →
geo24/100

Discovery of CD80 and CD86 as recent activation markers on regulatory T cells by protein-RNA single-cell analysis

GEO Series GSE150060. Homo sapiens. 6 samples. Type: Other.

openGEO-OpenJun 2020View details →
geo24/100

An in situ method for identification of transcriptome-wide protein-RNA interactions in cells [in_situ_STAMP - Long-Read]"

GEO Series GSE240325. Homo sapiens. 2 samples. Type: Other.

openGEO-OpenJun 2024View details →
geo24/100

An in situ method for identification of transcriptome-wide protein-RNA interactions in cells [isstamp_lr]

GEO Series GSE263870. Homo sapiens. 6 samples. Type: Other.

openGEO-OpenJun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record