Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
641
datasets available to search
ShareScore release 0.9.0
Dataset results
641 results for “Protein interaction”
Dataset to manuscript entitled "Non-fluorescent transient states of tyrosine - a basis for label-free protein conformation and interaction studies" submitted to Scientifc Reports
<p><strong>This folder contains all raw data underlying the results presented in a manuscript, submitted to</strong><strong> </strong><strong><em>Scientifc Reports, </em>and entitled</strong><strong>:</strong></p> <p><strong><em>Non-fluorescent transient states of tyrosine - a basis for label-free protein conformation and interaction studies</em></strong></p> <p> </p> <p><strong>Authored by:</strong></p> <p><strong>Niusha Bagheri <sup>1</sup>, Hongjian Chen <sup>1</sup>, Mihailo Rabasovic <sup>2</sup>, Jerker Widengren <sup>1,*</sup></strong></p> <p><sup>1 </sup>Royal Institute of Technology (KTH), Experimental Biomolecular Physics, Dept. Applied Physics, Albanova University Center 106 91 Stockholm, Sweden</p> <p><sup>2 </sup>Laboratory for Biophysics, Institute of Physics Belgrade, Pregrevica 118<br>11080 Zemun-Belgrade, Serbia</p> <p>* Corresponding author (jwideng@kth.se)</p> <p><strong> </strong></p> <p><strong>The data files are grouped into the different techniques used to generate them, and refer to the figures/tables in the manuscript where the extracted results are presented. </strong></p> <p> </p> <p><strong>ABSTRACT</strong></p> <p>The amino acids tryptophan, tyrosine, and phenylalanine have been extensively used for different label-free protein studies, based on the intensity, lifetime, wavelength and/or polarization of their emitted fluorescence. Like most fluorescent organic molecules, these amino acids also tend to undergo transitions into dark meta-stable states, such as triplet and photo-radical states. While this may be perceived as a problem, these transitions are also highly environment-sensitive and can be used as an additional set of parameters, reflecting interactions, folding states, and immediate environments around the proteins. In this work, we applied the transient state monitoring (TRAST) technique, analyzing the average intensity of tyrosine emission under different excitation modulations, to characterize the photo physics of tyrosine for such readout purposes. By investigating how the dark state transitions of tyrosine varied with excitation intensity and solvent conditions we established a photophysical model for tyrosine. Next, we studied Calmodulin (containing two tyrosines), which upon calcium binding takes a more folded conformation. From these TRAST experiments, performed with 280nm time-modulated excitation, we show that tyrosine dark state transitions clearly change with the calmodulin conformation, and may thus represent a useful source of information for (label-free) analyses of protein conformations and interactions.</p>
Inferred protein interactions between coronavirus and human proteins
<p>This repository contains the protein-protein interactions inferred by mimicINT (<a href="https://github.com/TAGC-NetworkBiology/mimicINT" target="_blank" rel="noopener">https://github.com/TAGC-NetworkBiology/mimicINT</a>) between the proteins of seven human coronaviruses (HCoV-229E, HCoV-HKU1, HCoV-NL63, HCoV-OC43, MERS-CoV, SARS-CoV and SARS-CoV-2) and substantial fraction of the human proteome. This dataset was generated in the context of the RiPCoN project (H2020-SC1-PHE-CORONAVIRUS-2020, <a href="https://cordis.europa.eu/project/id/101003633" target="_blank" rel="noopener">https://cordis.europa.eu/project/id/101003633</a>).</p>
Predictions of the SARS-CoV-2 B.1.1.529 Variant Spike Protein Receptor Binding Domain Structure and Neutralizing Antibody Interactions
<p>Using AlphaFold2 and HADDOCK, we have generated a predicted structure for the SARS-CoV-2 B.1.1.529 variant's Spike receptor binding domain and then predicted the binding interaction with neutralizing antibodies. This was performed to understand the potential structural changes in the receptor binding domain of B.1.1.529 and how this may affect vaccine efficacy through antibody interaction.</p>
Towards a reproducible interactome: semantic-based detection of redundancies to unify protein-protein interaction databases
<p>Protein-protein interactions (PPIs) play an ubiquitous and fundamental role in all biological processes. Information on PPIs described in the literature is annotated and made available by several protein-interaction databases. Because most databases have their own curation rules and priorities, they often annotate overlapping sets of publications, which leads to redundancies. We developed a semantic-based approach which enables to accurately detect redundancies within PPI datasets from multiple databases. We applied this approach to assemble a "reproducible interactome", with PPIs supported by at least two methods or publications.</p>
Neural relational inference to learn long-range allosteric interactions in proteins from molecular dynamics simulations
<p>MD simulations used in the studies of the publication "<strong>Neural relational inference to learn long-range allosteric interactions in proteins from molecular dynamics simulations</strong>"</p>
Diffraction images of a crystal of the F-BAR domain of PSTPIP1 (Proline-serine-threonine phosphatase-interacting protein 1) mutant G258A (PDB entry 7AAL)
<p>Diffraction images of a crystal of the F-BAR domain of PSTPIP1 (residues 1-289), mutant G258A.</p> <p>Data were collected on a single crystal at the beamline i03 of the Diamond Light Source synchrotron (Didcot, UK) using radiation of 0.9999 Å wavelength and a PILATUS3 6M detector. The dataset consists of 2400 images (0.15 degree oscillation per image). Crystals belong to the space group P2(1)2(1)2(1) with unit cell dimensions a=48.19 Å, b=73.02 Å, c=205.25 Å. The asymmetric unit contains an homodimer of the F-BAR domain (~53% solvent content), which is the biological unit.</p> <p>Diffraction data was notably anisotropic. The lowest resolution limit was 2.92 Å in the direction b* and the highest limits were 1.97 Å and 2.09 in the directions a* and c*, respectively.</p> <p> </p> <p>The structure derived form these data is published in:</p> <p>Manso, J.A., Marcos, T., Ruiz-Martín, V. Casas J, Alcón P, Sánchez Crespo M, Bayón Y, de Pereda JM, Alonso A <em>PSTPIP1-LYP phosphatase interaction: structural basis and implications for autoinflammatory disorders</em>. <strong>Cell. Mol. Life Sci</strong>. 79, 131 (2022). <a href="https://doi.org/10.1007/s00018-022-04173-w">https://doi.org/10.1007/s00018-022-04173-w</a></p> <p>The structure is available at the PDB under the code 7AAL:</p> <p><a href="https://www.ebi.ac.uk/pdbe/entry/pdb/7aal">https://www.ebi.ac.uk/pdbe/entry/pdb/7aal</a></p>
Diffraction images of a crystal of the F-BAR domain of PSTPIP1 (Proline-serine-threonine phosphatase-interacting protein 1) bound to the C-terminal homology (CTH) segment of the phosphatase LYP (PTPN22) (PDB entry 7AAM)
<p>Diffraction images of a crystal of the F-BAR domain of human PSTPIP1 (residues 1-289, Uniprot reference O43586-1) in complex with the CTH of LYP (residues 787-807, Uniprot Q9Y2R2-1).</p> <p>Data were collected on a single crystal at the beamline i03 of the Diamond Light Source synchrotron (Didcot, UK) using radiation of 0.99987 Å wavelength and a PILATUS3 6M detector. The dataset consists of 3 groups, each containing of 1800 images (0.1 degree oscillation per image), collected at three different positions of the same crystal. Crystal belongs to the space group P2(1)2(1)2(1) with unit cell dimensions a=48.0 Å, b=72.0 Å, c=205.0 Å. The asymmetric unit contains an homodimer of the F-BAR domain bound to a LYP-CTH (~53% solvent content), which is the biological complex.</p> <p>Diffraction data was notably anisotropic. The lowest resolution limit was 4.05 Å in the direction b* and the highest limits were 2.11 Å and 2.10 in the directions a* and c*, respectively.</p> <p> </p> <p>The structure derived form these data is published in:</p> <p>Manso, J.A., Marcos, T., Ruiz-Martín, V. Casas J, Alcón P, Sánchez Crespo M, Bayón Y, de Pereda JM, Alonso A <em>PSTPIP1-LYP phosphatase interaction: structural basis and implications for autoinflammatory disorders</em>. <strong>Cell. Mol. Life Sci</strong>. 79, 131 (2022). <a href="https://doi.org/10.1007/s00018-022-04173-w">https://doi.org/10.1007/s00018-022-04173-w</a></p> <p>The structure is available at the PDB under the code <strong>7AAM</strong>:</p> <p><a href="https://www.ebi.ac.uk/pdbe/entry/pdb/7aam">https://www.ebi.ac.uk/pdbe/entry/pdb/7aam</a></p>
Diffraction images of a crystal of the F-BAR domain of PSTPIP1 (Proline-serine-threonine phosphatase-interacting protein 1) (PDB entry 7AAN)
<p>Diffraction images of a crystal of the F-BAR domain of PSTPIP1 (residues 1-289).</p> <p>Data were collected on a single crystal at the beamline i03 of the Diamond Light Source synchrotron (Didcot, UK) using radiation of 0.99987 Å wavelength and a PILATUS3 6M detector. The dataset consists of 3600 images (0.15 degree oscillation per image) that were collected: 2400 at one position and the other 1200 at a second site in the same crystal. Crystal belongs to the space group P2(1)2(1)2(1) with unit cell dimensions a=48.3 Å, b=71.9 Å, c=204.6 Å. The asymmetric unit contains an homodimer of the F-BAR domain (~53% solvent content), which is the biological unit.</p> <p>Diffraction data was notably anisotropic. The lowest resolution limit was 4.32 Å in the direction b* and the highest limits were 2.12 Å and 2.17 in the directions a* and c*, respectively.</p> <p>The structure derived form these data is published in:</p> <p>Manso, J.A., Marcos, T., Ruiz-Martín, V. Casas J, Alcón P, Sánchez Crespo M, Bayón Y, de Pereda JM, Alonso A <em>PSTPIP1-LYP phosphatase interaction: structural basis and implications for autoinflammatory disorders</em>. <strong>Cell. Mol. Life Sci</strong>. 79, 131 (2022). <a href="https://doi.org/10.1007/s00018-022-04173-w">https://doi.org/10.1007/s00018-022-04173-w</a></p> <p>The structure is available at the PDB under the code <strong>7AAN</strong>:</p> <p><a href="https://www.ebi.ac.uk/pdbe/entry/pdb/7aan">https://www.ebi.ac.uk/pdbe/entry/pdb/7aan</a></p>
Simulation data for "Corrole-protein interactions in H-Nox and HasA"
<p>Molecular dynamics simulation trajectories, metadata, and analysis script for the manuscript "Corrode-Protein interactions in H-Nox and HasA" (<a href="https://pubs.rsc.org/en/content/articlehtml/2022/cb/d2cb00004k">article link</a>) by Christopher M. Lemon, Amos J. Nissley, Naomi R. Latorraca, Elizabeth C. Wittenborn, and Michael A. Marletta, RCS Chem Biol. 2022, 3, 571-581. </p>
Data for RAPPPID: Towards Generalisable Protein Interaction Prediction with AWD-LSTM Twin Networks
<p>Data for RAPPPID, a method for the Regularised Automative Prediction of Protein-Protein Interactions using Deep Learning.</p> <p>These datasets are in a format that RAPPPID is ready to read.<br> <br> <strong>Comparatives Dataset</strong><br> These datasets were derived from the STRING v11 <em>H. sapiens</em> dataset, according to the C1, C2, and C3 procedures outlined by Park and Marcotte, 2012. Negative samples are sampled randomly from the space of proteins not known to interact. See <a href="https://doi.org/10.1101/2021.08.13.456309">Szymborski & Emad</a> for details.<br> <br> <strong>Repeatability Datasets</strong><br> The following datasets are all derived from STRING in the manner as the comparatives dataset, but three different random seeds are used for drawing proteins.<br> <br> <strong>References</strong><br> Park,Y. and Marcotte,E.M. (2012) Flaws in evaluation schemes for pair-input computational predictions. Nat Methods, 9, 1134–1136.</p> <p>Szklarczyk, D., Gable, A. L., Lyon, D., Junge, A., Wyder, S., Huerta-Cepas, J., Simonovic, M., Doncheva, N. T., Morris, J. H., Bork, P., Jensen, L. J., and Mering, C. (2019). String v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets. Nucleic Acids Research, 47(D1), D607–D613.<br> <br> Szymborski,J. and Emad,A. (2021) RAPPPID: Towards Generalisable Protein Interaction Prediction with AWD-LSTM Twin Networks. bioRxiv https://doi.org/10.1101/2021.08.13.456309</p>
Profiling phage-host interactions between Skunavirus receptor binding proteins and lactococcal cell wall polysaccharide structures
Open the record for dataset details and reuse information.
Elution profiles and protein interaction data accompanying "Ancient eukaryotic protein interactions illuminate modern genetic disorders"
<div> <p> </p> <table> <tbody> <tr> <td> <h2><strong>DESCRIPTION</strong></h2> </td> <td> <h2><strong>FILENAME</strong></h2> </td> <td> <h2><strong>LOCATION</strong></h2> </td> </tr> <tr> <td> <p>LECA 10K OG set</p> </td> <td> <p>leca_ogs_annotated.xlsx</p> </td> <td> <p>Paper, Table S1</p> <p>Zenodo</p> </td> </tr> <tr> <td> <p>Summary of biological resources</p> </td> <td> <p>resource_summary.xlsx</p> </td> <td> <p>Paper, Table S2</p> <p>Zenodo</p> </td> </tr> <tr> <td> <p>LECA interactome (complexes)</p> </td> <td> <p>leca_ppis_fdr10_clustered_annotated.xlsx</p> </td> <td> <p>Paper, Table S3</p> <p>Zenodo</p> </td> </tr> <tr> <td> <p>CFMS - ref proteomes</p> </td> <td> <p>cfms_ref_proteomes.xlsx</p> </td> <td> <p>Paper, Table S4</p> <p>Zenodo</p> </td> </tr> <tr> <td> <p>ML - top algorithms</p> </td> <td> <p>tpot_top_algorithms.xlsx</p> </td> <td> <p>Paper, Table S5</p> <p>Zenodo</p> </td> </tr> <tr> <td> <p>LECA interactome (pairwise)</p> </td> <td> <p>leca_ppis_fdr10_pairwise.csv</p> </td> <td> <p>Zenodo</p> </td> </tr> <tr> <td> <p>UniProt Subcellular Localization IDs</p> </td> <td> <p>uniprot_localization_codes.xlsx</p> </td> <td> <p>Zenodo</p> </td> </tr> <tr> <td> <p>Dollo parsimony - ref proteomes</p> </td> <td> <p>dollo_parsimony_ref_proteomes.xlsx</p> </td> <td> <p>Zenodo</p> </td> </tr> <tr> <td> <p>Dollo parsimony - input trait matrix</p> </td> <td> <p>dollo_parsimony_count_matrix.tsv</p> </td> <td> <p>Zenodo</p> </td> </tr> <tr> <td> <p>CFMS - raw elution profiles</p> </td> <td> <p>amorphea_raw_elution_vectors.csv</p> <p>excavata_raw_elution_vectors.csv</p> <p>tsar_raw_elution_vectors.csv</p> <p>archaeplastida_raw_elution_vectors.csv</p> </td> <td> <p>Zenodo</p> </td> </tr> <tr> <td> <p>CFMS - normalized elution profiles</p> </td> <td> <p>amorphea_norm_elution_vectors.csv</p> <p>excavata_norm_elution_vectors.csv</p> <p>tsar_norm_elution_vectors.csv</p> <p>archaeplastida_norm_elution_vectors.csv</p> </td> <td> <p>Zenodo</p> </td> </tr> <tr> <td> <p>CFMS/APMS - complete feature matrix</p> </td> <td> <p>feature_matrix.csv</p> </td> <td> <p>Zenodo</p> </td> </tr> <tr> <td> <p>ML - top features</p> </td> <td> <p>linearsvc_top_100_features.xlsx</p> </td> <td> <p>Zenodo</p> </td> </tr> <tr> <td> <p>OMIM disease propagation, statistics</p> </td> <td> <p>omim_disease_propagation_stats.xlsx</p> </td> <td> <p>Zenodo</p> </td> </tr> <tr> <td> <p>OMIM disease propagation, top 20 hits per disease</p> </td> <td> <p>omim_disease_propagation_top20hits_per_disease.xlsx</p> </td> <td> <p>Zenodo</p> </td> </tr> <tr> <td> <p>Curated OMIM gene-disease relationships for LECA OGs</p> </td> <td> <p>omim_disease_network.tsv</p> </td> <td> <p>Zenodo</p> </td> </tr> <tr> <td>Curated OMIM gene-disease relationships for human UniProt IDs</td> <td>omim_disease_groups.csv</td> <td>Zenodo</td> </tr> </tbody> </table> </div> <p> </p>
Predicting metal-protein interactions using cofolding methods: Status quo
<p>Metals play important roles for enzyme function and many therapeutically relevant proteins. Despite the fact that the first drugs developed via computer aided drug design were metalloprotein inhibitors, many computational pipelines still discard metalloproteins due to the difficulties of modelling them computationally. New "cofolding" methods such as AlphaFold3 (AF3) and RoseTTAfold-AllAtom (RFAA) promise to improve this issue by being able to dock small molecules in presence of multiple complex cofactors including metals or covalent modifications. Here, we analyze the current status for metal ion prediction using these methods. We find that currently only AF3 provides realistic predictions for metal ions, RFAA in contrast does perform worse than more specialized models such as AllMetal3D in predicting the location of metal ions accurately. We find that AF3 predictions are consistent with expected physico-chemical trends/intuition whereas RFAA often also predicts unrealistic metal ion locations.</p>
Pleckstrin Homology domain Interacting Protein (PHIP); A Target Enabling Package
<p>SGC Oxford has expressed, purified and crystallized the second bromodomain of PHIP as part of the probe programme. Fragment screening and X-ray crystallography identified binders, some of which optimised to uM affinity. However, molecules with probe properties were not obtained. Consequently it has been decided to put the information generated into the public domain.</p>
The proteasome-interacting Ecm29 protein disassembles the 26S proteasome in response to oxidative stress
<p>This repository contains the modeling files and the analysis related to the article <a href="https://www.ncbi.nlm.nih.gov/pubmed/28821611">"The proteasome-interacting Ecm29 protein disassembles the 26S proteasome in response to oxidative stress"</a> by Wang et al. in J Biol Chem 2017.</p> <p><strong>For more information</strong> about how to reproduce this modeling, see the <a href="https://salilab.org/ecm29/">Sali lab website</a> or the README file.</p>
Data for the "Systematic mapping of protein-metabolite interactions in central metabolism of Escherichia coli"
<p>This dataset contains raw and processed NMR data used in the publication.</p>
Alphafold2 and AlphaFold-Multimer Predicted Interactions of Soybean Proteins with Macrophomina phaseolina Effectors reveals putative protease inhibitors and SUSS effectors.
<p> </p> <ul> <li> <p><strong>Kunitz Monomer Prediction</strong>:</p> <ul> <li><strong>Data</strong>: Analysis of soybean Kunitz proteins.</li> <li><strong>Details</strong>: Detected on the apoplast at 3 days post-infection with <em>Macrophomina phaseolina</em>.</li> <li><strong>File</strong>: <code>KUNITZ_monomers_outputdir.zip</code></li> </ul> </li> <li> <p><strong>Uncharacterized M. phaseolina Protein Monomer Prediction</strong>:</p> <ul> <li><strong>Data</strong>: Predictions for uncharacterized proteins.</li> <li><strong>Details</strong>: Detected on the apoplast at 3 days post-infection.</li> <li><strong>File</strong>: <code>uncharacterised_proteins_SUSS_effectoroutputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Positive Validation Set</strong>:</p> <ul> <li><strong>Data</strong>: Experimental verification of protein-inhibitor pairs.</li> <li><strong>Details</strong>: Pairs include experimentally verified interactions, specifically proteins and inhibitors, but lack resolved crystal structures.</li> <li><strong>File</strong>: <code>existing_non_existinpairs_Validation_outputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Soybean Serine Protease-Kunitz Interaction</strong>:</p> <ul> <li><strong>Data</strong>: Interactions between soybean serine proteases and Kunitz proteins.</li> <li><strong>Details</strong>: Analyzed in the apoplastic space at 3 days post-infection.</li> <li><strong>File</strong>: <code>glycine max_Serine protease_Vs_Gmaxkunitz_outputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Cysteine Protease without Pro-domain-MoErs-like effector Interaction</strong>:</p> <ul> <li><strong>Data</strong>: Interactions involving cysteine proteases.</li> <li><strong>Details</strong>: Rice RD21 and soybean cysteine proteases with pro-domains removed interacting with <em>MoErs1</em> and <em>MoErs1</em>-like M.phaseolina effectors.</li> <li><strong>File</strong>: <code>AF2-Multimer_RD21&GmaxCproteases_MoERS1_screening_outputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Fungal Serine Protease-Kunitz Interaction</strong>:</p> <ul> <li><strong>Data</strong>: Interactions between <em>Macrophomina phaseolina</em> serine proteases and soybean Kunitz proteins.</li> <li><strong>Details</strong>: Evaluated in the apoplastic space at 3 days post-infection.</li> <li><strong>File</strong>: <code>fungalSerineprotease_Vs_Gmax_kunitzoutputdir.zip</code></li> </ul> </li> <li> <p><strong>Negative Validation Set</strong>:</p> <ul> <li><strong>Data</strong>: Known non-interacting pairs.</li> <li><strong>Details</strong>: Non-interacting pairs of serine proteases-chitinases that are not resolved as crystal structures</li> <li><strong>File</strong>: <code>Gmax_Serineprotease_Vs_Gmaxchitinases_Validation_outputdir.zip</code></li> </ul> </li> </ul>
Deep learning model for characterizing protein-RNA interactions from sequence at single-base resolution
<p> </p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip.h5</a> - This file contains the training, validation, and test data for the Reformer model.</p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip_bc.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip_bc.h5</a> - This file contains the training, validation, and test data for the Reformer-BC model.</p> <p><a href="https://zenodo.org/api/records/14027315/draft/files/Reformer-code.zip/content" target="_blank" rel="noopener">Reformer-code.zip</a> - This file contains the training code of Reformer.</p>
DrugProt corpus: Biocreative VII Track 1 - Text mining drug and chemical-protein interactions
<p>Gold Standard annotations of the DrugProt corpus (training and development sets). Also, test and background sets.</p><p> </p><p><strong>Please cite if you use any DrugProt resource:</strong></p><p>Antonio Miranda-Escalada, Farrokh Mehryary, Jouni Luoma, Darryl Estrada-Zavala, Luis Gasco, Sampo Pyysalo, Alfonso Valencia, Martin Krallinger, Overview of DrugProt task at BioCreative VII: data and methods for large-scale text mining and knowledge graph generation of heterogenous chemical–protein relations, <i>Database</i>, Volume 2023, 2023, baad080</p><blockquote><p><i>@article{miranda2023overview, title={Overview of DrugProt task at BioCreative VII: data and methods for large-scale text mining and knowledge graph generation of heterogenous chemical--protein relations}, author={Miranda-Escalada, Antonio and Mehryary, Farrokh and Luoma, Jouni and Estrada-Zavala, Darryl and Gasco, Luis and Pyysalo, Sampo and Valencia, Alfonso and Krallinger, Martin}, journal={Database}, volume={2023}, pages={baad080}, year={2023}, publisher={Oxford University Press UK} }</i></p></blockquote><p>Miranda, Antonio, et al. "Overview of DrugProt BioCreative VII track: quality evaluation and large scale text mining of drug-gene/protein relations." <i>Proceedings of the seventh BioCreative challenge evaluation workshop</i>. 2021.</p><blockquote><p><i>@inproceedings{miranda2021overview, title={Overview of DrugProt BioCreative VII track: quality evaluation and large scale text mining of drug-gene/protein relations}, author={Miranda, Antonio and Mehryary, Farrokh and Luoma, Jouni and Pyysalo, Sampo and Valencia, Alfonso and Krallinger, Martin}, booktitle={Proceedings of the seventh BioCreative challenge evaluation workshop}, year={2021} }</i></p></blockquote><p> </p><p><strong>Introduction</strong></p><p>The aim of the DrugProt track (similar to the previous CHEMPROT task of BioCreative VI) is to promote the development and evaluation of systems that are able to automatically detect in relations between chemical compounds/drug and genes/proteins. We have therefore generated a manually annotated corpus, the <i>DrugProt corpus</i>, where domain experts have exhaustively labeled:(a) all chemical and gene mentions, and (b) all binary relationships between them corresponding to a specific set of biologically relevant relation types (<i>DrugProt relation classes</i>). There is also an increasing interested in the integration of chemical and biomedical data understood as curation of relationships between biological and chemical entities from text and storing such information in form of structured annotation databases. Such databases are of key relevance not only for biological but also for pharmacological and clinical research. A range of different types chemical-protein/gene interactions are of key relevance for biology, including metabolic relations (e.g. substrates, products) inhibition, binding or induction associations.</p><p>The DrugProt track aims to address these needs and to promote the development of systems able to extract chemical-protein interactions that might be of relevance for precision medicine as well as for drug discovery and basic biomedical research.</p><p>The DrugProt track in BioCreative VII (BC VII) will explore recognition of chemical-protein entity relations from abstracts.</p><p>Teams participating in this track are provided with:</p><ul><li>PubMed abstracts</li><li>Manually annotated chemical compound mentions</li><li>Manually annotated gene/protein mentions</li><li>Manually annotated chemical compound-protein relations</li></ul><p> </p><p><strong>Zip structure:</strong></p><ul><li>Training set folder with<ul><li>drugprot_training_abstracts.tsv: PubMed records</li><li>drugprot_training_entities.tsv: manually labeled mention annotations of chemical compounds and genes/proteins</li><li>drugprot_training_relations.tsv: chemical-protein relation annotations</li></ul></li><li>Development set folder with<ul><li>drugprot_development_abstracts.tsv</li><li>drugprot_development_entities.tsv</li><li>drugprot_development_relations.tsv</li></ul></li><li>Test+background set folder with<ul><li>test_background_abstracts.tsv</li><li>test_background_entities.tsv</li></ul></li></ul><p> </p><p><strong>Data format description</strong></p><p>The <strong>input text files</strong> for the DrugProt track are plain-text, UTF8-encoded PubMed records in a tab-separated format with the following three columns:</p><ol><li>Article identifier (PMID, PubMed identifier)</li><li>Title of the article</li><li>Abstract of the article</li></ol><p> </p><p>DrugProt <strong>entity mention annotation files</strong> contain manually labeled mention annotations of chemical compounds and genes/proteins. Such files consist of tab-separated fields containing the following six columns:</p><ol><li>Article identifier (PMID)</li><li>Term number (for this record)</li><li>Type of entity mention (CHEMICAL, GENE-Y, GENE-N)</li><li>Start character offset of the entity mention</li><li>End character offset of the entity mention</li><li>Text string of the entity mention</li></ol><p>Each line contains one entity, and <i>each entity is uniquely identified by its PMID and the Term Number</i>. Besides, each annotation contains an annotation type, the start-offset -the index of the first character of the annotated span in the text-, the end-offset -the index of the first character after the annotated span- and the text spanned by the annotation.</p><p>Example DrugProt <i>training</i> entity mention annotations:</p><p>11808879 T1 GENE-Y 1860 1866 KIR6.2 11808879 T2 GENE-N 1993 2016 glutamate dehydrogenase 11808879 T3 GENE-Y 2242 2253 glucokinase 23017395 T1 CHEMICAL 216 223 HMG-CoA 23017395 T2 CHEMICAL 258 261 EPA</p><p> </p><p>Example DrugProt <i>development</i> entity mention annotations (no distinction between GENE-Y and GENE-N):</p><p>11808879 T1 GENE 1860 1866 KIR6.2 11808879 T2 GENE 1993 2016 glutamate dehydrogenase 11808879 T3 GENE 2242 2253 glucokinase 23017395 T1 CHEMICAL 216 223 HMG-CoA 23017395 T2 CHEMICAL 258 261 EPA</p><p><br>DrugProt <strong>relation annotations</strong> are distributed as a file that contains the detailed chemical-protein relation annotations prepared for the DrugProt track. There are no relation annotations for the test+background set (the goal of the task is to predict them). It consists of tab-separated columns containing:</p><ol><li>Article identifier (PMID)</li><li>DrugProt relation</li><li>Interactor argument 1 (<i>of type CHEMICAL</i>)</li><li>Interactor argument 2 (<i>of type GENE</i>)</li></ol><p>Each line contains one relation, and <i>each relation is identified by the PMID, the relation type and the two related entities</i>. In the below example, to find the entities involved in the first relation, you must find the entities with Term Identifier T1 and T52 <i>within the PMID 12488248.</i></p><p>Example DrugProt relation annotations:</p><p>12488248 INHIBITOR Arg1:T1 Arg2:T52 12488248 INHIBITOR Arg1:T2 Arg2:T52 23220562 ACTIVATOR Arg1:T12 Arg2:T42 23220562 ACTIVATOR Arg1:T12 Arg2:T43 23220562 INDIRECT-DOWNREGULATOR Arg1:T1 Arg2:T14</p><p> </p><p>Please, cite:</p><p>@inproceedings{krallinger2017overview, title={Overview of the BioCreative VI chemical-protein interaction Track}, author={Krallinger, Martin and Rabal, Obdulia and Akhondi, Saber A and P{\'e}rez, Mart{\i}n P{\'e}rez and Santamar{\'\i}a, Jes{\'u}s and Rodr{\'\i}guez, Gael P{\'e}rez and others}, booktitle={Proceedings of the sixth BioCreative challenge evaluation workshop}, volume={1}, pages={141--146}, year={2017}}</p><p> </p><p><strong>Summary statistics:</strong></p><p>Training set Development set Documents 3500 750 Tokens 1001168 199620 Annotated Entities 89529 18858 Annotated Relations 17288 3765</p><p> </p><p>Annotated Entities:</p><p>Training Entities Development Entities CHEMICAL 46274 9853 GENE-Y [Normalizable] 28421 - GENE-N [Non-Normalizable] 14834 - Gene Total (N+Y) 43255 9005 Total 89529 18858</p><p> </p><p>Annotated Relations:</p><p>Training Relations Development Relations INDIRECT-DOWNREGULATOR 1330 332 INDIRECT-UPREGULATOR 1379 302 DIRECT-REGULATOR 2250 458 ACTIVATOR 1429 246 INHIBITOR 5392 1152 AGONIST 659 131 AGONIST-ACTIVATOR 29 10 AGONIST-INHIBITOR 13 2 ANTAGONIST 972 218 PRODUCT-OF 921 158 SUBSTRATE 2003 495 SUBSTRATE_PRODUCT-OF 25 3 PART-OF 886 258 Total 17288 3765</p><p> </p><p>For further information, please visit <a href="https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/">https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/</a> or email us at krallinger.martin@gmail.com and antoniomiresc@gmail.com</p><p> </p><p><strong>Related resources:</strong></p><ul><li><a href="https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/">Web</a></li><li><a href="https://github.com/tonifuc3m/drugprot-evaluation-library">Evaluation library</a></li><li><a href="https://codalab.lisn.upsaclay.fr/competitions/8293">Online evaluation (CodaLab)</a></li><li><a href="https://doi.org/10.5281/zenodo.4957137">Relation annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957576">Gene and protein annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957518">Chemicals and drugs annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.7252201">DrugProt Silver Standard Knowledge Graph</a></li><li><a href="https://doi.org/10.5281/zenodo.5042178">FAQ</a></li><li><a href="https://doi.org/10.5281/zenodo.5119878">DrugProt Large Scale Additional SubTrack</a></li><li><a href="https://doi.org/10.5281/zenodo.5656991">DrugProt Large Scale document collection protocol</a></li><li><a href="https://doi.org/10.5281/zenodo.8246229">DrugProt Complete PubMed Knowledge Graph</a><br> </li></ul>
Dataset for article: Co-evolutionary landscape at the interface and non-interface regions of protein-protein interaction complexes
<p>Proteins involved in interactions throughout the course of evolution tend to co-evolve and compensatory changes may occur in interacting proteins to maintain or refine such interactions. However, certain residue pair alterations may prove to be detrimental for functional interactions. Hence, determining co-evolutionary pairings that could be structurally or functionally relevant for maintaining the conservation of an inter-protein interaction is important. Inter-protein co-evolution analysis in several complexes utilizing multiple existing methodologies suggested that co-evolutionary pairings can occur in spatially proximal and distant regions in inter-protein interactions. Subsequently, the Co-Var (<b>Co</b>rrelated <b>Var</b>iation) method based on mutual information and Bhattacharyya coefficient was developed, validated, and found to perform relatively better than CAPS and EV-complex. Interestingly, while applying the Co-Var measure and EV-complex program on a set of protein-protein interaction complexes, co-evolutionary pairings were obtained in interface and non-interface regions in protein complexes. The Co-Var approach involves determining high degree co-evolutionary pairings that include multiple co-evolutionary connections between particular co-evolved residue positions in one protein with multiple residue positions in the binding partner. Detailed analyses of high degree co-evolutionary pairings in protein-protein complexes involved in cancer metastasis suggested that most of the residue positions forming such co-evolutionary connections mainly occurred within functional domains of constituent proteins and substitution mutations were also common among these positions. The physiological relevance of these predictions suggests that Co-Var can predict residues that could be crucial for preserving functional protein-protein interactions. Finally, <b>Co-Var </b>web server (<a href="http://www.hpppi.iicb.res.in/ishi/covar/index.html">http://www.hpppi.iicb.res.in/ishi/covar/index.html</a>) that implements this methodology identifies co-evolutionary pairings in intra and inter-protein interactions.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.