Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
347
datasets available to search
ShareScore release 0.9.0
Dataset results
347 results for “structural proteins”
Code and data for: Network-based protein structural classification
<p>Code and data related to the research article titled, "Network-based protein structural classification".</p> <p>More information about the code and the data is available at https://nd.edu/~cone/NETPCLASS/</p>
Snapshots, frequency contact maps analysis, Poisson Boltzmann calculations, and data scripts for characterization of structural and energetic differences between conformations of the SARS-CoV-2 spike protein
<p><strong>Molecular dynamics simulation</strong> trajectories, which have been performed using the Amber ff14SB force field running with the Amber18 package at the NSF-funded (OAC-1826915, OAC-1828163) ELSA high performance computing cluster at The College of New Jersey. Simulation methodology and further details are described in [1] and [2]. For further details on the trajectories, please contact Joseph Baker (bakerj@tcnj.edu).</p> <p>The <strong>Poisson Boltzmann </strong>energy calculations have been achieved by using the input_files.tar.xz found here and solving the Poisson Boltzmann equation with pygbe. A more detailed example and tutorial can be found at [4]. For further details contact Horacio V Guzman.</p> <p><strong>The dataset contains </strong></p> <ul> <li><strong>A total of 30 snapshots of the three trajectories (10 snapshots each system = two per replica x 5 replicas/system):</strong></li> </ul> <ol> <li>SARS-CoV-2002 spike protein with three RBD in the down positions: "COV2-DDD/PDB/" .</li> <li>SARS-CoV-2002 spike protein with one RBD in the up and two RBD in the down positions: "COV2-UDD/PDB/".</li> <li>SARS-CoV-2002 spike protein with two RBD in the up and one RBD in the down positions: "COV2-DUU/PDB/".</li> </ol> <ul> <li><strong>Input files for Poisson-Boltzmann analysis</strong>:</li> </ul> <ol> <li>PoissonBoltzmann/input_files.tar.xz</li> </ol> <ul> <li><strong>Data for the frequency contact map and processing scripts</strong>:</li> </ul> <ol> <li>cov2-ddd.pdb, cov2-udd.pdb, cov2-duu.pdb reference PDB files.</li> <li>Contact maps [3] at "COV2-DDD/CONTACT_MAP/", "COV2-UDD/CONTACT_MAP/", "COV2-DUU/CONTACT_MAP/".</li> <li>frequency.lua: get frequency of contacts from a set of contacts map files.</li> <li>diff_frequency.lua: get differential frequency of contacts from a set of frequency files.</li> <li>Frequency of contacts listed in frequency.data files at "COV2-DDD/", "COV2-UDD/" and "COV2-DUU/" directories.</li> </ol> <p>Read the "INFO" files for further informations.</p> <p>This dataset and the code is part of a collaboration between:</p> <ul> <li>The Institute of Fundamental Technological Research, Polish Academy of Sciences, Warsaw, Poland (supported by the National Science Centre, Poland, under grant No. 2017/26/D/NZ1/0046)</li> <li>Department of Chemistry, The College of New Jersey, New Jersey, United States (supported by National Science Foundation under grant numbers OAC-1826915 and OAC-1828163).</li> <li>Jozef Stefan Institute, Ljubljana, Slovenia (supported by the Slovenian Research Agency (Funding No. P1-0055)).</li> <li>School of engineering in bioinformatics, University of Talca, Talca, Chile.</li> </ul> <p>[1] Rodrigo A. Moreira, Mateusz Chwastyk, Joseph L. Baker, Horacio V Guzman, & Adolfo B. Poma. (2020). All-atom simulations snapshots and contact maps analysis scripts for SARS-CoV-2002 and SARS-CoV-2 spike proteins with and without ACE2 enzyme (Version 0.1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.3817447</p> <p>[2] Chad W. Hopkins, Scott Le Grand, Ross C. Walker, and Adrian E. Roitberg. Long-Time-Step Molecular Dynamics through Hydrogen Mass Repartitioning. Journal of Chemical Theory and Computation 2015 11 (4), 1864-1874. http://doi.org/10.1021/ct5010406</p> <p>[3] Rodrigo A. Moreira, Mateusz Chwastyk, Joseph L. Baker, Horacio V Guzman, & Adolfo B. Poma. Quantitative determination of mechanical stability in the novel coronavirus spike protein. Nanoscale, 2020,12, 16409-16413. <a href="https://doi.org/10.1039/D0NR03969A">https://doi.org/10.1039/D0NR03969A</a></p> <p>[4] https://github.com/pyF4all</p>
X-ray structure ensemble refinement of the second bromodomain of Pleckstrin homology domain interacting protein (PHIP) (space group P21212)
<p>X-ray structure ensemble refinement of the second bromodomain of Pleckstrin homology domain interacting protein (PHIP) (space group P21212). Raw diffraction images are available on Zenodo: 10.5281/zenodo.4086066. A single conformer model was deposited in the Protein Data Bank under accession code <a href="https://www.ebi.ac.uk/pdbe/entry/pdb/7AV8">7AV8</a>. Refinement was carried out with <a href="https://www.phenix-online.org/documentation/reference/ensemble_refinement.html">phenix.ensemble_refinement</a> and the repository contains all input and output files:</p> <p>Input:</p> <ul> <li>mx8421v63_xPHIPAx1521_free.mtz</li> <li>refine9.pdb</li> </ul> <p>Output:</p> <ul> <li>PHIPA-P21212_ensemble_refinement_ensemble.geo</li> <li>PHIPA-P21212_ensemble_refinement_ensemble.pdb</li> <li>PHIPA-P21212_ensemble_refinement_ensemble.mtz</li> <li>PHIPA-P21212_ensemble_refinement_ensemble.log</li> </ul>
Host-pathogen protein interactions predicted using structure
<p>This dataset accompanies a manuscript describing a method to predict host-pathogen protein interactions using structure:</p> <p>Host-pathogen protein interactions predicted by comparative modeling.<br /> Davis FP, Barkan DT, Eswar N, McKerrow JH, Sali A. Protein Sci (2007) 16:2585-2596.<br /> http://www.proteinscience.org/cgi/doi/10.1110/ps.073228407</p> <p>The files contain predictions made for 10 human pathogens including species of Mycobacterium, Apicomplexa, and Kinetoplastida. The species.all.zip files contain all interactions predictions for each species along with the filter criteria that each interaction passed. The species.filter.zip files contains the same information, but only for the subset of interactions that passed the biological and network-level filters. These files can be viewed in any spreadsheet program, such as Excel.</p> <p> </p> <p>Predictions were made for interactions between human and </p> <ol> <li>Mycobacterium tuberculosis: mtuber</li> <li>Mycobacterium leprae: mleprae</li> <li>Leishmania major: lmajor</li> <li>Trypanosoma brucei: tbrucei</li> <li>Trypanosoma cruzi: tcruzi</li> <li>Cryptosporidium hominis: chominis</li> <li>Cryptosporidium parvum: cparvum</li> <li>Plasmodium falciparum: pfalciparum</li> <li>Plasmodium vivax: pvivax</li> <li>Toxoplasma gondii: tgondi</li> </ol>
Protein Structure Initiative - TargetTrack 2000-2017 - all data files
<p><strong>Protein Structure Initiative - TargetTrack protein target registration database (795 MB, gzipped tarball)</strong></p> <p>The Protein Structure Initiative was a high-throughput structural genomics effort from 2000-2015 focused on developing technologies to enable greater coverage of protein structure space. Over its 15-year tenure, over 100 investigators at 35 centers (see ContributingCenters.xls) declared over 350,000 protein sequences (targets) that they would study using state-of-the-art protein production and structure determination methods. Many of these targets were selected through bioinformatics-based methods to serve as representatives for sequence and structure clusters. </p> <p>From 2003-2010, these selected sequences and some basic identifying metadata were kept in a database called TargetDB, created at the Research Collaboratory for Structural Bioinformatics at Rutgers University. In 2008, a second database named PepcDB was created to track detailed experimental trial history and the standard protocols used by the PSI centers. These two databases became the principal structural genomics target databases, and were rolled into the <strong>PSI Structural Biology Knowledgebase</strong> in 2008. </p> <p>As part of the third phase of the PSI, TargetDB and PepcDB were merged into a single resource, <strong>TargetTrack</strong>, to facilitate one-stop access to the data as well as expanding the schema to include new required data items. Participating centers deposited the latest status on their active targets and the protocols that were used (along with any deviations) on a weekly or quarterly basis. TargetTrack provided a variety of pre-computed data downloads on a weekly basis as well. </p> <p>In July 2017, the Structural Biology Knowledgebase ceased operations. The files provided in this tarball represent the final datafiles generated by TargetTrack (timestamp June 30, 2017). <strong>Please read the README included in this dataset for descriptions of each file. </strong></p> <p><strong>The entire TargetTrack datafile in XML format can be found in /TargetTrack XML files/tt.xml.gz</strong></p> <p>Key documentation can be found in the /Documentation folder.<br> TargetTrack schema: targetTrack-v1.4.1.pdf<br> Spreadsheet with TargetTrack enumerations for relevant fields: targetTrackEnumeratedDataItems-v1.4.1-1.xls<br> Image depicted the XML data schema: targetTrack-v1.4.1.jpg</p> <p>These files are 868 MB in total size, uncompressed. <br> To open the tarball, use the command 'tar -zxvf TargetTrack-1Jul2017.tar.gz'</p> <p>-- created by the PSI Structural Biology Knowledgebase, July 5, 2017</p>
Coordinate files from LipIDens: Simulation assisted interpretation of lipid densities in cryo-EM structures of membrane proteins.
<p>Coordinate files from the first and last frame of coarse-grained (CG) and atomistic (AT) molecular dynamics (MD) simulations used throughout the LipIDens pipeline.</p><p>CG simulations were run for HHAT, OTOP1, ELIC, MscS, TRPV6, ChRmine, Ste2, Connexin-50, NPC1 and the PAT complex. All CG simulations were run for 10 x 15 μs with the exception of NPC1 which was simulated for 10 x 30 μs.</p><p>AT simulations were run for HHAT (5 x 200 ns) and ELIC (3 x 200 ns) in apo configurations.</p><p><strong>File description:</strong></p><p>Directories for each protein are listed with the suffix CG or AT used to indicate the simulation resolution. </p><p>md_fit_firstframe_<i>X</i>.gro - GROMACS structure file for the first frame of replicate <i>X</i>. </p><p>md_fit_lastframe_<i>X</i>.gro - GROMACS structure file for the last frame of replicate <i>X</i>. </p>
Structural models of viral proteins described in Krupovic M, et al., Proc Natl Acad Sci U S A. 2024
<p>This archive contains structural models in PDB format described in Krupovic M, Kuhn JH, Fischer MG, Koonin EV. Natural history of eukaryotic DNA viruses with double jelly-roll major capsid proteins. Proc Natl Acad Sci U S A. 2024</p>
Predicted structures of SLC-protein complexes and controls
<p>Solute carrier (SLC) transporters form a major superfamily which transport a wide range of substrates across cellular and organellar membranes. To study their regulation on protein level and position SLCs in the human interactome, we conducted a large-scale interrogation of the protein-protein interactions (PPIs) of SLCs employing affinity purification combined with mass spectrometry (AP-MS). This study resulted in thousands of novel protein interactions of SLCs. For a subset of SLC protein complexes, we performed structural predictions using AlphaFold (v2.2, v2.3 and v3). The dataset attached contains the structures in PDB/CIF format. The structures were further discussed in the associated manuscript. In addition, an annotation table is provided, which summarizes the scores for each modelled complex. </p>
Cognition-Associated Protein Structural Changes in a Rat Model of Aging are Related to Reduced Refolding Capacity – Peptide Quantifications
<p>Cognitive decline during aging represents a major societal burden, causing both personal and economic hardship in an increasingly aging population. There are a few well-known proteins that can misfold and aggregate in an age-dependent manner, such as amyloid β and α-synuclein. However, many studies have found that the proteostasis network, which functions to keep proteins properly folded, is impaired with age, suggesting that there may be many more proteins that incur structural alterations with age. Here, we used limited-proteolysis mass spectrometry (LiP-MS), a structural proteomic method, to globally interrogate protein conformational changes in a rat model of cognitive aging. Specifically, we compared soluble hippocampal proteins from aged rats with preserved cognition to those from aged rats with impaired cognition. We identified several hundred proteins as having undergone cognition-associated structural changes (CASCs). We report that CASC proteins are substantially more likely to be nonrefoldable than non-CASC proteins, meaning they typically cannot spontaneously refold to their native conformations after being chemically denatured. The potentially cofounding variable of post-translational modifications is systematically addressed, and we find that oxidation and phosphorylation cannot significantly explain the limited proteolysis signal. These findings suggest that noncovalent, conformational alterations may be general features in cognitive decline, and more broadly, that proteins need not form amyloids for their misfolded states to be relevant to age-related deterioration in cognitive abilities.</p> <p>This deposition provides processed peptide quantifications for all LC-MS/MS proteomics experiments conducted for this study.</p>
Simulated complex structures of the h-FBP21 tandem WW domain with proline-rich ligand extracted from SmB/B' core-splicing protein
<p>The tandem WW domain of the human formin-binding protein 21 (h-FBP21 tWW) consists of two WW domains separated by a flexible linker. It can bind target sequences in two different orientations and the flexibility of the linker additionally allows the two WW domains to adopt various relative orientations to each other. As consequence, the elucidation of possible complex structures for the h-FBP21 tWW is very challenging.</p> <p>Here, we present two complex structures for the h-FBP21 tWW and a proline-rich sequence from its natural binding partner, the core-splicing protein SmB/B’. Showing parallel (‘6’) and antiparallel (’14’) binding orientation, the two structures also differ in the relative positioning of the WW domains.</p> <p>For further instructions regarding the files, please refer to ‘README’.</p>
Crystal structure of the tandem kinase & triphosphate tunnel metalloenzyme domain module of the TTM1 protein from Arabidoposis thaliana in complex with inorganic phosphate and citric acid - 3lambda SeMAD dataset
<p>bzip2ed tar archive containing the diffraction images (Pilatus 2M-F detector, SLS beamline PXIII, collected on 19.12.2016) for 3 wavelength Se MAD experiment (infl, inflection point, peak, peak, rem, high energy remote) and the associated data processing files (xds_inf, xds_peak, xds_rem) </p>
Crystal structure of the tandem kinase & triphosphate tunnel metalloenzyme domain module of the TTM1 protein from Arabidoposis thaliana in complex with an adenosine nucleotide analog.
<p>bzip2ed tar archive containing the diffraction images (Pilatus 2M-F detector, SLS beamline PXIII, collected on 19.12.2016) and the associated data processing files (xds) </p>
Crystal structure of the tandem kinase & triphosphate tunnel metalloenzyme domain module of the TTM1 protein from Arabidoposis thaliana in complex with inorganic phosphate and citric acid - native dataset
<p>bzip2ed tar archive containing the diffraction images (Pilatus 2M-F detector, SLS beamline PXIII, collected on 19.12.2016) and the associated data processing files (xds) </p>
Data of the protocol paper: Protein Structural Modeling for Electron Microscopy Maps Using VESPER and MAINMAST
<p>The archived file contains data of the test cases used used in the protocol paper. For each protocol, there's a folder containing input files the protocol takes, and a folder for the protocol output files.</p>
Conserved structural elements specialize ATAD1 as a membrane protein extraction machine
<p>The mitochondrial AAA protein ATAD1 (in humans; Msp1 in yeast) removes mislocalized membrane proteins, as well as stuck import substrates from the mitochondrial outer membrane, facilitating their re-insertion into their cognate organelles and maintaining mitochondria's protein import capacity. In doing so, it helps to maintain proteostasis in mitochondria. How ATAD1 tackles the energetic challenge to extract hydrophobic membrane proteins from the lipid bilayer and what structural features adapt ATAD1 for its particular function has remained a mystery. Previously, we determined the structure of Msp1 in complex with a peptide substrate (Wang et al., 2020). The structure showed that Msp1's mechanism follows the general principle established for AAA proteins while adopting several structural features that specialize it for its function. Among these features in Msp1 was the utilization of multiple aromatic amino acids to firmly grip the substrate in the central pore. However, it was not clear whether the aromatic nature of these amino acids were required, or if they could be functionally replaced by aliphatic amino acids. In this work, we determined the cryo-EM structures of the human ATAD1 in complex with a peptide substrate at near atomic resolution. The structures show that phylogenetically conserved structural elements adapt ATAD1 for its function while generally adopting a conserved mechanism shared by many AAA proteins. We developed a microscopy-based assay reporting on protein mislocalization, with which we directly assessed ATAD1's activity in live cells and showed that both aromatic amino acids in pore-loop 1 are required for ATAD1's function and cannot be substituted by aliphatic amino acids. A short α-helix at the C-terminus strongly facilitates ATAD1's oligomerization, a structural feature that distinguishes ATAD1 from its closely related proteins.</p>
FASST Structure Database files for "Tertiary motifs as building blocks for the design of protein-binding peptides"
<p>FASST Database files for use in the <a href="https://github.com/swanss/peptide_design">peptide design</a> pipeline.</p> <p>A complete list of structures provided in the databases is provided in the supplementary information of the Protein Science article.</p> <p>Singlechain structures: <a href="https://onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2Fpro.4322&file=pro4322-sup-0004-TableS6.txt">pro4322-sup-0004-TableS6.txt</a> (format: PDBID_CHAINID)</p> <p>Multichain structures: <a href="https://onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1002%2Fpro.4322&file=pro4322-sup-0005-TableS7.txt">pro4322-sup-0005-TableS7.txt</a> (format: PDBID)</p>
Results for blind docking Tocriscreen 2.0 compounds against AlphaFold2 structural modules of human and parasite proteins
<p>Docking (GNINA 1.0) of 1280 Tocriscreen 2.0 compounds against >4000 AlphaFold structures from humans and the parasitic nematode <em>Brugia malayi</em>. </p> <p>Results are provided as Pickle objects or GZipped CSVs.</p> <p>Tuned machine learning models (random forest and XGBoost) are also included for classifying actives/decoys.</p>
A computational study of the structure and function of human Zrt and Irt-like proteins metal transporters: An elevator-type transport mechanism predicted by AlphaFold2
<p>Data produced and analyzed in the manuscript "A computational study of the structure and function of human Zrt and Irt-like proteins metal transporters: An elevator-type transport mechanism predicted by AlphaFold2" by Pasquadibisceglie et al.</p> <p><br> If you include these data in your manuscript, please cite: Pasquadibisceglie A, Leccese A and Polticelli F (2022) A computational study of the structure and function of human Zrt and Irt-like proteins metal transporters: An elevator-type transport mechanism predicted by AlphaFold2. <em>Front. Chem.</em> 10:1004815. doi: 10.3389/fchem.2022.1004815</p>
Data for COLLAPSE: A representation learning framework for identification and characterization of protein structural sites
<p>Data for methods described in the paper "COLLAPSE: A representation learning framework for identification and characterization of protein structural sites" by Alexander Derry and Russ B. Altman (BioRxiv, 2022). https://www.biorxiv.org/content/10.1101/2022.07.20.500713v1</p>
Raw NGS Data for "Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins"
<p>This directory contains relevant fastq files used for deep sequencing analysis in the publication “Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins”. </p> <p>Fastq files are provided for presorted, uninduced and induced populations from DMS experiments of four homologs (TtgR, TetR, RolR, and MphR). Three replicates were performed for each sample.</p> <p>Data analysis of this deep sequencing data was performed using custom scripts, which are described in the methods section of the publication.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.