Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
53
datasets available to search
ShareScore release 0.9.0
Dataset results
53 results for “protein folding”
Temperature-dependent fold-switching mechanism of the circadian clock protein KaiB
<p>Derived data accompanying publication of <em>Temperature-dependent fold-switching mechanism of the circadian clock protein KaiB</em> (Zhang et al., PNAS 2024).</p> <p> </p> <p>This dataset contains data for fold-switching of KaiB from simulations performed using the Upside coarse-grained model (Jumper et al. PLoS Comput. Bio 2017). Files contained include collective variables, kinetic quantities (committors), and initial structures used to seed unbiased simulations. These data should be sufficient recreate the analysis shown in the associated publicaion. Raw trajectory files have not been deposited due to their size; contact the author (Spencer Guo) to request.</p>
PhasAGE Expert Seminar- Intrinsic Protein Disorder and Conditional Folding in AlphaFoldDB
<p>The PhasAGE <strong>Expert Seminars</strong> consist of a series of talks with speakers from PhasAGE partner’s institutions to promote a successful transfer of knowledge about PhasAGE topics – biomolecular phase separation, aging and age-related diseases.</p>
Data for Stabilization of non-native folds and programmable protein gelation in compositionally designed deep eutectic solvents
<div> <p>Full set of data related to the publication "Stabilization of non-native folds and programmable protein gelation in compositionally designed deep eutectic solvents", published in ACS Nano with DOI:<a title="https://doi.org/10.1021/acsnano.4c01950" href="https://doi.org/10.1021/acsnano.4c01950">10.1021/acsnano.4c01950</a></p> <p> Full details on data treatment and logging are included in the file "DataLogging.pdf". All data use ASCII encoding in delimited .txt files.</p> <p> </p> </div>
Data supporting "Slowest-first translation scheme: Structural asymmetry along protein sequences and co-translational folding"
<p>Contains data for a set of 16,200 non-redundant protein structures taken from the Protein Data Bank. Associated code can be found at https://github.com/jomimc/FoldAsymCode.</p>
Molecular dynamics trajectories of protein folding
<p>Molecular dynamics trajectories of protein folding are deposited for educational purposes.</p> <p>Currently, the following trajectories are available:</p> <ul> <li>Chignolin (five independent NVT simulations up to 1.5 micro-sec): <ul> <li>`movie.pse` is a PyMOL session file of MD trajectories. </li> <li>`movie.mp4` is a movie file that shows you how a protein folds during simulation.</li> <li>xtc files (Gromacs compressed format) and corresponding tpr files.</li> <li>gro files for movie making</li> </ul> </li> </ul> <p> </p> <p>Computational setting </p> <ul> <li>Amber ff99SB-ILDN for protein (Lindorff-Larsen, K. <em>et al.</em> Improved side-chain torsion potentials for the Amber ff99SB protein force field. <em>Proteins</em> <strong>78</strong>, 1950–1958 (2010))</li> <li>TIP3P water model</li> <li>0.1 M salt concentration </li> <li>NVT ensemble at 300 K with V-rescale thermostat (Bussi, G., Donadio, D. & Parrinello, M. Canonical sampling through velocity rescaling. <em>J. Chem. Phys.</em> <strong>126</strong>, 014101 (2007))</li> <li>Time step : 2 fs</li> <li>gromacs-2020.6</li> </ul>
Artificial intelligence method to design and fold alpha-helical structural proteins from the primary amino acid sequence
<p>Dataset for paper: Z. Qin, L. Wu, H. Sun, S. Huo, T. Ma, E. Lim, P.-Y. Chen, B. Marelli, M.J. Buehler, Artificial intelligence method to design and fold alpha-helical structural proteins from the primary amino acid sequence, Extreme Mechanics Letters, Vol. 36, 100652, 2020. <a href="https://doi.org/10.1016/j.eml.2020.100652">https://doi.org/10.1016/j.eml.2020.100652</a>.</p> <p>Code: https://github.com/lamm-mit/MNNN/ </p>
Memory kernel extraction and mean first-passage time for fast-folding proteins (Q - trajectories)
<p>Fraction of native contacts reaction coordinate (Q) trajectories for the 8 proteins that appear in the Dalton et. al. PNAS 2023. For the original all-atom data from which the Q(t) were calculated, contact the group of David E. Shaw at Shaw Research (see the paper - K. Lindorff-Larsen, S. Piana, R. O. Dror, D. E. Shaw, How fast-folding proteins fold. Science 334, 517–520 (2011))</p> <p>Data contains:<br> - Q(t) trajectories for 8 proteins <br> - Corresponding free energy profiles<br> <br> Example analysis codes (written in C++) are included for:<br> - Free energy calculation<br> - Mean first-passage times<br> - Velocity-velocity correlation function and position-force correlation function calculations, needed for memory kernel extraction<br> - Memory kernel extraction (see Ayaz et. al. PNAS 2021)</p>
Dataset for Peptide binder design with inverse folding and protein structure prediction
<p>Dataset for a paper on peptide design</p> <p> </p> <p><br> mutated_peptides - results for randomly intriduced mutations in protein-peptide complexes that can be predicted at 2 Å (Figure 1)<br> pdb_peptide - variation in the number of recycles (1-10) for 96 peptides (Figure 1)<br> minibinder - results for the minibinder set (Figure 2)<br> Pfam - results for the Pfam set (Figures 4+5)<br> protein_mpnn - results on protein_mpnn test set (Figure 6)</p> <p> </p> <p> </p>
Modulation of a protein-folding landscape revealed by AFM-based force spectroscopy notwithstanding instrumental limitations
Open the record for dataset details and reuse information.
Birth of new protein folds and functions in the virome - Structure Database
<p>This is the database of protein structures described in the manuscript "Birth of new protein folds and functions in the virome", by authors Jason Nomburg, Nate Price, and Jennifer A. Doudna.</p>
Validation of de novo designed water-soluble and transmembrane proteins by in silico folding and melting
<p>Here are all of the datasets generated and analysed during this study. </p> <p>Here is a breakdown of their content:</p> <ul> <li><strong>8_stranded_transmembrane_barrels.zip</strong> - raw data from Alphafold (3 and 48 recycles), ESMFold and raptor predictions of the 8 stranded TMBs. A file with all the sequences is also given</li> <li><strong>12_stranded_transmembrane_barrels.zip - </strong>raw data from the Alphafold and ESMfold predictions of the 12 stranded TMBs. A file with all the sequences is also given</li> <li><strong>water_soluble_barrels.zip</strong> - raw data from the Alphafold and ESMfold predictions of the water soluble beta barrels (designable and non-designable). A file with all the sequences is also given</li> <li><strong>all design models.zip</strong> - original design models for water-soluble (designable and non-designable), 8-stranded and 12-stranded TMBs</li> </ul> <p> </p> <ul> <li><strong>ESMfold_masking_exp.tar - </strong>this tar file contains all the ESMfold masking experiments performed to the water-soluble, 8 and 12-stranded transmembrane barrels. Inside there are zipped datasets for each masking experiment<br> </li> <li> <p><strong>ziped_raw_csv_files.zip - </strong>raw csv files with all the data necessary to analyse the figures </p> </li> <li> <p><strong>analysis_notebooks.zip </strong>- Jupyter notebooks used to analyse the output prediction data for all figures</p> </li> </ul> <p> </p> <p> </p>
Unraveling the Unfolding Mechanism of Pseudoazurin: Insights into Stabilizing Cupredoxin Fold as a Common Domain of Cu-Containing Proteins
<p>This dataset includes molecular dynamics (MD) simulation trajectories and experimental data used in the title named study. The MD trajectories cover simulations of pseudoazurin under various conditions: apo (pH 2, pH 3, pH 7), holo (pH 2, pH 3, pH 7), and explicit water simulations of holo at pH 7. Additionally, the dataset contains raw experimental data, including small-angle neutron scattering (SANS) curves, visible (Vis) absorption spectra, and circular dichroism (CD) spectra. This comprehensive dataset supports the investigation of unfolding mechanism of Pseudoazurin.</p>
On the role of native contact cooperativity in protein folding
<p>This repository provides supporting information related to the paper "On the role of contact cooperativity in protein folding" by Wang, Frechette and Best ( <a href="https://doi.org/10.1073/pnas.2319249121">https://doi.org/10.1073/pnas.2319249121</a> ) and contains: </p> <ol> <li>C++ code using Boltzmann machines to learn the parameters of an Ising model that best describes contact formation observed in an MD simulation. </li> <li>The fitted parameters and python scripts used for plotting several of the figures in the paper.</li> </ol>
Fueling ab initio folding with oceanic metagenomics enables structure and function predictions of new protein families
<p>Code and protein sequence database to construct multiple sequence alignment from Tara Ocean data.</p>
Supplemental Material for 'Deep Generative Models of Protein Structure Uncover Distant Relationships Across a Continuous Fold Space' and DeepUrfold
<p>Data provided for the paper Draizen, EJ, Veretnik, S, Mura, C, and Bourne, PE. "Deep Generative Models of Protein Structure Uncover Distant Relationships Across a Continuous Fold Space." <em>Nature Communications</em>, Aug. 2024.</p> <div> </div> <p> </p>
Folding pathway of a discontinuous two-domain protein_1
<p>This data set contains all the raw data collected for the preparation of the manuscript "Folding pathway of a discontinuous two-domain protein".</p> <p>Raw data for the main figures Figure 1B-C, Figure 2A, C-D, Figure 4B-C, Figure 5B, C-D and for the supplementary figures Figure S1A-B, Figure S3C, E, G, Figure S5, Figure S7A-B, Figure S8A, D, Figure S9, and Figure S15A-C are deposited. Data is given figure-wise in folders and figure panels in sub-folders. The recurrent raw data used for multiple figures is mentioned for respective figures. Document explaining in detail about the figures and respective data set is also provided.</p> <p>Data is given as measured single-molecule TCSPC data as well as the background measurements as buffer and IRF as dpbs measurements in each respective folder.</p> <p>Due to size, the repository is uploaded in two parts. This part I has all the data for above figures except Figure S3, Figure S8, Figure S9 and Figure S15 for which associated data are deposited in part II repository with Zenodo DOI https://doi.org/10.5281/zenodo.8136592.</p>
The insertase YidC chaperones the polytopic membrane protein MelB inserting and folding simultaneously from both termini
<p>The deposited data set contains data for the main figure 3 of the manuscript "The insertase YidC chaperones the polytopic membrane protein MelB inserting and folding simultaneously from both termini" by Blaimschein et al. published in Structure (2023).</p>
The insertase YidC chaperones the polytopic membrane protein MelB inserting and folding simultaneously from both termini
<p>The deposited data set contains data for the main figure 2 of the manuscript "The insertase YidC chaperones the polytopic membrane protein MelB inserting and folding simultaneously from both termini" by Blaimschein et al. published in Structure (2023).</p>
The insertase YidC chaperones the polytopic membrane protein MelB inserting and folding simultaneously from both termini
<p>The deposited data set contains data for the main figure 6 of the manuscript "The insertase YidC chaperones the polytopic membrane protein MelB inserting and folding simultaneously from both termini" by Blaimschein et al. published in Structure (2023).</p>
Computing free energies of fold-switching proteins using MELD x MD
<p>In this Zenodo repository, we provide the MELD script and data for a few representative systems in the DP-MELD.zip, GA_GB-MELD.zip, and RfaH-MELD.zip files. The contents of the repository are described below:</p> <ol> <li> <p>MELD Simulation:</p> <ul> <li>Filename: DP-MELD.zip, GA_GB-MELD.zip, and RfaH-MELD.zip</li> <li>Description: These archives contain the necessary files and scripts for the MELD simulation.</li> </ul> </li> <li> <p>Setup Script:</p> <ul> <li>Filename: setup.py</li> <li>Description: This script is used to set up the MELD simulation.</li> </ul> </li> <li> <p>Trajectory Analysis Script:</p> <ul> <li>Filename: Clustering.sh</li> <li>Description: This script analyzes the trajectories obtained from the MELD simulation.</li> </ul> </li> <li> <p>Protein Information:</p> <ul> <li>Location: TEMPLATES folder</li> <li>Files: <ul> <li>Protein topology file: .top</li> <li>Coordinate file: .crd</li> <li>PDB file: .PDB</li> </ul> </li> <li>Description: These files provide input information for a few representative proteins.</li> </ul> </li> <li> <p>Residue-Residue Contact Information:</p> <ul> <li>Files: <ul> <li>contact_model1.dat</li> <li>contact_model2.dat</li> </ul> </li> <li>Description: These files contain information about the contacts between residues.</li> </ul> </li> <li> <p>Replica Trajectory Files:</p> <ul> <li>Filename: trajectory.00.dcd</li> <li>Description: These files contain the trajectories obtained from the simulation for the corresponding bottom replica.</li> </ul> </li> <li>Clustering Output: <ul> <li>Folders: Cluster_6 or Cluster_3.5</li> <li>Description: These folders contain the results of the clustering analysis, including the computed population and the average conformers for each cluster.</li> </ul> </li> </ol> <p>Furthermore, we provide an additional archive called unfold.zip:</p> <ol> <li>Unfolded Ensemble: <ul> <li>Filename: unfold.zip</li> <li>Description: This archive contains the unfolded ensemble, which is used to determine the force required for rebalancing two group springs for the MELD run between two conformers (A and B) before executing the MELD simulation.</li> </ul> </li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.