Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.9.0
Dataset results
7 results for “Protein Data Bank”
List of the structures of S-protein in complex with ligands deposited in the Protein Data Bank until the 1st January 2021.
<p>All 131 structures of SARS-CoV-2 S-protein in complex with a ligand released on the PDB until the 1<sup>st</sup> January 2021 were categorised by ligand type: hACE2, antibody Fab fragments, VHH antibody fragments or <em>de novo</em> designed peptide scaffolds. The ligands’ amino acid sequences, the method by which the structures were determined and their resolution were retrieved from the PDB. Information regarding the ligands' production method, dissociation constants (K<sub>D</sub>), S-protein segment against which the K<sub>D</sub> were measured and the determination methods were retrieved from the respective references. The categorisation of ligands by S-protein binding site and listing of S-protein conformation in each structure were achieved by visual analysis of all the structures using molecular visualisation software PyMOL.</p>
Ligands detected by CheckMyBlob in the Protein Data Bank
<p>A dataset containing numerical descriptions all ligands that CheckMyBlob was capable of detecting automatically on the entire PDB as of May 1st, 2017. It is the "master" data set containing all ligands queried and detected as described in the Kowiel et al. paper "Automatic recognition of ligands in electron density by machine learning methods".</p> <p>The file is compressed using 7zip to allow for faster downloads. The compressed file weighs around 1.1 GB, whereas the uncompressed CSV will take close to 3.0 GB of disk space. The all_summary.csv file can be used to reproduce the filtered ligand data sets (CMB, TAMC, CL) described in the Kowiel et al. paper. Additionally, the data set can be used as a source to create data sets based on other filtering criteria (e.g. ligand subsets of your choice), or on its own a as source of knowledge about all ligands that CheckMyBlob was capable of detecting automatically on the entire PDB as of May 1st, 2017.</p> <p>For machine learning applications, please read the code at https://github.com/dabrze/CheckMyBlob, to see examples describing how to preprocess the data. In particular, the following attributes should not be used during model training or testing: "pdb_code", "res_id", "chain_id", "local_res_atom_count", "local_res_atom_non_h_count", "local_res_atom_non_h_occupancy_sum", "local_res_atom_non_h_electron_sum", "local_res_atom_non_h_electron_occupancy_sum", "local_res_atom_C_count", "local_res_atom_N_count", "local_res_atom_O_count", "local_res_atom_S_count", "dict_atom_non_h_count", "dict_atom_non_h_electron_sum", "dict_atom_C_count", "dict_atom_N_count", "dict_atom_O_count", "dict_atom_S_count", "fo_col", "fc_col", "weight_col", "grid_space", "solvent_radius", "solvent_opening_radius", "part_step_FoFc_std_min", "part_step_FoFc_std_max", "part_step_FoFc_std_step", "local_volume", "res_coverage", "blob_coverage", "blob_volume_coverage", "blob_volume_coverage_second", "res_volume_coverage", "res_volume_coverage_second", "skeleton_data", "resolution_max_limit", "part_step_FoFc_std_min", "part_step_FoFc_std_max", "part_step_FoFc_std_step".</p> <p>The target attribute for classification is: <strong>res_name</strong>.</p>
Dataset of Bnet values of 23 radiation damage series analysed in "Quantifying and comparing radiation damage in the Protein Data Bank"
<p>Dataset of Bnet values, plus associated metadata (resolution, Rwork, Rfree, temperature, etc.), for 23 radiation damage series deposited in the PDB that are analysed in the publication "Quantifying and comparing radiation damage in the Protein Data Bank"</p>
Dataset of Bnet values of 93,978 PDB-REDO structures analysed in "Quantifying and comparing radiation damage in the Protein Data Bank"
<p>Dataset of Bnet values, plus associated metadata (resolution, Rwork, Rfree, temperature, etc.), for 93,978 PDB-REDO structures analysed in the publication "Quantifying and comparing radiation damage in the Protein Data Bank"</p>
Data from: Exploring the universe of protein structures beyond the Protein Data Bank
It is currently believed that the atlas of existing protein structures is faithfully represented in the Protein Data Bank. However, whether this atlas covers the full universe of all possible protein structures is still a highly debated issue. By using a sophisticated numerical approach, we performed an exhaustive exploration of the conformational space of a 60 amino acid polypeptide chain described with an accurate all-atom interaction potential. We generated a database of around 30,000 compact folds with at least 30% of secondary structure corresponding to local minima of the potential energy. This ensemble plausibly represents the universe of protein folds of similar length; indeed, all the known folds are represented in the set with good accuracy. However, we discover that the known folds form a rather small subset, which cannot be reproduced by choosing random structures in the database. Rather, natural and possible folds differ by the contact order, on average significantly smaller in the former. This suggests the presence of an evolutionary bias, possibly related to kinetic accessibility, towards structures with shorter loops between contacting residues. Beside their conceptual relevance, the new structures open a range of practical applications such as the development of accurate structure prediction strategies, the optimization of force fields, and the identification and design of novel folds.
Data from: Exploring the universe of protein structures beyond the Protein Data Bank
Open the record for dataset details and reuse information.
Dataset comparing different skewness methods tested in "Quantifying and comparing radiation damage in the Protein Data Bank"
<p>Dataset comparing five different skewness metrics (including the Bnet metric) across 23 radiation damage datasets in the publication "Quantifying and comparing radiation damage in the Protein Data Bank"</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.