Skip to main content
zenodoopen

High quality protein residues: Top2018 all-atom-filtered residues - mmCIF

<p>Introduction<br> --------------------------------------------------------------------------------<br> This directory contains files from the Top2018 dataset by the Richardson Lab at Duke University.</p> <p>These are high-quality residues from high-quality, low redundancy protein chains in the PDB.</p> <p>This dataset is quality-filtered on all atoms in the residue.</p> <p>The accompanying publication is:<br> Williams, C. J., Richardson, D. C., &amp; Richardson, J. S. (2021). The importance of residue-level filtering, and the Top2018 best-parts dataset of high‐quality protein residues. Protein Science. http://doi.org/10.1002/pro.4239</p> <p>Usage recommendations<br> --------------------------------------------------------------------------------<br> Protein residues that fail the filtering criteria described below have been removed from the files.&nbsp; As a result, these files can be considered pre-filtered and will return only results for residues of good model quality with supporting experimental data.&nbsp; All protein atoms have been considered in filtering; these files should be usable for any protein question.&nbsp; If your work is strictly limited to mainchain atoms (plus CB), there is a separate version that has been filtered on only mainchain atoms.</p> <p>The Top2018 contains several different levels of homology clustering (30%, 50%, 70%, 90%) to ensure nonredundant datasets.&nbsp; The 70% homology level is a reliable default.&nbsp; These chains are listed in top2018_chains_hom70_fullfiltered_60pct_complete.txt and found in top2018_pdbs_full_filtered_hom70.tar.gz</p> <p>Files are organized in subdirectories based on the first two letters of their PDB ids.&nbsp; The included python script sample_file_loop.py may aid in accessing the directory structure.</p> <p>Files already contain hydrogens added by Reduce.&nbsp; NQH flips have been performed to ensure that these are the best versions of these structures.</p> <p>top2018_metadata_full_filtered.csv contains information on release date, resolution, and validation scores for each file.</p> <p>top2018_passrates_full_filtered.csv contains information on how many protein residues from the original chain passed the quality filters.</p> <p><br> Homology sets:<br> --------------------------------------------------------------------------------<br> Using sequence homology clusters provided by the RCSB PDB, for each homology cluster, the best chain was selected for inclusion in the dataset.&nbsp; This ensures minimal sequence/structural redundancy.</p> <p>The Top2018 is available at several different levels of homology clustering, which may be appropriate to different uses.&nbsp; Lists of the included chains at each homology level are included in this distribution.</p> <p>Lower homology numbers mean less redundancy, but fewer total chains in the dataset.</p> <p>For general use, ***we recommend the 70% homology set*** as a good balance between inclusivity and variety. This list is given in the file top2018_chains_hom70_fullfiltered_60pct_complete.txt</p> <p><br> Usage caveats:<br> --------------------------------------------------------------------------------<br> These files are incomplete.&nbsp; They are single chains from structures that may have had multiple chains.&nbsp; Residues that fail the filtering criteria have been removed.&nbsp; Programs with strong requirements for completeness or uninterrupted chains should be used with care.&nbsp; Chain completeness and fragmentation statistics are available in top2018_passrates_full_filtered.csv and as _top2018.percent_passrate in the .cif file.</p> <p>All ligands and waters associated with the chain have been preserved without filtering.&nbsp; Robust ligand filtering is beyond the scope of this dataset.&nbsp; Trust the ligands at your own discretion.</p> <p><br> Filtering criteria: Chain level<br> --------------------------------------------------------------------------------<br> Chain is protein<br> Released on or before Dec 31, 2018<br> Resolution &lt; 2.0<br> MolProbity Score &lt; 2.0<br> &lt;3% residues have cbeta deviations<br> &lt;2% residues have covalent bond length outliers<br> &lt;2% residues have covalent bond geometry outliers</p> <p>Using sequence homology clusters provided by the RCSB PDB, for each homology cluster, the chain with the best (lowest) average of Resolution and MolProbity Score was selected.</p> <p><br> Filtering criteria: Residue level<br> --------------------------------------------------------------------------------<br> Even excellent structures usually contain some poorly-resolved regions.&nbsp; Residue-level filtering helps avoid including these regions in otherwise high-quality data</p> <p>All atoms in a residue:<br> Bfactor &lt;= 40<br> Real-space correlation coefficient (rscc) &gt;= 0.7<br> 2Fo-Fc map value &gt;= 1.2</p> <p>Additionally, residues are not allowed to have:<br> Covalent geometry outliers<br> Steric overlaps or &quot;clashes&quot;, as per Probe<br> Alternate conformations</p> <p><br> Chain Completeness criteria<br> --------------------------------------------------------------------------------<br> Chains which lost &gt;40% of their residues during filtering were dropped from this dataset.&nbsp; All chains present here are at least 60% complete.</p> <p>Filtering documentation<br> --------------------------------------------------------------------------------<br> Each file documents its pruned residues and included segments in a cif data block named data_top2018_dataset. This block can be found at the end of the file.</p> <p>In the _top2018_deleted_residue loop, causes of pruning are documented. If a residues was removed due to failing the B-factor filter, a &quot;b&quot; will appear in the appropriate column. Otherwise, a &quot;.&quot; will appear. Other filtering criteria are treated similarly with the following codes:<br> b - B-factor<br> c - RSCC<br> m - map value<br> g - geometry outliers<br> o - steric overlaps<br> a - alternate conformations</p> <p>Version history<br> --------------------------------------------------------------------------------<br> Version 0.9<br> Initial upload to establish DOI</p> <p>Version 1.0<br> Initial full upload</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
20
Reuse readiness
8
Engagement
4