Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
56
datasets available to search
ShareScore release 0.9.0
Dataset results
56 results for “AlphaFold”
Alphafold and ColabFold models of E. coli and consensus Bcs complexes
<p>ColabFold and AlphaFold 3 models used for structure modeling and interpretation in Anso et al. 'Structural basis for synthase activation and cellulose modification in the <em>E. coli</em> Type II Bcs secretion system'. </p>
Collective Variable for Metadynamics Derived from AlphaFold Output
<p>AlphaFold is the state of the art method for prediction of 3D structures of proteins from the amino acid sequence by neural networks. One of the outputs of AlphaFold is a probability profile of inter-residue distances for all residue pairs. We used this profile to evaluate any conformation of the studied protein to express its compliance with the AlphaFold prediction. This value can be used as a collective variable in metadynamics or parallel tempering metadynamics to accelerate protein folding in a molecular simulation. We applied this approach on folding of mini-proteins Trp-cage and beta hairpin. See V. Spiwok, M. Krečka & A. Křenek: <a href="http://doi.org/10.3389/fmolb.2022.878133">Collective Variable for Metadynamics Derived from AlphaFold Output</a> <em>Frontiers in Molecular Biosciences</em> <strong>9</strong> 878133 (2022) DOI: 10.3389/fmolb.2022.878133.</p>
The Encyclopedia of Domains (TED) structural domains assignments for AlphaFold Database v4
<h3>Dataset description:</h3> <p>The Encyclopedia of Domains (TED) is a joint effort by CATH (Orengo group) and the Jones group at University College London to identify and classify protein domains in AlphaFold2 models from AlphaFold Database version 4, covering over 188 million unique sequences and 365 million domain assignments. </p> <p>In this data release, we will be making available to the community a table of domain boundaries and additional metadata on quality (pLDDT, globularity, number of secondary structures), taxonomy, and putative CATH SuperFamily or Fold assignments, for all 365 million domains (~324 million domains in TED100 and ~40 million domains in TED-redundant).</p> <p>For all chains in the chain-level TED-redundant files, the file contains boundary predictions, consensus level and information on the TED100 representative.</p> <p>For both TED100 and TED-redundant we provide domain boundary predictions outputted by each of the three methods employed in the project (Chainsaw, Merizo, UniDoc). </p> <p>We are making available 7,427 PDB files for potentially novel folds identified during the TED classification process, with an annotation table sorted by novelty, as well as 6,433 highly symmetrical folds representatives.</p> <p>Please use the gunzip command to extract files with a '.gz' extension and "tar -xzvf file.tar.gz" to open .tar.gz files .</p> <p>CATH annotations have been assigned using the Foldseek algorithm applied in various modes, and the Foldclass algorithm, both of which are used to report significant structural similarity to a known CATH domain. </p> <p><br><strong>Note: The TED protocol differs from that of the standard CATH Assignment protocol for superfamily assignment, which also involves HMM-based protocols and manual curation for classification into superfamilies.</strong></p> <h3> </h3> <h3><strong>Changelog Version 5:</strong></h3> <ul> <li><strong>Add</strong>: ted_365m.domain_summary.cath.globularity.taxid.tsv.tar.gz - This table, in the same format as the previous ted_100_324m.domain_summary.cath.globularity.taxid.tsv.tar.gz, contains per-domain annotations for the whole of TED, including metadata on domain quality metrics such as secondary structure elements counts, globularity scores, average pLDDT and taxonomical assignments.</li> <li><strong>Add:</strong> high_symmetry_folds_set.domain_summary.tsv.gz - subset of ted_365m.domain_summary.cath.globularity.taxid.tsv containing information on 6,433 high symmetry folds in TED. The entries are sorted in descending order by Z-score obtained from SymD.</li> <li><strong>Add:</strong> high_symmetry_folds_set_models.tar.gz - TED domain models in PDB format for 6,433 high symmetry folds in TED.</li> <li><strong>Add:</strong> ISP_data.tar.gz - Raw data for Interacting SuperFamily Pairs calculations used in the manuscript. A more detailed description of the ISP data is available below as well as within the tar.gz file. </li> <li><strong>Add:</strong> ted_redundant_40m_domain_id.list.gz - list of TED_domain_ID in TED redundant</li> <li><strong>Add:</strong> ted_100_324m_domain_id.list.gz - list of TED_domain_ID in TED100</li> <li><strong>Fix/Replace</strong>: A domain-level summary of TED, now consolidated into <strong>ted_365m.domain_summary.cath.globularity.taxid.tsv</strong>, is consistent with the protocol used in the manuscript. As Foldclass and Foldseek T-level hits provide all 4 CATH digits, we removed the H portion of the CATH code from each prediction at the T-level. <br>Previously, the following columns <br> 14. cath_label - CATH superfamily code if predicted, either a C.A.T.H. homologous superfamily or C.A.T. fold assignment. i.e. 3.40.50.300<br> 15. cath_assignment_level - H for homologous superfamily assignment, T for fold level assignment.<br> 16. cath_assignment_method - Method used to assign a CATH label, either Foldseek or Foldclass<br>sometimes showed an additional label with a T-level prediction by Foldclass in the case of T-level assignments obtained by Foldseek, e.g.<br>3.40.30,3.40.30 T foldseek,foldclass<br>This has now been corrected to reflect the TED protocol, with Foldclass T-level assignments applied only to domains where a T-level assignment could not be applied using Foldseek, e.g. <br>domain-x 3.40.30 T foldseek<br>domain-y 3.20.20 T foldclass<br><br>Thus, in the current version of the data, CATH assignments label can only be <br>H-level assignment by Foldseek (i.e. 3.40.50.300 H foldseek)<br>T-level assignment by Foldseek (i.e. 3.40.30 T foldseek)<br>T-level assignment by Foldclass (i.e. 3.40.30 T foldclass)<br>or no assignment (- - - )</li> </ul> <h3><br>This dataset contains:</h3> <ul> <li><strong>ted_214m_per_chain_segmentation.tsv</strong><br>The file contains all 214M protein chains in TED with consensus domain boundaries and proteome information in the following columns.<br>1. AFDB_model_ID: chain identifier from AFDB in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br>2. md5 hash for chain sequence<br>3. nres - number of residues in chain<br>4. n_high - number of high consensus domains predicted in chain<br>5. n_med - number of medium consensus domains predicted in chain<br>6. n_low - number of low consensus domains predicted in chain<br>7. high_consesnsus - boundaries of high consensus domains predicted in chain. If none, 'na'<br>8. med_consensus - boundaries of medium consensus domains predicted in chain. If none, 'na'<br>9. low_consensus - boundaries of low consensus domains predicted in chain. If none, 'na'<br>10. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4<br><br></li> <li><strong>ted_365m_domain_boundaries_consensus_level.tsv.gz</strong><br>The file contains all domain assignments in TED100 and TED-redundant (365M) in the format:<br>1. TED_ID: TED domain identifier in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED03<br>2. Boundaries: domain boundaries in the format <start>-<stop> or <start>-<stop>_<start>-<stop> for discontinuous domains.<br>3. Consensus: either high or medium.<br><br></li> <li><strong>ted_100_324m_domain_id.list.gz</strong> - list of ~324 million domain identifiers in TED100, one per line in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED0<br><br></li> <li><strong>ted_redundant_40m_domain_id.list.gz</strong> - list of ~40 million domain identifiers in TED redundant, one per line in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED0<br><br></li> <li><strong>ted_365m.domain_summary.cath.globularity.taxid.tsv, novel_folds_set.domain_summary.tsv</strong> and <strong>high_symmetry_folds_set.domain_summary.tsv</strong> are header-less with the following columns separated by tabs (.tsv). novel_folds_set.domain_summary.tsv is sorted by novelty<br><strong>Note: The TED protocol differs from that of the standard CATH Assignment protocol for superfamily assignment, which also involves HMM-based protocols and manual curation for classification into superfamilies.</strong><br><br> 1. ted_id - TED domain identifier in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED03<br> 2. md5_domain - md5 hash of domain sequence<br> 3. consensus_level - medium (2 methods agreement) or high (3 methods agreement)<br> 4. chopping - domain boundaries in the format <start>-<stop> or <start>-<stop>_<start>-<stop> for discontinuous domains<br> 5. nres_domain - number of residues in domain<br> 6. num_segments - number of individual segments in domain. <br> 7. plddt - average pLDDT for domain (range from 0 to 100)<br> 8. num_helix_strand_turn - number of helix strand turns predicted by STRIDE<br> 9. num_helix - number of helices predicted by STRIDE<br> 10. num_strand - number of strands predicted by STRIDE<br> 11. num_helix_strand - number of helices and strands predicted by STRIDE<br> 12. num_turn - number of turns predicted by STRIDE<br> 13. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4<br> 14. cath_label - CATH superfamily code if predicted, either a C.A.T.H. homologous superfamily or C.A.T. fold assignment. i.e. 3.40.50.300. Otherwise '-'<br> 15. cath_assignment_level - H for homologous superfamily assignment, T for fold level assignment. Otherwise '-'<br> 16. cath_assignment_method - Method used to assign a CATH label, either foldseek or foldclass. Otherwise '-'<br> 17. packing_density - metric used to determine globularity. A domain with packing_density >=10.333 and norm_rg below 0.356 is considered globular<br> 18. norm_rg - normalised radius of gyration. A domain with packing_density >=10.333 AND norm_rg below 0.356 is considered globular. <br> 19. tax_common_name - Common name for organism<br> 20. tax_scientific_name - Scientific name for organism<br> 21. tax_lineage - Full taxonomic lineage.<br><br></li> <li><strong>ted_324m_seq_clustering.cathlabels.tsv.gz</strong> <br>The file contains the results of the domain sequences clustering with MMseqs2. <br>Columns:<br>1. Cluster_representative<br>2. Cluster_member<br>3. CATH code assignment if available i.e. 3.40.50.300 for a domain with a homologous match or 3.20.20 for a domain matching at the fold level in the CATH classification<br>4. CATH assignment type - either Foldseek-T, Foldseek-H or Foldclass<br><br></li> <li><strong>Domain assignments for TED redundant using single-chain and multi-chain consensus in</strong> <strong>ted_redundant_39m.multichain.consensus_domain_summary.taxid.tsv.gz and ted_redundant_39m.singlechain.consensus_domain_summary.taxid.tsv.gz</strong> </li> <li> The file <strong>ted_redundant_39m.multichain.consensus_domain_summary.taxid.tsv.gz</strong> contains a header with the following fields. Each column is tab-separated (.tsv).<br> 1. TED_redundant_id - TED chain identifier in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br> 2. md5 - md5 hash for chain sequence<br> 3. nres - number of residues in chain<br> 4. n_high - number of high consensus domains predicted in chain<br> 5. n_med - number of medium consensus domains predicted in chain<br> 6. high_consensus - boundaries of high consensus domains predicted in chain<br> 7. med_consensus - boundaries of medium consensus domains predicted in chain<br> 8. ndom_consensus - number of consensus domains predicted in chain<br> 9. n_targets - number of chains considered for consensus calculation<br> 10. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4<br> 11. TED_redundant_species - Scientific name for organism the chain originally comes from.<br> 12. TED100_chain_rep - TED100 representative for chain <br> 13. TED100_chain_rep_species - Species of TED100 representative for chain.</li> <li>The file <strong>ted_redundant_39m.singlechain.consensus_domain_summary.taxid.tsv</strong> contains a header with the following fields. Each column is tab-separated (.tsv).<br> 1. TED_redundant_id - TED chain identifier in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br> 2. md5 - md5 hash for chain sequence<br> 3. nres - number of residues in chain<br> 4. n_high - number of high consensus domains predicted in chain<br> 5. n_med - number of medium consensus domains predicted in chain<br> 6. high_consensus - boundaries of high consensus domains predicted in chain<br> 7. med_consensus - boundaries of medium consensus domains predicted in chain<br> 8. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4 <br> 9. TED_redundant_species - Scientific name for organism the chain originally comes from<br> 10. TED100_chain_rep - TED100 representative for chain <br> 11. TED100_chain_rep_species - Species of TED100 representative for chain.</li> </ul> <p> </p> <ul> <li><strong>novel_folds_set_models.tar.gz</strong> contains PDB files of all novel folds representatives identified in TED100.<br><br></li> <li><strong>high_symmetry_folds_set_models.tar.gz</strong> contains PDB files of all highly symmetrical folds representatives identified in TED100.<br> </li> <li><strong>Per-tool domain boundaries_predictions</strong> - All per-tool domain boundaries predictions for TED100 and TED-redundant are in the same format with the following columns.<br> 1. TED_chainID - TED chain identifier in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br> 2. TED_chain_md5 - md5 hash for chain sequence<br> 3. TED_chain_length - number of residues in chain<br> 4. ndoms - number of domains predicted in chains<br> 5. Domain boundaries - domain boundaries in the format <start>-<stop> or <start>-<stop>_<start>-<stop> for discontinuous domains<br> 6. Prediction probability - probability of each per-chain prediction<br>Domain boundaries predictions share the same format, with each segment separated by '_' and segment boundaries (start,stop) separated by '-'<br> <br> i.e.domain prediction by Merizo for AF-A0A000-F1-model_v4<br> AF-A0A000-F1-model_v4 e8872c7a0261b9e88e6ff47eb34e4162 394 2 10-52_289-394,53-288 0.90077<br> <br> Merizo predicts one continuous domain and a discontinuous domain,<br> Domain1 (discontinuous): 10-52_289-394<br> segment1: 10-52<br> segment2: 289-394<br> Domain 2 (continuous):<br> segment 1: 53-288<br><br></li> <li><strong>ISP_data.tar.gz</strong> contains raw data for the Interacting Superfamily Pairs (ISP) calculations featured in the manuscript. The archive contains a README as well as :<br>all_ISP_data_cath.pkl: <br>ISP data for CATH 4.3 in Python pickle format. <br>A Python dictionary with the following contents: <br>Each key is an ISP, e.g. '3.40.640.10-3.90.1150.10' <br>Each value is a dictionary with the following contents:<br>key 'aligned_domain_pairs': value is a Python list of length N, each element is a string specifying the two TED domain IDs in contact, separated by a colon, e.g. "AF-A0A000-F1-model_v4_TED02:AF-A0A000-F1-model_v4_TED01" <br>key 'vectors': value is a numpy.ndarray of shape (N, 3). Each row is a raw unnormalized interaction vector for the corresponding domain pair, after aligning to the reference structure. <br>Any given index in each list or ndarray has the data for a single domain pair; the order is constant in each list/array. <br>---------------------------------- <br>all_ISP_data_afdb.pkl: <br>ISP data for TED100 in Python pickle format. <br>A Python dictionary with the following contents: <br>Each key is an ISP, e.g. '3.40.640.10-3.90.1150.10' <br>Each value is a dictionary with the following contents:<br>key 'aligned_domain_pairs': value is a Python list of length N, each element is a string specifying the two TED domain IDs in contact, separated by a colon, e.g. "AF-A0A000-F1-model_v4_TED02:AF-A0A000-F1-model_v4_TED01"<br>key 'vectors': value is a numpy.ndarray of shape (N, 3). Each row is a raw unnormalized interaction vector for the corresponding domain pair, after aligning to the reference structure. <br>key 'choppings': value is a Python list of length N, each element is a colon-separated string containing the TED chopping strings for the domains in contact, e.g. "54-288:11-41_290-389". The format for each 'chopping' follows that used in the main TED TSV files. <br>key 'pae_score': value is a numpy.ndarray of floats of shape (N,). Each element is the median PAE score between the domains in contact, computed across both relevant parts of the PAE matrix as described in the paper. <br>Any given index in each list or ndarray has the data for a single domain pair; the order is constant in each list/array. <br>NB: each list of domains has not been filtered by PAE score, so the pae_score values include values greater than 4.0, which was the threshold used to filter confident predictions in the manuscript.<br>------------------------------------- <br>isp_data_afdbonly_nopaefilter.csv:<br>A subset of the data in the TED100 .pkl file above, in CSV format. <br>Each row contains the following fields: AFDB ID, e.g. AF-A0A000-F1-model_v4 ISP, e.g. 3.40.640.10-3.90.1150.10<br>Colon-separated domain ID pair, e.g. AF-A0A000-F1-model_v4_TED02:AF-A0A000-F1-model_v4_TED01 PAE score, e.g. '4.0' <br>As the aforementioned files, this data has not been filtered by PAE score values.<br><br></li> <li><strong>ted-tools-main.zip</strong> - copy of the https://github.com/psipred/ted-tools repository, containing tools and software used to generate TED.<br><br></li> <li><strong>cath-alphaflow-main.zip</strong> - copy of CATH-AlphaFlow, used to generate globularity scores for TED domains.<br><br></li> <li><strong>ted-web-master.zip</strong> - copy of TED-web, containing code to generate the web interface of TED (https://ted.cathdb.info)<br><br></li> <li><strong>gofocus_data.tar.bz2</strong> - GOFocus model weights</li> </ul>
AlphaFold structures reported in "AlphaFold2 Can Predict Single-Mutation Effects"
<p>This contains AlphaFold predictions for X proteins that are found in the Protein Data Bank (PDB), that were used to evalluate AlphaFold's predictions of mutation effects. This includes one set of structures predicted by AlphaFold2.0, using default settings, and one structure for each of 5 models. This also includes structures predicted by the ColabFold version of AlphaFold (6 recycles, 5 models, no template, amber minimization, 4 repeats).</p><p>There are also additional predicted structures that are found in the PDB that were not analyzed in the paper.</p><p>There are AlphaFold predictions for three proteins (BFP / RFP, GFP, and PafA), covering either all (BFP/RFP, PafA) or a subset (GFP) of the sequences in three datasets of phenotype measurements from high-throughput experiments.</p><p>Results are separated into tar files based on whether DeepMind (AF2.0) or ColabFold implementation was used.</p><p>Folders under "ColabFold/PDB" are labelled according to a sequence ID, since multiple PDB structures can exist for a single sequence. These sequence IDs can be mapped back to PDB IDs using the information in "seq_id_pdb_id.json".</p><p>All PDB files have been compressed using Foldcomp (<a href="https://github.com/steineggerlab/foldcomp">https://github.com/steineggerlab/foldcomp</a>). Foldcomp is required to decompress the ".fcz" files in order to recover the ".pdb" files.</p>
Integrating AlphaFold pLDDT Scores into CABS-flex for Enhanced Protein Flexibility Simulations
<div>This dataset accompanies the publication "Integrating AlphaFold pLDDT Scores into CABS-flex for enhanced protein flexibility simulations".</div> <div>This project was funded by the OPUS grant from the National Science Centre, Poland [2020/39/B/NZ2/01301].</div> <div> </div> <div>Training_set_protein_chains.txt and Whole_set_protein_chains.txt have lists of all PDB ID and chain used.</div> <div>Description_of_runs.csv has a list of every run tested. Run number correponds to csv file in Results_run.tar.gz.</div> <div>Every csv file has following columns: </div> <div> <ul> <li>PDB - PDB ID and chain </li> <li>Total_residues</li> <li>%_C - Percent of secondary structure assigned as coil by DSSP</li> <li>%_H - Percent of secondary structure assigned as helix by DSSP</li> <li>%_E - Percent of secondary structure assigned as sheet by DSSP</li> <li>%_T - Percent of secondary structure assigned as turn by DSSP</li> <li>pLDDT_mean - Average pLDDT score across all residues</li> <li>pLDDT_std - Standard deviation of pLDDT scores across all residues</li> <li>Unique_restraints - Number of unique restraints created by CABS-flex</li> <li>RMSF_CABS_R1_corr - RMSF correlation between CABS-flex and first MD simulation from ATLAS</li> <li>RMSF_CABS_R2_corr - RMSF correlation between CABS-flex and second MD simulation from ATLAS</li> <li>RMSF_CABS_R3_corr - RMSF correlation between CABS-flex and third MD simulation from ATLAS</li> <li>Highest_RMSF_corr - Highest RMSF correlation out of three</li> </ul> </div> <div> </div>
Improved protein complex prediction with AlphaFold-multimer by denoising the MSA profile
<p>Supporting data for AFProfile</p> <p>casp15.tar.zst - predicted structures and MSAs for the CASP15 set<br>native_afm_2_6_bench.tar.zst - native cif files for all complexes with ranking confidence <0.75 in the AFM 2-6 chains benchmark ( https://doi.org/10.1093/bioinformatics/btad424)<br>pred_top_models_afm_2_6_bench.tar.zst - predicted top ranked models and scores for all 100 samples.<br>afm_opt_metrics.csv - the best models, confidences and MMscores for the AFProfile run on the AFM 2-6 chains benchmark (n=427 structures)<br>msa_shapes.csv - the shape of the MSA as input to AFM for each structure<br>directed.tar.zst - contains all models for the AFProfile run on the AFM 2-6 chains benchmark (n=42700 samples)</p> <p>The directories are compressed with zstd: https://github.com/facebook/zstd<br>Uncompress:<br>tar --use-compress-program full/path/to/zstd -xvf file.tar.zst</p>
Data generated for the publication: Keeping it in the family: Using protein family templates to rescue poor AlphaFold models unliked
<p>Data and manuscript of:</p> <p>Keeping it in the family: Using protein family templates to rescue low confidence AlphaFold2 models</p> <p>Francesco Costa1, Matthias Blum1 and Alex Bateman1</p> <ol> <li>European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton. CB10 1SD. UK</li> </ol> <ul> <li>results contains the workflow results;</li> <li>AF2_seed contains results of the comparison with AF2 run with multiple seeds;</li> </ul>
Supplementary data frames, AlphaFold models, Normal Mode Analysis (NMA) Data, and NMA of Corresponding NMR Ensembles in the S2RCI, MD, and S2 Datasets for "Gradations in protein dynamics captured by experimental NMR are not well represented by AlphaFold2 models and other computational metrics"
<h1><strong>Changes applied to V2</strong></h1> <p>In addition to the supplementary dataframes and AlphaFold models from each dataset in V1, V2 includes the additional data outlined below.</p> <p>The <strong>S2RCI</strong> and <strong>MD</strong> datasets include comprehensive analyses of AlphaFold2 models (both before and after truncation). These datasets feature: </p> <ul> <li><strong>AlphaFold2 Models</strong>: Both original and truncated structures. </li> <li><strong>WEBnma Modes</strong>: `modes.txt` files generated from WEBnma analysis, available for both non-truncated and truncated AF2 models. </li> <li><strong>Root-Mean-Square-Fluctuations (RMSF)</strong>: Profiles calculated before and after truncation of AF2 models. </li> <li><strong>NMR Data: Normal Mode Analysis (NMA)</strong>: Performed on corresponding NMR ensembles (see below). </li> </ul> <p> </p> <p>The <strong>NMR Data</strong> of NMA in these datasets includes: </p> <ul> <li>NMR ensembles </li> <li>Individual NMR models extracted from each ensemble </li> <li>STRIDE secondary structure calculations per-individual NMR models</li> <li>RMSF profiles per-individual NMR models</li> </ul> <p>For detailed information, please refer to the `Readme.txt` file within each corresponding folder. </p> <p>The <strong>S2 dataset</strong> includes all the features listed above, except for the NMR analysis.</p>
Prediction and Visualization of Human Transmembrane Proteins using AlphaFold and Protein Language Models
<p><strong>Description:</strong> <strong>TMvis</strong> ("TMvis496.tar.gz") is a dataset containing 496 3D-structures of predicted human transmembrane proteins (TMP) and their predicted membrane embedding. The method TMbed [1], based on the protein language model ProtT5 [2] predicted 4.967 TMP for the human proteome (20,375 proteins, UniProt [3] version April 2022; excluding TITIN_HUMAN due to length). For these proteins, we obtained AlphaFold [4] structures from AlphaFoldDB [5] with an average per-residue confidence score (pLDDT) of more than 90%. This resulted in the 496 proteins of TMvis, as can be found in "TMvis496.fasta". The membrane embedding was predicted using the methods ANVIL [6], PPM3 [7], and per-residue TMbed predictions. As the three methods are based on different approaches, we decided to publish results for all. The figure “TMvis_project_overview.png” provides a graphical overview for each step described above.</p> <p><strong>TMvis Folder Structure:</strong> TMvis is separated into “alpha” containing predicted alpha-helical TMPs, and “beta” containing predicted beta-barrel TMPs. Within these folders, each protein is assigned one folder, identifiable by the respective unique UniProt ID. Each protein folder consists of:<br> - “UniprotID.fasta” with UniProt ID, sequence, TMbed per-residue prediction<br> - “AF-UniprotID-F1-model_v2.pdb” with the AlphaFold structure<br> - “AF-UniprotID-F1-model_v2.cif” with the AlphaFold structure<br> - “AF-UniprotID-F1-model_v2_ANVIL.pdb” with predicted ANVIL membrane embedding<br> - “AF-UniprotID-F1-model_v2_ppm.pdb” predicted PPM3 membrane embedding</p> <p>TMvis <br> | <br> ├── alpha <br> │ │ <br> │ ├── A0A087X1C5 <br> │ │ ├── A0A087X1C5.fasta <br> │ │ ├── AF-A0A087X1C5-F1-model_v2.pdb <br> │ │ ├── AF-A0A087X1C5-F1-model_v2.cif <br> │ │ ├── AF-A0A087X1C5-F1-model_v2_ANVIL.pdb <br> │ │ └── AF-A0A087X1C5-F1-model_v2_ppm.PDB <br> │ └── ... <br> └── beta <br> └── P45880</p> <p><strong>TMvis visualization:</strong> The 3D-visualization of every protein in the dataset TMvis can be easily accessed using the Jupyter Notebook “TMvis.ipynb”. It contains detailed descriptions the different membrane prediction tools ANVIL, PPM3, and TMbed as well as the respective code. Additionally, it allows to visualize the per-residue confidence scores (pLDDT) of AlphaFold.</p> <p>——————————————————————————————————————————————————————————————————————————</p> <p><strong>References:</strong></p> <p>[1] TMbed - TMbed Bernhofer, Michael, and Burkhard Rost. 2022. “TMbed – Transmembrane Proteins Predicted through Language Model Embeddings.” bioRxiv.</p> <p>[2] ProtT5 - A. Elnaggar et al., "ProtTrans: Towards Cracking the Language of Lifes Code Through Self-Supervised Deep Learning and High Performance Computing," in IEEE Transactions on Pattern Analysis and Machine Intelligence, doi: 10.1109/TPAMI.2021.3095381.</p> <p>[3] UniProt - UniProt Consortium (2021). UniProt: the universal protein knowledgebase in 2021. Nucleic acids research, 49(D1), D480–D489.</p> <p>[4] AlphaFold - AlphaFold Jumper, John, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, et al. 2021. “Highly Accurate Protein Structure Prediction with AlphaFold.” Nature 596 (7873): 583–89.</p> <p>[5] Alphafold DB - Varadi, Mihaly, Stephen Anyango, Mandar Deshpande, Sreenath Nair, Cindy Natassia, Galabina Yordanova, David Yuan, et al. 2022. “AlphaFold Protein Structure Database: Massively Expanding the Structural Coverage of Protein-Sequence Space with High-Accuracy Models.” Nucleic Acids Research 50 (D1): D439–44.</p> <p>[6] ANVIL - ANVIL Postic, Guillaume, Yassine Ghouzam, Vincent Guiraud, and Jean-Christophe Gelly. 2016. “Membrane Positioning for High- and Low-Resolution Protein Structures through a Binary Classification Approach.” Protein Engineering, Design & Selection: PEDS 29 (3): 87–91.</p> <p>[7] PPM3 - PPM3 Lomize, Mikhail A., Irina D. Pogozheva, Hyeon Joo, Henry I. Mosberg, and Andrei L. Lomize. 2012. “OPM Database and PPM Web Server: Resources for Positioning of Proteins in Membranes.” Nucleic Acids Research 40 (Database issue): D370–76.</p> <p>——————————————————————————————————————————————————————————————————————————</p> <p><strong>License:</strong></p> <p>This work is licensed under a Creative Commons Attribution 4.0 International License (CC-BY 4.0).</p> <p> </p>
Structures extracted from the AlphaFold database for the set of sequences used in Rosenberg et al. Nat. Comm. 2022.
<p>Structures extracted from the AlphaFold database for the set of sequences used in Rosenberg, A.A., Marx, A. & Bronstein, A.M. Codon-specific Ramachandran plots show amino acid backbone conformation depends on identity of the translated codon. Nat Commum 13, 2815 (2022) <a href="https://doi.org/10.1038/s41467-022-30390-9">https://doi.org/10.1038/s41467-022-30390-9</a>.</p>
AlphaFold 3 models for all subcomplexes embedded in the Chaetomium thermophilum pyruvate dehydrogenase complex
<p>This datast includes all input and output files of the Alphafold 3 web server runs regarding the prediction of structures for the embedded proteins and their interactions for the eukaryotic pyruvate dehydrogenase complex.</p>
Alphafold2 and AlphaFold-Multimer Predicted Interactions of Soybean Proteins with Macrophomina phaseolina Effectors reveals putative protease inhibitors and SUSS effectors.
<p> </p> <ul> <li> <p><strong>Kunitz Monomer Prediction</strong>:</p> <ul> <li><strong>Data</strong>: Analysis of soybean Kunitz proteins.</li> <li><strong>Details</strong>: Detected on the apoplast at 3 days post-infection with <em>Macrophomina phaseolina</em>.</li> <li><strong>File</strong>: <code>KUNITZ_monomers_outputdir.zip</code></li> </ul> </li> <li> <p><strong>Uncharacterized M. phaseolina Protein Monomer Prediction</strong>:</p> <ul> <li><strong>Data</strong>: Predictions for uncharacterized proteins.</li> <li><strong>Details</strong>: Detected on the apoplast at 3 days post-infection.</li> <li><strong>File</strong>: <code>uncharacterised_proteins_SUSS_effectoroutputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Positive Validation Set</strong>:</p> <ul> <li><strong>Data</strong>: Experimental verification of protein-inhibitor pairs.</li> <li><strong>Details</strong>: Pairs include experimentally verified interactions, specifically proteins and inhibitors, but lack resolved crystal structures.</li> <li><strong>File</strong>: <code>existing_non_existinpairs_Validation_outputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Soybean Serine Protease-Kunitz Interaction</strong>:</p> <ul> <li><strong>Data</strong>: Interactions between soybean serine proteases and Kunitz proteins.</li> <li><strong>Details</strong>: Analyzed in the apoplastic space at 3 days post-infection.</li> <li><strong>File</strong>: <code>glycine max_Serine protease_Vs_Gmaxkunitz_outputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Cysteine Protease without Pro-domain-MoErs-like effector Interaction</strong>:</p> <ul> <li><strong>Data</strong>: Interactions involving cysteine proteases.</li> <li><strong>Details</strong>: Rice RD21 and soybean cysteine proteases with pro-domains removed interacting with <em>MoErs1</em> and <em>MoErs1</em>-like M.phaseolina effectors.</li> <li><strong>File</strong>: <code>AF2-Multimer_RD21&GmaxCproteases_MoERS1_screening_outputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Fungal Serine Protease-Kunitz Interaction</strong>:</p> <ul> <li><strong>Data</strong>: Interactions between <em>Macrophomina phaseolina</em> serine proteases and soybean Kunitz proteins.</li> <li><strong>Details</strong>: Evaluated in the apoplastic space at 3 days post-infection.</li> <li><strong>File</strong>: <code>fungalSerineprotease_Vs_Gmax_kunitzoutputdir.zip</code></li> </ul> </li> <li> <p><strong>Negative Validation Set</strong>:</p> <ul> <li><strong>Data</strong>: Known non-interacting pairs.</li> <li><strong>Details</strong>: Non-interacting pairs of serine proteases-chitinases that are not resolved as crystal structures</li> <li><strong>File</strong>: <code>Gmax_Serineprotease_Vs_Gmaxchitinases_Validation_outputdir.zip</code></li> </ul> </li> </ul>
Accurate Prediction of Enzyme Thermostabilization with Rosetta using AlphaFold Ensembles
<p>DT<sub>M</sub> vs DG<sub>f,mut</sub> values for scoring LovD, LipA, <em>p</em>-nitrobenzyl esterase, xylanase A and tryptophan 6-halogenase variants (<em>DTM_vs_DDGf_mut.xlsx</em>).</p> <p>AlphaFold predicted structures in PDB and Pymol sessions formats for top scoring LovD, LovD6, LovD9, LipA WT, LipA 6B, <em>p</em>-nitrobenzyl esterase WT, xylanase A WT and tryptophan 6-halogenase WT decoys (<em>mAF-min_ensembles.zip</em>).</p> <p>Rosetta energies for all calculations (<em>Rosetta_scores.zip</em>).</p>
Metalloprotein AlphaFold set with enzyme/non-enzyme labeled sites
<pre>The AlphaFold set contains computationally generated structures for metalloproteins that were used to test MAHOMES II's enzyme/non-enzyme predictive performance (Feehan et al. 2023). README.md - Detailed description of AlphaFold set generation. AF-...-model_v2.pdb - Files with the 3D atomic coordinates of a metalloprotein. MAHOMES-II_AlphaFold_set_site_data.csv - Contains the data used during the generation of the AlphaFold set for the final sites. Columns are - Entry: The UniProt accession number of the protein with the bound metal site. - struc_id: The structures AlphaFold DB name (Febuary 2022) and the name of the file in this directory with added metal site. - metal_resName: The two letter PDB residue abbreviation for the site's metal - metal_seqID: The residue index number for the added metal ion. - Enzyme: The enzyme (True) or non-enzyme (False) label. - Entry name: UniProt entry name. - Protein names: The UniProt provided metalloprotein name(s). - Number of homologs with solved structures (PDB): Number of protein sequences in the PDB (May 21, 2020) with an E-value < 1. - Number of homologs in MAHOMES II dataset and T-metal-sites10: Number of protein sequences used to train and evaluate MAHOMES II with an E-value < 1 (0 for all entries). - Metal binding note: UniProt metal binding note that includes information covering the metal’s identity and catalytic flag. - Metal coordinating residue seqIDs: The sequence indices for the metal coordinating residues included in the UniProt’s metal binding section.</pre>
AlphaFold-Multimer (v3) model of a complex of Protein Phosphatase 2A catalytic subunit (PP2A/C), PP2A scaffold subunit (PP2A/A), the B55alpha substrate binding subunit, and FAM122A, a PP2A inhibitor protein.
<p>AlphaFold-Multimer (v3) model structure of a complex of Protein Phosphatase 2A catalytic subunit (PP2A/C), PP2A scaffold subunit (PP2A/A), the B55alpha substrate binding subunit, and FAM122A, a PP2A inhibitor protein. The protein sequences were obtained from UniProt:</p> <p>P67775 · PP2AA_HUMAN</p> <p>P30153 · 2AAA_HUMAN</p> <p>P63151 · 2ABA_HUMAN</p> <p>Q96E09 · PBIR1_HUMAN</p> <p>Coordinates are in mmCIF format. A PyMOL session file is included.</p> <p>Modeling was performed with AlphaFold-Multimer v3, downloaded from DeepMind's github. Structure prediction was performed without templates. The model was relaxed with Amber.</p> <p>FAM122A binds using a short linear motif (SLIM) (residues 84-89) first identified in RBL1 (p107) (Fowle et al., eLife, <a href="https://doi.org/10.7554/eLife.63181">https://doi.org/10.7554/eLife.63181</a> in the form of a short alpha helix (residues 84-92). This is followed by a long alpha helix (residues 96-122) which blocks access to the active site of the catalytic subunit. There are further contacts of FAM122A (residues 150-170) with one of the beta sheets of the B55alpha beta propeller domain. FAM122A regulates PP2A activity during the cell cycle.</p> <p> </p>
UltraScan Solution Modeler (US-SOMO) hydrodynamic parameter, structural small angle scattering and SESCA circular dichroism (CD) calculations on AlphaFold predicted structures
Open the record for dataset details and reuse information.
AlphaFold models and supporting data for the annotation of Vairimorpha necatrix
<p>V_necatrix_alphafold.zip - Contains AlphaFold models and associated files for all V. necatrix proteins.</p><p>chimerax_annotater_plugin.zip - Contains the ChimeraX plugin we developed and used to annotate the V. necatrix proteome.</p><p>v_necatrix_annotation_data.zip - Contains all data used in the ChimeraX plugin to annotate the V. necatrix proteome.</p><p>V_necatrix_proteome.fasta - Protein fasta file containing all protein sequences</p><p>V_necatrix_haplotype[1-4].fasta - Nucleotide fasta file for each haplotype</p>
AlphaFold structures with AlphaMissense scores
<p>These repository provides:</p> <ol> <li>NEW: AFwAM-pdb-qb.tar file including pdb.gz files for human protein structures from the AlphaFoldDb with occupancy column set to residue-wise mean of all a.a. variations and temperature factor column set to residue-wise mean of single nucleotde variations; a PyMOL plugin file (coloram-qb.py) for coloring these structures (<code>coloram column=b</code> or <code>coloram column=q</code>; b is the default)</li> <li>AFwAM-pdb.tar file including pdb.gz files for human protein structures from the <a href="https://alphafold.ebi.ac.uk/">AlphaFoldDb</a> with occupancy and temperature factor columns set to residue-wise mean of <a href="../records/8208688">AlphaMissense</a> scores;</li> <li>A PyMOL plugin file (coloram.py) for coloring these structures;</li> <li>For data, Python scripts, and notebooks, please refer to the pub.zip file; detailed instructions are provided in the README.md within this archive and further explained in our manuscript.</li> </ol> <p><br>For alternative data access, please visit <a href="https://alphamissense.hegelab.org/">https://alphamissense.hegelab.org</a>.</p> <p> </p> <p>Disclaimer: The AlphaMissense Database and other information provided on or linked to this site is for theoretical modelling only, caution should be exercised in use. It is provided "as-is" without any warranty of any kind, whether express or implied. For clarity, no warranty is given that use of the information shall not infringe the rights of any third party (and this disclaimer takes precedence over any contrary provisions in the Google Cloud Platform Terms of Service). The information provided is not intended to be a substitute for professional medical advice, diagnosis, or treatment, and does not constitute medical or other professional advice.</p> <p>Data contained within the AlphaMissense Database is provided for non-commercial research use only under CC BY-NC-SA 4.0 license.</p> <p>DeepMind - AlphaMissense: <a href="https://doi.org/10.1126/science.adg7492">https://doi.org/10.1126/science.adg7492</a></p>
AlphaFold Structures for: A wheat tandem kinase activates an NLR to trigger immunity
<p>AlphaFold Predictions used for analysis in "A wheat tandem kinase activates an NLR to trigger immunity".</p>
Dataset for AlphaDesign: A de novo protein design framework based on AlphaFold
<p>This dataset consists of output data from the work reported in: </p> <p>Jendrusch, M., Korbel, J. O., & Sadiq, S. K. (2021). AlphaDesign: A de novo protein design framework based on AlphaFold. bioRxiv.</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.