Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
44
datasets available to search
ShareScore release 0.9.0
Dataset results
44 results for “AlphaFold2”
Overcoming Limitation of AlphaFold2 by Deep-mutational Scanning and Stability-Selection of Protein Sequences
<p>This repository contains the processed datasets and corresponding code used in our study. While AlphaFold2 revolutionizes protein structure prediction, its accuracy critically depends on evolutionary information from natural homologs—limiting applications for proteins with sparse sequence families. Here, we bypass this bottleneck by employing deep mutational scanning and stability-guided selection to generate artificial homologs. Fed into AlphaFold2, these synthetic sequences match the accuracy achieved on well-predicted proteins with rich natural homology, while providing highly accurate predictions for difficult targets—including orphan proteins previously deemed "unpredictable." Our approach achieves high accuracy (<3 Å RMSD for 5/8 and <2 Å RMSD for 7/8 targets after excluding intrinsically flexible regions). Thus, integrating simple, scalable molecular biology (mutagenesis/selection) with high-throughput sequencing can deliver the accuracy similar to but at a fraction of the cost and time of traditional experimental structure-determination methods. This hybrid framework could democratize high-resolution structural biology, opening avenues to determine structures of protein complexes, modified proteins, and condition-dependent conformations. </p>
AlphaFold2 models of human DNA polymerase epsilon catalytic subunit A
<p>Models of the structure of human DNA polymerase epsilon catalytic subunit A built with the program AlphaFold2. Files are in mmCIF format. Models include:</p> <p>1) <a href="https://zenodo.org/api/files/c8652c0b-1bc6-44d5-b11e-e8e107dba67e/DPOE1_HUMAN_AF2_NterminalLobe_4m8o.cif">DPOE1_HUMAN_AF2_NterminalLobe_4m8o.cif </a>: AlphaFold2 model built with template PDB:4M8O (yeast DNA polymerase epsilon N-terminal lobe with bound DNA)</p> <p>2) <a href="https://zenodo.org/api/files/c8652c0b-1bc6-44d5-b11e-e8e107dba67e/DPOE1_HUMAN_AF2_FullLength_6wjv.cif">DPOE1_HUMAN_AF2_FullLength_6wjv.cif </a>: AlphaFold2 model built with template PDB:6WJV (yeast DNA polymerase epsilon full-length protein, N and C terminal lobes, without DNA).</p> <p>3) <a href="https://zenodo.org/api/files/c8652c0b-1bc6-44d5-b11e-e8e107dba67e/DPOE1_HUMAN_AF2_NterminalLobe_4m8o_withDNA.cif">DPOE1_HUMAN_AF2_NterminalLobe_4m8o_withDNA.cif</a> : AlphaFold2 model built with template 4M8O (File #1 above) with DNA added from PDB entry 4M8O and subjected to energy minimization with the program AMBER using force field ff14SB.</p> <p>The models were used to estimate the change in free energy of mutations found in patients with ovarian cancer, colon cancer, and endometrial cancer. Paper to be submitted Dec 2022.</p>
Alphafold2_ab_initio iterative predictions for folding intermediate identification
<p>PDB ids starts from 1 and 8, rmsds, plddts, t-sne embeddings.</p> <p>Check related biorxiv preprint: AlphaFold2 knows some protein folding principles; DOI: https://doi.org/10.1101/2024.08.25.609581.</p>
Alphafold2 models of 47G4 (Hu19) scFv dimers and of VLVL dimers
<p><strong>Alphafold2 models of 47G4 (Hu19) scFv dimers show a variety of swapped VH-VL structures.</strong></p> <p>Alphafold2 dimers models of scFvs formed from VH and VL domains tend to swap to form diabodies with long linkers as for short linkers. The 47G4 scFv used in this study uses a VL-to-VH scFv construct with a long linker (see sequence in Figure 1)</p> <p>See Supplement files for:</p> <ul> <li>The 5 best Alphafold2 models of a predicted 47G4 (Hu19) scFv dimer, using the sequence in Figure 1. <ul> <li>In PDB format. The 5 models are ranked by AF2 in files: <ul> <li><strong>ranked_0...4.pdb </strong></li> </ul> </li> <li>The 5 models can be compared by superimposition of the VL domain of chain A: in iCn3D using an iCn3D PNG image file. To see the models and analyze them in 3D, invoke iCn3D with the link: <a href="https://www.ncbi.nlm.nih.gov/Structure/icn3d/full.html">https://www.ncbi.nlm.nih.gov/Structure/icn3d/full.html</a> ; open “iCn3D PNG image file” using the<strong> file:</strong> <ul> <li><strong>5models_Hu19_scFv_AF2_icn3d_loadable.png</strong></li> </ul> </li> </ul> </li> </ul> <p><strong>Alphafold2 models for 47G4 (Hu19) VL-VL dimers</strong></p> <ul> <li>In PDB format. The 25 models are ranked by AF2 in files: <ul> <li><strong>Hu19-VLVL-AF2-ranked_0…24.pdb </strong></li> <li>Only one out of 25 models offers an inverted interface but offers lower confidence. All other present the canonical interface.</li> </ul> </li> </ul> <p><strong>List of PDB structures containing VL-VL quaternary interfaces (PDB accessed on 10/11/2021)</strong></p> <ul> <li>See file: <strong>PDB-</strong><strong>Igs-VLVL-asu+bu-with-over-10contacts.xls</strong> <ul> <li>The file represents PDB structures that contain contacting VL domains with more than 10 arbitrary contacts in either the first biological Unit (BU) or asymmetric unit (ASU). Both BU and ASU are used as both may contain valid dimers, as in the the inverted VLVL dimer structure (7JO8)</li> <li>Some <strong>iCn3D links</strong> are provided in the excel file to look and analyze a number of structures. Some links compare 7JO8 to 1REI, or other VLVL parallel vs antiparallel dimer such as 1LVE vs 5LVE.</li> </ul> </li> </ul> <p><strong>VLVL based diabody models: canonical and inverted</strong></p> <ul> <li>A model predicted by Alphafold2 version 2.0 for the 47G4 (Hu19) scFv dimer, using the sequence in Figure 1. Supplement file in PDB format: <ul> <li><strong>Hu19-VLVL-diabody-model.pdb</strong></li> </ul> </li> <li>Model building of a scFv dimer using the observed VLVL dimer (7JO8) <ul> <li><strong>Hu19-inverted VLVL-diabody-model.pdb</strong></li> </ul> </li> <li>Both models compared:<strong> VLVLDIAB_VLINVDIAB_icn3d_loadable.png </strong></li> </ul> <p><em>Where no link is specified, PDBids can be directly loaded in iCn3D using the link</em> <a href="https://www.ncbi.nlm.nih.gov/Structure/icn3d/full.html">https://www.ncbi.nlm.nih.gov/Structure/icn3d/full.html</a></p> <p> </p>
AlphaFold structures reported in "AlphaFold2 Can Predict Single-Mutation Effects"
<p>This contains AlphaFold predictions for X proteins that are found in the Protein Data Bank (PDB), that were used to evalluate AlphaFold's predictions of mutation effects. This includes one set of structures predicted by AlphaFold2.0, using default settings, and one structure for each of 5 models. This also includes structures predicted by the ColabFold version of AlphaFold (6 recycles, 5 models, no template, amber minimization, 4 repeats).</p><p>There are also additional predicted structures that are found in the PDB that were not analyzed in the paper.</p><p>There are AlphaFold predictions for three proteins (BFP / RFP, GFP, and PafA), covering either all (BFP/RFP, PafA) or a subset (GFP) of the sequences in three datasets of phenotype measurements from high-throughput experiments.</p><p>Results are separated into tar files based on whether DeepMind (AF2.0) or ColabFold implementation was used.</p><p>Folders under "ColabFold/PDB" are labelled according to a sequence ID, since multiple PDB structures can exist for a single sequence. These sequence IDs can be mapped back to PDB IDs using the information in "seq_id_pdb_id.json".</p><p>All PDB files have been compressed using Foldcomp (<a href="https://github.com/steineggerlab/foldcomp">https://github.com/steineggerlab/foldcomp</a>). Foldcomp is required to decompress the ".fcz" files in order to recover the ".pdb" files.</p>
Supplementary data frames, AlphaFold models, Normal Mode Analysis (NMA) Data, and NMA of Corresponding NMR Ensembles in the S2RCI, MD, and S2 Datasets for "Gradations in protein dynamics captured by experimental NMR are not well represented by AlphaFold2 models and other computational metrics"
<h1><strong>Changes applied to V2</strong></h1> <p>In addition to the supplementary dataframes and AlphaFold models from each dataset in V1, V2 includes the additional data outlined below.</p> <p>The <strong>S2RCI</strong> and <strong>MD</strong> datasets include comprehensive analyses of AlphaFold2 models (both before and after truncation). These datasets feature: </p> <ul> <li><strong>AlphaFold2 Models</strong>: Both original and truncated structures. </li> <li><strong>WEBnma Modes</strong>: `modes.txt` files generated from WEBnma analysis, available for both non-truncated and truncated AF2 models. </li> <li><strong>Root-Mean-Square-Fluctuations (RMSF)</strong>: Profiles calculated before and after truncation of AF2 models. </li> <li><strong>NMR Data: Normal Mode Analysis (NMA)</strong>: Performed on corresponding NMR ensembles (see below). </li> </ul> <p> </p> <p>The <strong>NMR Data</strong> of NMA in these datasets includes: </p> <ul> <li>NMR ensembles </li> <li>Individual NMR models extracted from each ensemble </li> <li>STRIDE secondary structure calculations per-individual NMR models</li> <li>RMSF profiles per-individual NMR models</li> </ul> <p>For detailed information, please refer to the `Readme.txt` file within each corresponding folder. </p> <p>The <strong>S2 dataset</strong> includes all the features listed above, except for the NMR analysis.</p>
Data - AlphaFold2 Predicts Alternative Conformation Populations in Green Fluorescent Protein Variants
<p><strong>MSAs.zip </strong>Multiple sequences alignments generated by AlphaFold2 structure prediction of 7 engineered GFPs.</p> <p><strong>AF2_models_column_masking.zip </strong>AlphaFold2 models of the alternative conformations of 7 engineered GFPs.</p> <p><strong>MD_trajectories_PyMOL.zip</strong> Molecular dynamics trajectories (PyMOL sessions) of the alternative conformations of 7 engineered GFPs.</p> <p><strong>MD_analysis.zip </strong>Root mean square deviation and per-residue root mean square fluctuations along molecular dynamics simulations of 7 engineered GFPs.</p> <p><strong>rmsd_values.zip</strong> Root mean square deviation relative to crystallographic GFP structure for AlphaFold2 models and molecular dynamics frames (global and central alpha-helix)</p>
DPCstruct Classification of AlphaFold2-Predicted Protein Structures
<p>This dataset contains DPCstruct domain classifications for protein structures predicted by AlphaFold2, as presented in the paper "Unsupervised Domain Classification of AlphaFold2-Predicted Protein Structures."</p> <p>DPCstruct was applied to a non-redundant set of the AlphaFold Database v4.0, known as Foldseek Clusters, which includes approximately 15 million representative proteins, as described in the work by <a href="https://doi.org/10.1038/s41586-023-06510-w">Barrio-Hernandez et al.</a></p> <p>This repository provides the results of our classification, along with all the data related to the analyses presented in our study. DPCstruct algorithm can be found at <a href="https://github.com/RitAreaSciencePark/DPCstruct">https://github.com/RitAreaSciencePark/DPCstruct</a> together with examples on how to use it.</p> <p><strong>FILES DESCRIPTION:</strong></p> <ul> <li><strong>dpcstruct_classification.tsv: </strong>List of domains identified by DPCstruct and their corresponding metacluster. Columns: Metacluster ID, Protein Uniprot ID, domain start, domain end.</li> <li><strong>mcs_reps.fasta:</strong> For each metacluster, two representative domains were selected: one representing the center of the cluster and the other being the domain with the highest pLDDT score. If these are the same, only one domain is included as the representative. This file contains the list of representative domains and their sequences in FASTA format.</li> <li><strong><span>mcs_reps_pdbs.zip: </span></strong>Contains a PDB file for each representative domain. The filename is structured as 'proteinID_metacluster.pdb'.</li> <li><strong>mcs_properties.tsv:</strong> Set of properties per metacluster, including: <ul> <li><strong>mcID:</strong> Metacluster ID.</li> <li><strong>size:</strong> Number of domains.</li> <li><strong>len_aa:</strong> Average length of domains (number of amino acids).</li> <li><strong>len_std:</strong> Standard deviation of domain lengths.</li> <li><strong>len_ratio:</strong> Ratio of len_std to len_aa.</li> <li><strong>plddt:</strong> Average predicted LDDT as reported by AlphaFold2.</li> <li><strong>disorder:</strong> Average intrinsic disorder score calculated with AIUPred.</li> <li><strong>alntmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs.</li> <li><strong>tmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs, using the maximum between TM-score normalized by query or target.</li> <li><strong>lddt:</strong> Pairwise LDDT score, averaged over all pairs.</li> <li><strong>prob:</strong> Pairwise probability of homology according to SCOPe, as reported by Foldseek.</li> <li><strong>pident:</strong> Pairwise percentage identity, averaged over all pairs.</li> </ul> </li> <li><span><strong>annotated_[cath|scop]_qc[x]_t[x]_l[x].tsv:</strong> </span>For each fold in [CATH|SCOP], we provide the best matching DPCstruct domain, if available, along with the structural alignment information as reported by Foldseek. A fold is considered annotated if its alignment values meet or exceed the following thresholds: <ul> <li>qc: query coverage.</li> <li>t: template modelling score of the alignment.</li> <li>l: lddt score of the alignment.</li> </ul> </li> <li><strong>dpcstruct_consistency.tsv:</strong> Consistency of DPCstruct metaclusters with respect to Pfam 36.0 labels. Note that we consider a Pfam label to overlap with a DPCstruct domain even if it shares just one amino acid, which is why some metaclusters have many labels. In such cases, we only display 5 representative labels.</li> <li><strong>pfam_consistency.tsv:</strong> Consistency of Pfam Clans with respecto to DPCstruct labels.</li> </ul> <p><strong>Note:</strong> All 'tsv' files contain a header as the first row.</p> <p>If there is any doubt regarding the data or there is something missing please contact us: </p> <p>federico.barone@areasciencepark.it</p>
Alphafold2 and AlphaFold-Multimer Predicted Interactions of Soybean Proteins with Macrophomina phaseolina Effectors reveals putative protease inhibitors and SUSS effectors.
<p> </p> <ul> <li> <p><strong>Kunitz Monomer Prediction</strong>:</p> <ul> <li><strong>Data</strong>: Analysis of soybean Kunitz proteins.</li> <li><strong>Details</strong>: Detected on the apoplast at 3 days post-infection with <em>Macrophomina phaseolina</em>.</li> <li><strong>File</strong>: <code>KUNITZ_monomers_outputdir.zip</code></li> </ul> </li> <li> <p><strong>Uncharacterized M. phaseolina Protein Monomer Prediction</strong>:</p> <ul> <li><strong>Data</strong>: Predictions for uncharacterized proteins.</li> <li><strong>Details</strong>: Detected on the apoplast at 3 days post-infection.</li> <li><strong>File</strong>: <code>uncharacterised_proteins_SUSS_effectoroutputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Positive Validation Set</strong>:</p> <ul> <li><strong>Data</strong>: Experimental verification of protein-inhibitor pairs.</li> <li><strong>Details</strong>: Pairs include experimentally verified interactions, specifically proteins and inhibitors, but lack resolved crystal structures.</li> <li><strong>File</strong>: <code>existing_non_existinpairs_Validation_outputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Soybean Serine Protease-Kunitz Interaction</strong>:</p> <ul> <li><strong>Data</strong>: Interactions between soybean serine proteases and Kunitz proteins.</li> <li><strong>Details</strong>: Analyzed in the apoplastic space at 3 days post-infection.</li> <li><strong>File</strong>: <code>glycine max_Serine protease_Vs_Gmaxkunitz_outputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Cysteine Protease without Pro-domain-MoErs-like effector Interaction</strong>:</p> <ul> <li><strong>Data</strong>: Interactions involving cysteine proteases.</li> <li><strong>Details</strong>: Rice RD21 and soybean cysteine proteases with pro-domains removed interacting with <em>MoErs1</em> and <em>MoErs1</em>-like M.phaseolina effectors.</li> <li><strong>File</strong>: <code>AF2-Multimer_RD21&GmaxCproteases_MoERS1_screening_outputdir.zip</code></li> </ul> </li> </ul> <ul> <li> <p><strong>Fungal Serine Protease-Kunitz Interaction</strong>:</p> <ul> <li><strong>Data</strong>: Interactions between <em>Macrophomina phaseolina</em> serine proteases and soybean Kunitz proteins.</li> <li><strong>Details</strong>: Evaluated in the apoplastic space at 3 days post-infection.</li> <li><strong>File</strong>: <code>fungalSerineprotease_Vs_Gmax_kunitzoutputdir.zip</code></li> </ul> </li> <li> <p><strong>Negative Validation Set</strong>:</p> <ul> <li><strong>Data</strong>: Known non-interacting pairs.</li> <li><strong>Details</strong>: Non-interacting pairs of serine proteases-chitinases that are not resolved as crystal structures</li> <li><strong>File</strong>: <code>Gmax_Serineprotease_Vs_Gmaxchitinases_Validation_outputdir.zip</code></li> </ul> </li> </ul>
AlphaFold2-Based Characterization of Apo and Holo Protein Structures and Conformational Ensembles Using Randomized Alanine Sequence Scanning Adaptation: Capturing Shared Signature Dynamics and Ligand-Induced Conformational Changes
<p>Proteins often exist in multiple conformational states, influenced by the binding of ligands or substrates. The study of these states, particularly the apo (unbound) and holo (ligand-bound) forms, is crucial for understanding protein function, dynamics, and interactions. In the current study, we use AlphaFold2 that combines<span> randomized</span> <span><span> </span>alanine<span> </span>sequence masking<span> </span>with shallow multiple sequence alignment<span> </span>subsampling to expand the conformational diversity of the predicted structural<span> </span>ensembles and<span> </span>capture conformational changes between apo and holo protein forms. Using several well-established datasets of<span> </span>structurally diverse apo-holo protein pairs, the proposed approach </span><span>enables<span> </span>robust predictions of apo and holo structures and conformational ensembles, while also displaying notably similar dynamics distributions. These observations are consistent with<span> </span>the view </span><span> </span>that the intrinsic dynamics of allosteric proteins is defined by the structural topology of the fold and favors conserved conformational motions driven by soft modes among orthologs. We also found<span> </span>a significant <span>correlation </span>between conformational flexibility and <span> </span>AlphaFold2 metric of statistical significance pLDDT for the apo-holo pairs in which ligand binding induced local moderate conformational changes. For apo-holo pairs exhibiting larger structural changes, this relationship<span> </span>becomes nonlinear, reflecting inability of AlphaFold2 confidence metrics to identify high energy functional conformations. Our findings support the notion that AlphaFold2 approaches can yield reasonable accuracy in predicting minor conformational adjustments between apo and holo states, especially for proteins with <span> </span>moderate localized changes upon ligand binding. However, for large, hinge-like domain movements, AF2 tends to predict the most stable domain orientation which is typically the apo form rather than the full range of functional conformations characteristic of the holo ensemble. These results indicate that modeling of multiple functional states of proteins may require more accurate detection of flexible region conformations and cannot solely rely on the pLDDT metric as the major determinant of the prediction accuracy in reproducing functional conformational ensembles.<span> </span></p>
CATH Structural domains in AlphaFold2 models for 21 model organisms
<p>CATH structural domain assignments for AlphaFold2 models in 21 model organisms.</p> <p>The table cath-v4_3_0.alphafold-v2.2022-11-22.tsv contains the domain assignments with information on model quality, CATH superfamily and Class, organism, average pLDDT, percentage of residues not in secondary structure, globularity and domain origin (CATH-PDB, CATH-HMM,Pfam,newfams).</p> <p>Organisms included:</p> <p>Arabidopsis thaliana</p> <p>Caenorhabditis elegans</p> <p>Candida albicans</p> <p>Danio rerio</p> <p>Dictyostelium discoideum</p> <p>Drosophila melanogaster</p> <p>Escherichia coli</p> <p>Glycine max</p> <p>Homo sapiens</p> <p>Leishmania infantum</p> <p>Methanocaldococcus jannaschii</p> <p>Mus musculus</p> <p>Mycobacterium tuberculosis</p> <p>Oryza sativa</p> <p>Plasmodium falciparum</p> <p>Rattus norvegicus</p> <p>Saccharomyces cerevisiae</p> <p>Schizosaccharomyces pombe</p> <p>Staphylococcus aureus</p> <p>Trypanosoma cruzi</p> <p>Zea mays</p> <p> </p> <p>The archive</p> <pre>cath-v4_3_0.alphafold-v2.2022-11-22.by_superfamily.tgz</pre> <p>contains all domains assigned by CATH in the dataset as PDB files, divided by superfamily.</p> <p>Alternatively, if you're interested in a particular organism, an individual tarball containing all CATH domains as PDB files is available</p> <p>e.g. cath-v4_3_0.alphafold-v2.2022-11-22.arabidopsis_thaliana.tgz</p> <p>All domains included in this release are named as af_[UniProt_ID]_[start]_[stop].</p> <p> </p>
Exploring Host-Binding Machineries of Mycobacteriophages with AlphaFold2
<p>Coordinates of AlphaFold2 predicted structures presented in the publication "Exploring Host-Binding Machineries of Mycobacteriophages with AlphaFold2".</p>
Improving AlphaFold2-based Protein Tertiary Structure Prediction with MULTICOM in CASP15
<p>Improving AlphaFold2-based Protein Tertiary Structure Prediction with MULTICOM in CASP15</p>
Partial Atomic Model of the Tailed Lactococcal Phage TP901-1 as Predicted by AlphaFold2
Open the record for dataset details and reuse information.
All-atom molecular dynamics simulations of incomplete ATP synthase rotor rings with unusually high stoichiometry predicted by the AlphaFold2-based method
<p>The trajectories of all-atom MD simulations of <span>AlphaFold2 4, 11, 16 or 18-mer structures of the subunit <em>c</em> from<br></span><span><em>Candidatus Kryptonium thompsoni</em></span><span> (CKt_Nmer_lipid_mix_CHM36m_303K_500ns) and <br></span><span><em>Thalassoglobus polymorphus </em>(Tp_Nmer_lipid_mix_CHM36m_303K_500ns), and <br>AlphaFold2 11-mer structure of the subunit <em>c</em> from <em>Spinacia oleracea</em> (So_11mer-c20_POPC_CHM36m_303K_300ns) </span><span>in a lipid bilayer.</span></p> <p><span>Simulations have been performed using the CHARMM36m force field, running with the GROMACS 2022 package.</span></p>
Structural modeling of ion channels using AlphaFold2, RoseTTAFold2, and ESMFold
<p>Ion channels play key roles in human physiology and are important targets in drug discovery. The atomic-scale structures of ion channels provide invaluable insights into a fundamental understanding of the molecular mechanisms of channel gating and modulation. Recent breakthroughs in deep learning-based computational methods, such as AlphaFold, RoseTTAFold, and ESMFold have transformed research in protein structure prediction and design. We review the application of AlphaFold, RoseTTAFold, and ESMFold to structural modeling of ion channels using representative voltage-gated ion channels, including human voltage-gated sodium (Na<sub>V</sub>) channel - Na<sub>V</sub>1.8, human voltage-gated calcium (Ca<sub>V</sub>) channel – Ca<sub>V</sub>1.1, and human voltage-gated potassium (K<sub>V</sub>) channel – K<sub>V</sub>1.3. We compared AlphaFold, RoseTTAFold, and ESMFold structural models of Na<sub>V</sub>1.8, Ca<sub>V</sub>1.1, and K<sub>V</sub>1.3 with corresponding cryo-EM structures to assess details of their similarities and differences. Our findings shed light on the strengths and limitations of the current state-of-the-art deep learning-based computational methods for modeling ion channel structures, offering valuable insights to guide their future applications for ion channel research.</p>
Photoreceptor-induced LHL4 protects photosystem II monomer - XL-MS and Alphafold2 intergrative modelling
<p>This dataset refers to the following preprint: https://www.biorxiv.org/content/10.1101/2024.02.23.581703v1 </p> <p>It provides:</p> <ul> <li>readable XL_MS annotate fragmentattion spectra for the crosslinks detected involved in the LHL4 interaction with CP43 and CP47. These complements the raw files submited in PRIDE (accession PXD049352) and the Supplementary Table 1 annexed to teh manuscript supplementray materials</li> <li>The complete output files for the AlphaFold2 multimer prediciotn of the pairwise interactions between LHL4 and CP43/CP47</li> </ul> <p> </p>
Modeling Protein Conformations by Guiding AlphaFold2 with Distance Distributions. Application to Double Electron Electron Resonance (DEER) Spectroscopy.
<p>We describe a modified version of AlphaFold2 that incorporates experiential distance distributions into the network architecture for protein structure prediction. Harnessing the OpenFold platform, we fine-tuned AlphaFold2 on a small number of structurally dissimilar proteins to explicitly model distance distributions between spin labels determined from Double Electron-Electron Resonance (DEER) spectroscopy. We demonstrate the performance of the modified AlphaFold2, referred to as DEERFold, in switching the predicted conformations guided by experimental or simulated distance distributions. Remarkably, the intrinsic performance of AlphaFold2 substantially reduces the number and the accuracy of the widths of the distributions needed to drive conformational selection thereby increasing the experimental throughput. The blueprint of DEERFold can be generalized to other experimental methods where distance constraints can be represented by distributions. </p>
Ensemble Refinement of Stable-5-LOX and AlphaFold2 predictions including variants
<p>Ensemble Refinement of Stable-5-LOX in "closed" (PDB code: 7TTK) and "open" (PDB code: 7TTJ) conformations. AlphaFold2 models of Stable-5-LOX and variants from the research article "Helical remodeling augments 5-lipoxygenase activity."</p>
Results for blind docking Tocriscreen 2.0 compounds against AlphaFold2 structural modules of human and parasite proteins
<p>Docking (GNINA 1.0) of 1280 Tocriscreen 2.0 compounds against >4000 AlphaFold structures from humans and the parasitic nematode <em>Brugia malayi</em>. </p> <p>Results are provided as Pickle objects or GZipped CSVs.</p> <p>Tuned machine learning models (random forest and XGBoost) are also included for classifying actives/decoys.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.