Skip to main content
zenodoopen

Incorporating prior knowledge in the seeds of adaptive sampling molecular dynamics simulations of ligand transport in enzymes with buried active sites

<p><strong>00_Caver.tar.gz </strong>- Contains CAVER (https://caver.cz/fil/download/manual/caver_userguide.pdf) input and output files used for identification of transport pathways in LinB86(PDB ID: 5LKA).&nbsp;<br>final_clustering<br>├── tunnel_custers # contains caver output files for individual tunnels clusters&nbsp;<br>│ &nbsp; ├── ...<br>├── analysis # contains .csv output files for botttlenecks and tunnels charecteristics of individual tunnels clusters&nbsp;<br>│ &nbsp; ├── ...</p> <p><strong>01_CaverDock_Tunnels_Profile.tar.gz</strong> - Contains tunnel clusters consisting of the top 100 tunnels and the CaverDock analysis files obtained.&nbsp;<br># Each folder (tun_cluster_p1a, tun_cluster_p1b, tun_cluster_p2, tun_cluster_p3) contains input raw files used for CaverDock calculations for individual snapshots of the respective tunnel clusters named as stripped_system*. The details of those files are:<br>- <em>calculations/*/caverdock.conf</em> : &nbsp;The config file input for caverdock calculation.&nbsp;<br>-<em> calculations/*/DBE.pdbqt </em>: Input file for the substrate DBE.<br>- <em>calculations/*/stripped_system*.pdbqt </em>: Input file for the Protein/Receptor<br>- c<em>alculations/*/stripped_system*.dsd </em>: Tunnel discretization file. Notes:&nbsp;<br>- <em>calculations/*/stripped_system*.pdb </em>: PDB file for the tunnel.&nbsp;</p> <p><strong>02_Minimization_and_Equilibration.tar.gz</strong> - &nbsp;Contains input and output files used for AMBER minimization and equilibration of the molecular systems and seed conformations used for adaptive sampling simulations.</p> <p><strong>03_HTMD_Bulk</strong> - separate Zenodo repository, see below for the link. Contains input, output and restart files used for HTMD (High-throughput molecular dynamics) adaptive sampling simulations at 310K for Bulk schemes.&nbsp;<br><strong>04_HTMD_Cavity</strong> - separate Zenodo repository, see below for the link. Contains input, output and restart files used for HTMD (High-throughput molecular dynamics) adaptive sampling simulations at 310K for Cavity schemes.<br><strong>05_HTMD_Cavity_Bulk</strong> -<strong> </strong>separate Zenodo repository, see below for the link. Contains input, output and restart files used for HTMD (High-throughput molecular dynamics) adaptive sampling simulations at 310K for Cavity&amp;Bulk schemes.&nbsp;<br><strong>06_HTMD_Tunnels</strong> - separate Zenodo repository, see below for the link. Contains input, output and restart files used for HTMD (High-throughput molecular dynamics) adaptive sampling simulations at 310K for Tunnels schemes.&nbsp;</p> <p><strong>07_MD-Analysis.tar.gz</strong> - Contains MD analysis files obtained from 45 micro-seconds adaptive sampling simulations at 310K.&nbsp;<br># Each folder contains input raw files used to calculate epochs convergence, distances, RMSD, RMSF, kinetics, percentage, sample proportions and tunnel lengths. The details of those files are:<br>- <em>Epoch_Convergence/epochs_dist_counts.csv</em> : &nbsp;Contains the counts of DBE distances from the active-site (0-5 &Aring;), tunnel (5-19 &Aring;), and bulk (&gt;19 &Aring;) for the 30 epochs.&nbsp;<br>-&nbsp; <em>Distances/*/dist_s_r*.csv</em> : Contains .csv file for the &nbsp;frames wise for all studied schemes. The analysis were performed using the cpptraj program (https://amber-md.github.io/cpptraj/CPPTRAJ.xhtml). The following columns are present:<br>D107_OD1_DBE_C1 &nbsp;&nbsp;<br>D107_OD2_DBE_C1 &nbsp;&nbsp;<br>D107_OD1_DBE_C2 &nbsp;<br>D107_OD2_DBE_C2 &nbsp;&nbsp;<br>N37_ND2_DBE_Br1 &nbsp;<br>N37_ND2_DBE_Br2 &nbsp;<br>W108_NE1_DBE_Br1 &nbsp;<br>W108_NE1_DBE_Br2&nbsp;<br>D107_COM_DBE_COM &nbsp;&nbsp;<br>W108_COM_DBE_COM &nbsp; &nbsp;<br>N37_COM_DBE_COM &nbsp;&nbsp;<br>catal_COM_p1aCOM &nbsp; &nbsp;<br>catal_COM_p1bCOM &nbsp;&nbsp;<br>catal_COM_p2COM &nbsp; &nbsp;<br>catal_COM_p3COM &nbsp; &nbsp;<br>p1aCOM_DBE_COM &nbsp; &nbsp;<br>p1bCOM_DBE_COM &nbsp; &nbsp;<br>p2COM_DBE_COM &nbsp; &nbsp;<br>p3COM_DBE_COM &nbsp;<br>catal_COM_DBE_COM &nbsp; &nbsp;<br>p1aCOM_p1bCOM &nbsp;&nbsp;<br>p1aCOM_p2COM &nbsp;<br>p1aCOM_p3COM &nbsp;<br>p1bCOM_p2COM&nbsp;<br>p1bCOM_p3COM&nbsp;<br>p2COM_p3COM&nbsp;<br>- <em>RMSD_RMSF/*/*.csv</em> : Contains .csv files with RMSD and RMSF from the protein residues. For RMSF 1st row are residue number (1-295) and 2nd row are RMSF. For RMSD, 1st column are frame no. (0.1ns) and 2nd column are RMSD values respectively. The calcualtion were performed using pytraj (https://amber-md.github.io/pytraj/latest/index.html) program. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Example input: pytraj.rmsd(traj, mask='1-295@CA') &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Example input: pytraj.rmsf(traj, mask=':1-295', options='byres')<br>-&nbsp;<em>COM_RMSF/.xlsx</em> : Contains the center of mass (COM) distances calculated using the bottleneck residues for catalytic residues (N38, D108, W109), p1a (D147, F151, and V173), p1b (D147, W177, and L248), p2 (L211 and L248), and p3 (L143, F151, and I213)<br>- <em>Kinetics/kinetics.txt</em> : Contains .csv file with kinetic information from studied scheme: Cavity, Cavity&amp;Bulk and Tunnels for all the three replicates and calculated average kon, koff, koff/kon rates.<br>-&nbsp;<em>Percentages/.csv</em> : Contains csv files for the percentages of DBE localization and distances from the active-site (0-5 &Aring;), tunnel (5-19 &Aring;), and bulk (&gt;19 &Aring;).<br>- <em>Tunnels_lengths/.csv</em> : Contains <em>tunnel_lengths.csv</em>, <em>Summary of tunnel lengths.xlsx</em> files with lengths of top 100 tunnels snapshots for tunnel clusters p1a, p1b, p2 and p3 in <em>tunnel_lengths.csv</em> and summary of respective tunnel clusters generated from Caver output (for more details please check https://www.caver.cz/fil/download/manual/caver_userguide.pdf with keywork "summary.txt") in the <em>Summary of tunnel lengths.xlsx</em> file.&nbsp;<br>- <em>Sample_proportions/.csv</em> : Contains <em>sample_proportions.csv</em> file with proportions or fraction individual metastable states while performing the transition pathway analysis. For more details please check https://software.acellera.com/htmd/htmd.kinetics.html<br>&nbsp;or http://www.emma-project.org/v1.2.1/api/generated/pyemma.msm.flux.pathways.html?highlight=transition%20path<br>&nbsp;- <em>*.py</em> : Python scripts used to build and analysis Markov state models with use case and distances of ligand.<br>- <em>Generated_models/</em> : Contains <em>models_rep[].dat</em> files representating the matric data used to build MSM for respective schemes and replicates. The dirs are arranged as below:<br>├── Cavity<br>│ &nbsp; ├── model_rep1.dat<br>│ &nbsp; ├── model_rep2.dat<br>│ &nbsp; └── model_rep3.dat<br>├── Cavity_Bulk<br>│ &nbsp; ├── model_rep1.dat<br>│ &nbsp; ├── model_rep2.dat<br>│ &nbsp; └── model_rep3.dat<br>└── Tunnels<br>&nbsp; &nbsp; ├── model_rep1.dat<br>&nbsp; &nbsp; ├── model_rep2.dat<br>&nbsp; &nbsp; └── model_rep3.dat&nbsp;</p> <p><br><strong>08_TransportTools.tar.gz</strong> - Contains TransportTools (TT) analysis output, log and summary files for Cavity, Cavity&amp;Bulk and Tunnels schemes. For more details on the workflow of TT, please visit https://github.com/labbit-eu/transport_tools<br>results-*_rep0 # results for replicate 1 for given schemes for example cavity, cavity&amp;bulk or tunnels.<br>├── data<br>│ &nbsp; ├── super_clusters<br>├── &nbsp;_internal<br>│ &nbsp; ├── ...<br>├── statistics<br>│ &nbsp; ├── ...<br>results-*_rep1 # results for replicate 2 for given schemes for example cavity, cavity&amp;bulk or tunnels.<br>├── data<br>│ &nbsp; ├── super_clusters<br>├── &nbsp;_internal<br>│ &nbsp; ├── ...<br>├── statistics<br>│ &nbsp; ├── ...<br>results-*_rep2 # results for replicate 3 for given schemes for example cavity, cavity&amp;bulk or tunnels.<br>├── data<br>│ &nbsp; ├── super_clusters<br>├── &nbsp;_internal<br>│ &nbsp; ├── ...<br>├── statistics<br>│ &nbsp; ├── ...<br>- <em>event.csv</em> file contains the aggregated summary of events inferred from the&nbsp;<em>4-filtered_events_statistics.txt</em> files of each results of respective schemes</p> <p><strong>09_MSM_states.tar.gz</strong> - Contains the Markov state models (MSM) output files for Cavity, Cavity&amp;Bulk and Tunnels schemes and three replicates. The MSMs were generated using the pyEMMA program and HTMD framework, for further details please follow https://software.acellera.com/htmd/documentation.html. The directories looks as below:&nbsp;<br>├── Bulk<br>│ &nbsp; ├── rep1 # MSM states for replicate 1<br>│ &nbsp; ├── rep2 # MSM states for replicate 2<br>│ &nbsp; ├── rep3 # MSM states for replicate 3<br>├── Cavity<br>│ &nbsp; ├── rep1&nbsp;<br>│ &nbsp; ├── rep2&nbsp;<br>│ &nbsp; ├── rep3&nbsp;<br>├── Cavity&amp;Bulk<br>│ &nbsp; ├── rep1&nbsp;<br>│ &nbsp; ├── rep2<br>│ &nbsp; ├── rep3<br>├── Tunnels<br>│ &nbsp; ├── rep1&nbsp;<br>│ &nbsp; ├── rep2<br>│ &nbsp; ├── rep3</p> <p><br><strong>10_MSM_fingerprints.tar.gz</strong> - Contains the Markov state models (MSMs) distances generated from repository dir <strong>09_MSM_states</strong> consisting the model*.pdb files. The distances were calculated using the cpptraj program of AMBER18 package.<br>- <em>MSM_Distances/*/rep*/*.csv</em> : Contains .csv file for the generated MSM models (0,1,2..). The following columns (calculated distances) are present in the .csv files:<br>D107_OD1_DBE_C1 &nbsp;&nbsp;<br>D107_OD2_DBE_C1 &nbsp;&nbsp;<br>D107_OD1_DBE_C2 &nbsp;<br>D107_OD2_DBE_C2 &nbsp;&nbsp;<br>N37_ND2_DBE_Br1 &nbsp;<br>N37_ND2_DBE_Br2 &nbsp;<br>W108_NE1_DBE_Br1 &nbsp;<br>W108_NE1_DBE_Br2&nbsp;<br>D107_COM_DBE_COM &nbsp;&nbsp;<br>W108_COM_DBE_COM &nbsp; &nbsp;<br>N37_COM_DBE_COM &nbsp;&nbsp;<br>catal_COM_p1aCOM &nbsp; &nbsp;<br>catal_COM_p1bCOM &nbsp;&nbsp;<br>catal_COM_p2COM &nbsp; &nbsp;<br>catal_COM_p3COM &nbsp; &nbsp;<br>p1aCOM_DBE_COM &nbsp; &nbsp;<br>p1bCOM_DBE_COM &nbsp; &nbsp;<br>p2COM_DBE_COM &nbsp; &nbsp;<br>p3COM_DBE_COM &nbsp;<br>catal_COM_DBE_COM &nbsp; &nbsp;<br>p1aCOM_p1bCOM &nbsp;&nbsp;<br>p1aCOM_p2COM &nbsp;<br>p1aCOM_p3COM &nbsp;<br>p1bCOM_p2COM&nbsp;<br>p1bCOM_p3COM&nbsp;<br>p2COM_p3COM&nbsp;</p> <p><br><strong>11_ULS_clustering_and_transition_assignments.tar.gz</strong> - Contains files for analysis of utilization of the substrate DBE. Each folder contains two types of .csv files:<br>1. for the transition detection of DBE and the classification in &nbsp;Bulk (out_), Bottleneck (bt_), Unknown bottleneck (bt_unknown), Inside (in_) and&nbsp;<br>2. the second type as the charecterization on the tunnels utilization: Tunnel (p1a, p1b, p2, and p3), Mixed and Unknnown.<br># the details of the file arangements are as below for the studied schemes Bulk, Cavity, Cavity&amp;Bulk and Tunnels:<br>├── average_tunnel_utilization_per_scheme.png<br>├── average_tunnel_utilization.png<br>├── Bulk<br>│ &nbsp; ├── Bulk_run_htmd_0_combined_df.csv<br>│ &nbsp; ├── Bulk_run_htmd_0_transitions_counts.csv<br>│ &nbsp; ├── Bulk_run_htmd_1_combined_df.csv<br>│ &nbsp; ├── Bulk_run_htmd_1_transitions_counts.csv<br>│ &nbsp; ├── Bulk_run_htmd_2_combined_df.csv<br>│ &nbsp; └── Bulk_run_htmd_2_transitions_counts.csv<br>├── Bulk&amp;Cavity<br>│ &nbsp; ├── Cavity&amp;Bulk_run_htmd_0_combined_df.csv<br>│ &nbsp; ├── Cavity&amp;Bulk_run_htmd_0_transitions_counts.csv<br>│ &nbsp; ├── Cavity&amp;Bulk_run_htmd_1_combined_df.csv<br>│ &nbsp; ├── Cavity&amp;Bulk_run_htmd_1_transitions_counts.csv<br>│ &nbsp; ├── Cavity&amp;Bulk_run_htmd_2_combined_df.csv<br>│ &nbsp; └── Cavity&amp;Bulk_run_htmd_2_transitions_counts.csv<br>├── Cavity<br>│ &nbsp; ├── Cavity_run_htmd_0_combined_df.csv<br>│ &nbsp; ├── Cavity_run_htmd_0_transitions_counts.csv<br>│ &nbsp; ├── Cavity_run_htmd_1_combined_df.csv<br>│ &nbsp; ├── Cavity_run_htmd_1_transitions_counts.csv<br>│ &nbsp; ├── Cavity_run_htmd_2_combined_df.csv<br>│ &nbsp; └── Cavity_run_htmd_2_transitions_counts.csv<br>├── parse_distances_msm.py<br>├── schemes_comparison_piechart_per_scheme.png<br>└── Tunnels<br>&nbsp; &nbsp; ├── Tunnels_run_htmd_0_combined_df.csv<br>&nbsp; &nbsp; ├── Tunnels_run_htmd_0_transitions_counts.csv<br>&nbsp; &nbsp; ├── Tunnels_run_htmd_1_combined_df.csv<br>&nbsp; &nbsp; ├── Tunnels_run_htmd_1_transitions_counts.csv<br>&nbsp; &nbsp; ├── Tunnels_run_htmd_2_combined_df.csv<br>&nbsp; &nbsp; └── Tunnels_run_htmd_2_transitions_counts.csv</p> <p>&nbsp;</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics