Project files provided as supporting information to the manuscript "Information-theoretical measures identify accurate low-resolution representations of protein configurational space"
<p>The dataset contains the following compressed folder:</p> <p>-Notebooks.zip:</p> <p>This folder contains:<br> -python_script:<br> -RESREL.py: script performing the clusterization and computing the relevance resolution curves<br> -random_curves.py: script generating the random value and computing the corresponding RES-REV curves_s<br> -Cluster_distance_matrix.py: script returning the distance among clusters for a given partition.<br> -python_notebook:<br> -Exploratory_analysis.ipynb: Analysis performed on the 12-protein_dataset<br> -DMAPS_ANTI.ipynb: Diffusion Map for the Antibody<br> -DMAPS_COV_1ake.ipynb: Diffusion Map + Inter-Intra state decomposition of covariance for 1ake</p> <p>Packages required for the usage of these python scripts/notebooks:<br> -numpy<br> -pandas<br> -matplotlib<br> -seaborn<br> -multiprocessing<br> -scipy</p> <p> </p> <p>========<br> RAW DATA<br> ========</p> <p>The raw data produced and employed in this study are available on a Google Drive folder at the following address:</p> <p>https://drive.google.com/drive/folders/1PasAUCgpR5-gdzUVEdyusgZIayQN0Le9</p> <p>In this folder, together with the compressed Notebooks.zip folder, one can fin the compressed folder Data.zip, within which the following data are present:</p> <p>-12-protein_dataset:<br> -md.mdp: the .mdp file used in the MD simulations<br> -PROTEIN_PDB_CODE:<br> -Hk_{sel}.npy & Hs_{sel}.npy: the Rel & Res curves, sel=[all, CA, CB]<br> -RMSD_{sel}.npy: the RMSD matrix, sel=[all, CA, CB]<br> -npt.gro:protein+water+ions structure @TEO the equilibration (NVT+NPT)<br> -MSR_df.csv: a dataset containing the following columns<br> 'area' : area behind the Relevance-Resolution curve;<br> 'selection': the atomic selection (['all', 'CA', 'CB']) used to compute the RMSD matrix used for the clusterization (and consequently the Relevance-Resolution curves)<br> 'method': the linkage measure used in the clustering procedure, an integer in [0,6];<br> 'method_name': the linkage measure used in the clustering procedure, a string in ['average','ward','complete','single','centroid','median','weighted'];<br> 'rmsd_mean': the mean value of the rmsd vector along the trajectory computed wrt the first frame;<br> 'rmsd_var': the variance of the rmsd vector along the trajectory computed wrt the first frame;<br> 'rgy_mean': the mean value of the radius of gyration along the trajectory;<br> 'rgy_var': the variance of the radius of gyration along the trajectory;<br> 'rmsf_mean': the mean value of the rmsf;<br> 'rmsf_var': the variance of the rmsf;<br> 'RMSD_M_mean': the mean value of the RMSD matrix.<br> 'RMSD_M_var': the variance of the RMSD matrix.<br> -Random:<br> -curves.npy= 100K Relevance-Resolution Random curves for M=40001<br> -curves_s.npy= 100K Relevance-Resolution Random curves for M=15000<br> -validation_dataset:<br> -antibody:<br> -Hk_CB.npy & Hs_CB.npy: the Rel & Res curves<br> -RMSD_CB.npy: the RMSD matrix<br> -DIFF_{M}.npy: the eigenvalue/vector of the 10-D diffusion space<br> -Label_{method}.npy: the label vector for n_clusters<br> -1ake:<br> -Hk_{sel}.npy & Hs_{sel}.npy: the Rel & Res curves<br> -RMSD_{sel}.npy: the RMSD matrix<br> -DIFF_{M}.npy: the eigenvalue/vector of the 10-D diffusion space<br> -Label_{method}.npy: the label vector for n_clusters<br> -intra_{m}.npy: the intra-cluster covariance matrix<br> -inter_cov_{m}.npy: the inter-cluster correlation matrix</p> <p> </p> <p>NOTE<br> =====</p> <p>The matrices of the cluster distances for adenylate kinase and antibody have been computed through the script Cluster_distance_matrix.py.</p> <p>These matrices have not been included in the dataset because of their large size; the raw data are however available upon request.<br> </p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4