Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
19
datasets available to search
ShareScore release 0.9.0
Dataset results
19 results for “QM/MM”
Dataset: 800 QM/MM minimum energy pathway conformations for the acylation reactions of Toho-1/ampicillin and Toho-1/cefalexin
<p>This dataset consists of 800 coordinate files (in the CHARMM psf/cor format) for the QM/MM minimum energy pathways of the acylation reactions between a Class A beta-lactamases (Toho-1) and two beta-lactam antibiotic molecules (ampicillin and cefalexin).</p> <p>These files are:</p> <ul> <li>toho_amp.r1-ae.zip: The R1-AE acylation pathways for Toho-1/Ampicillin (200 pathways);</li> <li>toho_amp.r2-ae.zip: The R2-AE acylation pathways for Toho-1/Ampicillin (200 pathways);</li> <li>toho_cex.r1-ae.zip: The R1-AE acylation pathways for Toho-1/Cefalexin (200 pathways);</li> <li>toho_cex.r2-ae.zip: The R2-AE acylation pathways for Toho-1/Cefalexin (200 pathways);</li> <li>energies.zip: the replica energies at B3LYP-D3/6-31+G**/C36 level;</li> <li>chelpgs.zip: the ChElPG charges of all reactant replicas at B3LYP-D3/6-31+G**/C36 level;</li> <li>farrys.zip: the featurzied NumPy arrays for model training;</li> <li>peephole.zip: an example file for how the optimized MEPs look like; </li> <li>dftb3_benchmark.zip: the reference calculations to justify the use of DFTB3/3OB-F/C36 in MEP optimizations, the reference level of theory is B3LYP-D3/6-31G**/C36. </li> </ul> <p>The R1-AE pathways are the acylation uses Glu166 as the general base; the R2-AE pathways uses Lys73 and Glu166 as the concerted base. </p> <p>All QM/MM pathways are optimized at the DFTB3/3OB-f/CHARMM36 level of theory. </p> <p>Z. Song et al Mechanistic Insights into Enzyme Catalysis from Explaining Machine-Learned Quantum Mechanical and Molecular Mechanical Minimum Energy Pathways. <em>ACS Phys. Chem Au</em> 2022, <strong>2</strong>, 4, 316–330. DOI: <a href="https://doi.org/10.1021/acsphyschemau.2c00005">10.1021/acsphyschemau.2c00005</a></p>
Imipenem hydrolysis by the L1 and the NDM-1 metallo-beta-lactamases: QM/MM structures of the stationary points
<p>Stationary points (transition states and minima) for the hydrolysis reactions of the imipenem in the active sites of the L1 and NDM-1 metallo-beta-lactamases. The reactions occure in three elementary steps each. For the NDM-1 there are two compeeting pathways leading to the different products. Enamine and imine forms of the hydrolyzed imipenem may form and those are marked with "N" and "C", respectively, according to the atom being protonated at the last step. Structures are calculated at the QM(PBE0-D3/6-31G**)/MM(AMBER) level.</p>
Unveiling the Catalytic Mechanism of NADP+-Dependent Isocitrate Dehydrogenase with QM/MM Calculations
<p>Raw molecular dynamics simulations for the complexes simulated for 20 ns with the AMBER 12 software. Details can be found in the original manuscript (<a href="https://doi.org/10.1021/acscatal.5b01928">https://doi.org/10.1021/acscatal.5b01928</a>). Molecular topology in AMBER Parameter Topology format and Trajectories in NETCDF binary coordinate format.</p> <p>Two protonation states were simulated: non-charged state of catalytic Lys212/Asp275 (<strong>LYN626+ASH275+ASP666</strong>) and charged state of catalytic Lys212/Asp275 (<strong>LYS626+ASP275+ASP666</strong>)</p> <p>Production trajectories are included in the <strong>MM-MD</strong> folders. Input and restart coordinates files for reproducibility are located in folders <strong>inputs</strong> and <strong>rst</strong>.</p> <p>Instructions for the calculations are summarized in <em>amber_min-md.sh</em>.</p>
QM/MM MD simulations of the ES complexes of SARS-CoV-2 main protease and oligopeptide substrates
<p>qmdcd.7z : QM/MM MD trajectories for all considered systems in dcd format for QM parts without link atoms (QMpart_nolink.pdb)</p> <p>frames.7z : QM parts of the MD frames selected for the electron density analysis.</p> <p> </p> <p> </p>
Data Set "Systematic QM Region Construction in QM/MM Calculations Based on Uncertainty Quantification"
<p>Data set accompanying the publication "Systematic QM Region Construction in QM/MM Calculations Based on Uncertainty Quantification"</p> <p>This dataset contains:</p> <p>- PDB files of the reactant and product starting structure</p> <p>- modified AMBER95 force field file</p> <p>- AMS fragment files for the ligands and ions</p> <p>- AMS input files for all geometry optimizations and single point calculations</p>
QM/MM nonadiabatic trajectories for cytosine in argon, benzene and water
<p>The dataset includes all trajectories for the nonadiabatic (QM/MM) molecular dynamics simulations of cytosine in argon, benzene and water. Each tar.gz directory contains two directories: INPUT and RESULTS. In this latter one has the 50 trajectories for cytosine in argon and benzene and the 75 trajectories for cytosine in water. In the files named dyn.out one has the atomic coordinates, velocities and energies for each time step.</p>
Dataset: 1,000 QM/MM minimum energy pathway conformations for the deacylation reactions of GES-5/imipenem
<p>This dataset consists of 1,000 coordinate files (in the CHARMM psf/cor format) for the QM/MM minimum energy pathways of the deacylation reactions between a Class A beta-lactamases (GES-5) and the imipenem antibiotic molecules.</p> <p>All pathway conformations were optimized at DFTB3/3OB-f/CHARMM36 level with 36 replicas.</p> <p>All single point calculations and charge population analysis were done at B3LYP-D3/6-31+G(d,p)/CHARMM36 level.</p> <ul> <li>0.paths_ges_imi_d1.tar.gz: 500 pathway conformations for GES-5/IPM-Delta1 deacylation reactions.</li> <li>0.paths_ges_imi_d2.tar.gz: 500 pathway conformations for GES-5/IPM-Delta1 deacylation reactions.</li> <li>1.eners.zip: The single point replica energies along all GES-5/IPM pathways.</li> <li>1.chrgs.zip: The NBO charges of the QM region of all replica conformations along all GES-5/IPM pathways.</li> <li>2.datasets.zip: The Python codes to postprocess the molecular data and the featurized the NumPy arrays.</li> <li>3.gnn.zip: The Python codes that implements the edge-conditioned graph convolutional NN to predict the deacylation barriers.</li> <li>5.representative_conf.zip: The pathway conformations of all cluster centroids and an energetic representative (pathway id 22) pathway. Note: This file also serves as a peephole of how the pathway conformations from Reaction Path with Holonomic Constrains calculations looks like.</li> <li>6.benchmark.zip: The benchmark calculations that validates the DFTB3/3OB-f/CHARMM36 against DFTB3/3OB/CHARMM36 and B3LYP/6-31G(d,p)/CHARMM36 level of theory on the energetic representative (pathway id 22) pathway conformations. </li> <li>p.figures.zip: A series of Jupyter Notebooks that produces the visualizations in the work.</li> <li>README.md: A markdown file that contains additional descriptions.</li> <li>environment.yml: The Conda environment used for the graph-learning. </li> </ul> <p>Z. Song and P. Tao, Graph-Learning Guided Mechanistic Insights into Imipenem Hydrolysis in GES Carbapenemases. <strong><em>Electron. Struct.</em></strong> 2022, <strong>4</strong>, 034001. DOI: <a href="https://doi.org/10.1088/2516-1075/ac7993">10.1088/2516-1075/ac7993</a></p>
CP2K biomolecular QM/MM benchmarking data
<p>Dataset containing raw CP2K output logs produced by running benchmarks from the BioExcel QM/MM benchmark suite (<a href="https://doi.org/10.5281/zenodo.6591692">https://doi.org/10.5281/zenodo.6591692</a>), provided as part of BioExcel-2 project deliverable D1.6. </p> <p>An analysis of (a subset of) these benchmark results can be found in <a href="https://doi.org/10.5281/zenodo.6591574">https://doi.org/10.5281/zenodo.6591574</a> </p> <p> </p> <p><strong>Machines used for benchmarking:</strong></p> <p>Cirrus@EPCC (HPE SGI ICE XA):</p> <p><a href="https://www.cirrus.ac.uk/">https://www.cirrus.ac.uk</a></p> <ul> <li>Infiniband interconnect</li> <li>CPU compute nodes: <ul> <li>2 x 18-core Intel Xeon (Broadwell) E5-2695, 2.1 GHz</li> <li>256GB RAM</li> </ul> </li> <li>GPU compute nodes: <ul> <li>2 x NVIDIA Tesla V100 (Volta) SXM2-16GB</li> <li>2 x 20-core Intel Xeon (Cascade Lake) Gold 6248, 2.4 GHz</li> <li>384GB RAM</li> </ul> </li> </ul> <p>ARCHER2@EPCC (HPE CRAY EX):</p> <p><a href="https://www.archer2.ac.uk/">https://www.archer2.ac.uk/</a></p> <ul> <li>Interconnect: HPE Cray Slingshot</li> <li>Compute nodes: <ul> <li>2 x 64-core AMD EPYC (Zen2 Rome) 7742, 2.25GHz</li> <li>256GB RAM</li> </ul> </li> </ul> <p><strong>Benchmarking protocol</strong></p> <p>Repeated runs to rule out machine noise variability were performed for each benchmark on each machine for both 1 MD step and 6 MD steps. Subsequent analysis and visualisation of parallel scaling of average runtime per MD step and subroutine-level profiling was performed using our analysis script available from:</p> <p><a href="https://doi.org/10.5281/zenodo.6591681">https://doi.org/10.5281/zenodo.6591681</a></p> <p>Results on Cirrus CPU nodes are for CP2K release version 8.1, whilst results on Cirrus GPU nodes and ARCHER2 are with CP2K version 8.2.</p>
Hydrolysis of SeCN and SCN in the active site of TcDH: QM/MM stationary points
<p>Equilibrium geometry configurations of the enzyme-substrate complexes (ES), transition states (TS) and reaction products (EP) for the SeCN and SCN hydrolysis in the active site of the TcDH enzyme calculated at the QM(uPBE0-D3/6-31G**)/MM(AMBER99 level) using NWChem software. The "qm" in the file name denotes that only QM part is included.</p>
QM/MM models of 11 red fluorescent proteins
<p>Geometry configurations on the ground state potential energy surface are obtained at the QM(PBE0-D3/cc-pvdz)/MM(AMBER) level. The list of fluorescent proteins is as follows:</p> <ol> <li>mRojoA</li> <li>mRojo-THSL</li> <li>mRojo-VFAV</li> <li>mRojo-TFAL</li> <li>mRojo-VYGV</li> <li>mRojo-VYGL</li> <li>mRojo-TYGV</li> <li>RDSmCherry0.2</li> <li>RDSmCherry0.5</li> <li>mKeima</li> <li>eqFP670</li> </ol>
``HiPen'': a new dataset for validating (S)QM/MM free energy simulations
<p>Calculating free energy differences between levels of theory (i.e., <span class="math-tex">\(\Delta A^{low \to high}\)</span>) is integral to performing indirect (S)QM/MM free energy simulations. However, connecting levels of theory via free energy simulations has proved difficult due to (1) bond/angle degrees of freedom, (2) dihedral degrees of freedom, and (3) solvent arrangement differences between levels of theory, largely due to partial charge differences between levels of theory. In order to improve calculation of (S)QM/MM free energy simulations, the free energy simulation community should begin to compare methods based on convergence success relative to overall computational time and resource requirements. We have begun to compile such a dataset by calculating <span class="math-tex">\(\Delta A^{MM \to SCC-DFTB}\)</span> in gas phase for 22 drug-like molecules, as seen in our recent publication, Kearns, et al. <strong>2018</strong>, <em>Molecules</em>, Submitted, and we hope that future practitioners will do the same. With this work we hope to provide a standard for comparison for future FES methodologies; additionally, in the near future we hope to continue to add to this dataset including results in more complicated environments such as in solution and in enzyme. All data can be found in our publication and in the accompanying Supporting Information; raw data (such as simulation trajectories and raw data files) can be made available upon request. The purpose of this dataset publication is to make available all starting coordinates, topologies, parameter sets, and input files necessary to replicating the results published in our work.</p>
Data Set "Protein network centralities as descriptor for QM region construction in QM/MM simulations of enzymes"
<p>This data set accompanies the publication "Efficient automatic construction of atom-economical QM regions with point-charge variation analysis" by Felix Brandt and Christoph R. Jacob (TU Braunschweig, Germany) </p> <p>It contains:</p> <p>- PDB files of the starting structures</p> <p>- modified AMBER95 force field file</p> <p>- AMS fragment files for the substrates and ions</p> <p>- AMS input files for all geometry optimizations and single point calculations</p> <p>- Python script for WISP and centrality analysis</p>
D1.4 release of BioExcel QM/MM Benchmark Suite
<p>D1.4 release snapshot of the BioExcel QM/MM benchmark suite hosted at <a href="https://github.com/bioexcel/qmmm_benchmark_suite">https://github.com/bioexcel/qmmm_benchmark_suite</a></p> <p> </p>
Refining ligand poses in RNA/ligand complexes of pharmaceutical relevance: a perspective by QM/MM simulations and NMR measurements
<p>This directory contains files for the project:</p> <p>Title: Refining ligand poses in RNA/ligand complexes of pharmaceutical relevance: a perspective by QM/MM simulations and NMR measurements</p> <p>Authors: Gia Linh Hoang, Manuel Röck, Aldo Tancredi, Thomas Magauer, Davide Mandelli, Jörg B. Schulz, Sybille Krauss, Giulia Rossetti, Martin Tollinger, Paolo Carloni</p> <p>There are two folders containing input and parameter files of the classical MD simulations with GROMACS and QM/MM simulation with MiMiC, and an input file for NMR Chemical Shifts calculation with ORCA.</p>
Data for "Accelerating QM/MM simulations of electrochemical interfaces through machine learning of electronic charge densities"
Open the record for dataset details and reuse information.
Reference and ESPF-DRF QM/MM data for bimolecular interaction energies
<p>Reference QM data and ESPF-DRF QM/MM (generated with PyESPF + PySCF https://github.com/tomfay/PyESPF ) interaction energy data for a set of test bimolecular systems.</p>
D1.6 release of BioExcel QM/MM Benchmark Suite
<p>An update to the previous D1.4 release of the BioExcel QM/MM benchmark suite (see <a href="https://doi.org/10.5281/zenodo.3885124">https://doi.org/10.5281/zenodo.3885124</a>), provided as part of deliverable D1.6.</p> <p>This is a snapshot of the repository available at <a href="https://github.com/bioexcel/qmmm_benchmark_suite">https://github.com/bioexcel/qmmm_benchmark_suite</a></p> <p>Compared to the previous (D1.4) release, a number of additional benchmarks were included in order to facilitate systematic performance profiling, namely: </p> <ul> <li>Versions of the MQAE (solute-solvent), ClC (ion channel), and CBD_PHY (phytochrome) benchmarks that use the hybrid DFT functionals PBE0 and B3LYP to incorporate Hartree-Fock exchange energy (the initial release of the benchmark suite included only the GGA functionals PBE and BLYP).</li> <li>Versions of the MQAE and ClC benchmarks that make use of the Auxiliary Density Matrix Method (ADMM) to markedly reduce the high computational cost of the molecular environment optimized MOLOPT basis set when used in combination with hybrid functionals such as B3LYP and PBE0.</li> <li>Versions of the MQAE and ClC benchmarks that use the EMSL basis set in combination with the B3LYP hybrid functional.<strong> </strong></li> <li>A version of the ClC-19 benchmark that uses the HFX basis set in combination with the B3LYP hybrid functional<strong>.</strong></li> <li>A version of the MQAE benchmark in which the linear size of the cell defining the quantum region is increased by a factor 2, i.e. where the volume of the quantum region is enlarged by a factor 8</li> </ul> <p><br> </p>
Data for "Influence of Wobbling Tryptophan and Mutations on PET Degradation Explored by QM/MM Free Energy Calculations" paper
Open the record for dataset details and reuse information.
Tetrahedral Cu(I) Complexes for Thermally Activated Delayed Fluorescence: A Density Functional Benchmark Study with QM/MM Models
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.