Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

598

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

598 results for “Small molecules”

Learn how ShareScore rates datasets ↗
zenodo52/100

Bioactivity of small-molecule compounds against Haemonchus contortus

<div> <div> <div> <p>This dataset of small-molecule compounds and their effects on <em>H. contortus </em>was assembled based on the results obtained from screening two compound libraries (Medicines for Malaria Venture Pathogen Box, Compounds Australia Open Scaffolds set) to assess the effect of compounds on the motility of exsheathed third-stage larvae (xL3) of <em>H. contortus </em>(Preston et al., 2016, 2017). Additionally, select literature data were included to augment the in-house generated data.</p> </div> </div> </div>

opencc-by-4.0Apr 2024View details →
zenodo52/100

AIMEl-DB: Atomic Properties for 44K small organic molecules

<h3>AIMEl-DB: Atomic Properties for 44K small organic molecules</h3> <p>This dataset comprises atomic properties of 44K (44 470) molecules selected from the QM9 database. The file names are based on the same indexing system used for QM9.&nbsp;</p> <p>This dataset includes four types of files:</p> <ul> <li><strong>.com Files<br></strong>Input files for Gaussian 16. Simple-point energy calculations were carried out using the keywords<br><code># B3LYP/6-31G(2df,p) scf=(maxcycle=9999) nosymm output=wfx</code><br><br></li> <li><strong>.log Files<br></strong>Output files from Gaussian 16 calculation with the aformentioned parameters.<br><br></li> <li><strong>.wfx Files<br></strong>Wave function files from Gaussian 16 calculation. These files were used as inputs for QTAIM calculations.&nbsp;<br><br></li> <li><strong>.sumviz Files<br></strong>Output file from AIMAll software. The keywords used for the calculations were<br><code>aimqb -nogui -scp=false -nproc=8 -naat=4 input.wfx</code><br>Each .sumviz file contains more than 30 properties based on the Quantum Theory of Atoms in Molecules (QTAIM).<br><br></li> <li><strong>.csv Files<br></strong>These files contain the results of a in-house treament of .sumviz data. They cointain two calculated atomic properties:<br><br> <ol> <li>Total magnitude of the dipole moment, |mu|</li> <li>Total magnitude of the quadrupole moment, |Q|</li> </ol> </li> </ul> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; and two extracted atomic properties:<br><br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 3.&nbsp; Electronic Population, N<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 4. Atomic Energy, E</p> <p>&nbsp;</p> <p>The <code>aimel_merged_44k.csv</code> presents the concatenation of the 44 470 <strong>csv Files. </strong></p> <p>Additionaly, the <code>aimel_merged_38k.csv</code> presents the concatenation of the 38 876 <strong>csv Files.&nbsp;</strong>This file corresponds to the version 1.0 of the dataset.&nbsp;</p> <p><br>If you find this dataset useful, please cite the original paper:</p> <p>Meza-Gonz&aacute;lez, B., Ram&iacute;rez-Palma, D.I., Carpio-Mart&iacute;nez, P.&nbsp;<em>et al.</em>&nbsp;Quantum Topological Atomic Properties of 44K molecules.&nbsp;<em>Sci Data</em>&nbsp;<strong>11</strong>, 945 (2024). https://doi.org/10.1038/s41597-024-03723-0</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2024View details →
zenodo48/100

Clinically approved small molecule antibiotics since 2010

<p>Properties of clinically approved small molecule antibiotics since&nbsp;2010 are summarized.</p> <p>Properties include:&nbsp;Name&nbsp;&nbsp; &nbsp;Year approved&nbsp;&nbsp; &nbsp;Origin&nbsp;&nbsp; &nbsp;Mechanism of action&nbsp;&nbsp; &nbsp;Indication&nbsp;&nbsp; &nbsp;Administration&nbsp;&nbsp; &nbsp;Spectrum&nbsp;&nbsp; &nbsp;Protein binding&nbsp;&nbsp; &nbsp;Resistance&nbsp;&nbsp; &nbsp;Smiles&nbsp;&nbsp; &nbsp;Molweight&nbsp;&nbsp; &nbsp;cLogP&nbsp;&nbsp; &nbsp;cLogS&nbsp;&nbsp; &nbsp;H-Acceptors&nbsp;&nbsp; &nbsp;H-Donors&nbsp;&nbsp; &nbsp;Druglikeness&nbsp;&nbsp; &nbsp;DrugScore&nbsp;&nbsp; &nbsp;Total Molweight&nbsp;&nbsp; &nbsp;Monoisotopic Mass&nbsp;&nbsp; &nbsp;Total Surface Area&nbsp;&nbsp; &nbsp;Relative PSA&nbsp;&nbsp; &nbsp;Polar Surface Area&nbsp;&nbsp; &nbsp;LE from Molweight&nbsp;&nbsp; &nbsp;LLE from Molweight&nbsp;&nbsp; &nbsp;LELP from Molweight&nbsp;&nbsp; &nbsp;Shape Index&nbsp;&nbsp; &nbsp;Molecular Flexibility&nbsp;&nbsp; &nbsp;Molecular Complexity&nbsp;&nbsp; &nbsp;Structure of Smiles [idcode]</p> <p>&nbsp;</p> <p>Propeties in rows B to I were retrieved from the references cited in the dataset. &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;<br> Properties in rows K - AC were calculated with DataWarrior.&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;https://openmolecules.org/datawarrior/index.html</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Experimental n-octanol/water Partition/Distribution Coefficients Database for Small Molecules

<p>We critically compiled experimental values of&nbsp;log<em>P</em><sub>N,&nbsp;</sub>pK<sub>a</sub>, log<em>P</em><sub>I</sub>, and&nbsp;log<em>D</em><sub>pH</sub> of&nbsp;225&nbsp;entries based on earlier&nbsp;literature reports.&nbsp;The experimental techniques of log<em>P</em><sub>N</sub>&nbsp;,&nbsp;log<em>D</em><sub>pH</sub>, and&nbsp;log<em>P</em><sub>I</sub>&nbsp;for each molecule were thoroughly revised and added to the database.&nbsp;The molecules were classified between acids and bases according to their functional groups and&nbsp;pK<sub>a</sub>&nbsp;values.&nbsp;</p> <p>If you use this database please cite our&nbsp;<a href="https://doi.org/10.1002/cphc.202300548">ChemPhysChem</a> paper</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Database of small molecule X-ray absorption spectra, featurized structures, and neural network ensembles

<p>Companion data for arXiv preprint <em>Uncertainty-aware predictions of molecular X-ray absorption spectra using neural network ensembles</em>&nbsp;(<a href="https://arxiv.org/abs/2210.00336">https://arxiv.org/abs/2210.00336</a>), by&nbsp;Animesh Ghose, Mikhail Segal, Fanchen Meng, Zhu Liang, Mark S. Hybertsen, Xiaohui Qu, Eli Stavitski, Shinjae Yoo, Deyu Lu &amp;&nbsp;Matthew R. Carbone.</p> <p><strong>Included</strong></p> <ul> <li>*-XANES-*.tar.bz2: raw&nbsp;input/output files for all molecular simulations used in the work. These inputs and outputs correspond to the structural data in the QM9 dataset.</li> <li>ml_ready.tar.bz2: machine learning-ready data (featurized spectra). Used as input to the neural network ensembles.</li> <li>XANES-220712-ACSF-*.tar.bz2: neural network ensembles used in this work.</li> </ul> <p><strong>Notes</strong></p> <ul> <li>The FEFF9 code [J. J. Rehr, J. J. Kas, F. D. Vila, M. P. Prange, and&nbsp;K. Jorissen, <em>Phys. Chem. Chem. Phys.</em> <strong>12</strong>, 5503 (2010)]&nbsp;was used to generate all X-ray absorption near-edge structure (XANES) spectra.</li> <li>All molecular structures were sourced from the QM9 database [R. Ramakrishnan, P. O. Dral, M. Rupp, and O. A. Von Lilienfeld, <em>Sci. Data</em> <strong>1</strong>, 1 (2014)].</li> </ul> <p><strong>Funding</strong></p> <p>This research is based upon work supported by the U.S. Department of Energy, Office of Science, Office Basic Energy Sciences, under Award Number FWP PS-030. This research also used theory and computational resources of the Center for Functional Nanomaterials, which is a U.S. Department of Energy Office of Science User Facility, and the Scientific Data and Computing Center, a component of the Computational Science Initiative, at Brookhaven National Laboratory under Contract No. DE-SC0012704.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Structural Interaction Fingerprints and Machine Learning for predicting and explaining binding of small molecule ligands to RNA: a benchmark dataset

<p><b>Structural Interaction Fingerprints and Machine Learning for predicting and explaining binding of small molecule ligands to RNA: a benchmark dataset.</b></p><p>Ribonucleic acids (RNA) play crucial roles in living organisms as they are involved in key processes necessary for proper cell functioning. Some RNA molecules, such as bacterial ribosomes and precursor messenger RNA, are targets of small molecule drugs, while others, e.g., bacterial riboswitches or viral RNA motifs are considered as potential therapeutic targets. Thus, the continuous discovery of new functional RNA increases the demand for developing compounds targeting them and for methods for analyzing RNA—small molecule interactions. We recently developed fingeRNAt - a software for detecting non-covalent bonds formed within complexes of nucleic acids with different types of ligands. The program detects several non-covalent interactions, such as hydrogen and halogen bonds, ionic, Pi, inorganic ion- and water-mediated, lipophilic interactions, and encodes them as computational-friendly Structural Interaction Fingerprint (SIFt). Here we present the application of SIFts accompanied by machine learning methods for binding prediction of small molecules to RNA targets. We show that SIFt-based models outperform the classic, general-purpose scoring functions in virtual screening. We discuss the aid offered by Explainable Artificial Intelligence in the analysis of the binding prediction models, elucidating the decision-making process, and deciphering molecular recognition processes.</p>

opencc-zeroDec 2022View details →
zenodo44/100

Extreme dynamics of a small molecule in its bound state with an intrinsically disordered protein

<p>These data support the manuscript entitled "Extreme dynamics of a small molecule in its bound state with an intrinsically disordered protein" by Heller, Shukla, Figueiredo, and Hansen.</p><p>This data should be used with the code provided on GitHub at https://github.com/hansenlab-ucl/R2_IDP_small_mol. Once downloaded, this directory should be extracted using the following command:</p><p>&nbsp; &nbsp; tar -xzvf Data.tar.gz</p><p>The directory should be saved with the name 'Data' placed in the same directory as the GitHub README.md file.</p><p><strong>This dataset contains:&nbsp;</strong><br><i>Nuclear Magnetic Resonance (NMR) spectroscopy data files (.ft2 format) including:&nbsp;</i></p><p>* 1H 1D ligand-detected chemical shift titration of 5-fluoroindole (50 uM) with increasing concentrations of the protein, non-structural protein 5A, domains 2 and 3 (NS5A-D2D3), in 1H_1D_ft2_data/</p><p>* 1H pseudo-2D Diffusion Ordered SpectroscopY (DOSY) data of 5-fluoroindole (50 uM) with and without NS5A-D2D3 (75 uM) in 1H_DOSY_data/</p><p>* 1H-15N Heteronuclear Single Quantum Coherence (HSQC) measurements of NS5A-D2D3 (40 uM) in the absence and presence of 5-fluoroindole (160 and 320 uM) in 1H_15N_HSQC_ft2_and_metadata/</p><p>* 19F 1D ligand-detected chemical shift titration of 5-fluoroindole (50 uM) with increasing concentrations of NS5A-D2D3 in 19F_1D_ft2_data/</p><p>* 19F pseudo-2D ligand-detected longitudinal (spin-lattice, R1,eff) relaxation titration data of 5-fluoroindole (50 uM) with increasing concentrations of NS5A-D2D3 in 19F_R1eff_ft2_data/</p><p>* 19F pseudo-2D ligand-detected longitudinal (spin-spin, R2,eff) relaxation titration data of 5-fluoroindole (50 uM) with increasing concentrations of NS5A-D2D3 in 19F_R2eff_ft2_data/</p><p><i>Circular Dichroism (CD) data files (.txt format) including:&nbsp;</i></p><p>* CD measurements of NS5A-D2D3 at increasing concentrations in CD_data/no_molecule/</p><p>* CD measurements of NS5A-D2D3 with and without the small molecule, 5-fluoroindole CD_data/with_molecule/</p><p><i>Metadata&nbsp;</i></p><p>* Metadata from the Biological Magnetic Resonance Data Bank (https://bmrb.io/) used to determine scaling factors for the calculation of chemical shift perturbations in 1H_15N_HSQC_ft2_and_metadata/</p>

opencc-by-4.0May 2023View details →
zenodo44/100

RMG-DB-11: Enumerating Reaction Space for Small Molecule Chemistry

<p>This repository presents approximately 750 million atom-mapped reaction SMILES. Reactions are generated by applying templates from the Reaction Mechanism Generator (RMG) database to a subset of the species from GDB11. Thus, we refer to this dataset as RMG-DB-11 i.e., the Reaction Mechanism Generator Database whose species contain up to 11 heavy atoms. All SMILES have been canonicalized by RDKit. All reactions are labeled with their corresponding RMG template.</p> <p>This data serves as a crucial starting point for quantitative predictive chemistry. Many methods that search for transition state structures require atom-mapped SMILES, which this repository provides. This data is also well-suited for unsupervised pre-training of various machine learning models.</p> <p>To parse the data with Python, start with <em>import pandas as pd</em>. Reactions with 1-8 heavy atoms can be parsed using the following code snippet: <em>pd.read_csv(&lt;filepath&gt;)</em>. Reactions with 9 heavy atoms can be parsed using <em>pd.read_pickle(&lt;filepath&gt;, compression=&#39;zip&#39;)</em>. The file names below include the word &quot;zip&quot; as a helpful hint to use the compression argument. Due to the large number of reactions with 10 and 11 heavy atoms, these are split into smaller chunks. First untar the file using <em>tar -xvf &lt;tar_archive&gt;</em> to obtain several zipped pickle files that can each be parsed using the same method as with 9 heavy atoms.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

London Dispersion Governs the Interaction Mechanism of Small Polar and Non-Polar Molecules in Metal-Organic Frameworks

<p>Raw data set relating to publication.</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

COVID 19 SARS COV2 targets and small molecule data including insilico analysis

<p>Welcome to the repository for the COVID-19 research data.</p> <p>Corresponding Author: Girinath G. Pillai and few experts</p> <p>Co-authors: Team of experts, scholars and students</p> <p>To join dedicated Slack Discussion : <a href="https://join.slack.com/t/nyroindia/shared_invite/zt-ejes216c-QZzEK_G5tNKIjewbVj2IPA">https://join.slack.com/t/nyroindia/shared_invite/zt-ejes216c-QZzEK_G5tNKIjewbVj2IPA</a></p> <p>We commit to conduct research analysis and all the findings and data will be open and anyone can use or help us improve the data.</p> <p>The parameters for checkpoints are:</p> <p>A) Pharmacophore Modelling - i) generate pharmacophore reference maps from XRay crystal geometry, ii) Generate all possible conformers of the dataset molecules for screening.</p> <p>B) Virtual Screening - i) highest docking score within the dataset, ii) lowest clashes (interligand or intraligand), iii) interactions with key amino acid residues based on literature reports, PROSITE server and pocket finding algorithm like DoGSite or CASTp iv) optimal LE values and v) satisfactory interactions between small molecules and amino acids.</p> <p>C) Selection - i) binding affinity range predictions, lowest among the dataset, ii) free binding energy calculation considering desolvation terms, lowest among the dataset and iii) torsion analysis - coverage of bonds in CSD database.</p> <p>D) Optimization - i) pharmacokinetic properties to be generated from selected hits and an optimal balance of properties to be considered for candidate selection criteria.</p> <p>E) For novel lead molecules - i) chemical space exploration on building blocks could be carried out, ii) on-demand synthesis and procurement.</p> <p>If you prefer you could always cite <a href="https://github.com/giribio/COVID19">https://github.com/giribio/COVID19</a></p> <p>Feel free to create any issues in Github or feel free to contact me via Slack for any queries.</p> <p>Thanks and let us fight against COVID-19 in all possible ways.</p>

openother-openMay 2020View details →
zenodo40/100

QM7-X: A comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules

<p>Here, we introduce QM7-X, a comprehensive dataset of &gt; 40&nbsp;physicochemical properties for ~4.2&nbsp;M equilibrium and non-equilibrium structures of small organic molecules with up to seven non-hydrogen (C, N, O, S, Cl) atoms. To span this fundamentally important region of chemical compound space (CCS), QM7-X includes an exhaustive sampling of (meta-)stable equilibrium structures---comprised of constitutional/structural isomers and stereoisomers, e.g.,&nbsp;enantiomers and diastereomers (including cis-trans-and conformational isomers)---as well as 100&nbsp;non-equilibrium structural variations thereof to reach a total of ~4.2&nbsp;M molecular structures. Computed at the tightly converged quantum-mechanical PBE0+MBD level of theory, QM7-X contains global (molecular) and local (atom-in-a-molecule) properties ranging from ground state quantities (such as atomization energies and dipole moments) to response quantities (such as polarizability tensors and dispersion coefficients). By providing a systematic, extensive, and tightly converged dataset of quantum-mechanically computed physical and chemical properties, we expect that QM7-X will play a critical role in the development of next-generation machine-learning based models for exploring greater swaths of CCS and performing <em>in silico</em>&nbsp;design of molecules with targeted properties.</p> <p>The dataset is provided in eight HDF5 based files (compressed in .XZ files). One can also find here a README file with technical usage details and examples of how to access the information stored in the dataset (see createDB.py).&nbsp;</p> <p>*The paper explaining the generation of data stored in QM7-X can be found in <em>Sci Data</em>&nbsp;8,&nbsp;43 (2021). DOI: 10.1038/s41597-021-00812-2 . arXiv:&nbsp;https://arxiv.org/abs/2006.15139 .</p>

opencc-by-4.0Jun 2020View details →
zenodo40/100

Dataset: "Probabilistic Framework for Integration of Mass Spectrum and Retention Time Information in Small Molecule Identification"

<p>The SQLite database contains the pre-computed tandem mass spectra (MS2) and retention order scores used for the experiments in the publication: &quot;<a href="https://doi.org/10.1093/bioinformatics/btaa998">Probabilistic Framework for Integration of Mass Spectrum and Retention Time Information in Small Molecule Identification</a>&quot; by Bach et al. (2020).</p> <p>A detailed description of the database structure is given in the &#39;README.md&#39; and can also be found in the <a href="https://github.com/aalto-ics-kepaco/msms_rt_score_integration/tree/master/data">code-repository associated with the publication</a>. The database layout is illustrated in the &#39;db_layout.png&#39; file.</p> <p>The SQLite file &#39;ms_and_rt_score_DB_bach_etal_2020.db.gz&#39; is compressed using <a href="https://en.wikipedia.org/wiki/Gzip">gzip</a>.</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Optimized geometries for selected ions of ionic liquids and small molecules

<p>The geometries of these selected chemical entities were optimized at the ab initio or semiempirical levels of theory. They can be conveniently used to create more complicated systems through combining species like the free PACKMOL software offers.&nbsp;</p>

opencc-zeroAug 2016View details →
zenodo40/100

Allosteric Modulation of YAP/TAZ-TEAD Interaction by Palmitoylation and Small Molecule Inhibitors

<p>The tar file contains 3 directories, each of which contains all of the simulation input files (pdb, prmtop, inpcrd) and trajectory files (nc) for the simulations run.&nbsp;</p><ol><li>TEAD_Only: Simulations of the TEAD protein in the apo, inhibitor(inh)-bound, palmitate(plt)-bound, or palmitic acid(plm)-bound form&nbsp;</li><li>YAP-TEAD: Simulations of the YAP-TEAD heterodimer in the apo, inhibitor(inh)-bound, palmitate(plt)-bound, or palmitic acid(plm)-bound form&nbsp;</li><li>TAZ-TEAD: Simulations of the TAZ-TEAD heterodimer in the apo, inhibitor(inh)-bound, palmitate(plt)-bound, or palmitic acid(plm)-bound form&nbsp;</li></ol>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Small molecules targeting the structural dynamics of AR-V7 partially disordered protein using deep learning and physics based models.

<p>Partially disordered proteins can contain both stable and unstable secondary structure segments and &nbsp;are involved in various (mis)functions in the cell. The extensive conformational dynamics of partially disordered proteins scaling with extent of disorder and length of the protein hampers the efficiency of traditional experimental and in-silico structure-based drug discovery approaches. Therefore new efficient paradigms in drug discovery taking into account conformational ensembles of proteins need to emerge. In this study, using as a test case the AR-V7 transcription factor splicing variant related to prostate cancer, we present an automated &nbsp;methodology that can accelerate the screening of small molecule binders targeting partially disordered proteins. By swiftly identifying the conformational ensemble of AR-V7, and reducing the dimension of binding-sites by a factor of 90 by applying appropriate physicochemical filters, &nbsp;we combine physics based molecular docking and multi-objective classification machine learning models that speed up the screening of thousands of compounds targeting AR-V7 multiple binding sites. Our method not only identifies previously known binding sites of AR-V7, but also discovers new ones, as well as increases the multi-binding site hit-rate of small molecules by a factor of 17 compared to naive physics-based molecular docking.&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Supporting data for: "Data-driven discovery of cardiolipin-selective small molecules by computational active learning"

<p>This repository contains supporting data and code for the paper titled &quot;Data-driven discovery of cardiolipin-selective small molecules by computational active learning&quot; by Bernadette Mohr, Kirill Shmilovich, Isabel Kleinw&auml;chter, Dirk Schneider, Andrew L.Ferguson, and Tristan Bereau.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Regulatory spine RS3 residue of protein kinases: a lipophilic bystander or a decisive element in the small-molecule kinase inhibitor binding?

<p>Datasets related to publication:&nbsp;</p> <p>Shevchenko E, Pantsar T: Regulatory spine RS3 residue of protein kinases: a lipophilic bystander or a decisive element in the small-molecule kinase inhibitor binding?.&nbsp;<em><em>Biochem Soc Trans</em></em>&nbsp;28 February 2022; 50 (1): 633&ndash;648</p> <p>https://doi.org/10.1042/bst20210837</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Dataset: "Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data"

<p>Dataset used in the experiments of the publication: &quot;Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data&quot; by Bach et al.</p> <p><strong>File description:</strong></p> <ul> <li> <p>cfmid4.tar: MS&sup2; spectra simulated using <a href="https://bitbucket.org/wishartlab/cfm-id-code/src/CFM-ID_4.0.7/">CFM-ID (v4.0.7)</a> for all molecular candidate structures</p> </li> <li> <p>db_layout.png: Visualization of the SQLite database (DB) layout</p> </li> <li> <p>massbank.sqlite.gz: DB containing all needed data to (re-)run the experiments shown in the paper. Please read &quot;DB_README.md&quot; for further details. The database file can be unpacked using gzip.</p> </li> <li> <p>metfrag.tar: MetFrag input files and MS&sup2; scores for all candidate sets computed using the <a href="https://ipb-halle.github.io/MetFrag/projects/metfragcl/">MetFrag software</a>.</p> </li> <li> <p>sirius_scores.tar: MS&sup2; scores for all candidates and measured spectra using the <a href="https://bio.informatik.uni-jena.de/software/sirius/">SIRIUS software</a>.</p> </li> <li> <p>sirius_inputs.tar: Input (ms-files) for the SIRIUS software.</p> </li> <li> <p>DB_README.md: Description of each table in the &quot;massbank.sqlite&quot; SQLite DB.</p> </li> <li> <p>db_processing_scripts.tar: Scripts to re-produce the &quot;massbank.sqlite&quot; and a README.md providing further information on the process.</p> </li> <li> <p>massbank__2020.11__v0.6.1.sqlite: Base SQLite DB from which the &quot;massbank.sqlite&quot; was build up. It was created using the &quot;<a href="https://github.com/bachi55/massbank2db">massbank2db</a>&quot; (v0.6.1) Python package using the <a href="https://github.com/bachi55/MassBank-data/tree/2020.11-branch">MassBank release 2020.11</a>.</p> </li> <li> <p>substructure_fingerprints.tar: Pre-computed substructure counting fingerprints for all candidates related to our experiments.</p> </li> </ul> <p><strong>Instructions:</strong></p> <p>The &quot;massbank.sqlite&quot; can be directly used with the Structure Support Vector Machine Model (SSVM) described in the manuscript and implemented in the &quot;<a href="https://github.com/aalto-ics-kepaco/msms_rt_ssvm">ssvm</a>&quot; Python package.</p> <p>If desired, the database can be re-produced using the scripts provided in &quot;db_processing_scripts.tar&quot;:</p> <ol> <li>Create a directory for all data</li> <li>Download and extract the ... <ol> <li>Processing scripts</li> <li>MS&sup2; scorer outputs (e.g. metfrag.tar)</li> <li>Pre-computed substructure fingerprints</li> </ol> </li> <li>Follow the instructions given in the &quot;README.md&quot; of the &quot;db_processing_scripts.tar&quot;</li> </ol>

opencc-by-4.0Jan 2022View details →
zenodo40/100

ePharmaLib: A Versatile Library of e-Pharmacophores to Address Small-Molecule (Poly-)Pharmacology

<p><em><strong>The peer-reviewed publication for this dataset has now been published&nbsp;in&nbsp;Journal of Chemical Information and Modeling, and can be accessed here:&nbsp;<a href="https://doi.org/10.3390/epidemiologia2030024">https://doi.org/10.1021/acs.jcim.1c00135</a>. Please cite this when using the dataset.</strong></em></p> <p><em>Reverse pharmacophore screening (parallel screening) is an efficient and cost-effective computational method used to study the polypharmacology of drugs. To this end, we created ePharmaLib: a collection of 15,148 energetically optimized, structure-based pharmacophores (e-pharmacophores) , with 3 to 8 features constituting 12.6%, 17.9%, 20.9%, 17.1%, 10.6% and 20.9%, respectively. The pharmacophores were generated from the 3D macromolecular structures of druggable proteins in complex with diverse ligands, retrieved from the sc-PDB database (<a href="http://bioinfo-pharma.u-strasbg.fr/scPDB/">http://bioinfo-pharma.u-strasbg.fr/scPDB/</a>). ePharmaLib can either be used with the Schr&ouml;dinger&rsquo;s PHASE program (<a href="https://www.schrodinger.com/products/phase">https://schrodinger.com/products/phase</a>) or PHARAO, also known as Align-it (<a href="https://silicos-it.be.s3-website-eu-west-1.amazonaws.com/software/align-it/1.0.4/align-it.html">https://silicos-it.be.s3-website-eu-west-1.amazonaws.com/software/align-it/1.0.4/align-it.html</a>). &nbsp;Designed for drug discovery research, this ready-to-use library could dramatically expedite drug discovery by revealing novel molecular interactions of drugs in an efficient and cost-effective manner.</em></p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

A Data Resource for Prediction of Thermodynamic Properties of Small Molecules

<p>We developed a database of 2869 experimental values of enthalpy of formation and 1403 values for entropy for substances composed of stable small molecules, derived from the literature. We developed a model for predicting enthalpy of formation and entropy from semiempirical quantum mechanical calculations of energy and atom counts, and applied the model to a comprehensive database of 16,417 small molecules. The database of small-molecule thermodynamic properties will be useful for predicting the outcome of any process that might involve the generation or destruction of volatile products, such as atmospheric chemistry, volcanism, or waste pyrolysis. Additionally, the collected experimental thermodynamic values will be of value to others developing models to predict enthalpy and entropy.</p>

opencc-by-saMar 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record