Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
40
datasets available to search
ShareScore release 0.9.0
Dataset results
40 results for “Organic molecules”
AIMEl-DB: Atomic Properties for 44K small organic molecules
<h3>AIMEl-DB: Atomic Properties for 44K small organic molecules</h3> <p>This dataset comprises atomic properties of 44K (44 470) molecules selected from the QM9 database. The file names are based on the same indexing system used for QM9. </p> <p>This dataset includes four types of files:</p> <ul> <li><strong>.com Files<br></strong>Input files for Gaussian 16. Simple-point energy calculations were carried out using the keywords<br><code># B3LYP/6-31G(2df,p) scf=(maxcycle=9999) nosymm output=wfx</code><br><br></li> <li><strong>.log Files<br></strong>Output files from Gaussian 16 calculation with the aformentioned parameters.<br><br></li> <li><strong>.wfx Files<br></strong>Wave function files from Gaussian 16 calculation. These files were used as inputs for QTAIM calculations. <br><br></li> <li><strong>.sumviz Files<br></strong>Output file from AIMAll software. The keywords used for the calculations were<br><code>aimqb -nogui -scp=false -nproc=8 -naat=4 input.wfx</code><br>Each .sumviz file contains more than 30 properties based on the Quantum Theory of Atoms in Molecules (QTAIM).<br><br></li> <li><strong>.csv Files<br></strong>These files contain the results of a in-house treament of .sumviz data. They cointain two calculated atomic properties:<br><br> <ol> <li>Total magnitude of the dipole moment, |mu|</li> <li>Total magnitude of the quadrupole moment, |Q|</li> </ol> </li> </ul> <p> and two extracted atomic properties:<br><br> 3. Electronic Population, N<br> 4. Atomic Energy, E</p> <p> </p> <p>The <code>aimel_merged_44k.csv</code> presents the concatenation of the 44 470 <strong>csv Files. </strong></p> <p>Additionaly, the <code>aimel_merged_38k.csv</code> presents the concatenation of the 38 876 <strong>csv Files. </strong>This file corresponds to the version 1.0 of the dataset. </p> <p><br>If you find this dataset useful, please cite the original paper:</p> <p>Meza-González, B., Ramírez-Palma, D.I., Carpio-Martínez, P. <em>et al.</em> Quantum Topological Atomic Properties of 44K molecules. <em>Sci Data</em> <strong>11</strong>, 945 (2024). https://doi.org/10.1038/s41597-024-03723-0</p> <p> </p> <p> </p>
Dataset supporting the paper "Doublet-Singlet-Doublet Transition in a Single Organic Molecule Magnet On-Surface Constructed with up to 3 Aluminum Atoms. Nano Letters 21, 8317 (2021)"
<p>Dataset corresponding to theoretical calculations in the paper "Doublet-Singlet-Doublet Transition in a Single Organic Molecule Magnet On-Surface Constructed with up to 3 Aluminum Atoms" Nano Letters 21, 8317 (2021), <a href="https://doi.org/10.1021/acs.nanolett.1c02881">https://doi.org/10.1021/acs.nanolett.1c02881</a></p> <p>List of files:</p> <p>Several folders corresponding to the figures of the paper. They contain:</p> <ul> <li>.siesta files: STM images in WsXM format (http://www.wsxm.eu/) simulated using STMpw (<a href="https://doi.org/10.5281/zenodo.3581159">https://doi.org/10.5281/zenodo.3581159</a>).</li> <li>CONTCAR and POSCAR files: relaxed structures in VASP format. They can be visualized with VESTA (<a href="https://jp-minerals.org/vesta/en/">https://jp-minerals.org/vesta/en/</a>).</li> <li>.agr: grace files (<a href="https://plasma-gate.weizmann.ac.il/Grace/">https://plasma-gate.weizmann.ac.il/Grace/</a>).<br> </li> </ul>
Laboratory simulations of benzene oxidation and formation of highly oxygenated organic molecules (HOM)
<p>This dataset supplements the following manuscript:<br> Garmash, O., Rissanen, M. P., Pullinen, I., Schmitt, S., Kausiala, O., Tillmann, R., Percival, C., Bannan, T. J., Priestley, M., Hallquist, Å. M., Kleist, E., Kiendler-Scharr, A., Hallquist, M., Berndt, T., McFiggans, G., Wildt, J., Mentel, T., and Ehn, M.: Multi-generation OH oxidation as a source for highly oxygenated organic molecules from aromatics, Atmos. Chem. Phys. Discuss., https://doi.org/10.5194/acp-2019-582, in review, 2019.<br> It presents data from Table 1, Tables S1-S4 and Figures 5, A1 and A2, including model input data.</p>
London Dispersion Governs the Interaction Mechanism of Small Polar and Non-Polar Molecules in Metal-Organic Frameworks
<p>Raw data set relating to publication.</p>
QM7-X: A comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules
<p>Here, we introduce QM7-X, a comprehensive dataset of > 40 physicochemical properties for ~4.2 M equilibrium and non-equilibrium structures of small organic molecules with up to seven non-hydrogen (C, N, O, S, Cl) atoms. To span this fundamentally important region of chemical compound space (CCS), QM7-X includes an exhaustive sampling of (meta-)stable equilibrium structures---comprised of constitutional/structural isomers and stereoisomers, e.g., enantiomers and diastereomers (including cis-trans-and conformational isomers)---as well as 100 non-equilibrium structural variations thereof to reach a total of ~4.2 M molecular structures. Computed at the tightly converged quantum-mechanical PBE0+MBD level of theory, QM7-X contains global (molecular) and local (atom-in-a-molecule) properties ranging from ground state quantities (such as atomization energies and dipole moments) to response quantities (such as polarizability tensors and dispersion coefficients). By providing a systematic, extensive, and tightly converged dataset of quantum-mechanically computed physical and chemical properties, we expect that QM7-X will play a critical role in the development of next-generation machine-learning based models for exploring greater swaths of CCS and performing <em>in silico</em> design of molecules with targeted properties.</p> <p>The dataset is provided in eight HDF5 based files (compressed in .XZ files). One can also find here a README file with technical usage details and examples of how to access the information stored in the dataset (see createDB.py). </p> <p>*The paper explaining the generation of data stored in QM7-X can be found in <em>Sci Data</em> 8, 43 (2021). DOI: 10.1038/s41597-021-00812-2 . arXiv: https://arxiv.org/abs/2006.15139 .</p>
Calculated state-of-the art results for solvation and ionization energies of thousands of organic molecules relevant to battery design
<p>This dataset presents molecular properties critical for battery electrolyte design, specifically solvation energies, ionization potentials, and electron affinities. The dataset is intended for use in machine learning model testing and algorithm validation. The properties calculated include solvation energies using the COSMO-RS method [1] and ionization potentials and electron affinities using various high-accuracy computational methods as implemented in MOLPRO [2]. Computational details can be found in Ref. [3], with scripts used to generate the data mostly uploaded to our github repository [4].</p> <p>Molecular Datasets Considered:</p> <ul> <li> <p>QM9 Dataset: Contains small organic molecules broadly relevant for quantum chemistry [5]</p> </li> <li> <p>Electrolyte Genome Project (EGP): Focuses on materials relevant to electrolytes.[6]</p> </li> <li> <p>GDB17 and ZINC databases: Offer a broad chemical diversity with potential application in battery technologies. [7, 8]</p> </li> </ul> <h2>Data structure</h2> <p>How to Load the Data:</p> <p>All files can be loaded with</p> <p><br><code>import json</code></p> <p><code>with open("file.json", "r") as f:</code><br><code> data_dict = json.load(f)</code></p> <p><br>and the filestructure can be explored with</p> <p><code>data_dict.keys()</code></p> <p>We have also added an example script in python that shows how to extract all data from the JSON files following this link</p> <p><a href="https://github.com/chemspacelab/VienUppDa/blob/main/SolQuest/BIG_MAP_DATA/load_db.py">How to extract the data</a></p> <p>Note the file structure of the the AMONS JSON files is slightly different as explained below!</p> <h3>Solvation energies</h3> <p>The data is stored in two types of JSON archives: files for full molecules of GDB17 and ZINC and files for amons of GDB17 and ZINC. They are structured differently as amon entries are sorted by the number of heavy atoms in the amon (e.g., all amons with 3 heavy atoms are stored in <code>ni3</code>). Because of the large number of amons with 6 or 7 heavy atoms,they are further split into <code>ni6_1</code>, <code>ni6_2</code>, and so on. A sub dictionary of an amon dictionary or a full molecule dictionary contains the following keys:</p> <p><code>ECFP</code> - ECFP4 representation vector</p> <p><code>SMILES</code> - SMILES string</p> <p><code>SYMBOLS</code> - atomic symbols</p> <p><code>COORDS</code> - atomic positions in Angstrom</p> <p><code>ATOMIZATION</code> - atomization energy in [kcal/mol]</p> <p><code>DIPOLE</code> - dipole moment in Debye</p> <p><code>ENERGY</code> - energy in Hartree</p> <p><code>SOLVATION</code> - solvation energy in [kcal/mol] for different solvents at 300 K.</p> <p> </p> <p>Files:</p> <p> </p> <p><strong><em><code>GDB17.json.zip</code> </em></strong>(unpack with unzip first with unzip <strong><em><code>GDB17.json.zip</code></em></strong>) - subset of GDB17 random molecules</p> <p><strong><em><code>AMONS_ZINC.json</code> </em></strong>-<strong><em> </em></strong>all<strong><em> </em></strong>amons of ZINC up to 7 heavy atoms</p> <p><strong><em><code>EGP.json</code> </em></strong>- EGP molecules</p> <p><code><strong><em>AMONS_GDB17.json</em></strong></code> - all amons of GDB17 up to 7 heavy atoms</p> <p><code><strong>QM9IPEA_raw_molpro_output</strong>.zip</code> - compressed folder with raw Molpro input and output files</p> <table> <tbody> <tr> <td><strong>File Name</strong></td> <td><strong>Description </strong></td> <td><strong>Molecules</strong></td> </tr> <tr> <td>AMONS_GDB17.json</td> <td>GDB17 amons</td> <td>37860</td> </tr> <tr> <td>AMONS_ZINC.json</td> <td>ZINC amons </td> <td>88771</td> </tr> <tr> <td>GDB17.json</td> <td>Subset of GDB17</td> <td>309468</td> </tr> <tr> <td>EGP.json </td> <td>EGP molecules </td> <td>18362</td> </tr> </tbody> </table> <p>Atomic energies $E_{at}$ at BP and def2-TZVPD level in Hartree [Ha]</p> <table> <tbody> <tr> <td><strong>Element</strong></td> <td><strong>H</strong></td> <td><strong>C</strong></td> <td><strong>N</strong></td> <td><strong>O</strong></td> <td><strong>F</strong></td> <td><strong>Br</strong></td> <td><strong>Cl</strong></td> <td><strong>S</strong></td> <td><strong>P</strong></td> </tr> <tr> <td>Eat [Ha]</td> <td>-0.5</td> <td> -37.85</td> <td> -54.60</td> <td> -75.09</td> <td>-99.77</td> <td>-2574.40</td> <td> -460.20</td> <td> -398.16</td> <td>-341.30</td> </tr> </tbody> </table> <p> </p> <table> <tbody> <tr> <td><strong>B</strong></td> <td><strong>Si</strong></td> </tr> <tr> <td> -24.65</td> <td> -289.40</td> </tr> </tbody> </table> <p>We follow the convention of negative atomization energies for stablity compared to the isolated atoms:</p> <p>$E_{atomization} = E_{mol} - \sum_{i} E_{at,i}$</p> <p><br>Free energy of solvation at 300 K in [kcal/mol]:</p> <h3>Ionization potentials and electron affinities</h3> <p>The upload contains two JSON files, <strong><em>QM9IPEA.json</em></strong> and <strong><em>QM9IPEA_atom_ens.json</em></strong>. <strong><em>QM9IPEA.json </em></strong>summarizes MOLPRO calculation data grouping it along the following dictionary keys:</p> <p> </p> <p><strong>QM9IPEA.json</strong></p> <p><code>COORDS</code> atom coordinates in Angstroms<br><code>SYMBOLS</code> atom element symbols<br><code>ENERGY</code> total energies for each charge (0, -1, 1) and method considered<br><code>CPU_TIME</code> CPU times (in seconds) spent at each step of each part of the calculation<br><code>DISK_USAGE</code> highest total disk usage in GB<br><code>ATOMIZATION_ENERGY</code> atomization energy at charge 0 (all methods)<br><code>IONIZATION_ENERGY</code> ionization energy for all methods<br><code>ELECTRON_AFFINITY</code> electron affinity for all methods<br><code>HOMO_ENERGY</code> HOMO energy from DFHF calculations<br><code>LUMO_ENERGY</code> LUMO energy from DFHF calculations<br><code>QM9_ID</code> ID of the molecule in the QM9 dataset</p> <p><strong>QM9IPEA_atom_ens.json</strong></p> <p><code>SPINS</code> the spin assigned to elements during calculations of atomic energies<br><code>ENERGY</code> energies of atoms using different methods</p> <p> </p> <p> </p> <p>All energies are given in Hartrees with NaN indicating the calculation failed to converge. Ionization potentials and electron affinities can be recovered as energy differences between neutral and charged (+1 for ionization potentials, -1 for electron affinities) species.</p> <p>"CPU_time" entries contain steps corresponding to individual method calculations, as well as steps corresponding to program operation: "INT" (calculating integrals over basis functions relevant for the calculation), "FILE" (dumping intermediate data to restart file), and "RESTART" (importing restart data). The latter two steps appeared since we reused relevant integrals calculated for neutral species in charged species' calculations; we also used restart functionality to use HF density matrix obtained for the neutral species as the initial density matrix guess for the SCF-HF calculation for charged species. NaN CPU time value means the step was not present or that the calculation is invalid. Note that the CPU times were measured while parallelizing on 12 cores and were not adjusted to single-core.</p> <p><strong> </strong></p> <p><strong><em>QM9IPEA_atom_ens.json</em></strong> contains atomic energies used to calculate atomization energies in <strong><em>QM9IPEA.json</em></strong>, the dictionary keys are:</p> <p><code>SPINS</code> - the spin assigned to elements during calculations of atomic energies.</p> <p><code>ENERGY</code> - energies of atoms using different methods.</p> <p> </p> <p>(Note that H has only one electron and thus does not require a level of theory beyond Hartree-Fock.)</p> <p>NOTE: Additional calculations were performed between publication of arXiv:2308.11196 and creation of this upload. For the version of the dataset used in the manuscript, please refer to DOI:10.5281/zenodo.8252498.</p> <h3>Acknowledgement</h3> <p>This project has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No. 957189 (BIG-MAP) and No. 957213 (BATTERY 2030+). O.A.v.L. has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (grant agreement No. 772834). O.A.v.L. has received support as the Ed Clark Chair of Advanced Materials and as a Canada CIFAR AI Chair. O.A.v.L. acknowledges that this research is part of the University of Toronto’s Acceleration Consortium, which receives funding from the Canada First Research Excellence Fund (CFREF). Obtaining the presented computational results has been facilitated using the queueing system implemented at <a href="https://leruli.com">https://leruli.com</a>. The project has been supported by the Swedish Research Council (Vetenskapsrådet), and the Swedish National Strategic e-Science program eSSENCE as well as by computing resources from the Swedish National Infrastructure for Computing (SNIC/NAISS).</p> <p> </p> <h3>References</h3> <p>[1] Klamt, A.; Eckert, F. COSMO-RS: a novel and efficient method for the a priori prediction of thermophysical data of liquids. Fluid Phase Equilibria 2000, 172, 43–72</p> <p>[2] Werner, H.-J.; Knowles, P. J.; Knizia, G.; Manby, F. R.; Schutz, M. Molpro: a general-purpose quantum chemistry program package. WIREs Comput. Mol. Sci. 2012, 2, 242–253</p> <p>[3] arxiv link of draft</p> <p>[4] <a href="https://github.com/chemspacelab/ViennaUppDa">https://github.com/chemspacelab/ViennaUppDa</a></p> <p>[5] Ramakrishnan, R.; Dral, P. O.; Rupp, M.; von Lilienfeld, O. A. Quantum chemistry structures and properties of 134 kilo molecules. Sci. Data 2014, 1, 140022</p> <p>[6] Qu, X.; Jain, A.; Rajput, N. N.; Cheng, L.; Zhang, Y.; Ong, S. P.; Brafman, M.; Mag- inn, E.; Curtiss, L. A.; Persson, K. A. The Electrolyte Genome Project: A big data approach in battery materials discovery. Comput. Mater. Sci. 2015, 103, 56–67</p> <p><strong> </strong>[7] Ruddigkeit, L.; van Deursen, R.; Blum, L. C.; Reymond, J.-L. Enu- meration of 166 Billion Organic Small Molecules in the Chemical Universe Database GDB-17. Journal of Chemical Information and Modeling 2012, 52, 2864–2875</p> <p>[8] Irwin, J. J.; Shoichet, B. K. ZINC A Free Database of Commercially Available Compounds for Virtual Screening. Journal of Chemical Information and Modeling 2005, 45, 177–182.</p>
Data for "Experimental investigation into the volatilities of highly oxygenated organic molecules (HOM)"
<p>The data used in the preparation of the manuscript "Experimental investigation into the volatilities of highly oxygenated organic molecules (HOM)". The data consists of the data in each of the figures, as well as the time series for measured ozone, alpha-pinene, NOx, NO, condensation sink, temperature, relative humidity, aerosol mass concentration for organics, sulfate and ammonium, as well as the high resolution fitted compounds and unit mass resolution sticks from the CI-APi-TOF.</p>
Dataset from 'Influence of anthropogenic emissions on the composition of highly oxygenated organic molecules in Helsinki: a street canyon and urban background station comparison'
<p>This dataset supplements the following manuscript:</p> <p>Okuljar, M., Garmash, O., Olin, M., Kalliokoski, J., Timonen, H., Niemi, J. V., Paasonen, P., Kontkanen, J., Zhang, Y., Hellén, H., Kuuluvainen, H., Aurela, M., Manninen, H. E., Sipilä, M., Rönkkö, T., Petäjä, T., Kulmala, M., Dal Maso, M., and Ehn, M.: Influence of anthropogenic emissions on the composition of highly oxygenated organic molecules in Helsinki: a street canyon and urban background station comparison, EGUsphere [preprint], https://doi.org/10.5194/egusphere-2023-524, 2023.</p>
Data from: Pushing Raman spectroscopy over the edge: purported signatures of organic molecules in fossil animals are instrumental artefacts
<p>Widespread preservation of fossilized biomolecules in many fossil animals has recently been reported in six studies, based on Raman microspectroscopy. Here, we show that the putative Raman signatures of organic compounds in these fossils are actually instrumental artefacts resulting from intense background luminescence. Raman spectroscopy is based on the detection of photons scattered inelastically by matter upon its interaction with a laser beam. For many natural materials, this interaction also generates a luminescence signal that is often orders of magnitude more intense than the light produced by Raman scattering. Such luminescence, coupled with the transmission properties of the spectrometer, induced quasi-periodic ripples in the measured spectra that have been incorrectly interpreted as Raman signatures of organic molecules. Although several analytical strategies have been developed to overcome this common issue, Raman microspectroscopy as used in the studies questioned here cannot be used to identify fossil biomolecules.</p>
kefisher98/IP_EA_deltaSCF: Ionization Potential, Electron Affinity, and Delta SCF for Small Organic Molecules
<p>Ionization potential, electron affinity, and delta SCF for small organic molecules selected from the ANI-1 data set. Properties are calculated with 24 different density functional approximations as well as CCSD(T).</p>
Data-Driven Discovery of Carbonyl Organic Electrode Molecules: Machine Learning and Experiment
<p>Prediction model of organic molecular electrode</p>
Reduction of supported metal nanoparticles via oxidation of organic molecules: a strategy for nanoparticle redispersion
Open the record for dataset details and reuse information.
Effect of Guest Molecules on the Stacking Configuration of Covalent Organic Frameworks: A Periodic Energy Decomposition Analysis
<p>Illustrative videos that provide an in-depth understanding of the methodology and the result obtained from our recent study. A full description of the method is found in the paper. The potential energy surface was plotted using imshow Matplotlib and the 3d structures were rendered using povray. The videos are all free to use for educational purposes but please remember to cite the paper at anytime you use the videos. </p> <p> </p>
Organizing Structural Principles of the Interleukin-17 Ligand-Receptor Axis - Single molecule tracking - raw data
<p>This dataset contains the raw image data that was analyzed in the manuscript "Organizing Structural Principles of the Interleukin-17 Ligand-Receptor Axis"</p>
Organizing Structural Principles of the Interleukin-17 Ligand-Receptor Axis - Single molecule tracking - raw data - calibration images
<p>This dataset contains the images used for channel calibration for the single molecule data that was analyzed in the manuscript "Organizing Structural Principles of the Interleukin-17 Ligand-Receptor Axis"</p>
Particle-Phase Uptake and Chemistry of Highly Oxygenated Organic Molecules (HOMs) from α-Pinene OH Oxidation
<p>Secondary organic aerosol (SOA) forms a major part of the tropospheric submicron particle mass. Still, the exact formation mechanisms of SOA have remained elusive. It is now admitted that highly oxygenated organic molecules (HOMs) can contribute to a large fraction of SOA formation. In this study, we performed a set of chamber experiments to investigate the SOA formation, and the HOMs uptake and processing directly formed by OH-radical initiated oxidation of α-pinene under two different aerosol seed conditions. Numerous HOM compounds were identified using advanced online and offline analytical techniques and grouped into four classes according to their different uptake behaviors. For the first time, individual HOMs uptake coefficients ranging from 1.1×10<sup>-2</sup> to 1.5×10<sup>-1</sup> were experimentally determined and analyzed using a resistance model which considers uptake limitations by individual gas- and/or particle-phase processes. This study demonstrates that the uptake coefficients of HOMs strongly depend on their molar mass and their respective O/C ratio. Results show that aerosol seed composition and phase state affect the initial uptake of HOMs. Furthermore, the study demonstrates that the acidity and/or different seed phase-state can significantly enhance the subsequent uptake through occurring acidity-driven reactions reflected in a reactive behavior, particularly under (NH<sub>4</sub>)HSO<sub>4</sub> seed conditions, promoting up to 3 times a higher SOA mass formation including the formation of highly-oxidized organosulfates (HOOS). Overall, the present study implies that HOMs and their subsequent chemical processing can play an important role in both the early growth of newly formed particles and SOA formation when particle acidity is high.</p>
DFT torsiondrive data for: MACE-OFF23: Transferable Machine Learning Force Fields for Organic Molecules
<p>MACE-OFF23: Transferable Machine Learning Force Fields for Organic Molecules</p> <div><a href="https://arxiv.org/search/physics?searchtype=author&query=Kov%C3%A1cs,+D+P">Dávid Péter Kovács</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Moore,+J+H">J. Harry Moore</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Browning,+N+J">Nicholas J. Browning</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Batatia,+I">Ilyes Batatia</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Horton,+J+T">Joshua T. Horton</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Kapil,+V">Venkat Kapil</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Witt,+W+C">William C. Witt</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Magd%C4%83u,+I">Ioan-Bogdan Magdău</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Cole,+D+J">Daniel J. Cole</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Cs%C3%A1nyi,+G">Gábor Csányi </a><a href="https://doi.org/10.48550/arXiv.2312.15211">https://doi.org/10.48550/arXiv.2312.15211</a></div> <p> </p> <p>Supporting data including raw outputs from SPICE consistent torsion drives on the TorsionNet500 and OpenFF Biaryl datasets and HDF5 versions formated to be consistent with the rest of the SPICE dataset. See the <a href="../records/10975225">SPICE release</a> for more details. </p>
QM9-XAS database of 56k QM9 small organic molecules labeled with TDDFT X-ray absorption spectra
<p>Database for training graph neural network (GNN) models in <strong>Integrating Explainability into Graph Neural Network Models for the Prediction of X-ray Absorption Spectra, </strong>by Amir Kotobi, Kanishka Singh, Daniel Höche, Sadia Bari, Robert H.Meißner, and Annika Bande.</p> <p><strong>Included:</strong></p> <ul> <li>qm9_Cedge_xas_56k.npz: the TDDFT XAS spectra of 56k structures from the QM9 dataset, were employed to label the graph dataset. The dataset contains two pairs of key/value entries: <strong>spec_stk</strong>, which represents a 2D array containing energies and oscillator strengths of XAS spectra, and <strong>id</strong>, which consists of the indices of QM9 structures. This data was used to create the QM9-XAS graph dataset.</li> <li>qm9xas_orca_output.zip: the raw ORCA output of TDDFT calculations for the 56k QM9-XAS dataset consists of excitation energies, densities, molecular orbitals, and other relevant information. This unprocessed output serves as a source to derive ground truth data for explaining the predictions made by GNNs.</li> <li>qm9xas_spec_train_val.pt: processed graph train/validation dataset of 50k QM9 structures. It is used as input to GNN models for training and validation.</li> <li>qm9xas_spec_test.pt: processed graph test dataset of 6k QM9 structures. It is used to test the performance of trained GNN models.</li> </ul> <p><strong>Notes on the datasets:</strong></p> <ul> <li>The QM9-XAS dataset was created using ORCA electronic structure package [Neese, F., WIREs Computational Molecular Science 2012, 2, 73–78] to calculate carbon K-edge XAS spectra with the time-dependent density functional theory (TDDFT) method [Petersilka, M.; Gossmann, U. J.; Gross, E. K. U., Phys. Rev. Lett. 1996, 76, 1212–1215]</li> <li>The molecular structures of QM9-XAS datasets were sourced from the QM9 database [R. Ramakrishnan, P. O. Dral, M. Rupp, and O. A. Von Lilienfeld, <em>Sci. Data</em> 1, 1 (2014)].</li> </ul> <p><strong>Funding:</strong></p> <p>This<strong> </strong>research was funded by HIDA Trainee Network program, HAICU, Helmholtz AI-4-XAS, DASHH and HEIBRiDS graduate schools. For theoretical calculations and model training, computational resources at DESY and JFZ were used. </p>
Data from: Pushing Raman spectroscopy over the edge: purported signatures of organic molecules in fossil animals are instrumental artefacts
Open the record for dataset details and reuse information.
Fragment and torsion biasing algorithms for construction of small organic molecules in proteins using DOCK
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.