Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

598

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

598 results for “Small molecules”

Learn how ShareScore rates datasets ↗
zenodo36/100

Data supporting 'Small molecule and cell contact-inducible systems for controlling expression and differentiation in stem cells'

<p>Data supporting Soliman et al. 2024. Manuscript describes data collection practices and experimental design.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

An Integrated Machine Learning Approach Delineates an Entropic Expansion Mechanism for the Binding of a Small Molecule to α-Synuclein

<p>The <a href="https://zenodo.org/uploads/14177307" target="_blank" rel="noopener noreferrer">TRAJECTORY.zip</a> contains two trajectories. One for the apo simulation (traj_mw_c4.xtc) and the other for the water simulation in .gro format (traj.gro). The &nbsp;run1.tpr is the binary file for running the apo simulation in gromacs. For water simulation, the raw DOSPT files are also included to calculate the entropy.</p> <p>The <a href="https://zenodo.org/uploads/14177307" target="_blank" rel="noopener noreferrer">figure_raw_data.zip</a> file contains the raw data that was used for plotting the figures in the main text and in the supplementary material.</p> <p>For any further inquiries or requests for additional data, please contact&nbsp;<a rel="noopener">jmondal@tifrh.res.in</a>.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Automated identification of small molecules in cryo-electron microscopy data with density- and energy-guided evaluation

<p>Data and turorials for ligand identification with EMERALD-ID. Each Zip file contains ligand models and data used to generate their respective figures in the EMERALD-ID manuscript. For specific details, consult the README file in each directory.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Research Data Supporting "Electrostatic co-assembly of nanoparticles with oppositely charged small molecules into static and dynamic superstructures"

<p>This repository contains the set of data shown in the paper &quot;<strong>Electrostatic co-assembly of nanoparticles with oppositely charged small molecules into static and dynamic superstructures</strong>&quot;, published on <strong>Nature Chemistry </strong>(DOI: 10.1038/s41557-021-00752-9).</p> <p>The files in the folders are organized as follow:</p> <p><strong>AA_models/ : </strong>contains all the files needed to run the Atomistic simulations discussed in the paper (including the topologies and starting configurations).</p> <p><strong>CG_models/ : </strong>contains all the files needed to run the Coarse Grained simulations discussed in the paper (including the topologies and starting configurations).</p> <p><strong>citrate_analysis/ : </strong>contains all the files needed to reproduce the analysis of the CV and SOAP+PCA+PAMM (including input files and python scripts).</p> <p>Additional details are available in the methods section and in the Supporting Information of the main paper.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Result files (ONLYSTEREO): "Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data"

<p>Result files associated with the publication: &quot;<strong>Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data</strong>&quot; by Bach et al.</p> <p>The following files are included in the archive:</p> <ul> <li>Raw max-marginal predictions using LC-MS&sup2;Struct for all LC-MS&sup2; experiments of the ONLYSTEREO setup</li> <li>Averaged max-marginals for the LC-MS&sup2;Struct over all SSVM models</li> <li>Ranks for the ground-truth structures predicted by Only MS&sup2; and LC-MS&sup2;Struct (molecule class analysis)</li> </ul> <p>Instructions:</p> <ul> <li>clone the repository containing the experimental scripts and analysis notebooks: <a href="https://github.com/aalto-ics-kepaco/lcms2struct_exp">https://github.com/aalto-ics-kepaco/lcms2struct_exp</a></li> <li>download the archive in this repository</li> <li>unpack the archive in the git-repository root directory</li> <li>follow the instructions given in the <a href="https://github.com/aalto-ics-kepaco/lcms2struct_exp/blob/main/README.md">README.md</a> of the git-repository to reproduce the figures, etc.</li> </ul>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Result files (ALLDATA): "Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data with LC-MS²Struct"

<p>Result files associated with the publication: &quot;<strong>Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data</strong>&quot; by Bach et al.</p> <p>The following files are included in the archive:</p> <ul> <li>Raw max-marginal predictions using LC-MS&sup2;Struct for all LC-MS&sup2; experiments of the ALLDATA setup</li> <li>Averaged max-marginals for the LC-MS&sup2;Struct over all SSVM models</li> <li>Top-k accuracies for the comparison methods (MS&sup2;+RT, ...)</li> <li>Ranks for the ground-truth structures predicted by Only MS&sup2; and LC-MS&sup2;Struct (molecule class analysis)</li> </ul> <p>Instructions:</p> <ul> <li>clone the repository containing the experimental scripts and analysis notebooks: <a href="https://github.com/aalto-ics-kepaco/lcms2struct_exp">https://github.com/aalto-ics-kepaco/lcms2struct_exp</a></li> <li>download the archive in this repository</li> <li>unpack the archive in the git-repository root directory</li> <li>follow the instructions given in the <a href="https://github.com/aalto-ics-kepaco/lcms2struct_exp/blob/main/README.md">README.md</a> of the git-repository to reproduce the figures, etc.</li> </ul> <p>Version history:</p> <ul> <li><strong>Version 1</strong>: Experimental results for &quot;Method comparison&quot; and &quot;Molecule classification analysis&quot; where performed with <strong>2D fingerprints</strong> (<a href="https://www.biorxiv.org/content/10.1101/2022.02.11.480137v1">preprint v1</a>)</li> <li><strong>Version 2</strong> <em>(this version)</em>: Experimental results for &quot;Method comparison&quot; and &quot;Molecule classification analysis&quot; where performed with <strong>3D fingerprints</strong></li> </ul>

opencc-by-4.0Apr 2022View details →
dryad36/100

Data from: Towards high-resolution modeling of small molecule - ion channel interactions

<p>Ion channels are critical drug targets for a range of pathologies, such as epilepsy, pain, itch, autoimmunity, and cardiac arrhythmias. To develop effective and safe therapeutics, it is necessary to design small molecules with high potency and selectivity for specific ion channel subtypes. There has been increasing implementation of structure-guided drug design for the development of small molecules targeting ion channels. We evaluated the performance of two Rosetta ligand docking methods, RosettaLigand and GALigandDock, on structures of known ligand - cation channel complexes. Ligands were docked to voltage-gated sodium (Na<sub>V</sub>), voltage-gated calcium (Ca<sub>V</sub>), and transient receptor potential vanilloid (TRPV) channel families. For each test case, RosettaLigand and GALigandDock methods were able to frequently sample a ligand binding pose within 1-2 Å root mean square deviation (RMSD) relative to the experimental ligand coordinates. However, RosettaLigand and GALigandDock scoring functions cannot consistently identify experimental ligand coordinates as top-scoring models. Our study reveals that the proper scoring criteria for RosettaLigand and GALigandDock modeling of ligand - ion channel complexes should be assessed on a case-by-case basis using sufficient ligand and receptor interface sampling, knowledge about state specific interactions of the ion channel and inherent receptor site flexibility that could influence ligand binding.</p>

opencc-zeroMay 2024View details →
dryad36/100

Data from: Sustained intestinal epithelial monolayer wound closure after transient application of a FAK-activating small molecule

<p>M64HCl, which has drug-like properties, is a water-soluble Focal Adhesion Kinase (FAK) activator that promotes murine mucosal healing after ischemic or NSAID-induced injury. Since M64HCl has a short plasma half-life in vivo (less than two hours), it has been administered as a continuous infusion with osmotic minipumps in previous animal studies. However, the effects of more transient exposure to M64HCl on monolayer wound closure remained unclear. Herein, we compared the effects of shorter M64HCl treatment in vitro to continuous treatment for 24 hours on monolayer wound closure. We then investigated how long FAK activation and downstream ERK1/2 activation persist after two hours of M64HCl treatment in Caco-2 cells. M64HCl concentrations immediately after washing measured by mass spectrometry confirmed that M64HCl had been completely removed from the medium while intracellular concentrations had been reduced by 95%.<strong> </strong>Three-hour and four-hour<strong> </strong>M64HCl (100 nM) treatment promoted epithelial sheet migration over 24 hours similar to continuous 24-hour exposure. 100nM M64HCl did not increase cell number. Exposing cells twice with 2-hr exposures of M64HCl during a 24-hour period had a similar effect. Both FAK inhibitor PF-573228 (10 µM) and ERK kinase (MEK) inhibitor PD98059 (20 µM) reduced basal wound closure in the absence of M64HCl, and each completely prevented any stimulation of wound closure by M64HCl. Rho kinase inhibitor Y-27632 (20 µM) stimulated Caco-2 monolayer wound closure but no further increase was seen with M64HCl in the presence of Y-27632. M64HCl (100 nM) treatment for 3 hours stimulated Rho kinase activity.  M64HCl decreased F-actin in Caco-2 cells. Furthermore, a two-hour treatment with M64HCl (100 nM) stimulated sustained FAK activation and ERK1/2 activation for up to 16 and hours 24 hours, respectively. These results suggest that transient M64HCl treatment promotes prolonged intestinal epithelial monolayer wound closure by stimulating sustained activation of the FAK/ERK1/2 pathway. Such molecules may be useful to promote gastrointestinal mucosal repair even with a relatively short half-life.</p>

opencc-zeroJul 2024View details →
zenodo36/100

Supplemental Data for "Identifying novel variants of small molecules through database search of mass spectra"

<p>Supplemental Dataset 1 contains GNPS dataset information for the large scale search of GNPS vs Pubchem.</p> <p>Supplemental Dataset 2a contains the top scoring exact mode hit for each spectrum against PubChem and COCONUT. Files are split into "*chunk*" files of up to 10 million records each.</p> <p>Supplemental Dataset 2b contains the top scoring variable mode hit for each spectrum against COCONUT.</p> <p>Supplemental Dataset 3 contains mass spectra provided by Waters Corporation to analyze impurities of Imatinib.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Shutting Down Shigella Secretion: Characterizing Small Molecule Type Three Secretion System ATPase Inhibitors

<p>Spa47 inhibition data from &quot;Shutting Down <em>Shigella</em> Secretion: Characterizing Small Molecule Type Three Secretion System ATPase Inhibitors&quot;.</p>

opencc-by-4.0Oct 2018View details →
zenodo36/100

A Novel Integrated Reference-Counter Electrode for Electrochemical Measurements of HOMO and LUMO Levels in Small-Molecule Thin-Film Semiconductors for OLEDs

<p>This dataset <span>includes all raw data, processed data, and analysis scripts necessary to replicate the findings reported in the published paper.</span></p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Small Molecule Agonist Binding at human NOP Receptor Data

<p>This entry contains:</p> <ul> <li>MD input files and scripts to run the simulations of N/OFQ(1-13)-NH2 in complex with the human NOP receptor&nbsp; (PDB ID: 8F7X)</li> <li>Initial docking poses and&nbsp;MD input files and scripts to run the simulations of predicted small molecule agonist in complex with the human NOP receptor&nbsp;<br> <table> <tbody> <tr> <td>Ligand</td> <td>Binding Mode</td> <td>Receptor Structure</td> </tr> <tr> <td><strong>(<em>R</em>)-Ro 65-6570</strong></td> <td>RBM01&nbsp;</td> <td>8F7X</td> </tr> <tr> <td><strong>(<em>R</em>)-Ro 65-6570</strong></td> <td>RBM03 &nbsp;</td> <td>8F7X_mod</td> </tr> <tr> <td><strong>MCOPPB </strong></td> <td>MBM04</td> <td>8F7X_mod</td> </tr> <tr> <td><strong>MCOPPB </strong></td> <td>MBM05</td> <td>8F7X_mod</td> </tr> </tbody> </table> </li> <li>MD topology (.psf) and trajectory (.dcd) files of all five simulated systems (three replicas each)</li> </ul>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Transitive prediction of small molecule function through alignment of high-content screening resources

<p>This dataset supports the development of CLIP&lt;sup&gt;n&lt;/sup&gt;, a contrastive-learning framework designed to align heterogeneous high-content screening (HCS) profile datasets.</p> <p><strong>GitHub link</strong>: https://github.com/AltschulerWu-Lab/CLIPn</p> <h2>Directory Structure</h2> <h3>Data Files</h3> <ul> <li>HCS_datasets.pkl: Contains 13 high-content screening (HCS) datasets from multiple studies across 20 years.</li> <li>Hypoxia.pkl: Contains 8 profile datasets using different assays and treated under diverse hypoxia durations.</li> <li>Expression.pkl: Contains 2 transcriptional profile datasets and 6 image profile datasets for multimodal analysis.</li> </ul> <h3>Folders</h3> <p><strong>raw_profiles</strong>:<br>HCS13/<br>- Contains raw data from 13 high-content screening (HCS) datasets. Each dataset includes meta and feature files.&nbsp;</p> <p>L1000/<br>- CDRP_feature_exp.csv: Raw L1000 expression data from the CDRP dataset.<br>- CDRP_meta_exp.csv: Metadata associated with the CDRP expression data.<br>- LINCS_feature_exp.csv: Raw L1000 expression data from the LINCS dataset.<br>- LINCS_meta_exp.csv: Metadata associated with the LINCS expression data.</p> <p>RxRx3/<br>- RxRx3_feature_final.csv: Profile data from the RxRx3 dataset.<br>- RxRx3_meta_final.csv: Metadata from the RxRx3 dataset.</p> <p>Uncharacterized_compounds/<br>- NCI_cpnData.csv: Feature data for uncharacterized compounds from the NCI dataset.<br>- NCI_cpnInfo.csv: Information about uncharacterized compounds in the NCI dataset.<br>- Prestwick_UTSW_cpnData.csv: Feature data for uncharacterized compounds from the Prestwick UTSW dataset.<br>- Prestwick_UTSW_cpnInfo.csv: Information about uncharacterized compounds from the Prestwick UTSW dataset.</p> <h2><br>Usage</h2> <p><br><br><code>import pickle</code><br><code>with open('data.pkl', 'rb') as f:</code><br><code>&nbsp; &nbsp; data = pickle.load(f)</code></p> <p><code>X = data['X']</code><br><code>y = data['y']</code><br><br></p> <h2>Data Reference</h2> <p><br>For raw datasets from 13 HCS database, data and analysis pipeline for dataset 1 was obtained from https://www.science.org/doi/suppl/10.1126/science.1100709/suppl_file/perlman.som.zip; for datasets 2-3, data were shared by authors; For datasets 4-5, analysis code was downloaded from https://static-content.springer.com/esm/art%3A10.1038%2Fnbt.3419/MediaObjects/41587_2016_BFnbt3419_MOESM21_ESM.zip and data were shared by authors; For datasets 6-7, processed dataset was downloaded from AWS following instructions from https://github.com/carpenter-singh-lab/2022_Haghighi_NatureMethods, and replicate_level_cp_normalized.csv.gz features were used. For project datasets 8-13, datasets and analysis results were downloaded from https://zenodo.org/records/7352487. For RxRx3, dataset was obtained from https://www.rxrx.ai/rxrx3. L1000 transcript datasets were downloaded using the same link as datasets 6-7 and the processed transcript data files (named &ldquo;replicate_level_l1k.csv&rdquo;) were used.&nbsp;</p>

openmit-licenseOct 2024View details →
zenodo36/100

Exploring the Effects of LC Parameters on Retention Indices of Small Molecules using Amine Scaling - Supplementary Data

<p>In this document&nbsp;the retention times, calculated retention indices and calculated delta retention indices can be found&nbsp;that were used for the thesis on <strong>Exploring the Effects of LC Parameters on Retention Indices of Small Molecules using Amine Scaling</strong>, as well as the raw-data files of the performed measurements.</p>

opencc-by-4.0Dec 2022View details →
dryad36/100

Data from: Small molecule inhibitor of tau self-association in a mouse model of tauopathy: A preventive study in P301L tau JNPL3 mice

<p><span class="TextRun SCXW44549199 BCX0"><span class="NormalTextRun SCXW44549199 BCX0">Advances in </span><span class="NormalTextRun SCXW44549199 BCX0">tau biology and </span><span class="NormalTextRun SCXW44549199 BCX0">the </span><span class="NormalTextRun SCXW44549199 BCX0">difficulties of</span><span class="NormalTextRun SCXW44549199 BCX0"> amyloid-directed </span><span class="NormalTextRun SpellingErrorV2Themed SCXW44549199 BCX0">immuno</span><span class="NormalTextRun SpellingErrorV2Themed SCXW44549199 BCX0">therapeutics</span><span class="NormalTextRun SCXW44549199 BCX0"> have heightened interest in tau as a target for </span><span class="NormalTextRun SCXW44549199 BCX0">small molecule </span><span class="NormalTextRun SCXW44549199 BCX0">drug discovery for neurodegenerative diseases. </span><span class="NormalTextRun SCXW44549199 BCX0">Here</span><span class="NormalTextRun SCXW44549199 BCX0">,</span><span class="NormalTextRun SCXW44549199 BCX0"> we </span><span class="NormalTextRun SCXW44549199 BCX0">evaluate</span><span class="NormalTextRun SCXW44549199 BCX0">d</span> <span class="NormalTextRun SCXW44549199 BCX0">OLX-07010</span><span class="NormalTextRun SCXW44549199 BCX0">, a small molecule inhibitor of tau self-association,</span> <span class="NormalTextRun SCXW44549199 BCX0">for the prevention of </span><span class="NormalTextRun SCXW44549199 BCX0">tau aggregat</span><span class="NormalTextRun SCXW44549199 BCX0">ion</span><span class="NormalTextRun SCXW44549199 BCX0">. </span><span class="NormalTextRun SCXW44549199 BCX0">The primary endpoint of the study was </span><span class="NormalTextRun SCXW44549199 BCX0">statistically significant </span><span class="NormalTextRun SCXW44549199 BCX0">reduction of insoluble tau aggregates in treated </span><span class="NormalTextRun SCXW44549199 BCX0">JNPL3 </span><span class="NormalTextRun SCXW44549199 BCX0">mice compared </span><span class="NormalTextRun SCXW44549199 BCX0">with </span><span class="NormalTextRun SCXW44549199 BCX0">V</span><span class="NormalTextRun SCXW44549199 BCX0">ehicle-control</span><span class="NormalTextRun SCXW44549199 BCX0"> mice. </span><span class="NormalTextRun SCXW44549199 BCX0">S</span><span class="NormalTextRun SCXW44549199 BCX0">econdary endpoints were dose-dependent reduction of insoluble tau aggregates, reduction of phosphorylated tau, and reduction of soluble tau.</span> <span class="NormalTextRun SCXW44549199 BCX0">This study was performed in JNPL3 mice, which are representative of inherited forms of 4-repeat tauopathies with the P301L tau mutation</span><span class="NormalTextRun SCXW44549199 BCX0"> (</span><span class="NormalTextRun SpellingErrorV2Themed SCXW44549199 BCX0">eg</span><span class="NormalTextRun SCXW44549199 BCX0">, progressive supranuclear palsy</span><span class="NormalTextRun SCXW44549199 BCX0"> and</span><span class="NormalTextRun SCXW44549199 BCX0"> frontotemporal dementia</span><span class="NormalTextRun SCXW44549199 BCX0">)</span><span class="NormalTextRun SCXW44549199 BCX0">. The P301L mutation makes tau prone to aggregation; therefore, JNPL3 mice present a more challenging target than mouse models of human tau without mutations. </span><span class="NormalTextRun SCXW44549199 BCX0">JNPL3 mice </span><span class="NormalTextRun SCXW44549199 BCX0">were treated </span><span class="NormalTextRun SCXW44549199 BCX0">from 3 to 7 months</span> <span class="NormalTextRun SCXW44549199 BCX0">of</span> <span class="NormalTextRun SCXW44549199 BCX0">age with </span><span class="NormalTextRun SCXW44549199 BCX0">V</span><span class="NormalTextRun SCXW44549199 BCX0">ehicle</span><span class="NormalTextRun SCXW44549199 BCX0">, </span><span class="NormalTextRun AdvancedProofingIssueV2Themed SCXW44549199 BCX0">30 mg</span><span class="NormalTextRun SCXW44549199 BCX0">/kg compound</span><span class="NormalTextRun SCXW44549199 BCX0"> dose</span><span class="NormalTextRun SCXW44549199 BCX0">,</span> <span class="NormalTextRun SCXW44549199 BCX0">or </span><span class="NormalTextRun AdvancedProofingIssueV2Themed SCXW44549199 BCX0">40 mg</span><span class="NormalTextRun SCXW44549199 BCX0">/kg compound</span><span class="NormalTextRun SCXW44549199 BCX0"> dose</span><span class="NormalTextRun SCXW44549199 BCX0">. Biochemical </span><span class="NormalTextRun SCXW44549199 BCX0">methods were used to evaluate self-associated tau, insoluble tau aggregates, total tau</span><span class="NormalTextRun SCXW44549199 BCX0">,</span><span class="NormalTextRun SCXW44549199 BCX0"> and phosphorylated tau in the hindbrain</span><span class="NormalTextRun SCXW44549199 BCX0">,</span><span class="NormalTextRun SCXW44549199 BCX0"> cortex</span><span class="NormalTextRun SCXW44549199 BCX0">,</span><span class="NormalTextRun SCXW44549199 BCX0"> and hippocampus.</span> <span class="NormalTextRun SCXW44549199 BCX0">T</span><span class="NormalTextRun SCXW44549199 BCX0">he </span><span class="NormalTextRun SCXW44549199 BCX0">V</span><span class="NormalTextRun SCXW44549199 BCX0">ehicle group had higher levels of insoluble tau </span><span class="NormalTextRun SCXW44549199 BCX0">in the hindbrain </span><span class="NormalTextRun SCXW44549199 BCX0">than the </span><span class="NormalTextRun SCXW44549199 BCX0">B</span><span class="NormalTextRun SCXW44549199 BCX0">aseline group</span><span class="NormalTextRun SCXW44549199 BCX0">;</span> <span class="NormalTextRun SCXW44549199 BCX0">treatment with </span><span class="NormalTextRun SCXW44549199 BCX0">40 mg/kg</span> <span class="NormalTextRun SCXW44549199 BCX0">compound </span><span class="NormalTextRun SCXW44549199 BCX0">dose prevented this increase. </span><span class="NormalTextRun SCXW44549199 BCX0">In the cortex, t</span><span class="NormalTextRun SCXW44549199 BCX0">he levels of insoluble tau were similar in the </span><span class="NormalTextRun SCXW44549199 BCX0">B</span><span class="NormalTextRun SCXW44549199 BCX0">aseline and </span><span class="NormalTextRun SCXW44549199 BCX0">V</span><span class="NormalTextRun SCXW44549199 BCX0">ehicle </span><span class="NormalTextRun SCXW44549199 BCX0">groups</span><span class="NormalTextRun SCXW44549199 BCX0">,</span> <span class="NormalTextRun SCXW44549199 BCX0">indicating</span><span class="NormalTextRun SCXW44549199 BCX0"> that the pathological phenotype of these mice was beginning to </span><span class="NormalTextRun SCXW44549199 BCX0">emerge</span><span class="NormalTextRun SCXW44549199 BCX0"> at the </span><span class="NormalTextRun SCXW44549199 BCX0">study </span><span class="NormalTextRun SCXW44549199 BCX0">endpoint </span><span class="NormalTextRun SCXW44549199 BCX0">and that</span><span class="NormalTextRun SCXW44549199 BCX0"> the</span><span class="NormalTextRun SCXW44549199 BCX0">re was a delay in the</span><span class="NormalTextRun SCXW44549199 BCX0"> development of the phenotype of the model as originally characterized.</span> <span class="NormalTextRun SCXW44549199 BCX0">No drug-related adverse effects were </span><span class="NormalTextRun SCXW44549199 BCX0">observed</span><span class="NormalTextRun SCXW44549199 BCX0"> during the 4-month treatment period. </span></span><span class="EOP SCXW44549199 BCX0"> </span></p>

opencc-zeroJun 2023View details →
zenodo36/100

QM9-XAS database of 56k QM9 small organic molecules labeled with TDDFT X-ray absorption spectra

<p>Database for training graph neural network (GNN) models in&nbsp;<strong>Integrating Explainability into Graph Neural Network Models for the Prediction of X-ray Absorption Spectra,&nbsp;</strong>by&nbsp;Amir Kotobi,&nbsp;Kanishka Singh,&nbsp;Daniel H&ouml;che,&nbsp;Sadia Bari,&nbsp;Robert H.Mei&szlig;ner,&nbsp;and Annika Bande.</p> <p><strong>Included:</strong></p> <ul> <li>qm9_Cedge_xas_56k.npz: the TDDFT XAS spectra of 56k structures from the QM9 dataset, were employed to label the graph dataset. The dataset contains two pairs of key/value entries: <strong>spec_stk</strong>,&nbsp;which represents a 2D array containing energies and oscillator strengths of XAS spectra, and <strong>id</strong>,&nbsp;which consists of the indices of QM9 structures. This data was used to create the QM9-XAS graph dataset.</li> <li>qm9xas_orca_output.zip: the raw ORCA output of TDDFT calculations for the 56k QM9-XAS dataset consists of excitation energies, densities, molecular orbitals, and other relevant information. This unprocessed output serves as a source to derive ground truth data for explaining the predictions made&nbsp;by&nbsp;GNNs.</li> <li>qm9xas_spec_train_val.pt: processed graph train/validation dataset&nbsp;of 50k QM9 structures. It is used as input to GNN models for training and validation.</li> <li>qm9xas_spec_test.pt: processed graph test dataset&nbsp;of 6k QM9 structures. It is used to test&nbsp;the performance of trained GNN models.</li> </ul> <p><strong>Notes on the datasets:</strong></p> <ul> <li>The QM9-XAS dataset was created using&nbsp;ORCA electronic structure package [Neese, F.,&nbsp;WIREs Computational Molecular Science 2012, 2, 73&ndash;78] to calculate carbon K-edge XAS spectra with the time-dependent density functional theory (TDDFT)&nbsp;method [Petersilka, M.; Gossmann, U. J.; Gross, E. K. U.,&nbsp;Phys. Rev. Lett. 1996, 76, 1212&ndash;1215]</li> <li>The molecular structures of QM9-XAS datasets were sourced from the QM9 database [R. Ramakrishnan, P. O. Dral, M. Rupp, and O. A. Von Lilienfeld,&nbsp;<em>Sci. Data</em>&nbsp;1, 1 (2014)].</li> </ul> <p><strong>Funding:</strong></p> <p>This<strong>&nbsp;</strong>research was funded by&nbsp;HIDA Trainee Network program, HAICU, Helmholtz&nbsp;AI-4-XAS, DASHH and HEIBRiDS graduate schools. For theoretical calculations and model training, computational resources at DESY and JFZ&nbsp;were used.&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Enhanced Mapping of Small Molecule Binding Sites in Cells

<p>Selected docking poses for Wozniak et al manuscript. File name include the structure id (PDBID or Alphafold model) and the probe id, separated by an underscore (i.e., 6tjk_4.pdb, or AF-P10620-F1-model_6.pdb).</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

choderalab/geometry-benchmark-espaloma: Small molecule geometry benchmark dataset to validate espaloma-0.3

<p>This is a collection of preprocessed QM and MM optimized structures needed to perform the small molecule geometry benchmark study, described in the <strong>espaloma-0.3</strong> paper:</p> <p>Kenichiro Takaba, Iv&aacute;n Pulido,&nbsp;Pavan Kumar Behara, Mike Henry, Hugo MacDermott Opeskin, John D. Chodera, Yuanqing Wang.&nbsp;&quot;Machine-learned molecular mechanics force field for the simulation of protein-ligand systems and beyond&quot; (<a href="https://arxiv.org/abs/2307.07085">arXiv:2307.07085</a>)</p> <p>This benchmark study calculates and compares the RMSD, TFD, and ddE metrics for a specified set of MM force fields. The initial optimized structures were&nbsp;sourced from the&nbsp;<a href="https://github.com/openforcefield/qca-dataset-submission/tree/master/submissions/2021-06-04-OpenFF-Industry-Benchmark-Season-1-v1.1">OpenFF Industry Benchmark Season 1 v1.1</a>&nbsp;dataset, which is available through &nbsp;<a href="https://qcarchive.molssi.org/">QCArchive</a>.&nbsp;More details about the preprocessing steps is available at&nbsp;<a href="https://github.com/choderalab/geometry-benchmark-espaloma/tree/main/qc-opt-geo">https://github.com/choderalab/geometry-benchmark-espaloma/tree/main/qc-opt-geo</a>.</p> <ul> <li><strong>02-chunks.tar.gz</strong>:&nbsp;QM optimized structures chunked into small file sizes.</li> <li><strong>02-outputs-openff-2.0.0-espaloma-0.3.0rc1.tar.gz</strong>:&nbsp;MM optimized structures using openff-2.0.0 and espaloma-0.3.0rc1 force field (former release candidate of espaloma-0.3)</li> <li><strong>02-outputs-gaff2.11.tar.gz</strong>:&nbsp;MM optimized structures using gaff-2.11 force field</li> <li><strong>02-outputs-espaloma-0.3.0rc6.tar.gz</strong>:&nbsp;MM optimized structures using espaloma-0.3.0rc6 (espaloma-0.3) force field</li> <li><strong>02-outputs-openff-2.1.0.tar.gz</strong>: MM optimized structures using openff-2.1.0 force field</li> </ul>

opencc-by-4.0Sep 2023View details →
dryad36/100

Knime workflow and data files to calculate a lead-likeness evaluation function 'Λ' for small molecule compounds

<p class="RSCB01COMAbstract"><span class="06CHeading"><span>A simple and rational method to rank molecules' lead-likeness using continuous evaluation functions was developed to support epigenetic drug discovery. This strategy proved to be highly effective on model chemical libraries and finally helped driving synthetic efforts towards candidates of interest for epigenetic applications towards HDAC6, BRD4 and EZH2.</span></span></p> <p>Here we provide sample data and the knime workflow for generatating lead-like chemical probes for epigenetics. This tool and example data set will enable researchers to establish this workflow in their own laboratories and apply it to new synthetic scaffolds for medicinal chemistry.</p>

opencc-zeroOct 2023View details →
ClinicalTrials.gov36/100

Insulin-Sensitizing Anti-Inflammatory Small Molecule for Investigative Treatment of Dementia

ClinicalTrials.gov study NCT05227820. IPD Sharing: Not stated. Countries: 1. Publications: 36.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record