Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

598

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

598 results for “Small molecules”

Learn how ShareScore rates datasets ↗
zenodo40/100

Dataset supporting manuscript entitled 'Partitioning of Small Hydrophobic Molecules into Polydimethylsiloxane in Microfluidic Analytical Devices'

<p>This is the dataset supporting the manuscript &#39;Surface and bulk modifications of polydimethylsiloxane to reduce absorption/adsorption of small molecules in lab-on-a-chip&#39; published in Micromachines</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Trajectories and Code from "Small molecules targeting the disordered transactivation domain of the androgen receptor induce the formation of collapsed helical states" Zhu et al. 2022

<p>Trajectories, GROMACS&nbsp;input files, and analysis code from the manuscript &quot;Small molecules targeting the disordered transactivation domain of the androgen receptor induce the formation of collapsed helical states&quot; Zhu et al. 2022 (Nature Communications, In Press)</p> <p>https://www.biorxiv.org/content/10.1101/2021.12.23.474012v1.abstract</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Radiation Damage in Small Molecule Crystallography - Experiment 6

<p>Sixth data set in a series of experiments investigating the effect of radiation damage to a small molecule crystal structure.</p> <p>Sample: catena-(bis(m2-Glycyl-histidinato-N,N&#39;,O)-nickel(II) heptahydrate)</p> <p>sum formula: C16H36N8NiO13</p> <p>Wavelength: 0.9889&nbsp;Angstrom</p> <p>Temperature: 30 K</p> <p>Flux: 2.627&bull;10<sup>10&nbsp;</sup>ph/s</p> <p>Calculated dose (average DWD) per scan: 1.90&nbsp;MGy</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Radiation Damage in Small Molecule Crystallography - Experiment 5

<p>Fifth data set in a series of experiments investigating the effect of radiation damage to a small molecule crystal structure.</p> <p>Sample: catena-(bis(m2-Glycyl-histidinato-N,N&#39;,O)-nickel(II) heptahydrate)</p> <p>sum formula: C16H36N8NiO13</p> <p>Wavelength: 0.9028&nbsp;Angstrom</p> <p>Temperature: 100K</p> <p>Flux: 6.605&bull;10<sup>9&nbsp;</sup>ph/s</p> <p>Calculated dose (average DWD) per scan: 0.79&nbsp;MGy</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Radiation Damage in Small Molecule Crystallography - Experiment 3

<p>Third&nbsp;data set in a series of experiments investigating the effect of radiation damage to a small molecule crystal structure.</p> <p>Sample: catena-(bis(m2-Glycyl-histidinato-N,N&#39;,O)-nickel(II) heptahydrate)</p> <p>sum formula: C16H36N8NiO13</p> <p>Wavelength: 0.6889 Angstrom</p> <p>Temperature: 100K</p> <p>Flux:&nbsp;4.35&bull;109<sup>&nbsp;</sup>ph/s</p> <p>Calculated dose (average DWD) per scan: 0.32&nbsp;MGy</p>

opencc-by-4.0Jan 2019View details →
zenodo40/100

Radiation Damage in Small Molecule Crystallography - Experiment 8

<p>Eighth&nbsp;data set in a series of experiments investigating the effect of radiation damage to a small molecule crystal structure.</p> <p>Sample: catena-(bis(m2-Glycyl-histidinato-N,N&#39;,O)-nickel(II) heptahydrate)</p> <p>sum formula: C16H36N8NiO13</p> <p>Wavelength: 0.6889 Angstrom</p> <p>Temperature: 120 K</p> <p>Flux: 8.42&bull;109<sup>&nbsp;</sup>ph/s</p> <p>Crystal size: 0.050&nbsp;x 0.010 x 0.010 mm</p> <p>Calculated dose (DWD) per scan: 0.60&nbsp;MGy</p>

opencc-by-4.0Mar 2019View details →
zenodo40/100

Radiation damage in small molecule crystallography - experiment 1

<p>First data set in a series of experiments investigating the effect of radiation damage to a small molecule crystal structure.</p> <p>Sample: catena-(bis(m2-Glycyl-histidinato-N,N&#39;,O)-nickel(II) heptahydrate)</p> <p>sum formula: C16H36N8NiO13</p> <p>Wavelength: 0.6889 Angstrom</p> <p>Temperature: 100 K</p> <p>Flux: 8.8&bull;10<sup>9&nbsp;</sup>ph/s</p> <p>Crystal size: 0.050&nbsp;x 0.010 x 0.010 mm</p> <p>Calculated dose (DWD) per scan: 0.63&nbsp;MGy</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

Radiation Damage in Small Molecule Crystallography - Experiment 4

<p>Fourth data set in a series of experiments investigating the effect of radiation damage to a small molecule crystal structure.</p> <p>Sample: catena-(bis(m2-Glycyl-histidinato-N,N&#39;,O)-nickel(II) heptahydrate)</p> <p>sum formula: C16H36N8NiO13</p> <p>Wavelength: 0.6889 Angstrom</p> <p>Temperature: 100K</p> <p>Flux:&nbsp;8.162&bull;10<sup>9&nbsp;</sup>ph/s</p> <p>Calculated dose (average DWD) per scan: 0.57&nbsp;MGy</p>

opencc-by-4.0Jan 2019View details →
zenodo40/100

Radiation Damage in Small Molecule Crystallography - Experiment 2

<p>Second data set in a series of experiments investigating the effect of radiation damage to a small molecule crystal structure.</p> <p>Sample: catena-(bis(m2-Glycyl-histidinato-N,N&#39;,O)-nickel(II) heptahydrate)</p> <p>sum formula: C16H36N8NiO13</p> <p>Wavelength: 0.6889 Angstrom</p> <p>Temperature: 100 K</p> <p>Flux:&nbsp;1.73&bull;10<sup>10&nbsp;</sup>ph/s</p> <p>Crystal size: 0.050&nbsp;x 0.010 x 0.010 mm</p> <p>Calculated dose (average DWD) per scan: 1.27&nbsp;MGy</p>

opencc-by-4.0Jan 2019View details →
zenodo40/100

Integrated Omics-Based Discovery of Novel Genes, Secondary Metabolites Clusters, and Small Molecules in Penicillium spp. with Disparate Fungal Isolates

<p><em><span>Penicillium expansum</span></em><span> is a ubiquitous postharvest pathogen of pome fruit that causes blue mold decay of apple fruit while another member of the genus, <em>P. chrysogenum</em><span>,</span><em> </em>is a well-studied saprophyte used for antibiotic and small molecule production. While these two fungi have been investigated individually, the recent discovery of <em>P. chrysogenum </em>hindering <em>P. expansum</em> apple fruit infection has not been well studied. To shed light on this interaction between the two species, we conducted a comparative transcriptomic, metabolomic, and genomic study. Global transcriptional and metabolomic outputs were disparate between the species, nearly identical for the <em>P. chrysogenum </em>isolates, and different between <em>P. expansum </em>isolates. Further, the two <em>P. chrysogenum</em> genomes revealed secondary metabolite gene clusters that differed from <em>P. expansum</em>. This included the absence of an intact patulin gene cluster in <em>P. chrysogenum</em>, which corroborates the metabolomic data regarding the species&rsquo; inability to produce patulin. Additionally, <em>P. expansum </em>virulence gene homologues were identified in <em>P. chrysogenum </em>and were similarly transcriptionally regulated <em>in vitro</em>. Molecules with potential antimicrobial activity, and phytohormones like indole-3-acetic acid (IAA), were detected for the first time in <em>P. expansum</em> while pharmacological compounds like the well-studied antibiotic penicillin G were identified in <em>P. chrysogenum</em> culture supernatants. Our findings provide new omics-based resources that enable the study of small molecule production of interest, the potential of <em>Penicillium</em>-derived antimicrobials for postharvest decay control, and <em>P.</em> <em>expansum&rsquo;s</em> metabolites roles in host-pathogen interactions. </span></p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Single-molecule Fluorescent In Situ Hybridization (smFISH) for RNA detection in the fungal pathogen Candida albicans small example dataset

<p><strong>This small example dataset is connected to the protocol article titled:</strong></p> <p>Single-molecule Fluorescent <em>In Situ</em> Hybridization (smFISH) for RNA detection in the fungal pathogen <em>Candida albicans</em></p> <p><strong>Abstract:</strong></p> <p><em>Candida albicans</em> is the most prevalent human fungal pathogen. Its pathogenicity is linked to the ability of <em>C. albicans</em> to reversibly change morphology and to grow as yeast, pseudohyphal or hyphal cells in response to environmental stimuli. Understanding the molecular regulation controlling those morphological switches remains a challenge that, if solved, could help fight <em>C. albicans</em> infections.</p> <p>While numerous studies investigated gene expression changes occurring during <em>C. albicans</em> morphological switches using bulk approaches (e.g., RNA sequencing), here we describe a single-cell and single-molecule RNA imaging and analysis protocol to measure absolute mRNA counts in morphologically intact cells. To detect endogenous mRNAs in single fixed cells, we optimized a single molecule fluorescent <em>in situ</em> hybridization (smFISH) protocol for <em>C. albicans</em>, which allows one to quantify the differential expression of mRNAs in yeast, pseudohyphae or hyphal cells. We quantified the expression of two mRNAs, cell cycle-controlled mRNA (<em>CLB2)</em> and a transcription regulator (<em>EFG1</em>), which show differential expression in the different morphological cell types and in different nutrient conditions. In this protocol, we described in detail the major steps of this approach: growth and fixation, hybridization, imaging, cell-segmentation and mRNA spot analysis. Raw data is provided with the protocol to favour reproducibility. This approach could benefit the molecular characterization of <em>C. albicans</em> and other filamentous fungi, pathogenic or non-pathogenic.</p> <p><strong>Data description:</strong></p> <p>This dataset&nbsp;consists of a FISH experiment&nbsp;spanning two different mRNAs, EFG1 and CLB2, and one nutrient condition,&nbsp;SPIDER37, in Candida albicans. For culturing, the&nbsp;C. albicans wildtype strain SC5314&nbsp;was inoculated at 30 degrees overnight (~15 hours) in 10 mL of TSB medium in a 30 &deg;C&nbsp;shaking incubator. Next, samples were diluted to a density of 10^5 cells/ mL and inoculated for 6 hours in 30 mL&nbsp;Spider medium at 37&nbsp;&deg;C in falcon tubes on an orbital microplate shaker. Then, samples were fixated by adding PFA&nbsp;to a final concentration of 4% to the medium. For hybridization, both mRNAs were&nbsp;hybridized independently by specific DNA oligo labelled with a Quasar670 dye to enable the visualisation of single mRNA molecules. As both genes are labelled by the same dye, these oligos were not co-applied to the same sample but to independent samples.</p> <p><strong>Microscopy</strong></p> <p>For smFISH imaging we use an Olympus BX-63 epifluorescence microscope equipped with Ultrasonic stage and UPlanApo 100x 1.35NA oil-immersion objective (Olympus). Lumencore SOLA FISH light source, a Hamamatsu ORCA-Fusion sCMOS camera (6.5 &micro;m-pixel size) mounted using U-CMT C-Mount Adapter, and zero-pixel shift filter sets: F36-500 DAPI HC Brightline Bandpass Filter, F36-502 FITC HC BrightLine Filter, F36-542 Cy3 HC BrightLine Filter, and F36-523 Cy5 HC BrightLine Filter. Images are acquired across 61-81 optical sections (depending on the sample thickness) with a z-step size of 0.2 &mu;m. The CellSens software (Olympus) is used for instrument control and image acquisition. For the DAPI&nbsp; channel 10-50 ms of exposure was used. Whilst, for the CY5 channel, used for&nbsp;imaging the FISH probes, 750 ms was applied.&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

A large comprehensive curated dataset of small molecules and their activities covering three cardiac ion channels: hERG, Cav1.2, and Nav1.5

<p>The compressed data folder (dataset.rar) represents a data framework for researchers in the field of drug discovery to perform in depth analyses on a very large open-access unique and comprehensive hERG, Nav1.5, and Cav1.2 cardiotoxicity integrated database of small molecules and their activities. The database is organized as follows:</p> <ul> <li>Each sub-folder represents a cardiac ion channel target: hERG, Nav1.5, and Cav1.2</li> <li>Each target sub-folder consists of 3 files in CSV format: One file containing the development set (split into training and validation sets using an 80/20&nbsp;ratio&nbsp;for hyperparameter tuning). The other 2 files contain external evaluation sets. The first test dataset consists&nbsp;of compounds with a structural similarity of no more than 60% (Tanimoto similarity&nbsp; &le; 0.6) to the remaining development set, while the second test dataset comprises&nbsp;compounds with a structural similarity of no more than 70% (Tanimoto similarity &le; 0.7) to the remaining development set.</li> <li>Each file contains data with 7&nbsp;columns: "InChl Key" as a unique identifier of the chemical structure, "SMILES" as the string format of storage and exchange of the chemical structure, "Source" as the upstream data source from which the data was&nbsp;retrieved, "ChEMBL ID" as the&nbsp;ChEMBL identifier if the compound comes from&nbsp;ChEMBL database,&nbsp;"PubChem CID" as the&nbsp;PubChem compound&nbsp;identifier if the compound comes from&nbsp;PubChem database,&nbsp;"pIC50" as the&nbsp;negative logarithm of the half-maximal inhibitory concentration (IC50) to describe the potency of the compound, and "USED_AS" column specifying whether the compound was used for training or validation.</li> </ul> <p><strong>Upon usage, please cite this publication:</strong></p> <ul> <li>Issar Arab, Kristof Egghe, Kris Laukens, Ke Chen, Khaled Barakat, Wout Bittremieux, <strong>Benchmarking of Small Molecule Feature Representations for hERG, Nav1.5, and Cav1.2 Cardiotoxicity Prediction</strong>, <em>Journal of Chemical Information and Modeling</em>, (2023). doi:<a href="https://doi.org/10.1021/acs.jcim.3c01301">10.1021/acs.jcim.3c01301</a></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
dryad40/100

Data from: Pleomorphic effects of three small-molecule inhibitors on transcription elongation by <em>Mycobacterium tuberculosis</em> RNA polymerase

Open the record for dataset details and reuse information.

publicOct 2025View details →
dryad40/100

A selective small-molecule agonist of G protein-gated inwardly-rectifying potassium channels reduces epileptiform activity in a mouse model of tumor associated epilepsy - Thy1-GCaMP Tumor Electrophysiology

Open the record for dataset details and reuse information.

publicOct 2024View details →
zenodo36/100

Development of a high-throughput small molecule screening assay for phenotypical characterization of lysosomal storage disorder-affected cells, with infantile cystinosis as a proof of principle

<p>Together with the Pivot Park Screening Centre we performed a drug screen on CTNS-/- proximal tubule cells. For this we developed an assay to evaluate LC3-II positive puncta, and which may be applied for any disease in which autophagy plays an important role. The screen was optimized by the hotel for a 384 well format, making it useful for high throughput screening. The screen was performed with 1280 compounds from the Prestwick library.</p>

opencc-by-4.0Dec 2019View details →
zenodo36/100

Molecular geometries and energies from quantum mechanical calculations and small molecule force field evaluations.

<p>Force fields are used in a wide variety of contexts for classical molecular simulation, including studies on protein-ligand binding, membrane permeation, and thermophysical property prediction.<br> The quality of these studies relies on the quality of the force fields used to represent the systems.&nbsp;<br> Focusing on small molecules of fewer than 50 heavy atoms, this data compares nine force fields: GAFF, GAFF2, MMFF94, MMFF94S, OPLS3e, SMIRNOFF99Frosst, and the Open Force Field Parsley, versions 1.0, 1.1, and 1.2.<br> On a dataset comprising 22,675 molecular structures of 3,271 molecules, we analyzed force field-optimized geometries and conformer energies compared to reference quantum mechanical (QM) data.<br> <br> The data was created using scripts of the &nbsp;<a href="https://github.com/MobleyLab/benchmarkff/commit/fa45247aa9867f504c02eb6f62d8459a94a0a936">benchmarkff github repository</a>.</p> <p>A corresponding manuscript is submitted, a preprint is available on&nbsp;ChemRxiv:<br> <a href="https://doi.org/10.26434/chemrxiv.12551867.v2 ">Lim, Victoria T.; Hahn, David F.; Tresadern, Gary; Bayly, Christopher I.; Mobley, David (2020): Benchmark Assessment of Molecular Geometries and Energies from Small Molecule Force Fields. ChemRxiv. Preprint</a></p> <p>Read below or the file README.md for further information and description of the content:</p> <pre><code class="language-markdown"># README Version: 04 Nov 2020 For Python scripts that are NOT found in these directories, please check the [BenchmarkFF Github repo](https://github.com/MobleyLab/benchmarkff/tree/master/tools). ## Procedure 1. Prep OPLS3e file for analysis: standardize format by OpenEye in case of differences and convert from kJ/mol to kcal/mol. ``` cd prep python convert_extension.py -i opls3e_minimized.sd -o opls3e.sdf ``` 2. Remove mols that couldn't parameterize by ALL FFs. ``` python get_by_tag.py -i opls3e.sdf -s "SMILES QCArchive" -list trim3.txt -o trim3_full_opls3e.sdf ``` 3. Run analysis. ``` conda activate parsley # calc ddE, RMSD, and TFD distributions python compare_ffs.py -i match.in -t 'SMILES QCArchive' --plot &gt; metrics.out # match_minima, only in 01_analysis_all and 02_analysis_all_smaller_cutoff python match_minima.py -i match.in --plot --cutoff 1.0 --readpickle # look at specific subsets, only in 01_analysis_all python color_by_moiety.py -i match.in -p metrics.pickle -s N-N.dat azetidine.dat octahydrotetracene.dat -o scatter_tfd_3_ # look at outliers,only in 01_analysis_all and 02_analysis_all_smaller_cutoff python tailed_parameters.py -i refdata_trim_overlap_full_openff_unconstrained-1.2.0.sdf -f &lt;offxml file&gt; --metric 'TFD' --cutoff 0.12 --tag "TFD to trim_overlap_full_qcarchive.sdf" --tag_smiles "SMILES QCArchive" &gt; output_tfd.dat ``` ## Brief description of contents * High level: ``` . ├── 00_prep │   ├── convert_extension.py │   ├── opls3e_minimized.sd OPLS3e minimized structures from Schrodinger Maestro │   ├── opls3e.sdf standardized through OpenEye tools │   ├── opt_openff*.sdf OpenFF minimized conformations ├── 01_analysis_all compare all ffs (qm, GAFF(2), MMFF94(S), Smirnoff, OpenFF-X.X, OPLS3e) ├── 02_analysis_all_smaller_cutoff compare all ffs (qm, GAFF(2), MMFF94(S), Smirnoff, OpenFF-X.X, OPLS3e) with a smaller cutoff of .3 for match_minima ├── 03_analysis_latest_ffs compare only the latest versions of ffs (qm, GAFF2, MMFF94S, OpenFF-1.2, OPLS3e) ├── 04_analysis_openff_only compare only OpenFF ffs (qm, Smirnoff, OpenFF-X.X) └── README.md ``` * Inside an output directory: ``` YY_analysis_* various output files of above mentioned scripts, some are listed and described below: ├── bar*.png parameter coverage bar plots ├── ddE.dat relative energies data ├── fig_density_*.png scatter plots of ddE vs (RMSD or TFD) for each force field ├── match.in input file for compare_ffs.py ├── metrics.out output file for compare_ffs.py ├── metrics.pickle pickle file for compare_ffs.py -- you can read this into compare_ffs instead of rerunning the full analysis ├── refdata_*.sdf output SDF files with stored RMSD / TFD scores with reference to QM for each structure ├── relene_*.dat relative energies of matched conformers ├── ridge_dde.png compared energies plot ├── ridge_rmsd.svg compared rmsds plot ├── ridge_tfd.svg compared tfds plot ├── fig_scatter_*.png scatter plots of ddE vs (RMSD or TFD). these are noisy; I don't use these ├── trim3_*.sdf input SDF files for compare_ffs.py listed in match.in file ├── violin*.* violin plot showing ddE distributions ``` </code></pre>

opencc-by-4.0Nov 2020View details →
zenodo36/100

SMAdd-seq: Probing chromatin accessibility with small molecule DNA intercalation and nanopore sequencing

<p>Studies of in vivo chromatin organization have relied on the accessibility of the underlying DNA to nucleases or methyltransferases, which is limited by their requirement for purified nuclei and enzymatic treatment. Here, we introduce a nanopore-based sequencing technique called Small-Molecule Adduct sequencing (SMAdd-seq), where we profile chromatin accessibility by treating nuclei or intact cells with a small molecule, angelicin. Angelicin reacts with thymine bases in linker DNA not bound to core nucleosomes after UV light exposure, thereby labeling accessible DNA regions. By applying SMAdd-seq in Saccharomyces cerevisiae, we demonstrate that angelicin-modified DNA can be detected by its distinct nanopore current signals. To systematically identify angelicin modifications and analyze chromatin structure, we developed a neural network model, NEural network for mapping MOdifications in nanopore long-reads (NEMO). NEMO accurately called expected nucleosome occupancy patterns near transcription start sites at both bulk and single-molecule levels. We observe heterogeneity in chromatin structure and identify clusters of single-molecule reads with varying configurations at specific yeast loci. Furthermore, SMAdd-seq performs equivalently on purified yeast nuclei and intact cells, indicating the promise of this method for in vivo chromatin labeling on long single molecules to measure native chromatin dynamics and heterogeneity.</p>

opencc-by-4.0Mar 2024View details →
dryad36/100

Identification of a small molecule Tim-3 inhibitor to potentiate T cell-mediated antitumor immunotherapy in preclinical models

<p><span>T cell immunoglobulin and mucin-containing molecule 3 (Tim-3), expressed in dysfunctional and exhausted T cells, has been widely acknowledged as a promising immune checkpoint target for tumor immunotherapy. Here, using a strategy combining virtual and functional screening, we identified a compound named ML-T7 that targets the FG-CC' cleft of Tim-3, a highly conserved binding site of </span><span>phosphatidylserine</span><span> (PtdSer) and carcinoembryonic antigen-related cell adhesion molecule 1 (CEACAM1). ML-T7 enhanced the survival and antitumor activity of </span><span>primary CD8<sup>+</sup> cytotoxic T lymphocytes (CTLs) and human chimeric antigen receptor (CAR) T cells and reduced their exhaustion </span><span>in vitro and in vivo</span><span>. In addition, ML-T7 promoted NK cells' killing activity and DC antigen-presenting capacity, consistent with the reported activity of Tim-3.</span><span> Notably, ML-T7 strengthened DCs' functions through both Tim-3 and Tim-4, consistent with the hypothesis that Tim-4 contains a similar FG-CC' loop. Intraperitoneal dosing of ML-T7 showed comparable tumor inhibitory effects to Tim-3 blocking antibody. ML-T7 reduced syngeneic tumor progression in both wildtype and Tim-3 humanized mice and alleviated the immunosuppressive microenvironment. Furthermore, combined ML-T7 and anti-PD-1 therapy had greater therapeutic efficacy than monotherapy in mice, supporting further development of ML-T7 for tumor immunotherapy. Our study demonstrates a potential small molecule for selectively blocking Tim-3 and warrants further study.</span></p>

opencc-zeroNov 2023View details →
zenodo36/100

kefisher98/IP_EA_deltaSCF: Ionization Potential, Electron Affinity, and Delta SCF for Small Organic Molecules

<p>Ionization potential, electron affinity, and delta SCF for small organic molecules selected from the ANI-1 data set. Properties are calculated with 24 different density functional approximations as well as CCSD(T).</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Neural network ensembles and FEFF spectra for multi-modal small molecule chemical motif prediction

<p><strong>Data</strong></p> <ul> <li><strong>22-12-05-data</strong>: original molecular XANES data created from <a href="https://doi.org/10.1103/PhysRevResearch.5.013180">Ghose <em>et al.</em></a></li> <li><strong>23-04-26-ml-data</strong>: machine learning-ready data which is prepared in the format required by <a href="https://github.com/matthewcarbone/Crescendo">Crescendo</a>.</li> <li><strong>23-05-03-hp</strong>: hyper-parameter tuning results from 23-04-26-ml-data.</li> <li><strong>23-05-05-ensembles</strong>: ensemble results from 23-04-26-ml-data.</li> <li><strong>23-05-11-ml-data-CUTOFF8</strong>: a special machine learning-ready dataset constructed by a unique partitioning: only molecules with less than or equal to 8&nbsp;atoms/molecule are used for training/validation, the rest are used for testing.</li> <li><strong>23-12-06_torch_models</strong>: torch only models which can be easily used independently of our ML helper repository, Crescendo. Instead, it can be used with a few lines of code found in multimodal_molecules/core.py, in our <a href="https://github.com/AI-multimodal/multimodal-molecules">GitHub respository</a>.</li> </ul> <p><strong>Funding</strong></p> <p>This research is based upon work supported by the U.S. Department of Energy, Office of Science, Office Basic Energy Sciences, under Award Number FWP PS-030. This research also used theory and computational resources of the Center for Functional Nanomaterials, which is a U.S. Department of Energy Office of Science User Facility, and the Scientific Data and Computing Center, a component of the Computational Science Initiative, at Brookhaven National Laboratory under Contract No. DE-SC0012704.</p>

opencc-by-4.0Jun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record