Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
40
datasets available to search
ShareScore release 0.9.0
Dataset results
40 results for “Chemical Space”
Fig. 3 in Computational insight into the chemical space of plant growth regulators
Fig. 3. The representative examples of compounds from the reference dataset (the corresponding patent application numbers are provided).
Fig. 6 in Computational insight into the chemical space of plant growth regulators
Fig. 6. The Sammon map showing the distribution of the compounds within the related chemical space; outliers indicated by the ellipse are spread beyond the scope of the major set population. Axis per se do not make sense as they simply organize outputted 2D-lattice.
Fig. 7 in Computational insight into the chemical space of plant growth regulators
Fig. 7. The Kohonen map constructed for the whole reference dataset. The scale at the bottom identifies the number of compounds located in a node; the axes indicate the coordinates of neurons within the lattice; the contours are smoothed.
Enumerated Chemical Space and the parameter file
<p>Molecular formula for the study MTLBS1684 were enumerated using a set of formula filtering rules. </p>
Chemical compositions data for "Space weathering of the Chang'e-5 lunar sample from a mid-high latitude region on the Moon"
<p>Data for “Space weathering of the Chang’e-5 lunar sample from a mid-high latitude region on the Moon”</p>
Data set for Ligand additivity relationships enable efficient exploration of transition metal chemical space
<p>Dataset of transition metal complexes curated in pickle files and comma delimited format, scripts for CSD curation, and computed DFT properties for associated manuscript.</p>
Chemperium database for: Geometric Deep Learning for Molecular Property Predictions with Chemical Accuracy Across Chemical Space
<p>The dataset and trained models for the submitted manuscript "Geometric Deep Learning for Molecular Property Prediction with Chemical Accuracy Across Chemical Space"</p> <p>The trained models can be used in combination with the predict module in github.com/mrodobbe/chemperium. More information in README.md.</p> <p><em>When using these datasets, refer directly to the manuscript: https://doi.org/10.1186/s13321-024-00895-0 </em></p>
Supporting data of the publication: Conquering chemical spaces in the billion range: is docking a computational alternative to DNA-encoded libraries?
<p>Raw data of the publication: "Conquering chemical spaces in the billion range: is high-throughput docking a computational alternative to DNA-encoded libraries?" by Levente M. Mihalovits, Tibor V. Szalai, Dávid Bajusz and György Miklós Keserű. Figures of the manuscript and the supporting information were created using these data.</p>
Datasets used for chemical space visualization and the corresponding results.
<p>The datasets (.h5 format) used for dimensionality reduction (ChEMBL_datasets) and optimization results (DR_results) in the <a href="https://chemrxiv.org/engage/chemrxiv/article-details/66bb4da5f3f4b05290bccb6e">publication</a>.</p>
A CLOSE-UP LOOK AT THE CHEMICAL SPACE OF COMMERCIALLY AVAILABLE BUILDING BLOCKS FOR MEDICINAL CHEMISTRY
<p>The ability to efficiently synthesize desired compounds can be a limiting factor for chemical space exploration in drug discovery. This ability is conditioned not only by the existence of well-studied synthetic protocols but also by the availability of corresponding reagents, so-called building blocks (BB). In this work, we present a detailed analysis of the chemical space of 400K purchasable BB. The chemical space was defined by corresponding synthons – fragments contributed to the final molecules upon reaction. They allow an analysis of BB physicochemical properties and diversity, unbiased by the leaving and protective groups in actual reagents. The main classes of BB were analyzed in terms of their availability, rule-of-two-defined quality, and diversity. Available BBs were eventually compared to a reference set of biologically relevant synthons derived from ChEMBL fragmentation, in order to illustrate how well they cover the actual medicinal chemistry needs. This was performed on a newly constructed universal generative topographic map of synthon chemical space, allowing to visualize both libraries and analyze their overlapping and library-specific regions.</p> <p>The dataset includes annotated synthons extracted from ChEMBL and publicly available.</p>
Exploring the Chemical Space of Glycosylation in Noncovalent Protein Complexes: an Expedition along Different Structural Levels of Human Chorionic Gonadotropin Employing Mass Spectrometry
<p><strong>Supplementary files for "Exploring the Chemical Space of Glycosylation in Noncovalent Protein Complexes: an Expedition along Different Structural Levels of Human Chorionic Gonadotropin Employing Mass Spectrometry"</strong></p> <p><strong>Introduction</strong></p> <p>This data repository contains all previously unpublished raw data files for the manuscript “Exploring the Chemical Space of Glycosylation in Noncovalent Protein Complexes: an Expedition along Different Structural Levels of Human Chorionic Gonadotropin Employing Mass Spectrometry” by Maximilian Lebede<sup>||</sup>, Fiammetta Di Marco<sup>||</sup>, Wolfgang Esser-Skala, René Hennig, Therese Wohlschlager, Christian G. Huber.</p> <p><strong>Files</strong></p> <p>This repository contains 9 files:</p> <ul> <li><strong>Dimer Raw Files.zip</strong> folder containing 4 files of native-MS data (*.raw, Thermo RAW file format) of two batches of the drug product Ovitrelle® at native dimer level. </li> <li><strong>H11M9 Ovitrelle BA056714 Glycopeptide R1 230920_07.zip</strong> folder containing 1 file of HPLC-MS/MS glycopeptide data (*.raw, Thermo RAW file format) of one batch of the drug product Ovitrelle®.</li> <li><strong>H11M9 Ovitrelle BA056714 Glycopeptide R2 230920_08.zip</strong> folder containing 1 file of HPLC-MS/MS glycopeptide data (*.raw, Thermo RAW file format) of one batch of the drug product Ovitrelle®.</li> <li><strong>H11M9 Ovitrelle BA056714 Glycopeptide R3 230920_09.zip</strong> folder containing 1 file of HPLC-MS/MS glycopeptide data (*.raw, Thermo RAW file format) of one batch of the drug product Ovitrelle®.</li> <li><strong>H11M9 Ovitrelle BA059433 Glycopeptide R1 240920_15.zip</strong> folder containing 1 file of HPLC-MS/MS glycopeptide data (*.raw, Thermo RAW file format) of one batch of the drug product Ovitrelle®.</li> <li><strong>H11M9 Ovitrelle BA059433 Glycopeptide R2 240920_16.zip</strong> folder containing 1 file of HPLC-MS/MS glycopeptide data (*.raw, Thermo RAW file format) of one batch of the drug product Ovitrelle®.</li> <li><strong>H11M9 Ovitrelle BA059433 Glycopeptide R3 240920_17.zip</strong> folder containing 1 file of HPLC-MS/MS glycopeptide data (*.raw, Thermo RAW file format) of one batch of the drug product Ovitrelle®.</li> <li><strong>MoFi Settings.zip</strong> folder containing 12 files of MoFi settings (*.xml) to annotate deconvoluted spectra of hCG subunits and dimer of two Ovitrelle® batches, untreated and after desialylation. A typical MoFi setting file is build from protein sequence (*.FASTA), monosaccharide and frequent modification atomic composition (*.csv), glycan or glycoform library (*.csv) and deconvoluted spectrum in centroid (*.csv). Files are named as following: Settings_Ovitrelle_Batch number (BA056714 or BA059433)_Structural level (Alpha, Beta or Dimer)_Enzymatic treatement (Untreated or Sialidase).</li> <li><strong>Subunit Raw Files.zip</strong> folder containing 8 files of HPLC-MS data (*.raw, Thermo RAW file format) of two batches of the drug product Ovitrelle® at intact subunit level. </li> </ul> <p>Raw files are named as following: Instrument, Drug product (Ovitrelle), Batch number (BA056714 or BA059433), Structural level (Dimer, Subunits or Glycopeptides), Enzymatic treatment (untreated, Sialidase, PNGase F or PNGase F + Sialidase) and date. Glycopeptide data includes 3 replicates (R1-3).</p> <p><strong>License</strong></p> <p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this license, visit <a href="http://creativecommons.org/licenses/by/4.0/">http://creativecommons.org/licenses/by/4.0/</a> .</p> <p> </p>
DIY chemical space product file
<p>The chemical space made up from 1000 building blocks using Mcule's ARCHIE.</p>
Automated patent extraction powers generative modeling in focused chemical spaces: Training data and model checkpoints release
<p>Training data and model checkpoints accompanying paper on "Automated patent extraction powers generative modeling in focused chemical spaces". If you use this data, please cite the following manuscript:</p> <pre>@article{subramanian2023automated, title={Automated patent extraction powers generative modeling in focused chemical spaces}, author={Subramanian, Akshay and Greenman, Kevin P and Gervaix, Alexis and Yang, Tzuhsiung and G{\'o}mez-Bombarelli, Rafael}, journal={Digital Discovery}, year={2023}, publisher={Royal Society of Chemistry} }</pre>
Fig. 1 in Computational insight into the chemical space of plant growth regulators
Fig. 1. Representative examples of phytohormones.
Fig. 5 in Computational insight into the chemical space of plant growth regulators
Fig. 5. The representative examples of the reference scaffolds.
Fig. 2 in Computational insight into the chemical space of plant growth regulators
Fig. 2. The mechanisms of action for herbicides (examples).
Results from "Navigating a 10E+60" Chemical Space
Open the record for dataset details and reuse information.
Molecular Property Diagnostic Suite Compound Library (MPDS-CL): A Structure based Classification of the Chemical Space
<p>Molecular Property Diagnostic Suite-Compound Library (MPDS-CL), is an open-source galaxy-based cheminformatics web-portal which presents a structure-based classification of the molecules. A structure-based classification of nearly 150 million unique compounds, which are obtained from 42 publicly available databases were curated for redundancy removal through 97 hierarchically well-defined atom composition-based portions. These are further subjected to 56-bit fingerprint-based classification algorithm which led to a formation of 56 structurally well-defined classes. The classes thus obtained were further divided into clusters based on their molecular weight. Thus, the entire set of molecules was put in 56 different classes and 625 clusters. This led to the assignment of a unique ID, named as <em>MPDS AadharID</em>, for each of these 149 169 443 molecules. <em>MPDS AadharID</em> is akin to the unique number given to citizens in India (similar to the SSN in US, NINO in UK). MPDS-CL unique features are: a) several search options, such as exact structure search, substructure search, property-based search, fingerprint-based search, using SMILES, InChIKey and key-in; b) automatic generation of information for the processing for MPDS and other galaxy tools; c) providing the class and cluster of a molecule which makes it easier and fast to search for similar molecules and d) information related to the presence of the molecules in multiple databases. The MPDS-CL can be accessed at http://mpds.neist.res.in:8086/.</p>
Exploring the Chemical Design Space of Metal-Organic Frameworks for Photocatalysis
<p>In this work, we employ a chemical insights-based diversity-driven approach to search for metal-organic framework (MOF) photocatalysts. With an in silico design based on chemical insights, we populated areas in the chemical design space related to MOFs with photocatalytic potential. We selected a balanced dataset of DFT-based photocatalytic descriptors computed for 314 MOFs, comprising our in silico structures, a diverse subset of the QMOF database, and experimental MOF photocatalysts. With such a balanced dataset, we could fine-tune supervised machine-learning models from literature that allowed us to draw insights into relevant areas in the chemical design space for photocatalysis and potential bottlenecks.<br>Among our in silico MOFs, a few motifs stood out, such as Au-pyrazolate, Ti clusters and rod-shaped metal nodes, and a particular MOF designed with the Mn4Ca cluster, which mimics the OER center in the photosystem II of photosynthesis.<br>Overall, by combining three pillars --- the design of potential MOF photocatalysts guided by chemical insights, the DFT evaluation of photocatalytic descriptors, and the machine-learning approach --- we were able to gain insights into structure-property relationship, and identify trends in the chemical design space that can open new avenues for advancing the field of photocatalysis.</p>
The International Space Station Has a Unique and Extreme Microbial and Chemical Environment Driven by Use Patterns
Space habitation provides unique challenges in built environments isolated from Earth. We produced a 3D map of the microbes and metabolites throughout the International Space Station (ISS), with 803 samples collected during space flight, including controls. We find that the use of each of the nine sampled modules within the ISS strongly drives the microbiology and chemistry of the habitat. Relating the microbiology to other Earth habitats, we find that, as with human microbiomes, built environment microbiomes also align naturally along an axis of industrialization, with the ISS providing an extreme example of an industrialized environment. We demonstrate the utility of culture-independent sequencing for microbial risk monitoring, especially as the location of sequencing moves to space. The resulting resource of chemistry and microbiology in the space-built environment will guide long-term efforts to maintain human health in space for longer durations.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.