Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
18
datasets available to search
ShareScore release 0.9.0
Dataset results
18 results for “Compound library”
Morphing libraries, QSAR models, and compounds predicted to be active on the Glucocorticoid receptor (GR)
<p>This repository contains datasets and files related to the computational drug discovery project of the chemical space exploration of the Glucocorticoid receptor. The accompanying Python code is freely available in the GitHub repository (<a title="https://github.com/Iagea/GRML_analyses" href="https://github.com/Iagea/GRML_analyses" target="_blank" rel="noreferrer noopener">https://github.com/Iagea/GRML_analyses</a>).</p> <p><strong>Morphing Libraries:</strong></p> <ul> <li><strong>GRML_library.csv:</strong> The GRML library is the collection of 999,015 virtual compounds generated by Molpher [1-2] starting from GR ligands with unique Bemis-Murcko scaffolds collected from the ChEMBL17 and IMG libraries.</li> <li><strong>RML_library.csv:</strong> The RML library is the collection of 1,346,310 virtual compounds generated by Molpher starting from compounds with unique Bemis-Murcko scaffolds randomly selected from the ZINC database.</li> </ul> <p><strong>IMG library:</strong></p> <ul> <li><strong>IMG_non_proprietary.csv</strong>: The non-proprietary IMG library subset containing 12,956 compounds and their corresponding B-scores from the primary screen.</li> </ul> <p><strong>Molpher inputs:</strong></p> <ul> <li><strong>GR_inputs.csv</strong>: The GR inputs are the ligands used to create the GRML library, 204 compounds from ChEMBL17 (95 compounds) and the non-proprietary dataset from IMG (109 compounds).</li> <li><strong>Random_inputs.csv</strong>: The random inputs are 249 random ZINC compounds used to create the Random library.</li> </ul> <p><strong>Model's training sets:</strong></p> <ul> <li><strong>Model33_training_set.csv</strong>: Random forest classification model training set, it includes 865 compounds; known GR actives and inactives from ChEMBL33 (738 compounds) and non-proprietary active ligands from the IMG library (127 compounds).</li> <li><strong>Model17_training_set.csv</strong>: Random forest classification model training set, it includes 601 compounds; known GR actives and inactives from ChEMBL17 (474 compounds) and non-proprietary active ligands from the IMG library (127 compounds).</li> <li><strong>RFR_training_set.csv</strong>: Random forest regression model training set, it includes 89 compounds; known GR actives and inactives from ChEMBL33 that fit into the GR pharmacophore with the four features we describe in our paper.</li> </ul> <p><strong>Models:</strong></p> <ul> <li><strong>Model33.pkl:</strong> Python pickle file containing the trained Random forest classification models used along with Mondrian cross-conformal prediction to classify GR actives/inactives. This model was trained with ChEMBL33 and IMG libraries.</li> <li><strong>Model17.pkl</strong>: Python pickle file containing the trained Random forest classification models used along with Mondrian cross-conformal prediction to classify GR actives/inactives. This model was trained with ChEMBL17 and IMG libraries.</li> <li><strong>RFR_models.pkl</strong><em>: </em>Python pickle file containing the 100 trained random forest regression models used to rank the proposed active morphs. These models were trained with the RFR_training_set.csv.</li> </ul> <p><strong>Active predicted morphs:</strong></p> <ul> <li><strong>all_morphs_actives</strong><em><strong>_</strong></em><strong>predicted.xlsx:</strong> An Excel spreadsheet containing two sheets. 1) All 22,524 GRML active predicted morphs. 2) All 4,341 RML active predicted morphs. The QED, NIBR Severity Score, and Molskill Score are given for each morph.</li> </ul> <p><strong>Proposed GR active ligands:</strong></p> <ul> <li><strong>designed_ligands.xlsx</strong>: An Excel spreadsheet containing two sheets. 1) All 54 designed GR ligands with their QED, NIBR severity score, MolSkill score, consensus ranking from the 100 RFR models, and the result of the manual annotation and remarks, if available. 2) The structure of the 54 ligands based on their manual annotation and presence or not in ChEMBL33 database.</li> </ul> <p>Researchers and professionals in the field of drug discovery and cheminformatics may find these resources useful for further analysis and investigations.</p> <p><strong>Bibliography</strong></p> <p>[1] Hoksza, D., Škoda, P., Voršilák, M. <em>et al.</em> Molpher: a software framework for systematic chemical space exploration. <em>J Cheminform</em> <strong>6</strong>, 7 (2014). https://doi.org/10.1186/1758-2946-6-7</p> <p>[2] <a href="https://github.com/lich-uct/molpher-lib">https://github.com/lich-uct/molpher-lib</a></p>
BonMOLière: Small-Sized Libraries of Readily Purchasable Compounds, Optimized to Produce Genuine Hits in Biological Screens across the Protein Space
<p>BonMOLière is a set of small-sized libraries of readily purchasable compounds, optimized to produce genuine hits in biological screens across the protein space. The corresponding research article has been published in <em>Int. J. Mol. Sci.</em> 2021, 22(15), 7773, DOI: <a href="https://doi.org/10.3390/ijms22157773">https://doi.org/10.3390/ijms22157773</a></p>
Compounds of interest identified by screening a focused library against USP5 Zf-UBD with a 19F NMR assay
<p>Screening a set of commercial compounds selected from computational docking studies against USP5 zinc finger ubiquitin binding domain (Zf-UBD) using <sup>19</sup>F NMR spectroscopy. You can find details of preliminary <sup>19</sup>F NMR experiments <a href="https://zenodo.org/record/1246807#.WxV8HkgvzIV">here</a>.</p>
LC-MS/MS data of pharmaceutical compounds for lIbrary building
<p>Dataset used for MergeION package testing are provided by Janssen Pharmaceutica. It consists of known standard pharmaceutical compounds for which high quality Q-Exactive MS/MS data is provided. Details about these compounds can be found in Metadata_Library.txt. All datasets were acquired in positive ion mode through either DDA (data-dependent acquisition) or targeted MS/MS. Raw data in profile mode were converted into centroid-mode mzML or mzXML files using MSConvertGUI</p> <p> </p>
Metabolomics dataset relating to the pubblication: "Combining CRISPRi and metabolomics for functional annotation of compound libraries"
<p>Metabolomics dataset relating to the pubblication: "Combining CRISPRi and metabolomics for functional annotation of compound libraries"</p>
Screening the Sigma LOPAC®1280 library of compounds for protective effects against cisplatin-induced oto- and nephrotoxicity
<p>Dose-limiting toxicities for cisplatin administration, including ototoxicity and nephrotoxicity, impact the clinical utility of this effective chemotherapy agent and lead to lifelong complications, particularly in pediatric cancer survivors. Using a two-pronged drug screen employing the zebrafish lateral line as an <i>in vivo</i> readout for ototoxicity and kidney cell-based nephrotoxicity assay, we screened 1280 compounds and identified 22 that were both oto- and nephroprotective. Of these, dopamine and L-mimosine, a plant-based amino acid active in the dopamine pathway, were further investigated. Dopamine and L-mimosine protected the hair cells in the zebrafish otic vesicle from cisplatin-induced damage and preserved zebrafish larval glomerular filtration. Importantly, these compounds did not abrogate the cytotoxic effects of cisplatin on human cancer cells. This study provides insights into the mechanisms underlying cisplatin-induced oto- and nephrotoxicity and compelling preclinical evidence for the potential utility of dopamine and L-mimosine in the safer administration of cisplatin.</p>
Screening of 165,988 compounds reflecting biological and chemical diversity in the Janssen Library for anti-SARS-CoV-2 activity
<p>This report describes the most relevant results of screening a diversity set of the Janssen Pharmaceutica compound collection for potential activity against SARS-CoV-2 in a stable eGFP-expressing VeroE6 cell line infected with SARS-CoV-2.</p>
Fig. 4 in BMDMS-NP: A comprehensive ESI-MS/MS spectral library of natural compounds
Fig. 4. Evaluation of searching blocks against those of BMDMS-NP. Blocks in (A) and (C) are from NIST17 and those in (B) and (D) are from MoNA. Lines marked with 'x' in (C) and (D) show the performance after application of precursor m/z filter had been adopted prior to calculation of the similarity score between blocks.
Fig. 3 in BMDMS-NP: A comprehensive ESI-MS/MS spectral library of natural compounds
Fig. 3. Overview of BMDMS-NP. (A) The MS/MS library was established by two types of spectrometers with several conditions. (B) BMDMS-NP consists of information about compound spectra. The program provided by GitHub helped to search and build an in-house library. (C) Spectrum search is performed using the concept of data block. The spectra of the library are grouped as shown by the bold lines in the table. The user can group input spectral data into a block and search against BMDMS-NP and get ranking and scores for the searches.
Fig. 1 in BMDMS-NP: A comprehensive ESI-MS/MS spectral library of natural compounds
Fig. 1. Principal component analysis of molecular fingerprints between the 'natural product/for sale' subset of ZINC and BMDMS-NP. The graph describes that the structural diversity of BMDMS-NP is overlapped with that of ZINC. In order to compare intuitively, the ZINC points were randomly selected to be the same quantity as the BMDMS-NP points and the 95% confidence interval ellipses are added.
Fig. 2 in BMDMS-NP: A comprehensive ESI-MS/MS spectral library of natural compounds
Fig. 2. Venn diagram representing the number of overlapped metabolites between libraries. The number of each set was determined using InChIKey as the compound identifier.
Screening the Sigma LOPAC®1280 library of compounds for protective effects against cisplatin-induced oto- and nephrotoxicity
Open the record for dataset details and reuse information.
Molecular Property Diagnostic Suite Compound Library (MPDS-CL): A Structure based Classification of the Chemical Space
<p>Molecular Property Diagnostic Suite-Compound Library (MPDS-CL), is an open-source galaxy-based cheminformatics web-portal which presents a structure-based classification of the molecules. A structure-based classification of nearly 150 million unique compounds, which are obtained from 42 publicly available databases were curated for redundancy removal through 97 hierarchically well-defined atom composition-based portions. These are further subjected to 56-bit fingerprint-based classification algorithm which led to a formation of 56 structurally well-defined classes. The classes thus obtained were further divided into clusters based on their molecular weight. Thus, the entire set of molecules was put in 56 different classes and 625 clusters. This led to the assignment of a unique ID, named as <em>MPDS AadharID</em>, for each of these 149 169 443 molecules. <em>MPDS AadharID</em> is akin to the unique number given to citizens in India (similar to the SSN in US, NINO in UK). MPDS-CL unique features are: a) several search options, such as exact structure search, substructure search, property-based search, fingerprint-based search, using SMILES, InChIKey and key-in; b) automatic generation of information for the processing for MPDS and other galaxy tools; c) providing the class and cluster of a molecule which makes it easier and fast to search for similar molecules and d) information related to the presence of the molecules in multiple databases. The MPDS-CL can be accessed at http://mpds.neist.res.in:8086/.</p>
Development and validation of a high-throughput screening pipeline of compound libraries to target EMT
GEO Series GSE284656. Homo sapiens. 18 samples. Type: Expression profiling by high throughput sequencing.
Data from: Teratological and behavioral screening of the National Toxicology Program 91-compound library in zebrafish (Danio rerio)
To screen the tens of thousands of chemicals for which no toxicity data currently exists, it is necessary to move from in vivo rodent models to alternative models, such as zebrafish. Here, we used dechorionated Tropical 5D wildtype zebrafish embryos to screen a 91-compound library provided by the National Toxicology Program (NTP) for developmental toxicity. This library contained 86 unique chemicals that included negative controls, flame retardants, polycyclic aromatic hydrocarbons (PAHs), drugs, industrial chemicals and pesticides. Fish were exposed to five concentrations of each chemical or an equal amount of vehicle (0.5% DMSO) in embryo medium from 6 h post-fertilization (hpf) to 5 d post-fertilization (dpf). Fish were examined daily for mortality and teratogenic effects and photomotor behavior was assessed at 4 and 5 dpf. Of the five negative control compounds in the library, none caused mortality/teratogenesis, but two altered behavior. Chemicals provided in duplicate produced similar outcomes. Overall, 13 compounds caused mortality/teratology but not behavioral abnormalities, 24 only affected behavior, and 18 altered both endpoints, with behavior affected at concentrations that did not cause mortality/teratology (55/86 hits). Of the compounds that affected behavior, 52% caused behavioral abnormalities at either 4 or 5 dpf. Compounds within the same functional group caused different behavioral abnormalities, while similar behavioral patterns were caused by compounds from different groups. Our data suggest that behavior is a sensitive endpoint for developmental toxicity screening that integrates multiple modes of toxic action and is influenced by the age of the larval fish at the time of testing.
Data from: Multi-behavioral endpoint testing of an 87-chemical compound library in freshwater planarians
Open the record for dataset details and reuse information.
Data from: Teratological and behavioral screening of the National Toxicology Program 91-compound library in zebrafish (Danio rerio)
Open the record for dataset details and reuse information.
Tuebingen Compounds Library_Postcolumn_Reduction_ECell
<p>Q exactive HF</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.