Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
131
datasets available to search
ShareScore release 0.7.1
Dataset results
131 results for “Drug discovery”
High resolution deep mutational scanning of the melanocortin-4 receptor enables target characterization for drug discovery
<p>This record contains supplementary data files related to <a href="https://www.biorxiv.org/content/10.1101/2024.10.11.617882v1">High resolution deep mutational scanning of the melanocortin-4 receptor enables target characterization for drug discovery</a>, described below. For more information regarding statistical modeling and subsequent biological interpretation, please see <a href="https://github.com/octantbio/mc4r-dms">the mc4r-dms Github repository</a>.</p> <p> - <strong>MC4R-DMS-barcode-maps.tar.gz</strong> contains two large TSV files, each containing an oligonucleotide-barcode map for either the CRE or UAS system. Each map contains two columns: (i) the barcode sequence and (ii) the oligonucleotide (variant) identifier.</p> <p> - <strong>MC4R-DMS5-Gs-barcode-counts.tar.gz</strong>, <strong>MC4R-DMS8-Gq-barcode-counts.tar.gz</strong>, and <strong>MC4R-DMS11-Gs-barcode-counts.tar.gz</strong> each contain a set of TSV files representing the raw barcode counts for each sample, with one TSV file per sample. Each individual sample TSV file has two columns: (i) the read count and (ii) the barcode sequence</p> <p> - <strong>sample-properties.tar.gz</strong> contains three small TSV files each with three columns: (i) a sample ID matching the individual sample files in the barcode-counts directories, (ii) a condition or treatment label, and (iii) a condition or treatment dosage if applicable</p> <p> - <strong>MC4R-mapped-counts.tar.gz</strong> contains three TSV files, one per dataset (DMS5, DMS8, and DMS11), containing all barcodes joined to the relevant barcode map across all samples. Each file contains the following columns:</p> <table> <tbody> <tr> <td>sample</td> <td>the sample identifier corresponding to the file names in each barcode-counts directory</td> </tr> <tr> <td>barcode</td> <td>barcode sequence</td> </tr> <tr> <td>count</td> <td>read count</td> </tr> <tr> <td>lib</td> <td>library ID (always OCNT-DMSLIB-0)</td> </tr> <tr> <td>chunk</td> <td>chunk identifier</td> </tr> <tr> <td>wt_aa</td> <td>reference amino acid</td> </tr> <tr> <td>pos</td> <td>amino acid position</td> </tr> <tr> <td>mut_aa</td> <td>mutated amino acid</td> </tr> <tr> <td>wt_codon</td> <td>reference codon sequence</td> </tr> <tr> <td>mut_codon</td> <td>mutated codon sequence</td> </tr> <tr> <td>condition</td> <td>name of treatment condition</td> </tr> <tr> <td>condition_conc</td> <td>dosage of treatment compound, if relevnat</td> </tr> <tr> <td>stop_counts</td> <td>the natural log of the sum of stop-associated barcodes for the containing chunk</td> </tr> </tbody> </table> <p> </p>
Druglike molecule datasets for drug discovery
<p><strong>Background</strong><br> Trnasformer-based AI models have shown outstanding performance in identifying druggable candidate molecules. In most cases, models are trained on a massive amount of database of molecular information to capture the latent meaning of a given molecule. However, the desirable properties of candidate molecules include the feasibility of synthesizing them, low toxicity, and high druggability. In this study, we injected prior knowledge of the desirable properties of molecules during the training process.</p> <p><strong>Methods</strong><br> Using the PubChem database (100 M), we filtered druglike molecules based on the quantity of drug-likeliness (QED) score and the Pfizer rule. With this dataset of drug-like molecules, we trained both the molecular representation model (chemBERTa) and the molecular generation models (MolGPT). The molecular representation model was evaluated by fine-tuning the results on the MoleculeNet benchmark datasets, and the molecular generation model was evaluated based on the generated samples (10 K). </p> <p><strong>Results</strong><br> Training with druglike molecules enabled the generation of molecules with desirable properties without any conditioning. Although the molecular representation learning model was not remarkable, however, its performance in predicting clinical toxicology exceeded that of conventional molecular representation models.</p> <p><strong>Conclusion</strong><br> By training based on a dataset of druglike molecules, our approach enables molecular representation models to predict clinical toxicity more precisely. Furthermore, it enables the molecule generation model to generate molecules with desirable druglike properties without any conditional generation procedures.<br> </p> <p>-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p>import pickle</p> <p> </p> <p>with open("druglike_molecules_QED.pkl", "rb") as f:</p> <p> data = pickle.load(f)</p>
Deep learning and knowledge graph powered drug combination discovery against infectious diseases
<p>The Datasets and source codes for paper "<strong>Deep learning and knowledge graph powered drug combination discovery against infectious diseases</strong>".</p>
Fig. 1 in Design, synthesis and screening of a drug discovery library based on an Eremophila-derived serrulatane scaffold
Fig. 1. Chemical structures of the targeted NP diterpenoid scaffolds (1–2), methylated products (3–4) and the amide library (5–16).
Fig. 4 in Design, synthesis and screening of a drug discovery library based on an Eremophila-derived serrulatane scaffold
Fig. 4. Representative flow cytometry examples of GFP expression (i.e., HIV production) in J-Lat 10.6 cells in the absence of stimulation (A), treatment with 50 nM PMA (B) and PMA plus 100 μM compound 2 (C). Dose response profiles (D) of compounds on GFP expression induced by 50 nM PMA in J-Lat 10.6 cells. Data denote mean ± s.d. from three independent experiments.
Fig. 3 in Design, synthesis and screening of a drug discovery library based on an Eremophila-derived serrulatane scaffold
Fig. 3. Microscopic images of the skinny (Skn) phenotype in L4s induced by compound 3 compared with the wild-type (WT) phenotype control (cultured in culture medium + 0.25% DMSO); the bottom panels enlarged from the dashed/ boxed areas in the upper panels.
data used in study of "A universal programmable Gaussian Boson Sampler for drug discovery"
<p>1. Data used for analysis the average photon number distribution</p> <p>2. Data used for analysis the relation between squeezing level (r) and pump power (p)</p>
TrendyGenes, a computational pipeline for the detection of literature trends in academia and drug discovery
<p>TrendyGenes Literature Mining</p> <p>This repository contains the files and code to build the TrendyGenes pipeline described in the paper "TrendyGenes, a computational pipeline for the detection of literature trends in academia and drug discovery" (Serrano Nájera et al. 2021).</p> <p>Contents</p> <p>The folder contains the following files:</p> <ul> <li>PubMed_*.csv.gz: CSV files containing PubMed metadata (titles, abstracts etc.) split into multiple files</li> <li>CoCitations*.csv.gz: CSV files containing co-citation networks computed from PubMed</li> <li>MeSH2PMID.csv.gz: Map of MeSH terms to PMIDs</li> <li>Authorship_Neo4J_complete.csv.gz: Authorship information for PubMed papers</li> <li>Disease2PMID_Neo4J_complete.csv.gz: Map of disease terms to PMIDs after disambiguation</li> <li>Genes_Neo4J_complete_CCPU.csv.gz: Map of genes to PMIDs after disambiguation</li> <li>genes.csv.gz: List of human genes</li> <li>diseases.csv.gz: List of MeSH disease terms</li> <li>import_command*.txt: Commands to import data into Neo4j graph database</li> </ul> <p>Building the Knowledge Graph</p> <p>The various CSV files can be imported into a Neo4j graph database to build the knowledge graph containing publications, authors, genes, diseases etc. and their connections as described in the paper.</p> <p>The import_command*.txt files contain the Neo4J bulk import syntax needed to import the data into Neo4j:<br> https://neo4j.com/developer/guide-import-csv/</p> <p>Citation</p> <p>Serrano Nájera G, Narganes Carlón D, Crowther DJ. TrendyGenes, a computational pipeline for the detection of literature trends in academia and drug discovery. Scientific Reports. 2021 Aug 3;11(1):15747.</p> <p>License</p> <p>[MIT]</p> <p>This summarizes the key files provided and briefly explains how they can be used to build the knowledge graph database for the TrendyGenes pipeline. The citation provides a reference to the original paper.</p>
Supporting Information for Recommender Systems in Antiviral Drug Discovery
<p>Supporting Information for Recommender Systems in Antiviral Drug Discovery</p>
Small molecule sequestration of amyloid-β as a drug discovery strategy for Alzheimer's disease
<p>These data describe the bound and unbound ensembles of a disordered peptide (amyloid beta) in the presence and absence of a small, drug-like molecule. The data were produced by metadynamic metainference simulations restrained with NMR chemical shift data. We used PLUMED version 2.6.0 and GROMACS 2018.3. All the data and PLUMED input files required to reproduce the metadynamic metainference results are available on PLUMED-NEST (www.plumed-nest.org), the public repository of the PLUMED consortium, as plumID:20.014.</p> <p>A Jupyter Notebook describing the analysis of these results is available from GitHub at https://github.com/vendruscolo-lab/amyloid-beta_small_mol/ in Metadynamic_metainference/Analysis.</p> <p>These data support the findings of the manuscript entitled "Small molecule sequestration of amyloid-β as a drug discovery strategy for Alzheimer's disease" by Heller <em>et al</em>. DOI: 10.1101/729392</p>
Semantic text mining in early drug discovery for type 2 diabetes
<p>BACKGROUND: Surveying the scientific literature is an important part of early drug discovery; and with the ever-increasing amount of biomedical publications it is imperative to focus on the most interesting articles. Here we present a project that highlights new understanding (e.g.\ recently discovered modes of action) and identifies potential novel drug target, via a novel, data-driven text mining approach to score type 2 diabetes (T2D) relevance. We focused on monitoring trends and jumps in T2D relevance to help us be timely informed of important breakthroughs.<br> <br> METHODS: We extracted over 7 million <em>n</em>-grams from PubMed and then clustered around 240,000 linked to T2D into almost 50,000 T2D relevant `semantic concepts'. To score papers, these concepts were weighted depending on co-mentioning with core T2D proteins. A protein's current T2D relevance was determined by combining the scores of the papers mentioning it in the preceeding five years. The significance of a jump in a protein's rank was assessed by comparing it to previously observed jumps.<br> <br> RESULTS: We show that T2D relevant papers, also those not mentioning T2D explicitly, got assigned high scores by mentioning semantic concepts often used in connection with T2D, as shown by the enrichment of well known T2D proteins among the top scoring proteins. Our `high jumpers' identified important past developments in the apprehension of how certain key proteins relate to T2D, indicating that our method will make us aware of future breakthroughs. In summary, this project facilitated keeping up with current T2D research by repeatedly providing short lists of potential novel targets into our early drug discovery pipeline.</p>
Making Drug Discovery Data FAIR: The Yawning Gap between Aspiration and Implementation (recording)
<p>Recording from BioITWorld October 2020</p> <p>The FAIRification of data is gaining impetus. However, for drug discovery, the envisaged increased flow of structures and bioactivity into major public databases, such as PubChem, has not happened. Reasons will be reviewed, but a key impediment is that even when supplementary data from journal papers is submitted to open repositories, such as figshare, there is neither push nor pull into PubChem. Ways to ameliorate this major bottleneck will be discussed, including bypassing the entombment of chemistry in PDFs</p> <p> </p>
Caramel: an web-based QSAR tool for melanoma drug discovery - Bambu Models
Open the record for dataset details and reuse information.
Figure 7 from: Sivakumar B, Kaliappan I (2023) Lead drug discovery from imidazolinone derivatives with Aurora kinase inhibitors. Pharmacia 70(4): 1529-1540. https://doi.org/10.3897/pharmacia.70.e114935
Figure 7 Two-dimensional diagram of Compounds 1, 7, 9, 13, 19 interactions (top) during 100 ns MD simulation and hit molecule compound 19_1MQ4 interaction (bottom).
Figure 5 from: Sivakumar B, Kaliappan I (2023) Lead drug discovery from imidazolinone derivatives with Aurora kinase inhibitors. Pharmacia 70(4): 1529-1540. https://doi.org/10.3897/pharmacia.70.e114935
Figure 5 The root mean square fluctuation (RMSF) of 1MQ4 protein during 100 ns MD, representing local changes along the protein chain for Compounds 1, 7, 9, 13, 19.
Figure 4 from: Sivakumar B, Kaliappan I (2023) Lead drug discovery from imidazolinone derivatives with Aurora kinase inhibitors. Pharmacia 70(4): 1529-1540. https://doi.org/10.3897/pharmacia.70.e114935
Figure 4 The root mean square deviation (RMSD) of 1MQ4 protein during 100 ns MD, Compounds 1, 7, 9, 13, 19.
Figure 6 from: Sivakumar B, Kaliappan I (2023) Lead drug discovery from imidazolinone derivatives with Aurora kinase inhibitors. Pharmacia 70(4): 1529-1540. https://doi.org/10.3897/pharmacia.70.e114935
Figure 6 The plot represents the hydrogen bonding interactions of compounds 1, 7, 9, 13, 19 with esteem to deposits of 1MQ4 throughout 100 ns MD simulation.
Figure 3 from: Sivakumar B, Kaliappan I (2023) Lead drug discovery from imidazolinone derivatives with Aurora kinase inhibitors. Pharmacia 70(4): 1529-1540. https://doi.org/10.3897/pharmacia.70.e114935
Figure 3 Leads generated from QSAR and docking studies (compounds 1, 7, 9, 13, 19) structural stability analysis upon ligand docking.
Supplementary material 1 from: Sivakumar B, Kaliappan I (2023) Lead drug discovery from imidazolinone derivatives with Aurora kinase inhibitors. Pharmacia 70(4): 1529-1540. https://doi.org/10.3897/pharmacia.70.e114935
Supplementary data
Data for "The increase in prevalence of natural products through the drug discovery process"
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.