Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
131
datasets available to search
ShareScore release 0.7.1
Dataset results
131 results for “Drug discovery”
High-Throughput Drug Discovery for a Rare Neurological Disorder: Uncovering a Novel Therapeutic Opportunity for the 19q12 Autism Spectrum Disorder [19q12 neurons baseline data]
GEO Series GSE292758. Homo sapiens. 5 samples. Type: Expression profiling by high throughput sequencing.
In vivo intratumoral Heterogeneity in a dish: Scalable Forebrain Organoid Models of Embryonal Brain Tumors for High-Throughput Personalized Drug Discovery (human organoids subset 3)
GEO Series GSE270070. Homo sapiens; Mus musculus. 0 samples. Type: Expression profiling by high throughput sequencing.
Accelerating early anti-TB drug discovery by creating mycobacterial indicator strains that predict mode of action [Mycobacterium tuberculosis]
GEO Series GSE107831. Mycobacterium tuberculosis H37Rv. 36 samples. Type: Expression profiling by high throughput sequencing.
Figure 2 from: Sivakumar B, Kaliappan I (2023) Lead drug discovery from imidazolinone derivatives with Aurora kinase inhibitors. Pharmacia 70(4): 1529-1540. https://doi.org/10.3897/pharmacia.70.e114935
Figure 2
Figure 1 from: Sivakumar B, Kaliappan I (2023) Lead drug discovery from imidazolinone derivatives with Aurora kinase inhibitors. Pharmacia 70(4): 1529-1540. https://doi.org/10.3897/pharmacia.70.e114935
Figure 1
Dataset of AlphaFold's internal representations of 4,581 proteins relevant for drug discovery
<p>This dataset contains the outputs of the AlphaFold model for 4,581 proteins that are relevant targets in drug discovery.</p> <p>More information on the dataset can be found at the following repository: </p> <p><strong>Dataset structure:</strong></p> <p>↓ <strong>data/</strong>* -> main data directory</p> <blockquote> <p>↓ <strong>data/PID/</strong>* -> data of a single protein of length <strong>L</strong></p> <blockquote> <table> <tbody><tr> <th>Filename</th> <th>Description</th> <th>Tensor shape</th> <th>Lightweight</th> </tr> </tbody><tbody> <tr> <td><strong>single.npy</strong></td> <td>( s i ) evoformer single representation</td> <td><em>[<strong>L</strong> x 384]</em></td> <td>✔️</td> </tr> <tr> <td><strong>structure.npy</strong></td> <td>( a i ) output of the last layer of structure module</td> <td><em>[<strong>L</strong> x 384]</em></td> <td>✔️</td> </tr> <tr> <td><span><em><strong>msa.npy</strong>***</em></span></td> <td>( m s i ) processed MSA representation</td> <td><em>[<strong>N</strong> x <strong>L</strong> x 256]</em></td> <td> </td> </tr> <tr> <td><em><strong>pair.npy</strong>***</em></td> <td>( z i j ) evoformer pair representation</td> <td><em>[<strong>L</strong> x <strong>L</strong> x 128]</em></td> <td> </td> </tr> <tr> <td><strong>PID.pdb</strong></td> <td>3D protein structure prediction</td> <td> </td> <td>✔️</td> </tr> <tr> <td><strong>PID_unrelaxed.pdb</strong></td> <td>3D protein structure prediction w/o relaxation step (D)</td> <td> </td> <td>✔️</td> </tr> <tr> <td><strong>confidence.npy</strong>*</td> <td>confidence in structure prediction (0-100)</td> <td><em><a title="Béquignon OJM, Bongers BJ, Jespers W, IJzerman AP, van de Water B, van Westen GJP. Papyrus - A large scale curated dataset aimed at bioactivity predictions. ChemRxiv. Cambridge: Cambridge Open Engage; 2021; This content is a preprint and has not been peer-reviewed." href="https://chemrxiv.org/engage/chemrxiv/article-details/617aa2467a002162403d71f0" rel="nofollow">1</a></em></td> <td>✔️</td> </tr> <tr> <td><strong>plldt.npy</strong>*</td> <td>confidence in structure prediction per residue</td> <td><em>[<strong>L</strong>]</em></td> <td>✔️</td> </tr> <tr> <td><strong>PID.fasta</strong></td> <td>protein amino acid sequence and metadata</td> <td> </td> <td>✔️</td> </tr> <tr> <td><strong>timings.json</strong></td> <td>Processing log</td> <td> </td> <td>✔️</td> </tr> </tbody> </table> </blockquote> </blockquote> <blockquote> <p>↓ <strong>data/PID2/</strong>* -> data of protein #2</p> <blockquote> <p><strong>...</strong></p> </blockquote> </blockquote> <p>*<em>Note: <strong>L</strong>: sequence length, <strong>N</strong>: number of aligned sequences via MSA.</em></p>
Data for "Machine Learning Scoring Functions for Drug Discovery from Experimental and Computer-generated Protein-Ligand Structures: Towards Per-target Scoring Functions"
<p>Data used in "<em>Machine Learning Scoring Functions for Drug Discovery from Experimental and Computer-generated Protein-Ligand Structures: Towards Per-target Scoring Functions</em>"<br> by F. Pellicani, D. Dal Ben, A. Perali, S. Pilati</p> <p>If you use these data or the python script for your research or other activities, please cite the corresponding journal article.</p> <p> </p> <p>====================</p> <p>Uncompressing the zipped file <em>DataSFUnicam.zip</em> provies the following files and folders:</p> <p><br> <strong>DataSFUnicam/</strong></p> <p> </p> <p> ExperimentalDataPDBFiles/<br> <em>This folder contains 2408 .pdb files of experimental complex structures. The files are named with a univocal code corresponding to the protein-ligand complex.</em></p> <p> </p> <p> ExperimentalDataXLSXFile.xlsx<br> <em>This Excel file reports the experimental protein-ligand chemical information. In the sheet named “Foglio1”, the first column contains the univocal code of the protein-ligand complex, the second column contains the experimentally measured pK_d.</em></p> <p> </p> <p> SyntheticDataPDBFiles/<br> <em>This folder contains the .pdb files of the synthetic complex structures. The .pdb files are grouped in 17 folders according to just as many target proteins. The folders are named after the corresponding protein. Each folder contains the .pdb files for the best position of each protein-ligand pair according to the MOE docking score. The files are named with a univocal code.</em></p> <p> </p> <p> SyntheticDataXLSXFiles/<br> <em> The folder contains 17 Excel files with the chemical information of the synthetic protein-ligand complexes. The files are named after the corresponding target protein. In the sheet named “Foglio1” of each .xlsx file, the first column contains a univocal code of the protein-ligand complex in each conformation, the second column contains an auxiliary numerical code corresponding to the protein-ligand pair, the third column contains the experimentally measured pK_i, and the fourth column contains the docking score provided by the MOE software.</em></p> <p>====================</p> <p>USER GUIDE FOR THE PYTHON SCRIPT</p> <p>Download and uncompress the zipped file "<em>SFUnicam.zip</em>" with a command like "<em>unzip SFUnicam.zip</em>". </p> <p>The following file structure is created:</p> <p><em>SFUnicam/</em></p> <p> <em>ComplexToBePredictedFolder/4ey5_30.pdb <br> MaxAssMatrix.npy<br> my_model<br> devStndSynt.npy<br> mediaSynt.npy<br> UnicamSF13prot.py<br> README.txt</em><br> <br> The subfolder "<em>ComplexToBePredictedFolder/</em>" contains the example PDB file "<em>4ey5_30.pdb</em>".</p> <p>-) To execute the script "<em>UnicamSF13prot.py</em>", Python 3 should be installed with the following libraries and sublibraries:<br> <em>Keras:<br> Regularizers<br> Sequential (keras.models)<br> Conv1D, Dense, MaxPooling1D, GlobalMaxPooling1D, GlobalAveragePooling1D, AveragePooling1D (keras.layers)<br> Adam (keras.optimizers)<br> Numpy</em><br> <em>Tensorflow</em></p> <p>Operation:<br> -) Copy the .pdb file related to the protein-ligand complex whose affinity is to be predicted in the subfolder “<em>ComplexToBePredictedFolder/</em>”.<br> -) Make sure the following files are in the same folder where the python script is:<br> <em>MaxAssMatrix.npy<br> mediaSynt.npy<br> devStndSynt.npy<br> my_model</em><br> -) Run the code using Python 3 with a command like "<em>python3.x UnicamSF13prot.py</em>".<br> -) Enter the name of the protein-ligand PDB file whose affinity is to be predicted (excluding the extension ".pdb").<br> -) Read the predicted affinity from screen.<br> </p> <p> </p>
A Biospecimen and Clinical Data Study on Patients for Drug & Biomarker Discovery
ClinicalTrials.gov study NCT01592552. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Blood/Urine Markers for Drug Discovery for Renal Disease in Diabetes
ClinicalTrials.gov study NCT02580266. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Myeloma Novel Drug Discovery Ver 1.2
ClinicalTrials.gov study NCT05968417. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Drug Discovery for Parkinson's With Mutations in the GBA Gene
ClinicalTrials.gov study NCT05536388. IPD Sharing: NO. Countries: 1. Publications: 0.
Biomarker Discovery for Novel Drug Development in Idiopathic Pulmonary Fibrosis
ClinicalTrials.gov study NCT01718990. IPD Sharing: Not stated. Countries: 1. Publications: 0.
High fidelity patient-derived xenografts for accelerating prostate cancer discovery and drug development (aCGH)
GEO Series GSE41188. Homo sapiens. 24 samples. Type: Genome variation profiling by genome tiling array.
A multiplex single-cell RNA-seq pharmacotranscriptomics pipeline for drug discovery
GEO Series GSE274905. Homo sapiens. 2 samples. Type: Expression profiling by high throughput sequencing; Other.
Discovery of therapeutic agents targeting PKLR for NAFLD using drug repositioning
GEO Series GSE193627. Mus musculus. 64 samples. Type: Expression profiling by high throughput sequencing.
Human liver-derived organoids recapitulate Oropouche virus infection and manifestation, enabling antiviral drug discovery
GEO Series GSE316780. Homo sapiens. 32 samples. Type: Expression profiling by high throughput sequencing.
High fidelity patient-derived xenografts for accelerating prostate cancer discovery and drug development
GEO Series GSE41193. Homo sapiens. 56 samples. Type: Expression profiling by array; Genome variation profiling by genome tiling array.
High-Throughput Drug Discovery for a Rare Neurological Disorder: Uncovering a Novel Therapeutic Opportunity for the 19q12 Autism Spectrum Disorder [post treatment entrectinib patient data]
GEO Series GSE292760. Homo sapiens. 10 samples. Type: Expression profiling by high throughput sequencing.
Deriving Schwann Cells from hPSCs Enable Disease Modeling and Drug Discovery for Diabetic Peripheral Neuropathy
GEO Series GSE195730. Homo sapiens. 10 samples. Type: Expression profiling by high throughput sequencing.
Comparative analysis of glioblastoma initiating cells and patient-matched EPSC-derived neural stem cells as a discovery tool and drug matching strategy [Seq]
GEO Series GSE154367. Mus musculus. 30 samples. Type: Expression profiling by high throughput sequencing; Methylation profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.