Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25
datasets available to search
ShareScore release 0.9.0
Dataset results
25 results for “PubChem”
Transformations in PubChem - Full Dataset
<p>This is an archive of the data contained in the "Transformations" section in PubChem for integration into patRoon and other workflows.</p> <p>For further details see the ECI GitLab site: <a href="https://gitlab.lcsb.uni.lu/eci/pubchem/-/blob/master/annotations/tps/README.md">README</a> and main "<a href="https://gitlab.lcsb.uni.lu/eci/pubchem/-/tree/master/annotations/tps">tps</a>" folder.</p> <p>Credits:</p> <p>Concepts: E Schymanski, E Bolton, J Zhang, T Cheng;</p> <p>Code (in R): E Schymanski, R Helmus, P Thiessen</p> <p>Transformations: E Schymanski, J Zhang, T Cheng and many contributors to various lists!</p> <p>PubChem infrastructure: PubChem team</p> <p>Acknowledgements: ECI team who contributed to related efforts, especially: J. Krier, A. Lai, M. Narayanan, T. Kondic, P. Chirsir, E. Palm, Bashir Mayahi. All contributors to the NORMAN-SLE transformations!</p> <p>March 2025 released as v0.2.0 since the dataset grew by >3000 entries! The stats are: </p> <p># 10 Sept. 2025</p> <p>Unique Transformation Entries: 11082<br>Unique Reactions by CID: 9276<br>Unique Reactions by IK: 9263<br>Unique Reactions by IKFB: 8695<br>Unique NORMAN-SLE Compounds by CID: 8345<br>Unique ChEMBL Compounds by CID: 1419<br>Unique Compounds (all) by CID: 9404<br>Unique Predecessors (all) by CID: 3808<br>Unique Successors (all) by CID: 7423<br>Range of XlogP Differences: -12.5,10<br>Range of Mass Differences: -957.97490813,820.227106427</p>
PubChem OECD PFAS Larger PFAS Parts file for MetFrag
<p>This is a <a href="https://msbi.ipb-halle.de/MetFrag/">MetFrag</a> database file constructed from the "Molecule contains PFAS parts larger than CF<sub>2</sub>/CF<sub>3</sub>" subnode of the OECD PFAS Definition node in the <a href="https://pubchem.ncbi.nlm.nih.gov/classification/#hid=120"> PFAS and Fluorinated Organic Compounds in PubChem Tree</a> on the Classification Browser in PubChem.</p> <p>This file was constructed by downloading the node contents, selecting the columns of interest, changing the headers to MetFrag-compatible headers. Entries containing Xe, Pr, Po, Ru and W were removed; charges were also removed from formulas to avoid issues with MetFragCL.</p> <p>The construction of the tree is documented <a href="https://gitlab.com/uniluxembourg/lcsb/eci/pubchem-docs/-/raw/main/pfas-tree/PFAS_Tree.pdf?inline=false">here</a>.</p> <p>Note: PubChem authors have been removed from this version to comply with a presidential decree. This version was prepared exclusively by the remaining author. </p>
Experimental CCS Values in PubChem
<p>The collection of experimental collision cross section (CCS) values from ion mobility experiments in PubChem, retrieved via the code developed <a href="https://gitlab.lcsb.uni.lu/eci/pubchem/-/tree/master/annotations/CCS/CCS_retrieval">here</a>.</p> <ul> <li> "All_CCS_in_PubChem.csv" contains all CCS values extracted from PubChem (previous versions contained entries with no CIDs, current version has all entries mapped to CIDs).</li> <li>"All_CCS_in_PubChem_wInfo.csv" contains all CCS values, plus post-processing (splitting annotations, adducts and comments, adding chemical identifiers), but excludes all entries with no CID (and thus no structural information - not applicable in current version). </li> </ul> <p>Details how this file was produced are given <a href="https://gitlab.lcsb.uni.lu/eci/pubchem/-/tree/master/annotations/CCS/CCS_retrieval">here</a>.</p>
MassBank <=> PubChem Deposition/Annotation Repository
<p>This is a repository to exchange MassBank record and substance information to create deposition and annotation files in PubChem.</p> <p>Supporting code in: <a href="https://gitlab.com/uniluxembourg/lcsb/eci/pubchem/-/tree/master/massbank_eu" target="_blank" rel="noopener">https://gitlab.com/uniluxembourg/lcsb/eci/pubchem/-/tree/master/massbank_eu</a></p> <p>Credits:</p> <ul> <li>LCSB-ECI: Anjana Elapavalore, Todor Kondic, Emma Schymanski</li> <li>PubChem: Jeff Zhang, Paul Thiessen, Ben Shoemaker, Evan Bolton</li> <li>MassBank Consortium: Rene Meier, Steffen Neumann, Tobias Schulze</li> </ul> <p>Note: 20230419 files removes deprecated records from the 2022.12 release and still has InChIKeys added manually for ACES records (listed as NA) for annotation; all are present in the deposition. The unique substance file did not change, but close to 200 deprecated records were removed from the annotation set. 20230908, 20231129, 20240610, 20241126, 20250502: ACES InChIKeys added manually.</p>
S68 | HSDBTPS | Transformation Products Extracted from HSDB Content in PubChem
<p>This is the collection associated with list S68 HSDBTPS Transformation Products Extracted from HSDB Content in PubChem on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>HSDBTPS is a list of metabolites / transformation products extracted from the "Metabolites/Metabolism" section from HSDB (Hazardous Substance Data Bank) in PubChem (<a href="https://pubchem.ncbi.nlm.nih.gov/source/11933">https://pubchem.ncbi.nlm.nih.gov/source/11933</a>). Dataset DOI: <a href="https://doi.org/10.5281/zenodo.3827487">10.5281/zenodo.3827487</a>.</p> <p>Entries automatically extracted from HSDB are manually validated to remove mismatching CIDs, and add additional CIDs not captured using the descriptions provided. Files are created with a default to not import any data until this has been checked. Please report any mismatches, despite best efforts it is possible that errors are present, all files are under version control so that entries can be corrected/updated/enhanced over time and recommitted.</p> <p>Recent uploads (2022 and later) are handled with ShinyTPs (<a href="https://gitlab.lcsb.uni.lu/eci/shinytps">code</a> + article from Palm et al 2023 DOI:<a href="https://doi.org/10.1021/acs.estlett.3c00537">10.1021/acs.estlett.3c00537</a>). The original code associated with this deposit is located <a href="https://gitlab.lcsb.uni.lu/eci/pubchem/-/tree/master/annotations/tps/">here</a>. </p> <p>Updates:</p> <p>16 May 2020: Added the source file (extract of all HSDB Metabolites/Metabolism entries as is from JSON file), plus an updated S68_HSDBTPS_StructInfoOnly.csv with one new structure and renamed to include S68, plus the first draft of the Transformations table. 28 May 2020: added InChIKey and DIXSID files. 11 June 2020: added updated Transformations and Structure files to contain new CIDs and resulting reactions from new PubChem registrations. Dec 23, 2022: first dataset from Emma Palm added, using her TP curation app. DTXSID file dropped. 23 Mar 2023: added Biosystem and Enzyme columns, filling in biosystem where appropriate. 24 Mar 2023: new azo dye reactions from Emma Palm added. 1 Apr 2023: more azo dye reactions from Emma Palm. 4 Apr 2023: added new CIDs. 27 June 2023: added many new substances, including new CIDs. 28 June 2023 added transformations, 30 June 2023 added new CIDs. 11 July added missing CIDs to transformations table. 18 Nov 2023: added new reactions from Jolly Komolo (no new CIDs). 1 Dec 2023: new reactions from Marie, incl. two new CIDs. 19 Dec 2023: new reactions from Marie, incl. 1 new CID, updated CIDs from Dec 1. 29 April 2024: new reactions from Marie, CID updated. 16 July 2024: added reactions from Olga. 26 July 2024: added BPA=>MBP. 6 Aug 2024: adjusted many triazine names, added one new reaction. 27 Nov 2024: added Griseofulvin reactions. 8 Feb 2025: added Sertraline TP. 14 Mar 2025: added missing CID for O-Demethyl phosphamidon.</p>
Analog series from ChEMBL, PubChem, and DrugBank
<p>The datasets consist of analog series and key compounds extracted from ChEMBL, PubChem, and DrugBank. For each compound structural and activity information is provided. </p>
Highly promiscuous compounds from PubChem assays
<p>For the pool of 466 detected highly promiscuous compounds the PubChem ID and the corresponding ChEMBL ID(s) are provided. In addition, the detection status is set to "pains" or "aggregator" if the compound was detected as PAINS or an aggregator, respectively. Otherwise the status is set "passed". "ChEMBL analogues" lists the ChEMBL compound IDs of structural analogs of highly promiscuous compounds (if available). For the 466 compounds the number of targets and the corresponding PubChem target IDs are given in the last two columns. Compounds 1-26 (Compound No. 1-26) correspond to compounds shown in the publication.</p>
High-Priority Promiscuity Cliffs from PubChem
<p>A representative sample of 278 high-priority promiscuity cliffs identified from PubChem is provided. These promiscuity cliffs were formed by pairs of structurally analogous compounds with a difference in promiscuity degrees of at last 20 targets. These compounds were tested in at least 300 shared primary assays (SA) and showed assay similarity (AS) and assay frequency ratio (AFR) values of at least 0.85. A README file is also given.</p>
HSDB Metabolism / Metabolites Annotation Content from PubChem
<p>An archived version of annotation content extracted from PubChem for the <a href="https://pubchem.ncbi.nlm.nih.gov/source/11933">HSDB</a> section, as scripted up in <a href="https://gitlab.lcsb.uni.lu/eci/pubchem/-/blob/master/annotations/tps/extractAnnotations.R">extractAnnotations.R</a> (code available on the ECI <a href="https://gitlab.lcsb.uni.lu/eci/pubchem/">pubchem</a> repository)</p>
Collection of analog series-based (ASB) scaffolds shared between ZINC, ChEMBL, and PubChem
<p>Analog series-based (ASB) scaffolds shared between ZINC and ChEMBL (version 22), ZINC and PubChem and all the three databases are provided as three separate files. For each ASB scaffold, the SMILES representation of ZINC compounds is provided. In addition, the number of ZINC compounds, the number and the list of targets it was annotated with is reported. A README file is also given.</p>
PubChem Data Mining of OXPHOS inhibitors: scripts, data, and models
<p>README doc, source, and data files from PubChem data mining project to identify OXPHOS inhibitory chemotypes.</p>
Unfiltered Depositor-Provided Chemical Synonyms for Substance Records in PubChem
<p>This gzipped text file contains a list of all (live) substance records in PubChem with their "<strong>unfiltered"</strong> depositor-provided chemical synonyms, downloaded from PubChem in June 2017. Each line has a Substance ID (SID) and its chemical synonym, separated by a tab. The SID-synonym pairs in this file were used in the paper “<strong>PubChem Synonym Filtering Process Using Crowdsourcing</strong>” by Sunghwan Kim et al., published in the Journal of Cheminformatics (<a href="https://doi.org/10.1186/s13321-024-00868-3" target="_blank" rel="noopener">https://doi.org/10.1186/s13321-024-00868-3</a>). The up-to-date version of this file can be downloaded from the PubChem FTP Site (<a href="https://ftp.ncbi.nlm.nih.gov/pubchem/Substance/Extras/" target="_blank" rel="noopener">https://ftp.ncbi.nlm.nih.gov/pubchem/Substance/Extras/</a>).</p>
Matchms and PubChem cleaned MS/MS dataset from GNPS
<p>Dataset of MS/MS spectra retrieved from GNPS (https://gnps.ucsd.edu) on 25/01/2021, which underwent extensive metadata cleaning.</p> <p>Version v1 contained a bug (missing "NUM PEAKS" parameters), this was corrected in version v2.</p> <p>Metadata was cleaned and processed using matchms (https://github.com/matchms/matchms) and matchmsextras (https://github.com/matchms/matchmsextras). This largely consited of</p> <ul> <li>Empty spectra were removed.</li> <li>Compound names were cleaned</li> <li>charge, adduct, formula, ionmode fields were cleaned and corrected</li> <li>parent mass estimated were added (using precursor mz and adduct information)</li> <li>inchikey, inchi, and SMILES were checked and corrected</li> <li>Spectra which remained without inchi/inchikey/smiles were searched against pubchem based on their mass and name.</li> </ul> <p>This resulted in 210,400 spectra out of which 184,698 are annotated with InChIKey and SMILES and/or InChI.</p> <p>If you use this dataset for your research please cite the following:</p> <ul> <li>GNPS, e.g. [Wang, M. <em>et al.</em> Sharing and community curation of mass spectrometry data with GNPS. <em>Nat. Biotechnol. </em><strong>34</strong>, 828–837 (2016)]</li> <li>matchms: [ Huber, F. <em>et al.</em> matchms - processing and similarity evaluation of mass spectrometry data. <em>J. Open Source Softw. </em><strong>5</strong>, 2411 (2020) ]</li> <li>PubChem: [ Kim, S. <em>et al.</em> PubChem 2019 update: improved access to chemical data. <em>Nucleic Acids Res. </em><strong>47</strong>, D1102–D1109 (2019)]</li> </ul> <p>Many thanks!</p>
PubChem compounds tested in primary and confirmatory assays
<p>The set of 437,257 compounds that were tested in both primary and confirmatory assays was assembled from PubChem BioAssay collection and deposited in an EXCEL file. For each compound, its compound identifier in PubChem (i.e., cid), the number of primary and confirmatory assays it was tested in and activity annotations are reported.</p>
Compound activity data sets for 15 biological targets compiled from the ChEMBL and PubChem databases.
<p>Compound activity data sets for the 15 biological targets are deposited, along with structure-activity relationship matrices IDs. Active compounds were extracted from the ChEMBL database and inactive were from the PubChem database. Details of the data sets are described in the original publication. and the summary of the data sets is given in the readme.txt file. </p>
PubChem Compounds with MeSH Annotations
<p>PubChem database was searched for compounds that had a MeSH annotation. It is great that PubChem now allows a direct download for these compounds in a rich CSV file. </p>
10 Mio randomly selected PubChem SMILES strings with maximum length of 128
<p>This dataset contains SMILES strings of maximum 128 token randomly selected from the PubChem database.</p>
PubChem and ChEMBL-series processed dataset used in Exhaustive local chemical space exploration using a transformer model
<p>PubChem and ChEMBL-series processed dataset used in <span>Exhaustive local chemical space exploration using </span><span>a transformer model</span></p>
PubChem compounds before and after standardization
<p>This repository contains a subset of 200k PubChem compounds before and after standardization.<br> This dataset is the basis of the PubChem-pretrained model presented in "Standardizing chemical compounds with language models" (available on <a href="https://doi.org/10.26434/chemrxiv-2022-14ztf-v2">ChemRxiv</a>, see also the associated <a href="https://github.com/rxn4chemistry/rxn-standardization">GitHub repository</a>).</p> <p>The associated pretrained model is also provided here, along with the splits used for training.</p> <p>The data is provided under the <a href="https://cdla.dev/sharing-1-0/">CDLA-Sharing-1.0</a> license.</p> <p>Provided files:</p> <ul> <li>README.md: General README.</li> <li>src_and_tgt_all.csv: All the compounds before and after standardization, in CSV format.</li> <li>src-train.txt: The tokenized compounds before standardization in the train split.</li> <li>tgt-train.txt: The tokenized compounds after standardization in the train split.</li> <li>src-valid.txt: The tokenized compounds before standardization in the validation split.</li> <li>tgt-valid.txt: The tokenized compounds after standardization in the validation split.</li> <li>src-test.txt: The tokenized compounds before standardization in the test split.</li> <li>tgt-test.txt: The tokenized compounds after standardization in the test split.</li> <li>LICENSE.md: The details of the CDLA-Sharing-1.0 license.</li> <li>pretrained_pubchem_step_120000.pt: the pretrained model.</li> </ul>
WiP: "Keywords" on PubChem content
<p>Work in progress: exploring the creation of groups of chemicals based on PubChem annotation content in the form of "keywords" (the exact term is still a subject of debate ... for the moment keyword is the placeholder). This is a file deposition corresponding to the code base on the <a href="https://git-r3lab.uni.lu/eci/pubchem/-/tree/master/annotations/keywords">ECI GitLab</a> pages to create these files.</p> <p>Part of this work was performed at <a href="https://www.biohackathon-europe.org/">BioHackathon Europe</a> 2020 #BioHackEU20</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.