Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,923
datasets available to search
ShareScore release 0.9.0
Dataset results
1,923 results for “Compounds”
Compound database and subsets generated by the fragment network for stage 3 of the PHIP2 SAMPL7 Challenge
<p>The fragment network provides a convenient way to filter-out compounds that are dissimilar to the input hit(s). Overall, this search algorithm requires a compound input and 3 parameters: 1- the number of graph traversals (hops), 2- number of changes in heavy atom count (hac), 3- number of changes in ring atoms counts (rac). Please, read the reference (Hall, Murray and Verdonk, 2017) for the specifics of the methods.</p>
Research Data Supporting "Understanding Structural and Electronic Properties of Bismuth Trihalides and Related Compounds"
<p>Research Data Supporting "Understanding Structural and Electronic Properties of Bismuth Trihalides and Related Compounds"</p> <p>DOI: 10.1021/acs.inorgchem.9b03214</p>
Hydrogen storage properties of Mn and Cu for Fe substitution in TiFe0.9 intermetallic compound - Dataset related to publication.
<p>Data type: Experimental measurements, correlations and Van't Hoff plot. Date format: .opj. Origin of the data: Experimental pressure composition isotherm measurements. Data generated by a home-made Sieverts’ type apparatus from CNRS, ICMPE, Thiais, France. Software needed to plot the data: Origin.</p>
Systematic Data Analysis and Diagnostic Machine Learning Reveal Differences between Compounds with Single- and Multitarget Activity
<p>The deposited files contain balanced data sets of multi-target (MT) and single-target (ST) compounds (CPDs) used for machine learning studies (https://dx.doi.org/10.1021/acs.molpharmaceut.0c00901). The first file (st_mt_data.tsv) contains 15,142 MT- and 15,081 ST-CPDs and the second (st_dt_data.tsv) 1828 DT- and 1776 ST-CPDs. For each CPD, a nonstereo_aromatic_SMILES representation, the original ChEMBL_cid, UniProt (target) IDs, and CPD category (CPD_CAT) (i.e. DT/MT/ST) is provided. DT stands for 'diverse-target' and denotes a subset of MT-CPDs (as detailed in the publication). In addition, a CPD is tagged “Y” if it continued to be present in the data set after removal of 50% randomly selected CPDs or 50% CPD nearest neighbors (NN), respectively.</p>
VHP4Safety Compound Wikidata Triples
<p>Export as a CSV file with triples from the VHP4Safety Compound Wiki.</p>
SONAR -- experimental redox potentials for organic compounds undergoing 2-electron/2-proton transfer reactions
<p>reference data for the demo-compounds used as input for predicting redox potentials by a trained model </p><p>The file</p><ul><li>lists redox potentials and oxidized/reduced form for organic molecules undergoing a two-electron/two-proton reduction reaction (M + 2 e- + 2 H+ --> MH2)</li><li>contains data for 25 organic compounds compiled from various sources in literature</li><li>uses "|" as a separator</li><li>column names and explanations<ol><li><strong>ID</strong>: abbreviated trivial names e.g. for labelling</li><li><strong>orig redox potential [V]:</strong> original values reported in respective reference</li><li><strong>solvent</strong>: total formula, water (H2O) throughout</li><li><strong>pH</strong>: pH value of electrolyte solution. If not reported, inferred from the concentration of supporting electrolyte</li><li><strong>supporting_electrolyte</strong>: if spefified: total formula, if available; concentration</li><li><strong>SMILES_ox</strong>: molecular structures encoded as (manually assigned) SMILES strings for the oxidized species (M)</li><li><strong>SMILES_red</strong>: molecular structures encoded as (manually assigned) SMILES strings for the reduced species (MH2)</li><li><strong>ref_electrode:</strong> reference electrode the originally reported half cell potential refers to. If not specified, RHE was used as default</li><li><strong>redox potential vs SHE [V]</strong>:<ul><li>In case of missing information, reversible hydrogen electrode (RHE at pH = 0) was assumed, which corresponds to SHE</li><li>In case of conflicting entries (SHE and pH != 0), we assumed the pH should be accounted for and replaced "RHE" as reference electrode instead of "SHE". "NHE" was treated like "RHE".</li><li>In case the reference electrode was other than SHE, NHE or RHE, a respective offset was added. This was the case once for Ag/AgCl (assuming saturated solution, offset = 0.210, see respective reference)</li><li>Finally, the potential values were transferred to SHE according to: E(SHE) = E(RHE) + 0.05913 * pH</li><li>CAVEAT: Lacking information about individual pKa values, no other correction was made.</li></ul></li><li><strong>reference</strong>: orginal source</li></ol></li></ul>
NPClassifier predictions of COCONUT compounds
<p>Class predictions as returned by NPClassifier of COCONUT compounds.</p> <p>Used NPClassifier between 2024.01.03~2024.01.15. Used COCONUT version 2022.01.01</p> <p> </p> <p> </p> <p>References:</p> <ol> <li>Kim, H. W. et al. NPClassifier: A Deep Neural Network-Based Structural Classification Tool for Natural Products. J. Nat. Prod. 84, 2795–2807 (2021). 10.1021/acs.jnatprod.1c00399</li> <li>Sorokina, M., Merseburger, P., Rajan, K., Yirik, M. A. & Steinbeck, C. COCONUT online: Collection of Open Natural Products database. J. Cheminformatics 13, 2 (2021). https://doi.org/10.1186/s13321-020-00478-9 <p> </p> </li> </ol>
S109 | PARCEDC | List of 7074 potential endocrine disrupting compounds (EDCs) by PARC T4.2
<p>This is the collection associated with list S109 PARCEDC List of 7074 potential endocrine disrupting compounds (EDCs) by PARC T4.2 on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <div>A comprehensive list of 7074 endocrine disruptors compiled by PARC T4.2., integrating assessments by EU and national regulators, contributions from entities like the European Chemicals Agency (ECHA)and complemented by other potential endocrine disruptors (hormones, bisphenols, etc.).The list also includes potentially active endocrine disruptors screened from all substances in the NORMAN SusDat database (https://www.norman-network.com/nds/susdat/) using VEGA (QSAR) EDC prediction models (https://www.vegahub.eu/about-qsar), ToxCast database's in vitro assay data and ToxCast based machine learning predictions.</div> <div> </div> <div>List kindly provided by Sandrine Andres and Valeria Dulio, INERIS.</div>
volatile organic compounds that were measred in various sites in Israel, available by the Israel Ministry of Environmental Protection
<p>The dataset comprises measurements of volatile organic compounds sampled at multiple sites across Israel from 2010 to 2024, encompassing urban, rural, and suburban locations.</p>
Elevated increase in compound extreme heat-precipitation events over China
<p>This file contains the fractions (in percentage) of the compound extreme precipitation events that are preceded by an extreme heat event in China during 1961-2017. The compound events are identified based on the CN05.1 dataset at 0.5x0.5 resolution. Please contact us with any questions or concerns (email: luo.ming@hotmail.com).</p>
Strawberry volatile organic compounds metabolomic data and QTL study
<p>This dataset contains the supplementary materials of the publication "Multivariate QTL approach reveals a major regulator of terpenoid production and other volatiles in strawberry" of the same authors. In this study we extracted volatile organic compounds from several strawberry samples and analysed their identity and abundance. We used this volatile data to perform an extensive multivariate QTL study, the results of which can be found in this dataset.</p> <p>All analysis, results and figures can be reproduced using the folder included in the <strong>supplementary data 1</strong>. If you want to reproduce part or all of our analysis, only download sup data 1. </p> <p>If you only need one of our results or data table you can download them individually:</p> <ul> <li>Sup data 2: abundances of volatile organic compounds from a biparental and diverse panel (GWAS) population.</li> <li>Sup data 3 and 4: p-value tables for all metabolites as well as multivariate traits (see publication for more information).</li> <li>Sup table 1: Summary of metabolite abundances and heritabilities across both populations.</li> <li>Sup tables 2 and 3: significant QTL signals for each trait individually and summarised per QTL locus.</li> <li>Sup table 4: previously reported VOC QTLs in strawberry, with positions imputed in the Royal Royce genome.</li> <li>Sup table 5: metadata about all the identified compounds on this and previous studies.</li> <li>Sup table 6: SNP array positions imputed in the "Royal Royce" genome assembly.</li> <li>Sup table 7: Number of markers per chromosome in this analysis.</li> </ul> <h3>Update 2025</h3> <p>We updated the underlying code and datasets to reflect several revisions made to this work. Most notably, the QTL results have been reworked. They now do not include Blink or FarmCPU results (only mixed model results, obtained through statgenGWAS). Additionally, heritability estimations, QQ-plots and other figures have been added to the reproducible results code.</p>
SCALIBUR video 3: From sewage sludge to biofertilisers, bioplastics and compounds
<p>This video is part of a 3 part series explaining the innovative technologies being developed in the SCALIBUR project.</p> <p>The script is as follows:</p> <p>Ever wondered what happens next? Normally wastewater is cleaned up, contaminants removed, and returned to the water cycle. The sweet brown residue is known as sewage sludge. Yum. The SCALIBUR project is developing innovative technologies to convert urban sewage sludge into valuable products. Where we see waste SCALIBUR partners see a resource. Two systems are under development. A new start-to-end valorisation process will turn sludge into bio-fertilisers for agriculture, and biogas, which is converted into high value compounds through bioelectrochemical systems. Additionally a novel demo plant is being built for the production of PHA bioplastics from sludge. These technologies will help cities manage waste in a more sustainable and cost efficient way. And contribute to the creation of a truly circular bioeconomy in Europe.</p>
Metadata on EUbOPEN multiplex chemogenomic compound screen, wave 1
<p>This is the metadata about EUbOPEN multiplex chemogenomic compound screen, wave 1. The corresponding image data is found at <a href="https://www.ebi.ac.uk/biostudies/studies/S-BIAD145">https://www.ebi.ac.uk/biostudies/studies/S-BIAD145</a>.</p> <p>To compile the metadata Excel file into filelists, please use the Python scripts at: <a href="https://doi.org/10.5281/zenodo.6325622">https://doi.org/10.5281/zenodo.6325622</a>.</p> <p> </p>
Dimethylsulfoniopropionate-derived compound concentrations, volatile organic compound concentrations, and microorganism abundances around two corals and a seaweed in the reefs of Moorea (French Polynesia)
<p>These data belong in the paper: </p> <p>M. Masdeu-Navarro, J-F. Mangot, L. Xue, M. Cabrera-Brufau, S.G. Gardner, D.J. Kieber, J.M. González, R. Simó (2022). Spatial and diel patterns of volatile organic compounds, DMSP-derived compopunds and planktonic microorganisms around a tropical scleractinian coral colony. <em>Frontiers in Marine Science</em>.</p> <p>Concentrations of DMSP, acrylate, DMSO, DMS, DMDS, COS, CS2, isoprene, CH3I, CH2ClI, CH2Br2 and CHBr3 in seawater samples around colonies of the corals Acropora pulchra and Pocillopora sp., and the brown seaweed Turbinaria ornata. Abundances of high-DNA and low-DNA bacteria, Prochlorococcus, Synechococcus, picoeukaryotes and nanoeukaryotes in the same samples, as determined by flow cytometry. All samples were collected in April 2018 in the coral reefs of Mo'orea, French Polynesia. </p> <p>The upper set of data contains concentrations at the distance of 0.5 cm from the coral polyps on the branch tips or verrucae, as well as from the seaweed thalli (samples IN), and 2 m away, downcurrent (samples OUT). The second set of data corresponds to A. pulchra only, and contains seawater samples IN, OUT and AL, the latter being sampled at 0.5 cm from the base of the dead branches colonized by a turf alga. IN, OUT and AL samples were collected over an entire diel cycle, every 6 hours for a period of 30 hours.</p>
Dataset for publication: Compound parabolic collector solar disinfection system for the treatment of harvested rainwater, Strauss et al. (2018). DOI:10.1039/c8ew00152a.
<p>Datasets used for the publication: Strauss A, Reyneke B, Waso M and Khan W (2018) Compound parabolic collector solar disinfection system for the treatment of harvested rainwater. Environ Sci: Water Res. Technol. DOI: 10.1039/c8ew00152a. Please cite the article when using the datasets.</p> <p>Available datasets:</p> <ul> <li>WATERSPOUTT_688928_US_Environmental Conditions_01_1.0.0: Dataset describing the environmental conditions on sampling days while assessing a SODIS-CPC reactor for the treatment of roof-harvested rainwater.</li> <li>WATERSPOUTT_688928_US_SODIS-CPC Schematics_01_1.0.0: Schematic diagrams showing the design of the SODIS-CPC reactor.</li> <li>WATERSPOUTT_688928_US_SODIS-CPC-Microbiology_01_1.0.0: Dataset describing the results obtained while monitoring the microbiological quality of the roof-harvested rainwater before and after treatment with the SODIS-CPC reactor.</li> <li>WATERSPOUTT_688928_US_SODIS-CPC-Physicochemical_01_1.0.0: Dataset describing the physicochemical quality of the roof-harvested rainwater before and after treatment with the SODIS-CPC reactor.</li> <li>WATERSPOUTT_688928_US_UV-transmittance_01_1.0.0: Dataset describing the UV transmittance of polymethyl methacrylate and borosilicate glass.</li> </ul>
Compound annual growth rate for software: replication package
<p>This repository contains the reproducibility package (software and data) for the following paper.</p> <p>Les Hatton, Diomidis Spinellis, and Michiel van Genuchten. The long-term growth rate of evolving software: Empirical results and implications. <em>Journal of Software: Evolution and Process</em>, 29(5), May 2017. <a href="http://dx.doi.org/10.1002/smr.1847">doi:10.1002/smr.1847</a></p> <p>The amount of code in evolving software-intensive systems appears to be growing relentlessly, affecting products and entire businesses. Objective figures quantifying the software code growth rate bounds in systems over a large time scale can be used as a reliable predictive basis for the size of software assets. We analyze a reference base of over 404 million lines of open source and closed software systems to provide accurate bounds on source code growth rates. We find that software source code in systems doubles about every 42 months on average, corresponding to a median compound annual growth rate (CAGR) of 1.21±0.01. Software product and development managers can use our findings to bound estimates, to assess the trustworthiness of road maps, to recognise unsustainable growth, to judge the health of a software development project, and to predict a system’s hardware footprint.</p> <p> </p>
S52 | THSMOKE | Thirdhand Smoke (THS) Compounds
<p>This is the collection associated with list S52 THSMOKE on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/?q=suspect-list-exchange">https://www.norman-network.com/?q=suspect-list-exchange</a></p> <p>S52 </p> <p>THSMOKE</p> <p><strong>Thirdhand Smoke (THS) Compounds</strong></p> <p>THSMOKE XLSX , CSV (06/05/2019)<br> CompTox <a href="https://comptox.epa.gov/dashboard/chemical_lists/thsmoke">THSMOKE List</a></p> <p>THSMOKE InChIKeys (06/05/2019)</p> <p>Thirdhand Smoke (THS, the tobacco-related gases and particles that become embedded in materials), suspect list compiled by Sonia Torres and Noelia Ramirez (IISPV-URV) and Emma Schymanski (LCSB)</p>
Time-resolved compound repositioning predictions on a text-mined knowledge network
<p><strong>gs_positives.csv</strong>: The re-processed version of DrugCentral indications, utilized as training and testing positives in the analysis.</p> <p><strong>top_5000_predictions.csv</strong>: The top 5000 drug-disease pairs, by probability, produced by this analysis pipeline.</p> <p><strong>file_info.txt</strong>: Information about the column headings in of the two files.</p> <p> </p>
MeSCCon - Medical Spanish Chemical compound, drug and medication Name Lexicon (unfiltered version)
<p>The MeSCCon (Medical Spanish Chemical compound, drug and medication Name Lexicon) consists of a list or gazetteer of candidate names of chemicals, drugs, and medications mentioned in Spanish clinical texts. Thus MeSCCon serves as a lexical resource or dictionary for automatic detection of chemical/drug mentions, as well as indexing or classification of medical texts with such concept types.</p> <p>This collection was generated in a five step procedure:</p> <ol> <li>Automatic detection of mentions of chemicals and drugs in biomedical texts in English (including mapping/normalization to MeSH terms or ChEBI identifiers).</li> <li>Generation of a unique name list from the detected concept mentions.</li> <li>Basic filtering of non-chemical names or highly ambiguous mentions-abbreviations using basic characteristics like name morphology and length criteria.</li> <li>Automatic translation of name lists from English to Spanish using a medical machine translation system (see Soares, F. and Krallinger, M. BSC Participation in the WMT Translation of Biomedical Abstracts. In <em>Proceedings of the Fourth Conference on Machine Translation, Volume 3: Shared Task Papers, </em>pp. 175-178 2019; https://zenodo.org/record/3346802)</li> <li>Automatic mention lookup of translated names in a collection of 20 million Spanish clinical notes (primary care and pedriatrics).</li> </ol> <p>Every term in MeSCCon is identified by a text span (in Spanish), a target terminology namespace to which it was automatically mapped (MeSH or ChEBI) and the corresponding concept identifier in that terminology.</p> <p>Moreover, we provide for every text span the absolute term frequency, i.e. the number of matches in the corpus of 20 million clinical notes and the number of documents or notes in which it was found.</p> <p>Important note: no manual filtering of the MeSCCon was carried out, implying that some entries might comprise errors, either due to the initial name recognition and concept mapping in English or due to wrong automatic translations into Spanish.</p> <p>The MeSCCon resource is provided in two formats:</p> <ul> <li>TSV. Data is separated by tabs (\t). Every row of the file has the following fields:</li> </ul> <pre><code>terminology identifier translatedTerm termCount documentCount</code></pre> <ul> <li>JSON. Records are stored as a list of JSON objects. They have the following fields:</li> </ul> <pre><code class="language-javascript">{ "terminology":"MESH", "identifier":"D009020", "translatedTerm":"clorhidrato de morfina", "termFrequency":1, "documentFrequency":1 }</code></pre> <p> </p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital</p>
Data on respiration, substrate incorporation, and soil compound concentration in response to simulated root exudation
<p>In this study we used reverse microdialysis to release a mixture of <sup>13</sup>C-labeled substrates into intact meadow and forest soil cores (6-hour long) to simulate root exudation. We utilized three different artificial root exudates: sugars (glucose, fructose), organic acids (acetate, succinate), and a combination of sugars and organic acids (glucose, fructose, acetate, succinate); alongside a water-only control for comparison.</p> <p>We collected compounds from soil solutions and measured respiration. Due to <sup>13</sup>C-labeled substrate we could differentiate between substrate-derived respiration and SOM-derived respiration. Additionally, we extracted lipid fatty acids from soil and measured their <sup>13</sup>C incorporation.<br><br></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.