Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “cheminformatics”
Comparative Analysis of Anthraquinone and Chalcone Derivatives-Based Virtual Combinatorial Library. A Cheminformatics "Proof-of-Concept" Study
<p>This computational “proof-of-concept” study illustrated the combinatorial approach used to explain how the selected natural products' structures undergo molecular diversity analysis. A virtual combinatorial library (1.6M) based on 20 anthraquinones and 24 chalcones were enumerated. The resulting compounds were optimized to the near drug-likeness properties and the physicochemical descriptors were calculated for all datasets including FDA, Non-FDA, and natural products (NPs) datasets from ZINC 15. UMAP and principal component analysis (PCA) were applied to compare and represent the chemical space coverage of each dataset. Subsequently, the Laplacian score, and Gini coefficient, were applied to delineate feature selection, and selectivity among properties respectively. Finally, we demonstrated the diversity between the datasets by employing Murcko’s, and central scaffolds systems, calculated three fingerprint descriptors, and analyzed their diversity by PCA and self-organizing maps (SOM). The optimized enumeration resulted in 1,610,268 compounds with NP-Likeness, and synthetic feasibility mean scores close to FDA, Non-FDA, and NPs datasets. The overlap between the chemical space of 1.6M was more prominent with NPs. Laplacian score has prioritized NP-likeness and hydrogen bond acceptor properties (1.0 and 0.923) respectively, while the Gini coefficient showed that all properties have selective effects on datasets (0.81 to 0.93). Scaffold and fingerprint diversity indicated that the descending order for the tested datasets was FDA, Non-FDA, NPs, 1.6M. Virtual combinatorial libraries based on NPs can be considered as a source of the combinatorial compound with NP-likeness properties. Furthermore, measuring molecular diversity is supposed to be performed by different methods to allow for comparison and better judgment. </p> <p>This link provides an illustration of the whole virtual combinatorial library using the TMAP algorithm in addition to the complete dataset. TMAP is a recent algorithm applied to visualize ultra-large high-dimensional chemical libraries for structures and physicochemical properties (Probst & Reymond, 2020). This approach creates and distributes intuitive tree representations of big data sets with arbitrary dimensionality in the order of 10<sup>7</sup>.</p> <p><strong>To visualize the whole library of compounds, download the "index(2).rar", then extract the index.html that pop-up in the WinRAR application.</strong></p>
Workshop Material - 3D-e-Chem Structural Cheminformatics Workflows for Computer-Aided Drug Discovery
<p>The workshop at the KNIME user meeting (Berlin 9th of March 2018) is set up to stimulate participants with varying degrees of experience in cheminformatics to learn and apply the different structural cheminformatics tools and workflows developed within the context of the 3D-e-Chem project. You will learn how to construct and apply integrated cheminformatics workflows using the 3D-e-Chem KNIME nodes for the exploitation of G protein-coupled receptor and kinase data (two important pharmaceutical target classes) to obtain useful information for drug discovery.</p> <p>Information on the 3D-e-Chem KNIME nodes and workflows can be found online:</p> <p>3D-e-Chem GitHub website: <a href="http://3d-e-chem.github.io/">http://3d-e-chem.github.io/</a></p>
Exploiting Vector Pattern Diversity of Molecular Scaffolds for Cheminformatics Tasks in Drug Discovery
<p>Data and code to accompany the paper: <em>Exploiting Vector Pattern Diversity of Molecular Scaffolds for Cheminformatics Tasks in Drug Discovery. </em></p>
Interactive tmaps of the Global Chemicals Inventory (cheminformatic dataset)
<p>Interactive version of the tmaps described in the thesis "Cheminformatic techniques for screening chemicals on the globals market" (2024, ETH Zurich).</p> <p> </p> <p>Due to technical limitations, legends are numerical and are as follows:</p> <p>tmap - total inventories: Numbers correspond to the total number of inventories on which each chemical is registered.</p> <p>tmap - individual inventories: A number of 1 indicates that chemical is registered on the indicated inventory. A number of 0 indicates it is not.</p> <p>tmap - qsar: A number of 0 indicates the chemical is inside the applicability domain of the indicated inventory. A number of 2 indicates it is on the boundary (within 10% of the 95th percentile). A number of 6 indicates it is outside the domain.</p> <p>tmap - innovation: A number of 0 indicates the chemical is excluded from the analysis (coloured black). A number of 1 indicates it is considered existing on the global level. A number of 2 indicates it is mixed, and a number of 3 indicates it is considered new.</p>
PHD Thesis: Graph Set Data Mining - Clustering and Pattern Mining in the Context of Cheminformatics - Evaluation Data
<p>Evaluation data for the PHD thesis:</p> <p>Graph Set Data Mining<br> Clustering and Pattern Mining in the Context of Cheminformatics</p> <p>zur Erlangung des Grades eines<br> Doktors der Naturwissenschaften<br> der Technischen Universität Dortmund<br> an der Fakultät für Informatik</p> <p>von</p> <p>TIll Schäfer</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.