Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Pose Selector Workflow - Structure Input Files for Machine Learning (Set 2)
<p>Second set of structure input files for the docking poses of the remaining 2022 protein-ligand complexes. Together with the structure files and absolute binding free energy (ABFE) estimates shared in 10.5281/zenodo.11397017, this data can be used to train a machine-learning (ML) model predicting the ABFE of binding poses of protein-ligand complexes.</p> <p>More background on the workflow generating the structure files and ABFE estimates is provided in 10.5281/zenodo.11397017.</p>
Pose Selector Workflow - Docking Poses, Absolute Binding Free Energy Estimates and Structure Input Files for Machine Learning
<p>The Pose Selector (PS) workflow calculates absolute binding free energies (ABFEs) for binding poses of protein-ligand complexes. First, it converts the binding poses (both docking poses as well as experimentally observed ligand binding poses), which are provided as a combination of protein PDB file and ligand MOL2 file, into input files for molecular dynamics (MD) simulations with GROMACS after they have passed extensive quality checks and repair steps. Next, the PS workflow post-processes and analyses the last frame of the resulting eight 100 ps trajectories per binding pose with the Generalised Born model of implicit solvation as implemented in gmx_MMPBSA to obtain the ABFE estimates. The workflow was designed for soluble proteins without post-translational modifications, co-factors and non-standard amino acids, and it has limited support for coordinated ions.</p> <p>For the dataset published here, the PS workflow was run on docking poses generated for the PDBbind 2020 dataset (http://www.pdbbind.org.cn/index.php), shared in dockingPosesPDBBind2020.tar.gz. This entry and its partner entry 10.5281/zenodo.11397486 also share the intial coordinates used in the MD simulations of >800,000 docking poses of 4022 protein-ligand complexes (structureFiles_dockingPoses1.tar.gz in this entry and structureFiles_dockingPoses2.tar.gz in 10.5281/zenodo.11397486) and of the experimental ligand binding pose of 4549 complexes (structureFiles_experimentalStructures.tar.gz) as well as the corresponding ABFE estimates (absoluteBindingFreeEnergyEstimates.tar.gz). The MD simulations were run on the LUMI and MeluXina supercomputers while the implicit-solvent calculations were carried out on Galileo (Cineca).</p> <p>The README file describes the structure of the shared data in more detail and points out how to reproduce the MD trajectories and the subsequent implicit-solvent calculations yielding the free-energy estimates as well as how to use the data provided in this entry to train a machine-learning model predicting the ABFE of binding poses of protein-ligand complexes. The workflow scripts can be downloaded from GitHub (https://github.com/LigateProject/Pose-Selector-workflow). The MD simulations were run with GROMACS 2023.2 (https://manual.gromacs.org/2023.2/index.html), and the implicit-solvent calculations were carried out with gmx_MMPBSA 1.6.1 (https://valdes-tresanco-ms.github.io/gmx_MMPBSA/v1.6.1/).</p>
Raw data for: Use the 4S (Signal-Safe Speckle Subtraction): Explainable Machine Learning reveals the Giant Exoplanet AF Lep b in High-Contrast Imaging Data from 2011
<p>This collection of data contains all raw data needed to reproduce the results in the paper:</p> <p>Use the 4S (Signal-Safe Speckle Subtraction): Explainable Machine Learning reveals the Giant Exoplanet AF Lep b in High-Contrast Imaging Data from 2011</p> <p>It can also be used as a demonstration dataset for our Python package fours.</p> <p>More details can be found in the online documentation of our python package:<br><a href="https://fours.readthedocs.io/en/latest/">https://fours.readthedocs.io/en/latest/</a></p>
Intermediate results for: Use the 4S (Signal-Safe Speckle Subtraction): Explainable Machine Learning reveals the Giant Exoplanet AF Lep b in High-Contrast Imaging Data from 2011
<p>This collection contains all intermediate results needed to reproduce the results in the paper:</p> <p>Use the 4S (Signal-Safe Speckle Subtraction): Explainable Machine Learning reveals the Giant Exoplanet AF Lep b in High-Contrast Imaging Data from 2011</p> <p>You can use these intermediate results to create all plots in our paper without the need to run all experiments on a large cluster.</p> <p>More details can be found in the online documentation of our python package:<br><a href="https://fours.readthedocs.io/en/latest/">https://fours.readthedocs.io/en/latest/</a></p>
In situ conductometry for studying the homogenization of Al-Mg-Si alloys and predicting extrudate grain structure through machine learning
<p>This dataset includes the <em>in situ</em> impedance and time/temperature data from [1], grain structure data created by extrusion simulation coupled with physically-based microstructural simulation [2], and the predictions of the feed-forward neural network GRAINN-1/2 [1].</p> <p>[1] Österreicher, J. A., Zivanovic, D., Walenta, W., Maimone, S.,Hofbauer, M., Hovden, S., Tükör, Z., Arnoldt, A., Cerny, A. Kronsteiner, A., Antic, M., Zickler, G., Ehmeier, F., Mikulovic, M., Kunschert, G. (2024) . In situ conductometry for studying the homogenization of Al-Mg-Si alloys and predicting extrudate grain structure through machine learning. <em>Materials & Design</em>, 113070.</p> <p>[2] Hovden, S., Kronsteiner, J., Arnoldt, A., Horwatitsch, D., Kunschert, G., & Österreicher, J. A. (2024). Parameter study of extrusion simulation and grain structure prediction for 6xxx alloys with varied Fe content. <em>Materials Today Communications</em>, <em>38</em>, 108128.</p>
Auxiliary data files for replication of "Augmenting the availability of historical GDP per capita estimates through machine learning"
<p>This repository holds auxiliary data files needed for the replication "Augmenting the availability of historical GDP per capita estimates through machine learning". All further information and data is provided in the <a href="https://github.com/philmkoch/historicalGDPpc" target="_blank" rel="noopener">GitHub repository</a>.</p> <p>The data included in this auxiliary folder is based on the work by Laouenan et al. (https://www.nature.com/articles/s41597-022-01369-4). </p>
Dataset for: Adapting Explainable Machine Learning to Study Mechanical Properties of Two-Dimensional Hybrid Halide Perovskites
<p>This archive contains the in plane and out of plane Young's moduli (complete with respective VASP in and outputs) for 154 n=1 and 30 n>1 2D hybrid organic and inorganic perovskites. The data was used in the publication "Adapting Explainable Machine Learning to Study Mechanical Properties of Two-Dimensional Hybrid Halide Perovskites".</p> <p>Computational settings for the calculations were:</p> <p>Perdew-Burke-Ernzerhof (PBE) exchange-correlation with Tkatchenko-Scheffler (TS) van der Waals (vdW) corrections<br>Projector augmented-wave (PAW) method for the description of interactions between core and valence electrons.<br>A plane wave cutoff energy of 520 eV<br>A Γ-centered Monkhorst-Pack k-point mesh with a grid spacing of 2π × 0.040 Å−1 <br>Geometry optimizations were performed until energy and residual forces fell below 10−6 eV and 0.001 eV/ Å, respectively. <br><br></p>
Atmospheric rivers dataset for machine learning training
<p>A thorough description of the data and how it was created can be found: <a title="http://climate-cms.org/CNN-Atmospheric-Rivers/" href="http://climate-cms.org/CNN-Atmospheric-Rivers/" target="_blank" rel="noopener">http://climate-cms.org/CNN-Atmospheric-Rivers/</a></p> <p>A Jupyter notebook has also been created where we'll show you how to use this data to train a deep learning model to identify whether an Integrated Vapor Transport map contains an atmosperhic river. it can be found: <a title="CNN_AR_tutorial.ipynb" href="https://github.com/coecms/CNN-Atmospheric-Rivers/blob/main/CNN_AR_tutorial.ipynb" target="_blank" rel="noopener">CNN_AR_tutorial.ipynb</a> and had been published here: </p> <p>Mesto, M., Hobeichi, S., & Green, S. (2024). CNN-Atmospheric-Rivers (v1.0.0). Zenodo. <a href="https://doi.org/10.5281/zenodo.12538779" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.12538779</a></p> <p>The data is organised in three folders:</p> <ul> <li>IVT_ERA5_2Deg: Contains global IVT data.</li> <li>AR_Global: Contains polygons representing AR objects identified and hand-labelled in each IVT map.</li> <li>Training_Testing_tiles: Contains tiles of IVT data with an annotation file that classifies each tile as one of the following: ‘Atmospheric River’, ‘Ambiguous’. The ‘Ambiguous’ class refers to objects that are not clearly identifiable as atmospheric rivers.</li> </ul> <p><strong>Integrated Vapor Transport maps:</strong></p> <p>These maps were computed using the magnitude of the vertical integral of northward and eastward water vapour flux variables from ERA5. The IVT values are expressed in units of kg m^-1 s^-1. All the IVT TIFF files were loaded into ArcGIS software and displayed using a colour scheme that allows for the visual identification of atmospheric rivers. Data details:</p> <ul> <li>Folder: <strong><em>IVT_ERA5_2Deg</em></strong></li> <li>File format: TIFF</li> <li>Spatial resolution: 2 degrees</li> <li>Spatial coverage: Global (longitude: -180 to 180 , latitude: -90 to 90)</li> <li>Geographic Coordinate System: GCS_WGS_1984</li> <li>Temporal coverage: 1<sup>st</sup> – 5<sup>th</sup> day of January, April, July, October for 2010, 2013, 2015; these years correspond to La Niña, neutral, and El Niño year respectively</li> <li>Temporal resolution: Daily</li> <li>Naming of files: ivt_2deg_ddmmyyyy.tif</li> <li>Number of files: 60 (5 days × 4 months × 3 years)</li> <li>Number of channels in each file: 1</li> </ul> <p> </p> <p><strong>Atmospheric Rivers in IVT maps:</strong></p> <p>The annotation tool ‘Label Objects for Deep Learning’ was used to draw polygons to cover the shape of atmospheric rivers on each IVT map. Each polygon was assigned one of two labels: 'Atmospheric Rivers' or 'Ambiguous'. The polygons were drawn based on visual identification of the shape of atmospheric rivers, guided by IVT values close to 500kg m^-1 s^-1 as in Reid et al (2020). The 'Ambiguous' label was assigned to objects that were unclear in their classification as ARs. This ambiguity arose from objects that were shorter, wider, had slightly lower IVT values, or it was hard to tell if they were ARs of tropical cyclones during the early stages of their formation. Data details:</p> <ul> <li>Folder: <strong><em>AR_Global</em></strong></li> <li>File format: SHP (shapefile)</li> <li>Spatial coverage: Global</li> <li>Geographic Coordinate System: GCS_WGS_1984</li> <li>Temporal coverage: 1<sup>st</sup> – 5<sup>th</sup> day of January, April, July, October for 2010, 2013, 2015 (corresponding to La Niña, neutral, and El Niño year respectively)</li> <li>Temporal resolution: Daily</li> <li>Naming of files: ivt_2deg_ddmmyyyy_labelled.shp</li> <li>Number of files: 60 (5 days × 4 months × 3 years)</li> </ul> <p> </p> <p><strong>Dataset for deep learning training:</strong></p> <p>The tool 'Export Training Data for Deep Learning' uses the IVT maps in the 'AR_Global' folder and the shapefiles in the 'IVT_ERA5_2Deg' folder to create labelled tiles for deep learning training. Each tile in the map is assigned a label: 'Atmospheric River', 'Ambiguous', or no label if it doesn’t contain any AR or ambiguous shape. The generated map chips are stored in folder <em>Training_Testing_tiles/ RCNN_Masks_All_Tiles (no 3-10-15)/images</em>, and the labels are provided in the textfile 'map.txt' file in folder Training_Testing_tiles/ RCNN_Masks_All_Tiles (no 3-10-15)/. Data details:</p> <ul> <li>Folder: <strong><em>Training_Testing_tiles/</em> RCNN_Masks_All_Tiles (no 3-10-15)/images</strong></li> <li>File format: TIFF</li> <li>Spatial resolution: 2 degrees</li> <li>Spatial coverage: varies. Width of tile = 40 gridcells. Height of tile = 20grid cells</li> <li>Geographic Coordinate System: GCS_WGS_1984</li> <li>Temporal coverage: 1<sup>st</sup> – 5<sup>th</sup> day of January, April, July, October for 2010, 2013, 2015. These years correspond to La Niña, neutral, and El Niño year respectively. Please note that data for certain days are missing; these omissions correspond to days with no or only a single atmospheric river detected.</li> <li>Temporal resolution: Daily</li> <li>Number of files: varies</li> <li>Number of channels in each file: 1</li> </ul>
Machine learning and bioinformatics analysis of diagnostic biomarkers associated with the occurrence and development of lung adenocarcinoma
Open the record for dataset details and reuse information.
Data from: Machine learning without a processor: Emergent learning in a nonlinear analog network
<p>The capabilities of digital artificial neural networks grow rapidly with their size, however the time and energy required to train them does as well. The tradeoff is far better for Brains, where the constituent parts (neurons) update their analog connections in ignorance of the actions of other neurons, eschewing centralized processing. Recently introduced analog electronic <em>contrastive local learning networks </em>(CLLNs) share this important decentralized property. However their capabilities were limited because existing implementations are linear. In this dataset we include experimental demonstrations of a nonlinear CLLN, establishing a new paradigm for scalable learning. Included here are data and scripts required to generate figures 2-6 of the manuscript titled "Machine learning without a processor: Emergent learning in a nonlinear analog network".</p>
Nissl_4, Raw images for Machine learning for histological annotation and quantification of cortical layers.
<p>This dataset contains some images (TIFF image data) of <strong>brain juvenile rats Wistar Han (P14)</strong> scanned immunostained slides by using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel.</p> <p>These raw images are part of another Zenodo dataset <span><a href="https://doi.org/10.5281/zenodo.11544829" target="_blank" rel="noopener noreferrer">https://doi.org/10.5281/zenodo.11544829</a></span>, that contains the QuPath projects that uses this dataset and 5 others (from Nissl_1 to Nissl_6).</p>
Nissl_3, Raw images for Machine learning for histological annotation and quantification of cortical layers.
<p>This dataset contains images (TIFF image data) of <strong>brain juvenile rats Wistar Han (P14)</strong> scanned immunostained slides by using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel.</p> <p>These raw images are part of another Zenodo dataset <span><a href="https://doi.org/10.5281/zenodo.11544829" target="_blank" rel="noopener noreferrer">https://doi.org/10.5281/zenodo.11544829</a></span>, that contains the QuPath projects that uses this dataset and 5 others (from Nissl_1 to Nissl_6).</p>
Nissl_2, Raw images for Machine learning for histological annotation and quantification of cortical layers.
<p>This dataset contains images (TIFF image data) of <strong>brain juvenile rats Wistar Han (P14)</strong> scanned immunostained slides by using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel.</p> <p>These raw images are part of another Zenodo dataset <span><a href="https://doi.org/10.5281/zenodo.11544829" target="_blank" rel="noopener noreferrer">https://doi.org/10.5281/zenodo.11544829</a></span>, that contains the QuPath projects that uses this dataset and 5 others (from Nissl_1 to Nissl_6).</p>
Nissl_1, Raw images for Machine learning for histological annotation and quantification of cortical layers.
<p>This dataset contains images (TIFF image data) of <strong>brain juvenile rats Wistar Han (P14)</strong> scanned immunostained slides by using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel.</p> <p>These raw images are part of another Zenodo dataset <span><a href="https://doi.org/10.5281/zenodo.11544829" target="_blank" rel="noopener noreferrer">https://doi.org/10.5281/zenodo.11544829</a></span>, that contains the QuPath projects that uses this dataset and 5 others (from Nissl_1 to Nissl_6).</p> <p> </p>
Nissl_6, Raw images for Machine learning for histological annotation and quantification of cortical layers.
<p>This dataset contains images (TIFF image data) of <strong>brain juvenile rats Wistar Han (P14)</strong> scanned immunostained slides by using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel.</p> <p>These raw images are part of another Zenodo dataset <span><a href="https://doi.org/10.5281/zenodo.11544829">https://doi.org/10.5281/zenodo.11544829</a></span>, that contains the QuPath projects that uses this dataset and 5 others (from Nissl_1 to Nissl_6).</p> <p> </p> <p> </p>
Nissl_5, Raw images for Machine learning for histological annotation and quantification of cortical layers
<p>This dataset contains images (TIFF image data) of <strong>brain juvenile rats Wistar Han (P14)</strong> scanned immunostained slides by using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel.</p> <p>These raw images are part of another Zenodo dataset <span><a href="https://doi.org/10.5281/zenodo.11544829" target="_blank" rel="noopener noreferrer">https://doi.org/10.5281/zenodo.11544829</a></span>, that contains the QuPath projects that uses this dataset and 5 others (from Nissl_1 to Nissl_6).</p>
The state-of-the-art machine learning model for Plasma Protein Binding Prediction: computational modeling with OCHEM and experimental validation
<p><span>Institute of Materia Medica, Chinese Academy of Medical Sciences purchased 10,000 ChemDiv databases.</span></p>
FIGURE 1 in Numerical taxonomy and genus-species identification of Czekanowskiales in China based on machine learning
FIGURE 1. Distribution of Mesozoic Czekanowskiales fossils in China (Chinese basemap from the Standard Map Ser- vice, plan approval number: GS (2020) 4619).
FIGURE 4 in Numerical taxonomy and genus-species identification of Czekanowskiales in China based on machine learning
FIGURE 4. Trait importance scores for genus and species identification and heatmap of correlations between traits. (A) Importance scores for traits identified genus and species, with the solid symbols indicating key traits and the hollow symbols indicating nonkey traits; (B) proportion of importance scores for macro traits and cuticular traits identified at the genus and species; (C) heatmap of correlations between traits; serial numbers correspond to traits in Table 2.
FIGURE 5 in Numerical taxonomy and genus-species identification of Czekanowskiales in China based on machine learning
FIGURE 5. Accuracy distributions of five supervised learning algorithms and confusion matrix of the best algorithm. (A) Accuracy distribution of the five algorithms in genus-level identification (using mixed-key traits); (B) CART algorithm confusion matrix (using mixed-key traits); (C) accuracy distribution of five algorithms in genus-level identification (using only macro-key traits); (D) CART algorithm confusion matrix (using only macro-key traits); (E) accuracy distribution of five algorithms for species-level identification (using mixed-key traits); (F) LR algorithm confusion matrix (using mixed-key traits); (G) accuracy distribution of five algorithms for species-level identification (using only macro-key traits); (H) CART algorithm confusion matrix (using only macro-key traits). Genus abbreviations: Cz., Czekanowskia; Ph., Phoenicopsis; Sp., Sphenarion; So., Solenites.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.