Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Provably efficient machine learning for quantum many-body problems
<p>Raw data for the manuscript "Provably efficient machine learning for quantum many-body problems".</p>
KUALA: A Machine Learning-driven framework for kinase inhibitors repositioning
<p>The complete list of all predicted compounds for each kinase accompanied by RT thresholds.</p>
Supplementary information for: A continuous-score occupancy modeling framework for incorporating uncertain machine learning output in autonomous biodiversity surveys
<p><span>Ecologists often study biodiversity by evaluating species occupancy and the relationship between occupancy and other covariates. Occupancy models are now widely used to account for false absences in field surveys and to reduce bias in estimates of covariate relationships. Existing occupancy models take as inputs binary detection/non-detection observations of species at each visit to each site. However, autonomous sensing devices and machine learning models are increasingly used to survey biodiversity, generating a new type of observation record (i.e., continuous-score data) that reflects the model's confidence a species is present in each autonomously sensed file, instead of binary detection/non-detection data. These data are not directly compatible with traditional binary occupancy modeling methods.</span></p> <p><span>Here, we develop a new occupancy model that models continuous scores on a visit level as a Gaussian mixture, combining a distribution of scores for files that do contain the species of interest and a distribution of scores for files that do not. The model takes as input continuous scores for each autonomously sensed and classified file, along with an optional small number of binary, manually verified detection and non-detection annotations.</span></p> <p><span>We present a simulation study that shows that over a range of empirically realistic parameters, our model outperforms traditional occupancy models that are based on binary annotation alone. We also apply this new model to an empirical case study using data generated from five machine learning classifiers applied to autonomous acoustic recordings gathered in the eastern United States.</span></p> <p><span>Because our occupancy model generalizes allowable input data beyond binary observations, it is particularly well-suited to the increasing volume of machine learning classified data in ecology and conservation.</span></p>
Development of classifier of engagement in occupation with machine learning (CEOML) for quantifying context
<p>These are the raw data, code, and development model from the development of the classifier of engagement in occupation with machine learning (CEOML).</p>
Modeling Crash Severity and Collision Types Using Machine Learning
<p>Traffic safety analysis is the fundamental step for reducing economic, social, and environmental cost incurred due to traffic accidents. The essence of traffic safety is understanding the factors affecting crash occurrence, injury severity and collision type and their underlying relationships and predict-prevent future crash instances. Crash injury severity studies in past have utilized numerous statistical, econometric and Machine Learning (ML) and Artificial Intelligence (AI) tools to extract the underlying relationship between the crash causal factors and the consequent severity or collision type. The study aims to explore the Multi-Label Classification (MLC) tool from the domain of Artificial Intelligence (AI) for classification problems in the setting of traffic safety. MLC finds its application primarily in protein function, semantic scene, and music categorization problems. In the real world, multiple heterogenous subjective factors decide the extent of damage/severity of a particular crash instance. Theoretically, the traffic collision type and crash severity type can be correlated, and thus, it is intuitive to model them simultaneously. The ability of MLC to categorize an entity under analysis to more than one labels, correlated or uncorrelated, provides the approach an edge over the single-class (binary) or multi-class classification approach. The MLC based classification model was calibrated and tested using the historical crash data extracted for the state of Texas. The selection of study area was based on a link-level unsupervised principal component analysis-based clustering approach. Similar clustering approach was also tested at the county-level to understand the spatial behavior and thus transferability of the MLC approach to other key cities in the state. The performance of the proposed approach was tested, compared, and quantified with the conventional binary/multi-class classification tools used in the traffic safety domain. Inferences from the preliminary numerical analysis indicates that the proposed multi-label classification approach has promising performance compared to the traditional classification approaches, specifically found in traffic safety literatures.</p>
Data for "Quantum-corrected thickness-dependent thermal conductivity in amorphous silicon predicted by machine learning molecular dynamics simulations"
<p>This is the data set for the preprint <a href="https://arxiv.org/abs/2206.07605">arXiv:2206.07605</a> [cond-mat.mtrl-sci], obtained by the GPUMD code.</p> <p>Here are 6 directories.<br> 1). NEMD<br> 2). NEPpotential<br> 3). PDOS<br> 4). kappa-quenchRate<br> 5). kappa-size<br> 6). kappa-temperature<br> <br> 1). NEMD directory contains calculations of ballistic conductance using NEMD method, where 6 independent cycles are run to average.</p> <p>2). NEPpotential directory is the trained NEP potential.</p> <p>3). PDOS directory contains phonon density of states of a-Si samples generated by the quench rate of 10^{11} K/s.</p> <p>4). kappa-quenchRate directory contains HNEMD calculations of a-Si samples which are prepared using melt-quench temperature protocols with the quench rates covering from 10^{11} to 5x10^{12} K/s. In each case, 3 independent cycles are run.</p> <p>5). kappa-size directory contains HNEMD calculations based on different supercells. 6 independent cycles are run.</p> <p>6). kappa-temperature directory contains HNEMD calculations of a-Si samples which are prepared for different targeted temperatures using slow quench rate of 10^{11} K/s.</p> <p> </p>
Direct prediction for carbapenemase-producing and colistin-resistant Klebsiella pneumoniae isolates from routine MALDI-TOF mass spectrum using machine learning
<p>The emergence of carbapenem-nonsusceptible K. pneumoniae (CnSKP) leads a serious threat to patient survival and colistin resistance makes the treatment of CnSKP more difficultly. To make treatment strategy properly and quickly, we aimed to develop a rapid prediction method for CnSKP and colistin-resistant K. pneumoniae (ColRKP) based on the spectra of routine matrix-assisted laser desorption/ionization-time-of-flight mass spectrometry (MALDI–TOF MS). The machine learning (ML) model for differentiating CnSKP and carbapenem-susceptible K. pneumoniae (CSKP) showed accuracy of 0.8869 and AUC of 0.9551; the model for ColRKP and colistin-intermediate K. pneumoniae (ColIKP) showed accuracy of 0.8361 and the AUC of 0.8447.</p>
Using photodiodes and supervised Machine Learning for automatic classification of weld defects in laser welding of thin foils copper-to-steel battery tabs
<p>In this folder, excel files are stored with the results of signal processing that supported findings in the following paper:</p> <p>"Using photodiodes and supervised Machine Learning for automatic classification of weld defects in laser welding of thin foils copper-to-steell battery tabs".</p> <p>Matlab scripts and orginal signals will be uploaded soon with more detailed description.</p> <p> </p>
Machine learning techniques with code
<p>This dataset contains data from Paperswithcode.com obtained on January of 2022. The first file, 'Papers_with_abstracts' includes information about different research papers like their title, abstract, authors, etc. 'Links_between_papers_and_code' includes each paper connection to their corresponding github repository. Finally, 'Methods' includes a categorization of the aforementioned papers done by the community of Paperswithcode in different areas. These files were used in the Master Thesis Topic Modeling for Research Software done by María Ayuso in Universidad Politécnica de Madrid.</p>
Research data supporting: "Classifying soft self-assembled materials via unsupervised machine learning of defects"
<p>Research data supporting: "Classifying soft self-assembled materials via unsupervised machine learning of defects".</p> <p>The root folder contains 5 folders:</p> <ol> <li>FIBERS</li> <li>MEMBRANES_and_MICELLES</li> <li>NANOPARTICLES</li> <li>COMPARISON</li> <li>paper_images</li> </ol> <p>The folders 1. to 3. contain the data for every soft-matters architecture used to produce the results discussed in the main paper. Each of these folders contain additional sub-fordels: TRAJ, SOAP, PCA, CLUSTERING, containing the files discussed in the main paper.</p> <p>Folder 4. contains the data of the comparison between different classes of materials (SOAP, PCA, and CLUSTERING sub-folders).</p> <p>Folder 5. contains the images that are showed in the main paper and in the Supporting Information.</p>
Surrogate-modelling & machine learning dataset : finite element stress analysis of biaxial specimen with random elastic properties - 1000 samples
<p>Dataset finite element stress analysis of biaxial specimen with random elastic properties</p> <p>Unzip and execute dataset.py to visualise data samples. PyVista is needed.</p>
Surrogate-modelling & machine learning dataset : finite element stress analysis of biaxial specimen with random elastic properties - 100 samples
<p>Dataset finite element stress analysis of biaxial specimen with random elastic properties</p> <p>Unzip and execute dataset.py to visualise data samples. PyVista is needed</p>
Dynamics of the euphotic zone in the Black Sea: The synergy of data from profiling floats, machine learning and numerical modeling
<p>The datasets contain input data and data emulated by Neural networks (NN) used in the study 'Dynamics of the euphotic zone in the Black Sea: The synergy of data from profiling floats, machine learning and numerical modeling'</p> <p>- <strong>NN2018_CMEMS_ARGO.tar.gz</strong>: archive contains Matlab binary files consisting of NN-derived BGC variables (Chlorophyll-a, Oxygen and backscatter at 700nm) along ARGO float paths in 2018 with vertical resolution taken from CMEMS (13 depth levels in the depth range studied here); NN was applied either on CMEMS physics ('C') or on ARGO physics ('A'); additionally, CMEMS BGC model data (Chlorophyll-a and Oxygen) along these paths are included; (filenames follow the names of floats given in Table 1:<em> floatname</em>_2018_NNARGOCMEMS.mat)</p> <p>- <strong>NNalongARGO.tar.gz</strong>: archive contains Matlab binary files consisting of NN-derived BGC variables (Chlorophyll-a, Oxygen and backscatter at 700 nm) along ARGO float paths together with input ARGO data (time, latitude, longitude, salinity, temperature, sigma_T and BGC variables); all variables are mapped with 1m vertical resolution; depth range is [1m 150m] ; additionally, float ogs7 data include NO3, float hzg1 data does not contain Chlorophyll-a and backscatter at 700 nm; (filenames follow the names of floats given in Table 1:<em> floatname</em>_euph_1mRes_150mALLINCLNN.mat)</p> <p>- <strong>NNReconBlackSea.tar.gz</strong>: archive contains Matlab binary files consisting of basin wide NN derived BGC variables (Chlorophyll-a, Oxygen and backscatter at 700nm) for the years 2015-19 and 2010/11 (weekly mean data); additionally CMEMS data (time, latitude, longitude, salinity, temperature, sea surface height) are provided for the photic zone; (filenames are reconNNCMEMS_2015_2019.mat and reconNNCMEMS_2010_2011.mat, respectively)</p>
Statistical and machine learning methods for evaluating trends in air quality under changing meteorological conditions
<p>This repo includes the GEOS-Chem simulations and R scripts that are needed to replicate and evaluate the conclusions from Qiu, Zigler, and Selin, ACP, 2022 "Statistical and machine learning methods for evaluating trends in air quality under changing meteorological conditions".</p> <p><strong>The GEOS-Chem simulations</strong></p> <ul> <li>For the US (2011-2017): <ul> <li><em>observational_o3_pm_2011_2017_us.rds:</em> the simulated daily PM2.5 and O3 concentrations, and MERRA-2 meteorological features in the observational scenarios (<strong>changing</strong> meteorology, <strong>changing </strong>emissions).</li> <li><em>counterfactual_o3_pm_2011_2017_us.rds</em><em>:</em> the simulated daily PM2.5 and O3 concentrations, and MERRA-2 meteorological features in the counterfactual scenarios (<strong>constant</strong> meteorology, <strong>changing</strong> emissions).</li> <li><em>constant_emis_o3_pm_2012_2017_us.rds: </em>the simulated daily PM2.5 and O3 concentrations in the constant-emission scenarios (<strong>constant</strong> meteorology, <strong>constant</strong> emissions).</li> <li><em>regional_features_2011_2017_4x5_us.rds: </em>the MERRA-2 meteorological features in the observational scenarios (aggregated to 4x5 degrees), inputs for the "RF-regional" model.</li> </ul> </li> <li>For China (2013-2017): <ul> <li><em>observational_o3_pm_2013_2017_china.rds:</em> the simulated daily PM2.5 and O3 concentrations, and MERRA-2 meteorological features in the observational scenarios (<strong>changing</strong> meteorology, <strong>changing </strong>emissions).</li> <li><em>counterfactual_o3_pm_2013_2017_china.rds:</em> the simulated daily PM2.5 and O3 concentrations, and MERRA-2 meteorological features in the counterfactual scenarios (<strong>constant</strong> meteorology, <strong>changing</strong> emissions).</li> <li><em>constant_emis_o3_pm_2014_2017_china.rds</em><em>: </em>the simulated daily PM2.5 and O3 concentrations in the constant-emission scenarios (<strong>constant</strong> meteorology, <strong>constant</strong> emissions).</li> <li><em>regional_features_2013_2017_4x5_china.rds: </em>the MERRA-2 meteorological features in the observational scenarios (aggregated to 4x5 degrees), inputs for the "RF-regional" model.</li> </ul> </li> </ul> <p><strong>R scripts:</strong></p> <ul> <li><a href="https://zenodo.org/api/files/065be469-ef8d-4c9b-9bd8-a6a808275237/main.r">main.r</a>: the main script to perform statistical correction of meteorological variability.</li> <li>main.r uses functions from the other R script files (see below) which perform different statistical correction methods, respectively. </li> <li><a href="https://zenodo.org/api/files/065be469-ef8d-4c9b-9bd8-a6a808275237/parametric_regression_methods.r">parametric_regression_methods.r</a>: performs meteorological correction with parametric regression methods (MLR, polynomial, spline, GAM)</li> <li><a href="https://zenodo.org/api/files/065be469-ef8d-4c9b-9bd8-a6a808275237/tune_RF_regional.r">tune_RF_regional.r</a> and <a href="https://zenodo.org/api/files/065be469-ef8d-4c9b-9bd8-a6a808275237/RF_regional.r">RF_regional.r</a>: perform the meteorological correction with the "RF-regional" model</li> <li><a href="https://zenodo.org/api/files/065be469-ef8d-4c9b-9bd8-a6a808275237/GEOS_Chem_constant_emis.r">GEOS_Chem_constant_emis.r</a>: performs the meteorological correction using the simulations from the constant emission scenarios from the GEOS-Chem model</li> </ul> <p> </p>
Predicting peak daily maximum 8-hour ozone, and linkages to emissions and meteorology, in Southern California using machine learning methods
<p>crestlinetop30ozone19902019.csv includes the data used to build the modes for the annual 30 highest MDA8 concentrations at the Crestline site.</p>
Machine learning methods detect arm movement impairments in a patient with parieto-occipital lesion using only early kinematic information.
<p>This depository contains the data and codes of the paper: </p> <p>Bosco, A., Bertini C., Filippini M., Foglino C., Fattori P (2022). Machine learning methods detect arm movement impairments in a patient with parieto-occipital lesion using only early kinematic information. <em>Journal of Vision</em>, in press. </p> <p>Data are provided in mat-file, codes is provided in m-files (Matlab).</p>
Machine learning on syngeneic mouse tumor profiles to model clinical immunotherapy response
<p><span>Most cancer patients are refractory to immune checkpoint blockade (ICB) therapy, and proper patient stratification remains an open question. Primary patient data suffer from high heterogeneity, low accessibility, and lack of proper controls. In contrast, syngeneic mouse tumor models enable controlled experiments with ICB treatments. Using transcriptomic and experimental variables from >700 ICB-treated/control syngeneic mouse tumors, we developed a novel machine learning framework to model tumor immunity and identify factors influencing ICB response. Projected on human immunotherapy trial data, we found that the model can predict clinical ICB response. We further applied the model to predicting ICB-responsive/resistant cancer types in TCGA, which agreed well with existing clinical reports. Finally, feature analysis implicated factors associated with ICB response. In summary, our novel computational framework based on mouse tumor data reliably stratified patients regarding ICB response, informed resistance mechanisms, and has the potential for wide applications in disease treatment studies.</span></p>
Early Detection of Nucleation Events from Solution in LC-TEM by Machine Learning
<p>These data are images taken with a liquid cell transmission electron microscope and annotation data of particles.</p> <p>See <a href="https://github.com/hiroyasukatsuno/Early-Detection-of-Nucleation-Events-LC-TEM">this page (GitHub)</a>.</p> <p> </p>
Data for NHESS manuscript by Biass et al. (2022): Insights into the vulnerability of vegetation to tephra fallouts from interpretable machine learning and big Earth observation data
<p>This repository contains the data produced in the context of the following paper:</p> <blockquote> <p>Biass S, Jenkins SF, Aberhard WH, Delmelle P, Wilson T (2022): Insights into the vulnerability of vegetation to tephra fallouts from interpretable machine learning and big Earth observation data, Accepted in NHESS</p> </blockquote> <p>Naming convention is: `run_date`_`landcover`_`impact_metrics`_`VI`_`anomaly`_test.pkl, where:</p> <ul> <li>Landcover is either crops, shrubs, herbaceous vegetation (grass), forests (trees) or all together</li> <li>Impact metrics is either minV (impact magnitude) or minT (impact duration)</li> <li>VI is the vegetation index (here, EVI)</li> <li>Anomaly is the impact indicator (here, cumulative difference index)</li> </ul> <p>Refer to the associated paper for more information on the methodology.</p> <p>Files are saved as .pkl and were generated by the <a href="https://explainerdashboard.readthedocs.io/en/latest/">explainerdashboard</a> library. They are the result of <a href="https://xgboost.readthedocs.io/en/stable/">XGBoost</a> runs that were optimised with <a href="https://optuna.org">Optuna</a> and analysed with the <a href="https://shap.readthedocs.io/en/latest/">SHAP</a> library. They contain:</p> <ol> <li>The explanatory variables and observed and computed target variables for all features</li> <li>The SHAP values</li> </ol> <p>To load the files, use <a href="https://explainerdashboard.readthedocs.io/en/latest/cli.html?highlight=load#explainerdashboard.explainers.BaseExplainer.from_file">this method</a>.</p> <p> </p>
Mapping Phyllosilicates on the Asteroid Bennu Using Thermal Emission Spectra and Machine Learning Model Applications Datasets
<p>We provide laboratory spectra of pure minerals and mineral mixtures and the corresponding metadata. Derived from the laboratory data, we include the PLS model through coefficients. Through the application of the model, we provide the prediction values in volume% for Mg-rich serpentine, cronstedtite, and saponite for the BBD1, EQ3, and TAG datasets.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.