Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Integrated machine learning and GIS-based bathtub models to assess the future flood risk in the Kapuas River Delta, Indonesia
<p>To use the data and the code, please cite the following article: </p> <p>Joko Sampurno, Randy Ardianto, Emmanuel Hanert; Integrated machine learning and GIS-based bathtub models to assess the future flood risk in the Kapuas River Delta, Indonesia. <em><em>Journal of Hydroinformatics</em></em> 2022; jh2022106. DOI: <a href="https://doi.org/10.2166/hydro.2022.106">https://doi.org/10.2166/hydro.2022.106</a></p>
Label-Free Identification of Exosomes Using Raman Spectroscopy and Machine Learning
<p>Figures and datasets used in the manuscript "Label-Free Identification of Exosomes Using Raman Spectroscopy and Machine Learning". Codes that produce the figures are also included.</p> <p> </p> <p> </p> <div> </div>
Data Repository for: Machine-learning the spectral function of a hole in a quantum antiferromagnet
<p>The machine-learning dataset of 51^3 ~1.3 × 10^5 density of states (DOS) of a mobile hole in the t-t'-t''-J model theoretically generated by using the self-consistent Born approximation in the three-dimensional parameter space of t′ ∈ [−0.5, 0.5], t′′ ∈ [−0.5, 0.5] and J ∈ [0.2, 1.0], with each parameter sampled on a 51-point uniform grid. The dataset is randomly partitioned into an 80/10/10 training (T), validation (V), and testing T split. Note that each DOS A(ω) was calculated on a 1201-point uniform grid of ω ∈ [−6t, 6t], then it was resampled on a 301-point uniform grid for the forward problem and on a 354-point uniform grid for the inverse problem. The dataset used in the inverse problem is limited to the parameter space of t′ ∈ [−0.5, 0], t′′ ∈ [0, 0.5] and J ∈ [0.2, 1.0]. To open the enclosed .npz files, use numpy.load() in python3.</p>
PYTHIA6 Dataset: Machine learning-based jet and event classification at the Electron-Ion Collider with applications to hadron structure and spin physics
<p>Dataset corresponding to: https://inspirehep.net/literature/2164495</p> <p>The data set contains PYTHIA6 jets in ep collisions, with separate files for LO DIS and photoproduction samples. For further details about the data set, see https://inspirehep.net/literature/2164495. Examples of how to analyze the data set can be found at: https://github.com/jdmulligan/ml-eic-flavor.</p> <p>Please contact james.mulligan@berkeley.edu with any questions.</p>
Optimizing Performance and Energy Across Problem Sizes Through a Search Space Exploration and Machine Learning - dataset
<p>Dataset used for a journal submission at JPDC.</p>
Supplementary data: "Physics-informed machine learning for power grid frequency modelling"
<p>This repository contains result files for the paper "Physics-informed machine learning for power grid frequency modelling" <a href="https://doi.org/10.48550/arXiv.2211.01481">(Preprint)</a>. The code for producing the processed data and the results is <a href="https://github.com/johkruse/PIML-for-grid-frequency-modelling">available at github</a>.</p> <p><strong>Results</strong></p> <p>The result folder comprises the results of hyper-parameter optimisation, scaling variation and interpretation via SHAP. In particular, it contains these sub-folders and files:</p> <ul> <li><em>tuning </em>: Results of hyper-parameter tuning.</li> <li><em>best_model </em>: Weights of the trained model with best hyper-parameters.</li> <li><em>best_model_<scaling-variation> </em>: Weights of the trained models with best hyper-parameters but with a variation of the parameter scaling.</li> <li><em>fixed_model_hps.pkl </em>: Hyper-parameters that are not optimised.</li> <li><em>shap_values_<parameter>_long.h5</em> : SHAP values for the prediction of the system parameters.</li> </ul>
GMSK comb Spectrum Sensing for Machine Learning
<p>These files contains measurement results achievied in the following scenario:<br> Single PC with GNU Radio and connected USRP transmits GMSK comb signal (6x GMSK signal at center frequncies 2097.5 MHz, 2098.5 MHz, 2099.5 MHz, 2100.5 MHz, 2101.5 MHz, and 2102.5 MHz) with different amplifier gain.<br> Single PC with GNU Radio and connected USRP receives signal with central frequency 2100 MHz and bandwidth 40 MHz (treated like 40 x 1 MHz channels). It uses local osciallator offset equal 10 MHz and due to non-linear characteristics of amplifiers and filters in USRP the extreme 12 (on both sides) are removed. However to create (this) dataset only 6 channels where signal was transmitted were included.<br> <br> Received is set in three different possitions (similar height and distance from transmitter) - creating 3 files - and collects data to detect presence of transmitted signal. Data that can be found in the files is stored in CSV format to be easly analyzed in ML models. </p> <p>Data collected at during this experiment contains:<br> six columns of average received power in the analyzed channels (in dBm)<br> six columns of autocorrelation function kurtosis in the analyzed channels (in linear scale)<br> six columns of autocorrelation function skewness in the analyzed channels (in linear scale)<br> one column with information about signal transmission (label; 0 - noise, 1 - signal transmitted)<br> one column with transmitter gain (in dB; set to -100 in case of lack of transmission)</p>
Advancing Thermal Management with Machine Learning Potentials on Boron Nitride (BN) and Other Group 13 Nitrides
<p>Machine Learning Interatomic Potentials for Bulk and Bilayer Materials (M = B, Al, Ga, and In): Extraction of Phonons and 3rd Order Interatomic Force Constants</p>
INFLAMeR: a machine learning algorithm based on large-scale perturbation screening identified new lncRNAs regulating differentiation and survival of leukaemia cells
<p><strong>Abstract</strong></p> <p>Long non-coding RNAs (lncRNAs) are a diverse group of transcripts with poorly understood<br> functionality. To address this gap, we developed INFLAMeR, an advanced machine learning<br> model trained on CRISPRi screening data, to predict functional lncRNAs using comprehensive<br> genetic features. We experimentally validated the predictions by assessing their impact on cell<br> proliferation and anticancer drug resistance. Among the selected lncRNAs, 85% showed<br> significant effects upon knockdown, while low-scoring lncRNAs had no discernible impact.<br> Notably, our study elucidated the functional role of SNHG6 in hematopoietic differentiation.<br> INFLAMeR greatly enhances the prediction of functional lncRNAs, providing valuable insights<br> into their regulatory landscape. By integrating INFLAMeR with experimental validation, we can<br> identify and characterize functional lncRNAs in a cell-type-specific manner, contributing to a<br> deeper understanding of their involvement in cellular processes. Our findings revealed insights<br> into lncRNA biology and a framework for improving the identification of functional lncRNAs.</p>
Reproducing surface water isoscapes of δ18O and δ2H across China: A machine learning approach
<p>This dataset support this submission.</p>
Iterative Machine Learning for Classification and Discovery of Single-molecule Unfolding Trajectories from Force Spectroscopy Data (Raw Data)
<p>Raw data used for the testing of the FUSION Learning algorithm available at <a href="https://github.com/Nash-Lab/Fusion-Learning">https://github.com/Nash-Lab/Fusion-Learning</a>.</p> <p> </p> <ol> </ol>
Status Quo and Problems of Requirements Engineering for Machine Learning: Results from an International Survey
<p>This folder contains the current data used in the paper named 'Status Quo and Problems of Requirements Engineering for Machine Learning: Results from an International Survey'. We make available a ZIP file containing the survey used, the data collected from it and the Jupyter Notebooks used to build our analysis.</p>
Instrumentation Neutron Activation Analysis & Proton Induced X-RAY Emission techniques supported with Machine learning analysis for rare earth/macro/micro elements correlation from O. Sativa Rice varieties in Senegal River valley
<p>data sheet INAA;results</p>
Data set for Exploring Machine Learning-Based Methods for anomalies detection: Evidence from cryptocurrencies
<p><strong>Exploring Machine Learning-Based Methods for anomalies detection: Evidence from cryptocurrencies</strong></p>
Nomogram Built Based on Machine Learning to Predict Recurrence in Early-stage Hepatocellular Carcinoma Patients Treated with Ablation
Open the record for dataset details and reuse information.
Machine Learning aids rapid assessment of aftershocks: Application to the 2022-2023 Peace River earthquake sequence, Alberta, Canada
<p>Two catalogs are provided as part of the publication "Machine Learning aids rapid assessment of aftershocks: Application to the 2022-2023 Peace River earthquake sequence, Alberta, Canada", in The Seismic Record.</p> <p>Catalog_PR_EQTransformer_TSR.csv was produced using PhaseNet (<a title="PhaseNet: a deep-neural-network-based seismic arrival-time picking method" href="https://doi.org/10.1093/gji/ggy423" target="_blank" rel="noopener">Zhu and Beroza, 2018</a>); Catalog_PR_PhaseNet_TSR.csv was produced using EQTransformer (<a title="Earthquake transformer&mdash;an attentive deep-learning model for simultaneous earthquake detection and phase picking" href="https://doi.org/10.1038/s41467-020-17591-w" target="_blank" rel="noopener">Mousavi et al., 2020</a>). The files contain the following columns: </p> <ul> <li><strong>UTCDateTime:</strong> UTC Date and Time of detected event</li> <li><strong>NLL_Ev_Lat_Deg</strong>: Event Latitude in decimal degrees</li> <li><strong>NLL_Ev_Long_Deg</strong>: Event Longitude in decimal degrees</li> <li><strong>NLL_Ev_Z_km:</strong> Event Depth in kilometers</li> <li><strong>NLL_Ev_Lat_err: </strong>Event Latitude Error in kilometers</li> <li><strong>NLL_Ev_Long_err:</strong> Event Longitude Error in kilometers</li> <li><strong>NLL_Ev_Z_err</strong>: Event Depth Error in kilometers</li> <li><strong>Calibrated_Ml:</strong> Calibrated local magnitude</li> </ul> <p>Hypocenter locations and associated errors were calculated using NonLinLon (<a title="Probabilistic Earthquake Location in 3D and Layered Models" href="https://doi.org/10.1007/978-94-015-9536-0_5" target="_blank" rel="noopener">Lomax et al.. 2000</a>; <a title="Earthquake Location, Direct, Global-Search Methods" href="https://doi.org/10.1007/978-3-642-27737-5_150-2" target="_blank" rel="noopener">2009</a>) using the velocity model provided in <a title="Disposal From In Situ Bitumen Recovery Induced the ML 5.6 Peace River Earthquake" href="https://doi.org/10.1029/2023GL102940" target="_blank" rel="noopener">Schultz et al., 2023</a>. Magnitudes were calculated using <a title="Determination of Local Magnitude for Induced Earthquakes in the Western Canada Sedimentary Basin: An Update" href="https://csegrecorder.com/articles/view/determination-of-local-magnitude-for\-induced-earthquakes-in-the-wcsb" target="_blank" rel="noopener">Babaie-Mahani and Kao (2020)</a> and then calibrated (due to incorrect instrument response files) using <a title="Induced or Natural? Toward Rapid Expert Assessment, with Application to the Mw 5.2 Peace River Earthquake Sequence" href="https://doi.org/10.1785/0220230289" target="_blank" rel="noopener">Salvage et al., 2023.</a> </p>
Dopant network processing units as tuneable extreme learning machines - Data Sheet 1
<p>Inspired by the highly efficient information processing of the brain, which is based on the chemistry and physics of biological tissue, any material system and its physical properties could in principle be exploited for computation. However, it is not always obvious how to use a material system’s computational potential to the fullest. Here, we operate a dopant network processing unit (DNPU) as a tuneable extreme learning machine (ELM) and combine the principles of artificial evolution and ELM to optimise its computational performance on a non-linear classification benchmark task. We find that, for this task, there is an optimal, hybrid operation mode (“tuneable ELM mode”) in between the traditional ELM computing regime with a fixed DNPU and linearly weighted outputs (“fixed-ELM mode”) and the regime where the outputs of the non-linear system are directly tuned to generate the desired output (“direct-output mode”). We show that the tuneable ELM mode reduces the number of parameters needed to perform a formant-based vowel recognition benchmark task. Our results emphasise the power of analog in-matter computing and underline the importance of designing specialised material systems to optimally utilise their physical properties for computation.</p>
Using Machine Learning to Identify Responders to TACE or HAIC for uHCC
ClinicalTrials.gov study NCT07368530. IPD Sharing: NO. Countries: 1. Publications: 0.
A Feasibility Study to Improve Colorectal Cancer Screening Among Racially Diverse Zip Codes in a Persistent Poverty County Using Navigation and Machine Learning Predictive Algorithms
ClinicalTrials.gov study NCT05383976. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
"Machine Learning Analysis of Lingual Colorimetry and MADRS Anxiety-Depression Score in Acupuncture Patients"
ClinicalTrials.gov study NCT06899490. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.