Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Predicting diagnosis and survival of bone metastasis in breast cancer using machine learning: a SEER-based study
<p>The data for this article: predicting diagnosis and survival of bone metastasis in breast cancer using machine learning: a SEER-based study.</p>
Uncertainty quantification in cerebral circulation simulations focusing on the collateral flow: Surrogate model approach with machine learning
<p>Data and code underlying the findings reported in the paper titled "Uncertainty quantification in cerebral circulation simulations focusing on the collateral flow: Surrogate model approach with machine learning."</p>
Radiomics and machine learning analysis based on magnetic resonance imaging in the assessment of liver mucinous colorectal metastases
<p>I uploaded the images of the manuscript Radiomics and machine learning analysis based on magnetic resonance imaging in the assessment of liver mucinous colorectal metastases.</p>
On the application of an observations-based machine learning parameterization of surface layer fluxes within an atmospheric large-eddy simulation model: Article Data
<p>NetCDF datatset of presented results from the publication titled "On the application of an observations-based machine learning parameterization of surface layer fluxes within an atmospheric large-eddy simulation model" in the Journal of Geophysical Research - Atmospheres, Paper #2021JD036214R.</p>
Machine Learning Constructs Color Features to Accelerate Development of Long-Term Continuous Water Quality Monitoring
<p>This is a machine learning method for predicting the concentration of colored pollutants based on RGB and kmeans methods. This dataset includes raw images of pollutants as well as characteristic data of pollutants, as well as code for the model. You can see the contents of the zip file for details.</p>
Training data for ship track detection machine learning algorithms
<p>The training data and labels used to train the linked machine learning algorithm</p>
Application of Machine Learning in Chinese Medicine Differentiation of Dampness-heat Pattern in Patients with Type 2 Diabetes Mellitus
<p>Contains a compressed file of algorithm code and raw Excel data</p>
Climate data for Machine Learning based 100-year flood flow prediction model
<p>This study evaluates the application of ML technique over northeast United States regions and compares its performance to the U.S. Geological Survey (USGS) Streamflow Statistics (StreamStats)</p>
Application of machine learning to improve appropriateness of treatment in an orthopaedic setting of personalized medicine
<p>Raw Data</p>
Combination of whole genome sequencing and Supervised Machine Learning provides unambiguous identification of enterohemorrhagic Escherichia coli in raw milk
<p>These dataset are used in the "rename_list_of_groups.ipynb" notebook</p>
Blinded Predictions and Post-hoc Analysis of the Second Solubility Challenge Data: Exploring Training Data and Feature Set Selection for Machine and Deep Learning Models
<p>Training and test datasets and scripts for training models.</p>
Machine learning for Material Science 2022 - Neural Networks Assignment Dataset
<p>Data for the neural network assignment</p> <p> </p>
OQM9HK: A Large-scale Graph Dataset for Machine Learning in Materials Science
<p>This is a large-scale graph dataset of materials science based on the Open Quantum Materials Database (OQMD) v1.5 .</p> <p><a href="https://storage.googleapis.com/rimcs_cgnn/oqm9hk_dataset_Sep_30_2022.pdf" target="_blank" rel="noopener">Technical Report</a></p> <p><a href="https://www.rimcs.co.jp" target="_blank" rel="noopener">RIMCS Website</a></p> <p><strong>Data Loading</strong></p> <p>A Python code example:</p> <pre><code>import sys sys.path.append('/your/path/to/data/OQM9HK_BEL') import OQM9HK bel_path='/your/path/to/data/OQM9HK_BEL' config = OQM9HK.load_config(path=bel_path) print(config['atomic_numbers']) split = OQM9HK.load_split(path=bel_path) print(len(split['train']), len(split['val']), len(split['test'])) graph_data = OQM9HK.load_graph_data(path=bel_path) name = next(iter(graph_data)) # Frist entry's name graph = graph_data[name] # Graph object print(graph.nodes) print(graph.edge_sources) print(graph.edge_targets) dataset = OQM9HK.load_targets(path=bel_path) # Pandas dataframe print(dataset) train_set = dataset.iloc[split['train']] val_set = dataset.iloc[split['val']] test_set = dataset.iloc[split['test']]</code></pre> <p> </p>
Machine learning on large-scale proteomics data identifies tissue- and cell type-specific proteins
<p>Using data from 183 public human data sets from PRIDE, a machine learning model was trained to identify tissue and cell-type specific protein patterns. PRIDE projects were searched with ionbot and tissue/cell type annotation was manually added. Data from physiological samples were used to train a Random Forest model on protein abundances to classify samples into tissues and cell types. Subsequently, a one-vs-all classification and feature importance were used to analyse the most discriminating protein abundances per class. Based on protein abundance alone, the model was able to predict tissues with 98% accuracy, and cell types with 99% accuracy. The F-scores describe a clear view on tissue-specific proteins and tissue-specific protein expression patterns. In-depth feature analysis shows slight confusion between physiologically similar tissues, demonstrating the capacity of the algorithm to detect biologically relevant patterns. These results can in turn inform downstream uses, from identification of the tissue of origin of proteins in complex samples such as liquid biopsies, to studying the proteome of tissue-like samples such as organoids and cell lines</p>
Applications of machine learning tools for ultra-sensitive detection of lipoarabinomannan with plasmonic grating biosensors in clinical samples of tuberculosis
Background <p>Tuberculosis is one of the top ten causes of death globally and the leading cause of death from a single infectious agent. Eradicating the Tuberculosis epidemic by 2030 is one of the top United Nations Sustainable Development Goals. Early diagnosis is essential to achieving this goal because it improves individual prognosis and reduces transmission rates of asymptomatic infected. We aim to support this goal by developing rapid and sensitive diagnostics using machine learning algorithms to minimize the need for expert intervention. </p> Methods and Findings <p>A single-molecule fluorescence immunosorbent assay was used to detect the Tuberculosis biomarker lipoarabinomannan from a set of twenty clinical patient samples and a control set of spiked human urine. Tuberculosis status was separately confirmed by GeneXpert MTB/RIF and cell culture. Two machine learning algorithms, an automatic and a semiautomatic model, were developed and trained by the calibrated lipoarabinomannan titration assay data and then tested against the ground truth patient data. The semiautomatic model differed from the automatic model by an expert review step in the former, which calibrated the lower threshold to determine single molecules from background noise. The semiautomatic model was found to provide 88.89% clinical sensitivity, while the automatic model resulted in 77.78% clinical sensitivity.</p> Conclusions <p>The semiautomatic model outperformed the automatic model in clinical sensitivity as a result of the expert intervention applied during calibration and both models vastly outperformed manual expert counting in terms of time-to-detection and completion of analysis. Meanwhile, the clinical sensitivity of the automatic model could be improved significantly with a larger training dataset. In short, semiautomatic, and automatic Gaussian Mixture Models have a place in supporting rapid detection of Tuberculosis in resource-limited settings without sacrificing clinical sensitivity.</p>
Dataset and results for "Comparing machine learning and deep learning models for probabilistic post-processing of satellite precipitation-driven streamflow simulation"
<p>Dataset and results for "Comparing machine learning and deep learning models for probabilistic post-processing of satellite precipitation-driven streamflow simulation"</p> <p>Yuhang Zhang1, Aizhong Ye1*, Phu Nguyen2, Bita Analui2, Soroosh Sorooshian2, Kuolin Hsu2</p> <p>1 State Key Laboratory of Earth Surface Processes and Resource Ecology, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China.</p> <p>2 Center for Hydrometeorology and Remote Sensing, Department of Civil and Environmental Engineering, University of California, Irvine, Irvine, California, CA 92697, USA.</p> <p>## Dataset </p> <p>Streamflow simulations from one observed precipitation (CMA) and three satellite precipitation products (PDIR, IMERG-F, and GSMaP) for 522 sub-basins.</p> <p>- Q-CMA (streamflow reference)<br> - Q-PDIR (uncorrected)<br> - Q-IMERGF (uncorrected)<br> - Q-GSMAP (uncorrected)</p> <p>### Data structure</p> <p>- Head section (row1-row5)<br> - SubNO: 522 <br> - BeginT: 2003-01-01 00:00 <br> - EndT: 2019-12-31 00:00 <br> - Interval: 1440s (daily)<br> - Revise: 10 (scaling factor to keep int datatype)<br> - Point1 Point2 ... (Subbasin No.)<br> - Data section<br> - 6209 rows, 522 cols</p> <p>## Results</p> <p>Two post-processing model results for test period (2015-1-1 to 2018-12-31).</p> <p>### Data structure</p> <p>- 1462 rows, every row denotes each day from 2015-1-1 to 2018-12-31</p> <p>- 100 columns, every column denotes each quantile from 0.005 to 0.995, total 100 quantiles.</p> <p>### qrf-output</p> <p>- pdir (single input)<br> - imergf (single input)<br> - gsmap (single input)<br> - all (multiple inputs)</p> <p>### lstm-output</p> <p>- pdir (single input)<br> - imergf (single input)<br> - gsmap (single input)<br> - all (multiple inputs)</p> <p> </p>
UCI Machine Learning- Adult Dataset
https://archive.ics.uci.edu/ml/machine-learning-databases/adult/
Graph Machine Learning Dataset LinkedMDB (LinkedMDB-GML)
<p>LinkedMDB-GML is a heterogeneous graph machine learning dataset based on the full RDF knowledge graph <a href="https://ceur-ws.org/Vol-538/ldow2009\_paper12.pdf" target="_blank" rel="noopener">LinkedMDB</a> (6.1 million RDF triples).</p> <p>LinkedMDB-GML contains node types, such as movies and related entities, such as actors and directors.</p> <p>LinkedMDB-GML was created with <a href="https://github.com/davidlamprecht/AutoRDF2GML" target="_blank" rel="noopener">AutoRDF2GML</a>.</p>
Machine Learning GAN Deconvolution
<p>We initially benchmarked our GAN against the LUCYD network using a Z-stack of a selection U2OS cells acquired by widefield microscopy from the Li et al dataset (Li et al., 2022) . Following this we ran another benchmark against Deconwolf to gain a direct comparison of the improvements achieved using a Z-stack image from the ChrX-36plex OligoFISSEQ dataset (Nguyen et al., 2020).</p> <p><br>Data is organised as follow:</p> <p>├───Input<br>└───Output<br> ├───Deconwolf<br> ├───GAN<br> └───LUCYD</p> <p></p>
Data for: Interpolation and differentiation of alchemical degrees of freedom in machine learning interatomic potentials
<p>This is the dataset for the article "Interpolation and differentiation of alchemical degrees of freedom in machine learning interatomic potentials" by Juno Nam and Rafael Gomez-Bombarelli.</p> <ul> <li>Preprint: <a href="https://arxiv.org/abs/2404.10746">https://arxiv.org/abs/2404.10746</a></li> <li>Code: <a href="https://github.com/learningmatter-mit/alchemical-mlip">https://github.com/learningmatter-mit/alchemical-mlip</a></li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.