Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

16

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

16 results for “machine learning results”

Learn how ShareScore rates datasets ↗
zenodo44/100

Experiment on the performance of different machine learning algorithms for classification - Results

<h2>Results of a short performance study of machine learning algorithms</h2> <h3>Context and methodology</h3> <ul> <li>This data was produced while performing a university project to examine the performance of various machine learning algorithms on different prediction datasets</li> <li>The data serves the purpose of comparing the metrics of performing the different tasks</li> <li>The dataset contains a number of matrices for every classifier and every dataset</li> <li>The data was produced with python scripts provided further down and with the usage of the external datasets: <ul> <li>Membership Woes Dataset (OpenML): <a href="https://api.openml.org/d/44224">https://api.openml.org/d/44224</a></li> <li>Zoo dataset (UCI): <a href="https://doi.org/10.24432/C5R59V">https://doi.org/10.24432/C5R59V</a></li> <li>Breast Cancer Dataset: <a href="https://github.com/moritx/performance-experiment-machine-learning/tree/main/data">https://github.com/moritx/performance-experiment-machine-learning/tree/main/data</a></li> <li>Loan Dataset: <a href="https://github.com/moritx/performance-experiment-machine-learning/tree/main/data">https://github.com/moritx/performance-experiment-machine-learning/tree/main/data</a></li> </ul> </li> </ul> <h3>Technical details</h3> <ul> <li>The data consists of one JSON file</li> <li>The source code for producing this data is available at&nbsp;<a href="https://doi.org/10.5281/zenodo.11085222">https://doi.org/10.5281/zenodo.11085222</a></li> </ul> <h3>Structure of the data</h3> <p>[ {"classifier": ...,<br>"dataset": ...,<br>"hyper_parameters": ...,<br>"cross_validation_results": {<br>&nbsp; &nbsp; "fit_time": {} ,<br>&nbsp; &nbsp; "score_time": ...,<br>&nbsp; &nbsp; "metrics": {}<br>},&nbsp;<br>"holdout_test_results": ...}, ]</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Results: Towards Realistic SATD Identification Through Machine Learning Models: Ongoing Research and Preliminary Results

<p>Automated identification of self-admitted technical debt (SATD) has been crucial for advancements in managing such debt.&nbsp;<br>However, state-of-the-arts studies often overlook chronological factors, leading to experiments that do not faithfully replicate the conditions developers face in their daily routines.<br>This study initiates a chronological analysis of SATD identification through machine learning models, emphasizing the significance of temporal factors in automated SATD detection.&nbsp;<br>The research is in its preliminary phase, divided into two stages: evaluating model performance trained on historical data and tested in prospective contexts, and examining model generalization across various projects. Preliminary results reveal that the chronological factor can positively or negatively influence model performance and that some models are not sufficiently general when trained and tested on different projects.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Intermediate results for: Use the 4S (Signal-Safe Speckle Subtraction): Explainable Machine Learning reveals the Giant Exoplanet AF Lep b in High-Contrast Imaging Data from 2011

<p>This collection contains all intermediate results needed to reproduce the results in the paper:</p> <p>Use the 4S (Signal-Safe Speckle Subtraction): Explainable Machine Learning reveals the Giant Exoplanet AF Lep b in High-Contrast Imaging Data from 2011</p> <p>You can use these intermediate results to create all plots in our paper without the need to run all experiments on a large cluster.</p> <p>More details can be found in the online documentation of our python package:<br><a href="https://fours.readthedocs.io/en/latest/">https://fours.readthedocs.io/en/latest/</a></p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

BRAIN Journal-Automatic Anthropometric System Development Using Machine Learning-Figure 3. (3.a) – The flowchart of Graph cuts method; (3.b)- the result of Graph cuts image segmentation.

<p>Figure 3 describes the steps implemented Graph cuts algorithm for the segmentation of human body parts. The results obtained are 5 main sections that include the hands, the legs, the center of the body (chest, waist, hips), and the head. The result of the display image is taken from the human image database, which was collected by us (Нгуен, 2016).&nbsp;</p>

opencc-by-4.0Aug 2016View details →
zenodo40/100

BRAIN Journal-Automatic Anthropometric System Development Using Machine Learning-Figure 6. The result of building a 3D model based on RF and SVM classification with "Important features".

<p>From the chart of figure 6, we found that &quot;Important Features&quot; gave the best 3D model, which fits with the object in the image. The pattern is close to 90% compared with the true size. Apply classification algorithm RF increases the accuracy of the results and reduces computing time for the program. There are many methods for data classifying. One of them is the method of the support vector machine (SVM). The SVM method is represented by Vladimir N. Vapnik (1995) in Support Vector Machines (SVM) - a set of learning algorithms similar with the supervisor has two main tasks: the classification and the regression analysis. In this article we use the method of the SVM classification problem for the size of the human body with 5 classes to compare the performance between SVM methods and Random Forest algorithm.&nbsp;</p>

opencc-by-4.0Aug 2016View details →
zenodo40/100

BRAIN Journal-Automatic Anthropometric System Development Using Machine Learning-Figure 4. Flowchart and results of ICP algorithm

<p>The key concept of the standard ICP algorithm can be summarized in two steps: - Compute correspondences between the two scans. - Compute a transformation which minimizes the distance between corresponding points. It is forced to add a maximum matching threshold dmax. In most implementations of ICP, the choice of dmax represents a tradeoff between convergence and accuracy. A low-value result in bad convergence, a large value causes incorrect correspondences to pull the final alignment away from the correct value. Figure 4 describes the steps of the algorithm which determines the point features closest to object boundary. The result of the algorithm is described by images cut from the program (Нгуен, 2016)</p>

opencc-by-4.0Aug 2016View details →
zenodo40/100

Datasets, trained models and supporting results for machine learning tensorial properties of atomic systems via XPaiNN model.

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo40/100

Rating results obtained during a review of original articles on radiomics and machine learning for outcome prediction based on PET

<p>This upload provides Open Data associated with the publication&nbsp;&quot;Methodological evaluation of original articles on radiomics and machine learning for outcome prediction based on positron emission tomography (PET)&quot;&nbsp;by Rogasch JMM&nbsp;<em>et al.</em>&nbsp;(2023).</p> <p>The upload contains the item-by-item results of rating for all criteria and all 100 original articles. PubMed IDs are also included.</p> <p>Furthermore, a&nbsp;description of all variable names and how the rating categories were encoded in the data tables can be found in the PDF file &quot;ML_prediction_Dictionary_2023_08_27.pdf&quot;.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Stereoisomers are not Machine Learning's Best Friends: Experimental results of the prediction of the association constant between a cyclodextrin and a guest with Stereo2vec

<p>This study addresses the challenge of accurately identifying stereoisomers in cheminformatics which originates from our objective to apply machine learning to predict association constant between a cyclodextrin and a guest. Identifying stereoisomers is indeed crucial for machine learning applications. Current tools offer various molecular descriptors, including their textual representation as Isomeric SMILES which can distinguish stereoisomers. But such representation is text-based and does not have a fixed size, so a conversion is needed to make it usable to machine learning approaches. Word embedding techniques can be used to solve this problem. Mol2vec, a word embedding approach for molecules, offers such a conversion. Unfortunately, it cannot distinguish between stereoisomers due to its inability to capture the spatial configuration of molecular structures. This study proposes several approaches that use word embedding techniques to handle molecular discrimination using stereochemical information of molecules or considering Isomeric SMILES notation as a text in Natural Language Processing. Our aim is to generate a distinct vector for each unique molecule, correctly identifying stereoisomer information in cheminformatics. The proposed approaches are then compared on our original machine learning task: predicting the association constant between a cyclodextrin and a guest molecule.</p>

openbsd-3-clauseMay 2024View details →
zenodo36/100

A Comparative Analysis of Machine Learning Approaches to Gap Filling Meteorological Datasets (Results Only)

<p>This dataset contains the results from our evaluation of methodologies for filling gaps in meterological data. Variables were chosen to represent a standard set of measurements for purposes typically performed using Weather Stations installed in urban and rural areas.&nbsp;There are 4 dimensions by which we measure and validate each of the gap filling models across a large set of experimental configurations: meteorological variables <strong>TargetVar </strong>(dewpoint, humidity, leaf wetness, temp); <strong>feature_set</strong> (AWS, AWS-ERA5, ERA5, ERA5_Debias, Spatial); 3 types of machine learning algorithms <strong>ML</strong> (linear regression, random forests, LightGBM) combined with 2 non-ML algorithms; and <strong>gap_length</strong>: 1, 4, 36 and 288. &nbsp;A total of 1,720 experiments were conducted: 1,440 machine learning experiments; 160 using a spatial algorithm and 120 experiments using ERA5. For the 3 machine learning experiments, the average result was selected for each of 10 sites for 4 gap &nbsp;sizes (40 results) with the 3 ML models using 3 different feature sets (120 results).</p> <p>Data is provided in both CSV format and as a MySQL dump.</p> <p>Resultsets are accompanied with the SQL expression used to generate the result.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Result dataset for the paper "Determining Research Priorities Using Machine Learning"

<p>The dataset needed by the paper software to run the notebooks.</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Dataset and results for "Comparing machine learning and deep learning models for probabilistic post-processing of satellite precipitation-driven streamflow simulation"

<p>Dataset and results for &quot;Comparing machine learning and deep learning models for probabilistic post-processing of satellite precipitation-driven streamflow simulation&quot;</p> <p>Yuhang Zhang1, Aizhong Ye1*, Phu Nguyen2, Bita Analui2, Soroosh Sorooshian2, Kuolin Hsu2</p> <p>1 State Key Laboratory of Earth Surface Processes and Resource Ecology, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China.</p> <p>2 Center for Hydrometeorology and Remote Sensing, Department of Civil and Environmental Engineering, University of California, Irvine, Irvine, California, CA 92697, USA.</p> <p>## Dataset&nbsp;&nbsp; &nbsp;</p> <p>Streamflow simulations from one observed precipitation (CMA) and three satellite precipitation products (PDIR, IMERG-F, and GSMaP) for 522 sub-basins.</p> <p>- Q-CMA (streamflow reference)<br> - Q-PDIR (uncorrected)<br> - Q-IMERGF (uncorrected)<br> - Q-GSMAP (uncorrected)</p> <p>### Data structure</p> <p>- Head section (row1-row5)<br> &nbsp; - SubNO:&nbsp;&nbsp; &nbsp;522&nbsp;<br> &nbsp; - BeginT:&nbsp;&nbsp; &nbsp;2003-01-01 00:00&nbsp;<br> &nbsp; - EndT:&nbsp;&nbsp; &nbsp;2019-12-31 00:00&nbsp;<br> &nbsp; - Interval:&nbsp;&nbsp; &nbsp;1440s (daily)<br> &nbsp; - Revise:&nbsp;&nbsp; &nbsp;10 (scaling factor to keep int datatype)<br> &nbsp; - Point1&nbsp;&nbsp; &nbsp;Point2&nbsp;&nbsp; &nbsp;... (Subbasin No.)<br> - Data section<br> &nbsp; - 6209 rows, 522 cols</p> <p>## Results</p> <p>Two post-processing model results for test period (2015-1-1 to 2018-12-31).</p> <p>### Data structure</p> <p>- 1462 rows, every row denotes each day from 2015-1-1 to 2018-12-31</p> <p>- 100 columns, every column denotes each quantile from 0.005 to 0.995, total 100 quantiles.</p> <p>### qrf-output</p> <p>- pdir (single input)<br> - imergf (single input)<br> - gsmap (single input)<br> - all (multiple inputs)</p> <p>### lstm-output</p> <p>- pdir (single input)<br> - imergf (single input)<br> - gsmap (single input)<br> - all (multiple inputs)</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo28/100

Accuracy of EEG Biomarkers in the Detection of Clinical Outcome in Disorders of Consciousness after Severe Acquired Brain Injury: Preliminary Results of a Pilot Study Using a Machine Learning Approach

<p>Dataset for the accepted publication: &quot;Accuracy of EEG Biomarkers in the Detection of Clinical Outcome in Disorders of Consciousness after Severe Acquired Brain Injury: Preliminary Results of a Pilot Study Using a Machine Learning Approach&quot;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo28/100

Status Quo and Problems of Requirements Engineering for Machine Learning: Results from an International Survey

<p>This folder contains the current data used in the paper named &#39;Status Quo and Problems of Requirements Engineering for Machine Learning: Results from an International Survey&#39;. We make available a ZIP file containing the survey used, the data collected from it and the Jupyter Notebooks used to build our analysis.</p>

opencc-by-4.0Aug 2023View details →
zenodo24/100

Results of experiments in paper Simulated Hyperparameter Optimization for Statistical Tests in Machine Learning Benchmarks

<p>Results of the experiments&nbsp;in paper Simulated Hyperparameter Optimization for Statistical Tests in Machine Learning Benchmarks</p>

opencc-by-4.0Jun 2020View details →
zenodo24/100

Kolberger Heide community compositions and machine learning results

<p>Bacteria are ubiquitous and live in complex microbial communities, which can react rapidly to changing environmental conditions. Their physiological variety enables communities to respond in specific ways to environmental drivers, potentially resulting in distinct microbial fingerprints for a given environmental state. Our goal was to assess the opportunities and limitations of machine learning to detect fingerprints indicating the presence of the munition compound 2,4,6-trinitrotoluene (TNT) in southwestern Baltic Sea sediments.</p> <p>Over 40 environmental variables including grain size distribution, elemental composition and concentration of munition compounds (mostly at pmol g<sup>-1</sup> levels) from 150 sediments collected at the near-to-shore munition dumpsite Kolberger Heide by the German city of Kiel were combined with 16S rRNA gene amplicon sequencing libraries. Prediction was achieved using Random Forests; the robustness of predictions was validated using Artificial Neural Networks. To facilitate machine learning with microbiome data we developed the R package phyloseq2ML.</p> <p>Using the most classification-relevant 25 bacterial genera exclusively, potentially representing a TNT-indicative fingerprint, TNT was predicted correctly with up to 81.5&nbsp;% balanced accuracy. False positive classifications indicated that this approach has also the potential to identify samples where the original TNT contamination was no longer detectable. The sensitivity of this approach can be deduced from the fact that TNT presence was neither identified among the main drivers of the microbial community composition, nor did it correlate with sediment metal content, demonstrated by decreased prediction rates using environmental variables.</p> <p>Our results suggest that microbial communities can predict even minor influencing factors in complex environments, demonstrating the potential of this approach for the discovery of contamination events over an integrated period of time and for environmental monitoring in general.</p>

opencc-by-4.0Oct 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record