Skip to main content
zenodoopen

Machine Learning Process

<p><span>Image illustrating the essential steps for applying machine learning (ML) to your data.</span></p> <p><span>The learning process of an ML model starts with&nbsp;</span><strong><span>data selection</span></strong><span>, which involves collecting the necessary data from various places like databases, online repositories, or real-time systems. The second step is the&nbsp;</span><strong><span>preprocessing</span></strong><span>&nbsp;or&nbsp;</span><strong><span>data preparation</span></strong><span>. It includes cleaning the data to remove errors or inconsistencies and handling missing data. After that, the specific attributes from the&nbsp;</span><strong><span>structured data</span></strong><span>&nbsp;</span><span>are selected</span><span>&nbsp;</span><span>that will</span><span>&nbsp;help the ML algorithm learn. Then, the data&nbsp;</span><span>is split</span><span>&nbsp;into&nbsp;</span><strong><span>training and testing sets</span></strong><span>. The training set&nbsp;</span><span>is used</span><span>&nbsp;to build and train the ML model, while the testing set&nbsp;</span><span>is used</span><span>&nbsp;to evaluate its performance. Then, the ML model is chosen based on the problem&nbsp;</span><span>and</span><span>&nbsp;the training data is fed into the model to create the&nbsp;</span><strong><span>candidate model</span></strong><span>. Once satisfied with the model's performance,&nbsp;</span><span>the final ML model can be deployed</span><span>.</span></p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
0