Machine Learning Process
<p><span>Image illustrating the essential steps for applying machine learning (ML) to your data.</span></p> <p><span>The learning process of an ML model starts with </span><strong><span>data selection</span></strong><span>, which involves collecting the necessary data from various places like databases, online repositories, or real-time systems. The second step is the </span><strong><span>preprocessing</span></strong><span> or </span><strong><span>data preparation</span></strong><span>. It includes cleaning the data to remove errors or inconsistencies and handling missing data. After that, the specific attributes from the </span><strong><span>structured data</span></strong><span> </span><span>are selected</span><span> </span><span>that will</span><span> help the ML algorithm learn. Then, the data </span><span>is split</span><span> into </span><strong><span>training and testing sets</span></strong><span>. The training set </span><span>is used</span><span> to build and train the ML model, while the testing set </span><span>is used</span><span> to evaluate its performance. Then, the ML model is chosen based on the problem </span><span>and</span><span> the training data is fed into the model to create the </span><strong><span>candidate model</span></strong><span>. Once satisfied with the model's performance, </span><span>the final ML model can be deployed</span><span>.</span></p>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0