Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
geo20/100

Machine Learning Models Predict the Primary Sites of Head and Neck Squamous Cell Carcinoma Metastases Based on DNA Methylation

GEO Series GSE171994. Homo sapiens. 49 samples. Type: Methylation profiling by genome tiling array.

openGEO-OpenDec 2021View details →
geo20/100

Multi-omics and machine learning reveal context-specific gene regulatory activities of PML-RARA in Acute Promyelocytic Leukemia

GEO Series GSE173755. Homo sapiens. 28 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing; Other.

openGEO-OpenDec 2022View details →
geo20/100

Multi-omics and Machine Learning Accurately Predicts Clinical Response to Adalimumab and Etanercept Therapy in Patients with Rheumatoid Arthritis

GEO Series GSE138747. Homo sapiens. 320 samples. Type: Methylation profiling by genome tiling array; Expression profiling by high throughput sequencing.

openGEO-OpenAug 2020View details →
geo20/100

A machine learning approach to integrate big data for precision medicine in acute myeloid leukemia

GEO Series GSE108004. Homo sapiens. 54 samples. Type: Expression profiling by array; Expression profiling by high throughput sequencing.

openGEO-OpenDec 2017View details →
geo20/100

A machine learning approach to integrate big data for precision medicine in acute myeloid leukemia [RNA-Seq]

GEO Series GSE108003. Homo sapiens. 24 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenDec 2017View details →
geo20/100

Evaluation via Supervised Machine Learning of the Broiler Pectoralis Major and Liver Transcriptome in Association with the Muscle Myopathy Wooden Breast

GEO Series GSE144000. Gallus gallus. 35 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2020View details →
geo20/100

Machine learning accelerates the dissection of mitostasis as a cross-species central biological hub for leaf senescence

GEO Series GSE201607. Arabidopsis thaliana; Solanum lycopersicum; Oryza sativa Japonica Group. 18 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2022View details →
geo20/100

Model-to-crop conserved NUE Regulons enhance machine learning predictions of nitrogen use efficiency

GEO Series GSE280353. Arabidopsis thaliana; Zea mays. 354 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenMay 2025View details →
geo20/100

Integrative Analysis and Machine Learning based Characterization of Single Circulating Tumor Cells

GEO Series GSE129474. Homo sapiens. 15 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenApr 2020View details →
geo20/100

FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster

GEO Series GSE282899. Drosophila melanogaster. 4 samples. Type: Other.

openGEO-OpenJan 2026View details →
geo20/100

Machine learning-based prediction of the activity and specificity of Cas9 variants in gene editing

GEO Series GSE231840. Homo sapiens. 64 samples. Type: Other.

openGEO-OpenAug 2024View details →
zenodo20/100

Exploring Context-Free Languages via Planning: The Case for Automating Machine Learning -- Additional Plots

<p>A comparison to SOTA AutoML approach on various datasets</p>

opencc-by-4.0Jan 2020View details →
zenodo20/100

Long-term trends of ambient nitrate (NO3-) concentrations across China based on ensemble machine-learning models

<p>The monthly NO3- concentrations across China during 2005-2015</p>

opencc-by-4.0Aug 2020View details →
zenodo20/100

Data from the manuscript: Machine learning reveals structural characteristics of stereochemistry-specific interdigitation of synthetic monomycoloyl glycerol analogs

<h2>Four membrane configurations:</h2> <p>1. single = single bilayer (512 MMG molecules and 25600 water molecule)&nbsp;<br>2. db_large = large double bilayer (1024 MMG molecules and 51200 water molecules)<br>3. id_small = interdigitated double bilayer (256 MMG molecules and 12800 water molecules)<br>4. db_small = small double bilayer (256 MMG molecules and 12800 water molecules)</p> <p>Each membrane configuration has 2 different MMG analogs:</p> <p><br>1. MMG1 = MMG-1, 1:1 racemic mixture of stereoisomers with (2R,3S)/(2S,3R) configurations<br>2. MMG6 = MMG-6, 1:1 racemic mixture of stereoisomers with (2R,3R)/(2S,3S) configurations</p> <h2>File names:</h2> <p><br>Each trajectory folder contains starting structures, input structure files, trajectory, energy file and system topology file.<br>Additionally, the .mdp files and topologies are located in their own folders.</p> <p>-&nbsp; start.gro &nbsp; &nbsp;= starting structure before energy minimization**<br>-&nbsp; eq.gro &nbsp; &nbsp;= input structure file for production<br>-&nbsp; eq.cpt &nbsp; &nbsp;= equilibration checkpoint file<br>-&nbsp; run.gro &nbsp; &nbsp;= final frame structure file<br>-&nbsp; run.cpt &nbsp; &nbsp;= production checkpoint file<br>-&nbsp; run.tpr &nbsp; &nbsp;= production tpr file&nbsp;<br>- run.xtc &nbsp; &nbsp;= production trajectory file, final 500 ns<br>-&nbsp; run.edr &nbsp; &nbsp;= production energy file, final 500 ns</p> <p>the <em>db_large</em> system is built with equilibrated <em>db_small</em> system, hence the starting structure is before the equilibration.</p> <h2>MDP-files:</h2> <p><br>Molecular dynamics parameter files can be found in their own folder (05_mdps/). Files for energy minimization,<br>equilibration and production run are provided.</p> <p>- em1_charmm36.mdp &nbsp; &nbsp; &nbsp; &nbsp;= 1st energy minimization (for all systems)<br>- em_charmm36.mdp &nbsp; &nbsp; &nbsp; &nbsp;= 2nd energy minimization (for all systems)<br>- eq1_noposres_charmm36.mdp &nbsp; &nbsp;= equilibration without position restraints (for single)<br>- eq1_posresin_charmm36.mdp &nbsp; &nbsp;= equilibration with O1 atoms constrained in Z-direction (for db_large, id_small, db_small)<br>- production_charmm36.mdp &nbsp; &nbsp;= production run (for all systems)</p> <h2>Topologies and position restraint files:</h2> <p><br>System topologies for each MMG analog can be found in their own folders (06_top/MMG1/, 06_top/MMG6/).&nbsp;<br>Molecular topologies, forcefield parameters and position restraint files can be found from their respective folders (06_top/MMG1/topol/, 06_top/MMG6/topol/).</p> <h3>MMG-1:</h3> <p><br>- db_large-MMG1.top &nbsp; &nbsp; &nbsp; &nbsp;= MMG-1 large double-bilayer system<br>- db_small-MMG1.top &nbsp; &nbsp; &nbsp; &nbsp;= MMG-1 small systems: interdigitated and non-interdigitated double bilayers &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;<br>- single-MMG1.top &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;= MMG-1 single bilayer system<br>- forcefield.itp &nbsp; &nbsp; &nbsp; &nbsp; = force field parameters<br>- TIP3_CHARMM36.itp &nbsp; &nbsp; &nbsp; &nbsp;= water model<br>- MMG1_2R3S_outer.itp &nbsp; &nbsp; &nbsp; &nbsp; = outer leaflet 2R3S MMG-1 &nbsp; &nbsp; &nbsp;<br>- MMG1_2S3R_outer.itp &nbsp; &nbsp; &nbsp; &nbsp; = outer leaflet 2S3R MMG-1 &nbsp;&nbsp;<br>- MMG1_2R3S_inner.itp &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;= inner leaflet 2R3S MMG-1<br>- MMG1_2S3R_inner.itp &nbsp; &nbsp; &nbsp; &nbsp; = inner leaflet 2S3R MMG-1<br>- posres_MMG1_tail_z.itp &nbsp; &nbsp; &nbsp; &nbsp;= position restraints on MMG tails in z-direction<br>- posres_MMG1_O1_z.itp &nbsp; &nbsp; &nbsp; &nbsp; = position restraints on MMG headgroup O1 atom in z-direction</p> <h3>MMG-6:</h3> <p><br>- db_large-MMG6.top &nbsp; &nbsp; &nbsp; &nbsp;= MMG-6 large double-bilayer system<br>- db_small-MMG6.top &nbsp; &nbsp; &nbsp; &nbsp;= MMG-6 small systems: interdigitated and non-interdigitated double bilayers<br>- single-MMG6.top &nbsp; &nbsp; &nbsp; &nbsp;= MMG-6 single bilayer system<br>- forcefield.itp &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;= force field parameters<br>- TIP3_CHARMM36.itp &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; = water model<br>- MMG6_2S3S_outer.itp &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; = outer leaflet 2S3S MMG-6 &nbsp;<br>- MMG6_2R3R_outer.itp &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; = outer leaflet 2R3R MMG-6<br>- MMG6_2S3S_inner.itp &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; = inner leaflet 2S3S MMG-6<br>- MMG6_2R3R_inner.itp &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; = inner leaflet 2R3R MMG-6&nbsp;<br>- posres_MMG6_tail_z_2R3R.itp &nbsp; &nbsp;= position restraints on tails in z-direction &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;<br>- posres_MMG6_tail_z_2S3S.itp &nbsp; &nbsp;= position restraints on tails in z-direction<br>- posres_MMG6_O1_z.itp &nbsp; &nbsp; &nbsp; &nbsp;= position restraints on MMG headgroup O1 atom in z-direction</p>

restrictedcc-by-4.0Apr 2024View details →
zenodo20/100

MACHINE LEARNING IN THE FINANCIAL INDUSTRY: A BIBLIOMETRIC APPROACH TO EVIDENCING APPLICATIONS

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
zenodo20/100

Deciphering complex antibiotic resistance patterns in Helicobacter pylori through whole genome sequencing and machine learning

<p>Whole Genome Sequencing of <em>Helicobacter pylori</em> (<em>Hp</em>).</p>

opencc-by-4.0Dec 2023View details →
zenodo20/100

Dataset fot the study of Documentation Practices of Machine Learning Resources

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo20/100

Approaches for machine learning intermolecular interaction energies and application to energy components from symmetry adapted perturbation theory

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo20/100

Restricted Data Set for the article Analysing the impact of renewables on Iberian wholesale electricity market prices using machine learning techniques. Green Finance, 2024, 6 (2), 363-382

<p><span>Two of the series contained in the databases used in the article [</span><em><span>Analysing the impact of renewables on Iberian wholesale electricity market prices using machine learning techniques. Green Finance, 6 (2), 363-382</span></em><span>] are not authorized for public sharing. Specifically, these include the data corresponding to the Dutch TTF futures price and the API2 index. Consequently, they are deposited under restricted access.</span></p>

restrictedcc-by-4.0Nov 2024View details →
zenodo20/100

Machine Learning-Based Estimation of Experimental Artifacts and Image Quality in Fluorescence Microscopy - Supporting Data

<p>Supporting data for the reproduction of the results reported in Corbetta, E., Bocklitz, T., Machine learning based estimation of experimental artifacts and image quality in fluorescence microscopy (2024) [1].</p> <h2><strong>MM-IQA_Images_png and </strong><strong>MM-IQA_Images_tif</strong></h2> <p>Folder containing all the supporting datasets of the publication.</p> <ul> <li>images_manual_inspection: semisynthetic dataset used for the manual inspection of the quality metrics.</li> <li>images_lda_training: semisynthetic dataset used to train the Linear Discriminant Analysis (LDA) model.</li> <li>images_lda_prediction: datasets predicted by the LDA model. <ul> <li>images_experimental: every subfolder is a dataset composed of measurements of a different sample. Images are from publicly available datasets from [2] and [3].</li> <li>images_known_semisynthetic: knwon semisynthetic dataset used for prediction, assessment and interpretation of the trained model.</li> </ul> </li> </ul> <p>Tif files are the original data used for the study.</p> <h2><strong>MM-IQA_Source_data</strong></h2> <p>Folder containing all the supporting metadata of the publication.</p> <ul> <li>manual_inspection: quality metrics computed for the semisynthetic dataset used for the manual inspection. <ul> <li>manual_inspection_bg: indices for the selection of the background region in each sample.</li> <li>manual_inspection_free_parameters: parameters used for the generation of the simulated artifacts.</li> <li>manual_inspection_metrics: quality metrics computed for the dataset, used for the manual inspection.</li> <li>manual_inspection_samples: free parameters associated to each image of the dataset for manual inspection.</li> </ul> </li> <li>LDA_training: semisynthetic dataset used to train the Linear Discriminant Analysis (LDA) model. <ul> <li>lda_metrics_synthetic+semisynthetic_uniform_max: quality metrics computed for the training dataset, with maximum normalization of the images. (Not used in the manuscript)</li> <li>lda_metrics_synthetic+semisynthetic_uniform_rescale01: quality metrics computed for the training dataset, with image values rescaled between 0 and 1.</li> <li>parameters_all_degradations: parameters used for the generation of the simulated artifacts.</li> </ul> </li> <li>LDA_prediction: datasets predicted by the LDA model. <ul> <li>experimental: quality metrics computed for measurements of different samples. Images are from publicly available datasets from [2] and [3].</li> <li>known_semisynthetic: metadata for the knwon semi-synthetic dataset used for prediction, assessment and interpretation of the trained model: <ul> <li>known_semisynthetic_free_parameters: parameters used for the generation of the simulated artifacts.</li> <li>known_semisynthetic_metrics_rescale01: quality metrics computed for the training dataset, with image values rescaled between 0 and 1.</li> <li>known_semisynthetic_maxnorm_lda_results: lda prediction results, when metrics are maximum normalized to the training dataset.</li> <li>known_semisynthetic_znorm_lda_results: lda prediction results, when metrics are z-score normalized to the training dataset.</li> </ul> </li> </ul> </li> </ul> <p>Source data can be used to reproduce the results of the manuscript, using the codes shared in the public GitLab repository&nbsp;<em><a href="https://git.photonicdata.science/elena.corbetta/multi-marker-iqa" target="_blank" rel="noopener">multi-marker-IQA</a>.</em></p> <h3><em>How to use the source data</em></h3> <p>The following table describes which data can be used in the scripts provided in the public GitLab repository&nbsp;<em><a href="https://git.photonicdata.science/elena.corbetta/multi-marker-iqa" target="_blank" rel="noopener">multi-marker-IQA</a>.</em></p> <table> <tbody> <tr> <td><strong>Script</strong></td> <td><strong>Data to use</strong></td> <td><strong>Details</strong></td> </tr> <tr> <td>01_quality_metrics</td> <td>Subfolders of MM-IQA_Images_png</td> <td>Include all the images to evaluate in a single subfolder in&nbsp;<code>/test_images</code></td> </tr> <tr> <td>&nbsp;</td> <td>background_idx.xlsx</td> <td>The indices for the samples to evaluate must be included in the table</td> </tr> <tr> <td> <p>01_quality_metrics_visualization</p> <p>01_quality_metrics_visualization_notebook</p> </td> <td>manual_inspection_metrics.xlsx</td> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> <td>known_semisynthetic_metrics_rescale01</td> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> <td>Every metadata included in LDA_predcition/experimental/</td> <td>&nbsp;</td> </tr> <tr> <td> <p>02_lda_training+prediction</p> </td> <td>lda_metrics_synthetic+semisynthetic_uniform_rescale01</td> <td>As training dataset</td> </tr> <tr> <td>&nbsp;</td> <td>known_semisynthetic_metrics_rescale01</td> <td>As prediction dataset</td> </tr> <tr> <td>&nbsp;</td> <td>Every metadata included in LDA_predcition/experimental/</td> <td>As prediction dataset</td> </tr> <tr> <td> <p>02_lda_visualization</p> <p>02_lda_visualization_notebook</p> </td> <td>known_semisynthetic_maxnorm_lda_results</td> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> <td>known_semisynthetic_znorm_lda_results</td> <td>&nbsp;</td> </tr> <tr> <td>Notebook_test-iqa</td> <td>A small dataset with image data and the relative background index, if available.</td> <td>Use a limited number of images.</td> </tr> <tr> <td>Notebook_mm-iqa_workflow</td> <td>Any image dataset with the relative background indices</td> <td>For quality assessment and as prediction dataset</td> </tr> <tr> <td>&nbsp;</td> <td>lda_metrics_synthetic+semisynthetic_uniform_rescale01</td> <td>As training dataset</td> </tr> </tbody> </table> <p>&nbsp;</p> <h2>MM-IQA_Scripts</h2> <ul> <li><strong>multi-marker-iqa-main</strong>: original GitLab repository for MM-IQA, version available at the date of manuscript publication.</li> <li><strong>Notebooks_peer_review</strong>: additional notebooks generated during the peer-review process with the computation of metrics for natural images and correlation measures.</li> </ul>

restrictedcc-by-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record