Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Machine Learning Models Predict the Primary Sites of Head and Neck Squamous Cell Carcinoma Metastases Based on DNA Methylation
GEO Series GSE171994. Homo sapiens. 49 samples. Type: Methylation profiling by genome tiling array.
Multi-omics and machine learning reveal context-specific gene regulatory activities of PML-RARA in Acute Promyelocytic Leukemia
GEO Series GSE173755. Homo sapiens. 28 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing; Other.
Multi-omics and Machine Learning Accurately Predicts Clinical Response to Adalimumab and Etanercept Therapy in Patients with Rheumatoid Arthritis
GEO Series GSE138747. Homo sapiens. 320 samples. Type: Methylation profiling by genome tiling array; Expression profiling by high throughput sequencing.
A machine learning approach to integrate big data for precision medicine in acute myeloid leukemia
GEO Series GSE108004. Homo sapiens. 54 samples. Type: Expression profiling by array; Expression profiling by high throughput sequencing.
A machine learning approach to integrate big data for precision medicine in acute myeloid leukemia [RNA-Seq]
GEO Series GSE108003. Homo sapiens. 24 samples. Type: Expression profiling by high throughput sequencing.
Evaluation via Supervised Machine Learning of the Broiler Pectoralis Major and Liver Transcriptome in Association with the Muscle Myopathy Wooden Breast
GEO Series GSE144000. Gallus gallus. 35 samples. Type: Expression profiling by high throughput sequencing.
Machine learning accelerates the dissection of mitostasis as a cross-species central biological hub for leaf senescence
GEO Series GSE201607. Arabidopsis thaliana; Solanum lycopersicum; Oryza sativa Japonica Group. 18 samples. Type: Expression profiling by high throughput sequencing.
Model-to-crop conserved NUE Regulons enhance machine learning predictions of nitrogen use efficiency
GEO Series GSE280353. Arabidopsis thaliana; Zea mays. 354 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
Integrative Analysis and Machine Learning based Characterization of Single Circulating Tumor Cells
GEO Series GSE129474. Homo sapiens. 15 samples. Type: Expression profiling by high throughput sequencing.
FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster
GEO Series GSE282899. Drosophila melanogaster. 4 samples. Type: Other.
Machine learning-based prediction of the activity and specificity of Cas9 variants in gene editing
GEO Series GSE231840. Homo sapiens. 64 samples. Type: Other.
Exploring Context-Free Languages via Planning: The Case for Automating Machine Learning -- Additional Plots
<p>A comparison to SOTA AutoML approach on various datasets</p>
Long-term trends of ambient nitrate (NO3-) concentrations across China based on ensemble machine-learning models
<p>The monthly NO3- concentrations across China during 2005-2015</p>
Data from the manuscript: Machine learning reveals structural characteristics of stereochemistry-specific interdigitation of synthetic monomycoloyl glycerol analogs
<h2>Four membrane configurations:</h2> <p>1. single = single bilayer (512 MMG molecules and 25600 water molecule) <br>2. db_large = large double bilayer (1024 MMG molecules and 51200 water molecules)<br>3. id_small = interdigitated double bilayer (256 MMG molecules and 12800 water molecules)<br>4. db_small = small double bilayer (256 MMG molecules and 12800 water molecules)</p> <p>Each membrane configuration has 2 different MMG analogs:</p> <p><br>1. MMG1 = MMG-1, 1:1 racemic mixture of stereoisomers with (2R,3S)/(2S,3R) configurations<br>2. MMG6 = MMG-6, 1:1 racemic mixture of stereoisomers with (2R,3R)/(2S,3S) configurations</p> <h2>File names:</h2> <p><br>Each trajectory folder contains starting structures, input structure files, trajectory, energy file and system topology file.<br>Additionally, the .mdp files and topologies are located in their own folders.</p> <p>- start.gro = starting structure before energy minimization**<br>- eq.gro = input structure file for production<br>- eq.cpt = equilibration checkpoint file<br>- run.gro = final frame structure file<br>- run.cpt = production checkpoint file<br>- run.tpr = production tpr file <br>- run.xtc = production trajectory file, final 500 ns<br>- run.edr = production energy file, final 500 ns</p> <p>the <em>db_large</em> system is built with equilibrated <em>db_small</em> system, hence the starting structure is before the equilibration.</p> <h2>MDP-files:</h2> <p><br>Molecular dynamics parameter files can be found in their own folder (05_mdps/). Files for energy minimization,<br>equilibration and production run are provided.</p> <p>- em1_charmm36.mdp = 1st energy minimization (for all systems)<br>- em_charmm36.mdp = 2nd energy minimization (for all systems)<br>- eq1_noposres_charmm36.mdp = equilibration without position restraints (for single)<br>- eq1_posresin_charmm36.mdp = equilibration with O1 atoms constrained in Z-direction (for db_large, id_small, db_small)<br>- production_charmm36.mdp = production run (for all systems)</p> <h2>Topologies and position restraint files:</h2> <p><br>System topologies for each MMG analog can be found in their own folders (06_top/MMG1/, 06_top/MMG6/). <br>Molecular topologies, forcefield parameters and position restraint files can be found from their respective folders (06_top/MMG1/topol/, 06_top/MMG6/topol/).</p> <h3>MMG-1:</h3> <p><br>- db_large-MMG1.top = MMG-1 large double-bilayer system<br>- db_small-MMG1.top = MMG-1 small systems: interdigitated and non-interdigitated double bilayers <br>- single-MMG1.top = MMG-1 single bilayer system<br>- forcefield.itp = force field parameters<br>- TIP3_CHARMM36.itp = water model<br>- MMG1_2R3S_outer.itp = outer leaflet 2R3S MMG-1 <br>- MMG1_2S3R_outer.itp = outer leaflet 2S3R MMG-1 <br>- MMG1_2R3S_inner.itp = inner leaflet 2R3S MMG-1<br>- MMG1_2S3R_inner.itp = inner leaflet 2S3R MMG-1<br>- posres_MMG1_tail_z.itp = position restraints on MMG tails in z-direction<br>- posres_MMG1_O1_z.itp = position restraints on MMG headgroup O1 atom in z-direction</p> <h3>MMG-6:</h3> <p><br>- db_large-MMG6.top = MMG-6 large double-bilayer system<br>- db_small-MMG6.top = MMG-6 small systems: interdigitated and non-interdigitated double bilayers<br>- single-MMG6.top = MMG-6 single bilayer system<br>- forcefield.itp = force field parameters<br>- TIP3_CHARMM36.itp = water model<br>- MMG6_2S3S_outer.itp = outer leaflet 2S3S MMG-6 <br>- MMG6_2R3R_outer.itp = outer leaflet 2R3R MMG-6<br>- MMG6_2S3S_inner.itp = inner leaflet 2S3S MMG-6<br>- MMG6_2R3R_inner.itp = inner leaflet 2R3R MMG-6 <br>- posres_MMG6_tail_z_2R3R.itp = position restraints on tails in z-direction <br>- posres_MMG6_tail_z_2S3S.itp = position restraints on tails in z-direction<br>- posres_MMG6_O1_z.itp = position restraints on MMG headgroup O1 atom in z-direction</p>
MACHINE LEARNING IN THE FINANCIAL INDUSTRY: A BIBLIOMETRIC APPROACH TO EVIDENCING APPLICATIONS
Open the record for dataset details and reuse information.
Deciphering complex antibiotic resistance patterns in Helicobacter pylori through whole genome sequencing and machine learning
<p>Whole Genome Sequencing of <em>Helicobacter pylori</em> (<em>Hp</em>).</p>
Dataset fot the study of Documentation Practices of Machine Learning Resources
Open the record for dataset details and reuse information.
Approaches for machine learning intermolecular interaction energies and application to energy components from symmetry adapted perturbation theory
Open the record for dataset details and reuse information.
Restricted Data Set for the article Analysing the impact of renewables on Iberian wholesale electricity market prices using machine learning techniques. Green Finance, 2024, 6 (2), 363-382
<p><span>Two of the series contained in the databases used in the article [</span><em><span>Analysing the impact of renewables on Iberian wholesale electricity market prices using machine learning techniques. Green Finance, 6 (2), 363-382</span></em><span>] are not authorized for public sharing. Specifically, these include the data corresponding to the Dutch TTF futures price and the API2 index. Consequently, they are deposited under restricted access.</span></p>
Machine Learning-Based Estimation of Experimental Artifacts and Image Quality in Fluorescence Microscopy - Supporting Data
<p>Supporting data for the reproduction of the results reported in Corbetta, E., Bocklitz, T., Machine learning based estimation of experimental artifacts and image quality in fluorescence microscopy (2024) [1].</p> <h2><strong>MM-IQA_Images_png and </strong><strong>MM-IQA_Images_tif</strong></h2> <p>Folder containing all the supporting datasets of the publication.</p> <ul> <li>images_manual_inspection: semisynthetic dataset used for the manual inspection of the quality metrics.</li> <li>images_lda_training: semisynthetic dataset used to train the Linear Discriminant Analysis (LDA) model.</li> <li>images_lda_prediction: datasets predicted by the LDA model. <ul> <li>images_experimental: every subfolder is a dataset composed of measurements of a different sample. Images are from publicly available datasets from [2] and [3].</li> <li>images_known_semisynthetic: knwon semisynthetic dataset used for prediction, assessment and interpretation of the trained model.</li> </ul> </li> </ul> <p>Tif files are the original data used for the study.</p> <h2><strong>MM-IQA_Source_data</strong></h2> <p>Folder containing all the supporting metadata of the publication.</p> <ul> <li>manual_inspection: quality metrics computed for the semisynthetic dataset used for the manual inspection. <ul> <li>manual_inspection_bg: indices for the selection of the background region in each sample.</li> <li>manual_inspection_free_parameters: parameters used for the generation of the simulated artifacts.</li> <li>manual_inspection_metrics: quality metrics computed for the dataset, used for the manual inspection.</li> <li>manual_inspection_samples: free parameters associated to each image of the dataset for manual inspection.</li> </ul> </li> <li>LDA_training: semisynthetic dataset used to train the Linear Discriminant Analysis (LDA) model. <ul> <li>lda_metrics_synthetic+semisynthetic_uniform_max: quality metrics computed for the training dataset, with maximum normalization of the images. (Not used in the manuscript)</li> <li>lda_metrics_synthetic+semisynthetic_uniform_rescale01: quality metrics computed for the training dataset, with image values rescaled between 0 and 1.</li> <li>parameters_all_degradations: parameters used for the generation of the simulated artifacts.</li> </ul> </li> <li>LDA_prediction: datasets predicted by the LDA model. <ul> <li>experimental: quality metrics computed for measurements of different samples. Images are from publicly available datasets from [2] and [3].</li> <li>known_semisynthetic: metadata for the knwon semi-synthetic dataset used for prediction, assessment and interpretation of the trained model: <ul> <li>known_semisynthetic_free_parameters: parameters used for the generation of the simulated artifacts.</li> <li>known_semisynthetic_metrics_rescale01: quality metrics computed for the training dataset, with image values rescaled between 0 and 1.</li> <li>known_semisynthetic_maxnorm_lda_results: lda prediction results, when metrics are maximum normalized to the training dataset.</li> <li>known_semisynthetic_znorm_lda_results: lda prediction results, when metrics are z-score normalized to the training dataset.</li> </ul> </li> </ul> </li> </ul> <p>Source data can be used to reproduce the results of the manuscript, using the codes shared in the public GitLab repository <em><a href="https://git.photonicdata.science/elena.corbetta/multi-marker-iqa" target="_blank" rel="noopener">multi-marker-IQA</a>.</em></p> <h3><em>How to use the source data</em></h3> <p>The following table describes which data can be used in the scripts provided in the public GitLab repository <em><a href="https://git.photonicdata.science/elena.corbetta/multi-marker-iqa" target="_blank" rel="noopener">multi-marker-IQA</a>.</em></p> <table> <tbody> <tr> <td><strong>Script</strong></td> <td><strong>Data to use</strong></td> <td><strong>Details</strong></td> </tr> <tr> <td>01_quality_metrics</td> <td>Subfolders of MM-IQA_Images_png</td> <td>Include all the images to evaluate in a single subfolder in <code>/test_images</code></td> </tr> <tr> <td> </td> <td>background_idx.xlsx</td> <td>The indices for the samples to evaluate must be included in the table</td> </tr> <tr> <td> <p>01_quality_metrics_visualization</p> <p>01_quality_metrics_visualization_notebook</p> </td> <td>manual_inspection_metrics.xlsx</td> <td> </td> </tr> <tr> <td> </td> <td>known_semisynthetic_metrics_rescale01</td> <td> </td> </tr> <tr> <td> </td> <td>Every metadata included in LDA_predcition/experimental/</td> <td> </td> </tr> <tr> <td> <p>02_lda_training+prediction</p> </td> <td>lda_metrics_synthetic+semisynthetic_uniform_rescale01</td> <td>As training dataset</td> </tr> <tr> <td> </td> <td>known_semisynthetic_metrics_rescale01</td> <td>As prediction dataset</td> </tr> <tr> <td> </td> <td>Every metadata included in LDA_predcition/experimental/</td> <td>As prediction dataset</td> </tr> <tr> <td> <p>02_lda_visualization</p> <p>02_lda_visualization_notebook</p> </td> <td>known_semisynthetic_maxnorm_lda_results</td> <td> </td> </tr> <tr> <td> </td> <td>known_semisynthetic_znorm_lda_results</td> <td> </td> </tr> <tr> <td>Notebook_test-iqa</td> <td>A small dataset with image data and the relative background index, if available.</td> <td>Use a limited number of images.</td> </tr> <tr> <td>Notebook_mm-iqa_workflow</td> <td>Any image dataset with the relative background indices</td> <td>For quality assessment and as prediction dataset</td> </tr> <tr> <td> </td> <td>lda_metrics_synthetic+semisynthetic_uniform_rescale01</td> <td>As training dataset</td> </tr> </tbody> </table> <p> </p> <h2>MM-IQA_Scripts</h2> <ul> <li><strong>multi-marker-iqa-main</strong>: original GitLab repository for MM-IQA, version available at the date of manuscript publication.</li> <li><strong>Notebooks_peer_review</strong>: additional notebooks generated during the peer-review process with the computation of metrics for natural images and correlation measures.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.