Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

78

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

78 results for “Machine Learning Classification”

Learn how ShareScore rates datasets ↗
geo24/100

Radiogenomics of glioblastoma: Machine-learning based classification of molecular characteristics using multiparametric and multiregional MRI features

GEO Series GSE85539. Homo sapiens. 152 samples. Type: Methylation profiling by array.

openGEO-OpenNov 2016View details →
geo24/100

Machine-learning Classification Identifies Early Systemic Sclerosis Patients that Improve with Abatacept treatment by Modulating CD28-Pathways

GEO Series GSE217067. Homo sapiens. 233 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2023View details →
zenodo24/100

Star formation and morphological properties of galaxies in the P\lowercase{an}-STARRS 3$\pi$ survey- I.\\ A machine learning approach to galaxy and supernova classification

<pre>This is the catalog presented in Baldeschi, et al (2020). The catalog is subdivided in 13 csv files. Files description: First column: Panstar ID (integer) Second column: Right ascension [deg] (float) Third column: Declination [deg] (float) Fourth column: Probability for a source of being a star (P_star) (float) Fifth column: Probability for a galaxy of being higly star-forming (P_HSFF) (float) Sixth column: Probability for a galaxy of being spiral (P_spiral) (float) seventh column: Compleatness flag (string) If using this catalog for publications, please cite Baldeschi, et al (2020). Fourth column values are from Tachibana &amp; Miller (2018). For a detailed description of the columns we refer to the Appendix A of Baldeschi, et al (2020).</pre>

opencc-by-4.0Jul 2020View details →
zenodo24/100

Identifying galaxies, quasars and stars with machine learning: a new catalogue of classifications for 111 million SDSS sources without spectra - parquet format

<p>This is the same as the published data available under&nbsp;10.5281/zenodo.3768398, but in the format of parquet files. This means you can access it using Dask for convenience when using cloud compute facilities.&nbsp;</p> <p>Abstract: We used 3.1 million spectroscopically labelled sources from the Sloan Digital Sky Survey (SDSS) to train an optimised random forest classifier using photometry from the SDSS and the Widefield Infrared Survey Explorer (WISE). We applied this machine learning model to 111 million previously unlabelled sources from the SDSS photometric catalogue which did not have existing spectroscopic observations. Our new catalogue contains 50.4 million galaxies, 2.1 million quasars, and 58.8 million stars. We provide individual classification probabilities for each source, with 6.7 million galaxies (13%), 0.33 million quasars (15%), and 41.3 million stars (70%) having classification probabilities greater than 0.99; and 35.1 million galaxies (70%), 0.72 million quasars (34%), and 54.7 million stars (93%) having classification probabilities greater than 0.9. Precision, Recall, and F1 score were determined as a function of selected features and magnitude error. We investigate the effect of class imbalance on our machine learning model and discuss the implications of transfer learning for populations of sources at fainter magnitudes than the training set. We used a non-linear dimension reduction technique (Uniform Manifold Approximation and Projection: UMAP) in unsupervised, semi-supervised, and fully-supervised schemes to visualise the separation of galaxies, quasars, and stars in a two-dimensional space. When applying this algorithm to the 111 million sources without spectra, it is in strong agreement with the class labels applied by our random forest model.</p> <p>When using this dataset, please reference our paper via the journal (<a href="https://arxiv.org/abs/1909.10963">https://arxiv.org/abs/1909.10963</a>) and this DOI (10.5281/zenodo.4060257). If you make use of our scripts please reference our Github repository DOI (10.5281/zenodo.3855160).</p> <p>File descriptions:</p> <p>All of these files are Pandas Dataframes, saved as uncompressed parquet files for ease of access when using cloud compute such as Dask. df_spec_classprobs.parquet&nbsp;contains the spectroscopically observed sources used for training and testing. This has been cleaned, and has the results of the random forest classifier added as additional columns (sources used for training have NaNs in the class_pred column). SDSS-ML-all.parquet contains the 111 million photometrically observed sources, with our class labels and probabilities added.</p>

opencc-by-4.0Sep 2020View details →
zenodo24/100

Machine learning-based pulse wave analysis for classification of circle of Willis topology: an in silico study with 30,618 virtual subjects (database: Missing PCA P1)

<p>This repository contains the dataset for the Missing PCA P1 described in the article with the same name. MATLAB and Python codes for post-processing the dataset and the code for training and testing all machine learning models using the open-source library TensorFlow 2.12, the Keras application programming interface, and the Scikit-learn Python package can be found in here (<a href="https://zenodo.org/records/12519322" target="_blank" rel="noopener">https://zenodo.org/records/12519322</a>).</p>

opencc-by-4.0Jun 2024View details →
ClinicalTrials.gov24/100

Machine Learning to Analyze Facial Imaging, Voice and Spoken Language for the Capture and Classification of Cancer/Tumor Pain

ClinicalTrials.gov study NCT04442425. IPD Sharing: YES. Countries: 1. Publications: 0.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov24/100

Machine Learning-based Classification of Symptom Clusters and Online CBT

ClinicalTrials.gov study NCT06350201. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov24/100

Prediction of Antidepressant Treatment Response Using Machine Learning Classification Analysis

ClinicalTrials.gov study NCT02330679. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
geo24/100

EpiGe: A cytosine methyl-genotyping PCR-based machine-learning tool for rapid classification of medulloblastoma

GEO Series GSE210723. Homo sapiens. 74 samples. Type: Methylation profiling by array.

openGEO-OpenAug 2023View details →
geo16/100

The brain tumor classifications based on machine learning models trained with older version of methylation microarray chip compatible with the new EPIC v2 illumina’s chip.

GEO Series GSE229715. Homo sapiens. 32 samples. Type: Methylation profiling by genome tiling array.

openGEO-OpenSep 2023View details →
geo16/100

Machine learning-assisted classification of responders and non-responders to bortezomib treatment regimens PAD and VCD using new experimental dataset of 58 multiple myeloma RNA sequencing profiles

GEO Series GSE159426. Homo sapiens. 58 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2021View details →
geo16/100

DNA methylation-based machine learning classification distinguishes pleural mesothelioma from chronic pleuritis, pleural carcinosis, and pleomorphic lung carcinomas

GEO Series GSE203061. Homo sapiens. 76 samples. Type: Methylation profiling by genome tiling array.

openGEO-OpenMay 2022View details →
zenodo12/100

Replication Package for the Paper: "A Machine Learning Based Ensemble Method for Automatic Classification of Decisions"

<p>This is the replication package for the paper: &quot;A Machine Learning Based Ensemble Method for Automatic Classification of Decisions&quot;.&nbsp;It contains the source code and dataset of our experiment for the&nbsp;replication&nbsp;by&nbsp;other&nbsp;researchers. In the meanwhile, we provide brief description of the files in the replication&nbsp;package in the following.</p> <p><strong>1. code folder</strong></p> <ul> <li><em>experiment.py&nbsp;&nbsp;</em>contains the source code for our experiment, which is conducted on Windows 10 and Python 3.7.0.&nbsp;<strong>Note that you may&nbsp;get slightly</strong>&nbsp;<strong>different experiment&nbsp;results when conducting the experiments&nbsp;on different environment configurations.</strong></li> <li><em>requirements.txt</em>&nbsp; records all the installation packages and their version numbers needed for the current program to run.&nbsp;You&nbsp;can use &quot;<em>pip install -r requirements.txt</em>&quot; to rebuild the project and install all dependencies. <strong>Note that you may&nbsp;get slightly different experiment&nbsp;results when using different packages or versions.&nbsp;</strong></li> </ul> <p><strong>2. dataset folder</strong></p> <ul> <li><em>decisions.xlsx&nbsp;&nbsp;</em>contains 848 labelled sentence-level decisions from the Hibernate developer mailing list.</li> </ul>

restrictedMay 2020View details →
zenodo12/100

Replication Package for the Paper: "A Machine Learning Based Ensemble Method for Automatic Classification of Decisions: A Study of the Hibernate Developer Mailing List"

<p>This is the replication package for the paper: &quot;A Machine Learning Based Ensemble Method for Automatic Classification of Decisions: A Study of the Hibernate Developer Mailing List&quot;.&nbsp;It contains the source code and dataset of our experiment for the&nbsp;replication&nbsp;by&nbsp;other&nbsp;researchers. In the meanwhile, we provide brief description of the files in the replication&nbsp;package below.</p> <p><strong>1. code folder</strong></p> <ul> <li><em>experiment.py&nbsp;&nbsp;</em>contains the source code for our experiment, which is conducted on Windows 10 and Python 3.7.0.&nbsp;<strong>Note that you may&nbsp;get slightly</strong>&nbsp;<strong>different experiment&nbsp;results when conducting the experiments&nbsp;on different environment configurations.</strong></li> <li><em>requirement.txt</em>&nbsp; records all the installation packages and their version numbers needed for the current program to run.&nbsp;You&nbsp;can use &quot;<em>pip install -r requirement.txt</em>&quot; to rebuild the project and install all dependencies. <strong>Note that you may&nbsp;get slightly different experiment&nbsp;results when using different packages or versions.&nbsp;</strong></li> </ul> <p><strong>2. dataset folder</strong></p> <ul> <li><em>decisions.xlsx&nbsp;&nbsp;</em>contains 844&nbsp;labelled sentence-level decisions from the Hibernate developer mailing list.</li> </ul>

restrictedJul 2020View details →
zenodo12/100

Data for the manuscript "Classification of Stream, Hyperconcentrated, and Debris Flow Using Dimensional Analysis and Machine Learning"

<p>The excel file&nbsp;contains&nbsp;hydrological and sediment data.&nbsp;Also included are dimensional analysis data in our dataset.<br> The rar file contains the codes and data for SVM classification work.</p>

restrictedJul 2022View details →
zenodo12/100

Dataset for "Classification of Stream, Hyperconcentrated, and Debris Flow Using Dimensional Analysis and Machine Learning"

<p>Du J. et al., (2022). Dataset for &quot;Classification of Stream, Hyperconcentrated, and Debris Flow Using Dimensional Analysis and Machine Learning&quot;, Water Resources Research</p> <p>Table S1:&nbsp;Hydrologic Paramteres and Dimensionless Numbers of Debris Flows</p> <p>Table S2:&nbsp;Hydrologic Paramteres and Dimensionless Numbers of Hyperconcentrated Flows</p> <p>Table S3:&nbsp;Hydrologic Paramteres and Dimensionless Numbers of Stream Flows</p> <p>Table S4:&nbsp;Hydrologic Paramteres and Dimensionless Numbers of Lahars</p>

restrictedNov 2022View details →
zenodo12/100

Machine Learning approach to Classification of Resting-State EEG Microstates in stroke survivors

<p>Dataset -&nbsp;Machine Learning approach to Classification of Resting-State EEG Microstates in stroke survivors</p> <p>https://docs.google.com/spreadsheets/d/1MeEx9ysEC_tWyohqEmvc8Ey5mc5Z9NV2/edit#gid=160681293</p>

restrictedJan 2023View details →
zenodo8/100

MRI radiomics-based machine-learning classification of bone chondrosarcoma

<p><strong>Purpose:&nbsp;</strong>To evaluate the diagnostic performance of machine learning for discrimination between low-grade and high-grade cartilaginous bone tumors based on radiomic parameters extracted from unenhanced magnetic resonance imaging (MRI).</p> <p><strong>Methods:&nbsp;</strong>We retrospectively enrolled 58 patients with histologically-proven low-grade/atypical cartilaginous tumor of the appendicular skeleton (n = 26) or higher-grade chondrosarcoma (n = 32, including 16 appendicular and 16 axial lesions). They were randomly divided into training (n = 42) and test (n = 16) groups for model tuning and testing, respectively. All tumors were manually segmented on T1-weighted and T2-weighted images by drawing bidimensional regions of interest, which were used for first order and texture feature extraction. A Random Forest wrapper was employed for feature selection. The resulting dataset was used to train a locally weighted ensemble classifier (AdaboostM1). Its performance was assessed via 10-fold cross-validation on the training data and then on the previously unseen test set. Thereafter, an experienced musculoskeletal radiologist blinded to histological and radiomic data qualitatively evaluated the cartilaginous tumors in the test group.</p> <p><strong>Results:&nbsp;</strong>After feature selection, the dataset was reduced to 4 features extracted from T1-weighted images. AdaboostM1 correctly classified 85.7 % and 75 % of the lesions in the training and test groups, respectively. The corresponding areas under the receiver operating characteristic curve were 0.85 and 0.78. The radiologist correctly graded 81.3 % of the lesions. There was no significant difference in performance between the radiologist and machine learning classifier (P = 0.453).</p> <p><strong>Conclusions:&nbsp;</strong>Our machine learning approach showed good diagnostic performance for classification of low-to-high grade cartilaginous bone tumors and could prove a valuable aid in preoperative tumor characterization.</p>

restrictedNov 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record