Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

80

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

80 results for “feature learning”

Learn how ShareScore rates datasets ↗
zenodo32/100

Basal cell carcinoma diagnosis with fusion of deep learning and telangiectasia features

<p>Telangiectasia masks dataset created on a subset of the ISIC18, ISIC19 training datasets and the NIH study dataset R43 CA153927-01 and CA101639-02A2. All annotations are for Basal Cell Carcinoma lesions. This is an expanded dataset that was initially used in &ldquo;<a href="https://onlinelibrary.wiley.com/doi/10.1111/srt.13150">A Deep Learning Approach to Detect Blood Vessels in Basal Cell Carcinoma</a>&rdquo;.</p> <p>A sample lesion image and the corresponding mask is provided for preview. Lesion images and masks have been uploaded as separate zipped folders that can be downloaded.</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Mechanism of building a machine learning model and its main features

<p>This figure is a flowchart summarizing the steps required to build a machine learning model in general way.</p>

opencc-by-4.0May 2023View details →
zenodo28/100

Monitoring forest health using hyperspectral imagery: Does feature selection improve the performance of machine-learning techniques?

<p>This is a research compendium (RC) for the publication</p> <blockquote> <p>Monitoring forest health using hyperspectral imagery: Does feature selection improve the performance of machine-learning techniques?</p> </blockquote> <p>Code, figures, appendices and the manuscript can be found in the corresponding <a href="https://github.com/pat-s/2019-feature-selection">GitHub repository</a>.</p> <p>This RC is a static snapshot at the time of submission. The GitHub repository holds the latest version and may see changes after the publication was accepted.</p> <p><strong>Data sources and description</strong></p> <ul> <li><em>aoi.gpkg</em><strong>:</strong> Area of interest for downloading Sentinel-2 images. <em>Not used in the publication<strong>. </strong></em>Source:<em><strong> </strong></em>Custom.</li> <li><em>forest_mask.gpkg</em>: A forest/non-forest mask of the Basque Country. <em>Not used in the publication</em>. Source:<em><strong> </strong></em>Custom.</li> <li><em>hyperspectral.zip: </em>Hyperspectral remote sensing data used to extract reflectance values on the tree level. Source:<em><strong> </strong></em>Custom.</li> <li><em>plot-locations.gpkg: </em>Spatial location of the plots used in the study. Source:<em><strong> </strong></em>Custom.</li> <li><em>tree-in-situ-data-corrected.zip: </em>Corrected in-situ data containing defoliation information on the tree level. A correction of the spatial location was applied by the creators of the data. Source:<em><strong> </strong></em>Custom.</li> <li><em>tree-in-situ-data.zip: </em>First version of in-situ data containing defoliation information on the tree level.<em> Not used in the publication. </em>Source:<em><strong> </strong></em>Custom.</li> </ul> <p><strong>Licenses</strong></p> <p>All files are licensed under CC BY 4.0.</p>

opencc-by-4.0Jan 2020View details →
zenodo28/100

Transfer Learning for leveraging computer vision in infrastructure maintenance [extracted features]

<p>Dataset containing features extracted from images taken on single case of&nbsp;infrastructure facility for the purpose of training Transfer Learned CNN classifier. It is meant to be used with KrakN framework (https://github.com/MatZar01/KrakN), published with the research paper.</p>

opencc-by-4.0Apr 2020View details →
zenodo28/100

Dataset to replicate experiments in "Image Feature Learning with Genetic Programming" paper

<p>The zip file contains&nbsp; Dataset to replicate experiments in&nbsp; &quot;Image Feature Learning with Genetic Programming&quot; paper published at the PPSN 2020 conference.</p> <p>This package also contains a version of Lenet5 to classify the MNIST digits.</p> <p>MNIST dataset has been corrupted with salt noise.&nbsp; We made white a different proportion of the pixel at random (5% 10% 30% 40%).</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

Data set for "Dynamic perceptual feature selectivity in primary somatosensory cortex upon reversal learning"

<p>This repository contains the data used to generate the figures and well as the main codes that were used for analyses.</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

Investigating Non-Usually Employed Features in the Identification of Architectural Smells: A Machine Learning-Based Approach

<p>Architectural smells (ASs) negatively affect the maintenance and evolution of software at the architectural level. Most of the current approaches for ASs identification rely on the same small and well-known set of usually employed metrics (UE-Ms) with fixed thresholds. Machine learning (ML) is a promising technique for smell identification as algorithms can learn from a rich set of metrics/features, covering several characteristics of the software and incorporating a certain degree of subjectivity. This has been explored by building datasets with a robust and rich set of features, including not only the UE-Ms but also other non-usually employed metrics (NUE-Ms). However, usually the UE-Ms determine the output of the algorithms, obfuscating other metrics that have the potential to improve the classification. This also leads to inflated and difficult to maintain datasets.&nbsp;<br> In this paper, we investigate the accuracy of some ML algorithms employing only NUE-Ms. We scoped our study in the classification of two smells: God Component and Unstable Dependency. This investigation revealed a set of NUE-Ms that can be also used to identify these smells and the contribution of each one for the classification. Based on this information, software engineers can then build a final dataset just with the potential features. We also briefly present our tool, called InSet, that was used by academics and practitioners to identify smells in their systems. The feedback of them was used as the oracle to compare our tool to other approaches and good results were reached.&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo28/100

Predicting Hydrophobicity by Learning Spatiotemporal Features of Interfacial Water Structure: Combining Molecular Dynamics Simulations with Convolutional Neural Networks

<p>Files for reproducing results from Kelkar et al. (JPCB 2020) -&nbsp;Predicting Hydrophobicity by Learning Spatiotemporal Features of Interfacial Water Structure: Combining Molecular Dynamics Simulations with Convolutional Neural Networks</p> <p>&nbsp;</p> <p>This folder contains simulations starter files and also plug-and-play datasets to test ML algorithms on molecular dynamics (MD) simulation data.</p> <p>&nbsp;</p> <p>All analysis scripts can also be found on GitLab on this link:&nbsp;https://gitlab.com/atharva-kelkar/kelkar_et_al_jpcb_2020</p>

opencc-by-4.0Sep 2020View details →
zenodo28/100

3D-MSNet: A point cloud based deep learning model for untargeted feature detection and quantification in profile LC-HRMS data

<p>Supplementary data of 3D-MSNet</p>

opencc-by-4.0May 2022View details →
zenodo28/100

Data for EMO2023 Paper "Feature-based Benchmarking of Distance-based Multi/Many-objective Optimisation Problems: A Machine Learning Perspective"

<p><strong>Data for Paper &quot;Feature-based Benchmarking of Distance-based Multi/Many-objective Optimisation Problems: A Machine Learning Perspective&quot;</strong></p> <p><br> The file <strong>dbmopp_dataset_perf.csv</strong> contains results from the 945 x 30 instances, with the following columns:</p> <ul> <li><em>design_id</em>: problem identifier</li> <li><em>n_var</em>: number of variables {2, ..., 20}</li> <li><em>n_obj</em>: number of objectives {2, ..., 10}</li> <li><em>nonident_ps</em>: non-identical Pareto sets {0 (no), 1 (yes)}</li> <li><em>var_density</em>: varying density {0 (no), 1 (yes)}</li> <li><em>n_discon_ps</em>: number of disconnected Pareto sets {0, ..., 6}</li> <li><em>n_local_fronts</em>: number of local fronts {0, ..., 6}</li> <li><em>n_resist_regions</em>: number of dominance resistance regions {0, ..., 6}</li> <li><em>instance_id</em>: instance (fold) identifier {1, ..., 30}</li> <li><em>budget</em>: number of evaluations performed by the algorithm {5000, 10000, 30000, 50000}</li> <li><em>algo</em>: multi-objective evolutionary algorithm {NSGAII, IBEA, MOEAD, Random}</li> <li><em>hypervolume</em>: hypervolume reached by the algorithm [0.0, 1.0]</li> </ul> <p>&nbsp;</p> <p>The file <strong>dbmopp_dataset_perf_aggregated.csv</strong> contains average results from the 945 problems, with the following columns:</p> <ul> <li><em>design_id</em>: problem identifier</li> <li><em>n_var</em>: number of variables {2, ..., 20}</li> <li><em>n_obj</em>: number of objectives {2, ..., 10}</li> <li><em>nonident_ps</em>: non-identical Pareto sets {0 (no), 1 (yes)}</li> <li><em>var_density</em>: varying density {0 (no), 1 (yes)}</li> <li><em>n_discon_ps</em>: number of disconnected Pareto sets {0, ..., 6}</li> <li><em>n_local_fronts</em>: number of local fronts {0, ..., 6}</li> <li><em>n_resist_regions</em>: number of dominance resistance regions {0, ..., 6}</li> <li><em>budget</em>: number of evaluations performed by the algorithm {5000, 10000, 30000, 50000}</li> <li><em>algo</em>: multi-objective evolutionary algorithm {NSGAII, IBEA, MOEAD, Random}</li> <li><em>hypervolume_avg</em>: average hypervolume reached by the algorithm [0.0, 1.0]</li> <li><em>best</em>: 1 if the corresponding algorithm obtains the best average hypervolume, 0 otherwise</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo28/100

Advanced Iterative Model for Lumpy Skin Disease Prediction Using Fine-grained Feature Fusion and Adaptive Transfer Learning

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo28/100

Impact of Usability on Continuance Usage Intention in Language Learning Apps with Gamification Features

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo28/100

The Best of Both Worlds: Combining Learned Embeddings with Engineered Features for Accurate Prediction of Correct Patches

<p>Dataset for Panther</p>

opencc-by-4.0Nov 2022View details →
zenodo28/100

Contrastive learning-based histopathological feature infers molecular subtypes and clinical outcomes of breast cancer from unannotated whole slide images

<p>The breast cancer cohort&nbsp;came from the Changzhou Second&nbsp;People&#39;s Hospital&nbsp;(CZSPH)&nbsp;in&nbsp;Jiangsu, China.&nbsp;This cohort&nbsp;collected 91 FFPE WSIs from 90 breast cancer&nbsp;patients, including 15&nbsp;recurrence cases within 5 years.</p>

opencc-by-4.0May 2023View details →
ClinicalTrials.gov28/100

Deep Learning-Based Confocal Laser Microendoscopy Feature Atlas Construction and Its Application in Intelligent Diagnosis of Irritable Bowel Syndrome

ClinicalTrials.gov study NCT07051226. IPD Sharing: Not stated. Countries: 0. Publications: 32.

restrictedIPD-UNDECIDEDFeb 2026View details →
geo24/100

Multi-task learning uncovers robust translation cis-regulatory features

GEO Series GSE201766. Homo sapiens. 1 samples. Type: Other.

openGEO-OpenApr 2022View details →
geo24/100

ALS molecular subtypes are a combination of cellular and pathological features learned by deep multiomics classifiers (postmortem spinal cord)

GEO Series GSE272626. Homo sapiens. 266 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2024View details →
geo24/100

Radiogenomics of glioblastoma: Machine-learning based classification of molecular characteristics using multiparametric and multiregional MRI features

GEO Series GSE85539. Homo sapiens. 152 samples. Type: Methylation profiling by array.

openGEO-OpenNov 2016View details →
geo24/100

Machine learning analysis of the T cell receptor repertoire identifies sequence features that predict self-reactivity

GEO Series GSE221703. Mus musculus. 20 samples. Type: Other.

openGEO-OpenJan 2023View details →
geo24/100

ALS molecular subtypes are a combination of cellular and pathological features learned by deep multiomics classifiers (postmortem cortex)

GEO Series GSE272624. Homo sapiens. 289 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record