Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

80

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

80 results for “feature learning”

Learn how ShareScore rates datasets ↗
zenodo36/100

Machine learning classification of archaea and bacteria identifies novel predictive genomic features

<p>Dataset used for the classification analysis of archaea and bacteria based on 77 genomic features calculated using GBRAP (GenBank Retrieving, Analyzing and Parsing) tool (Vischioni, C. et al. Gbrap: a tool to retrieve, parse and analyze genbank files of viral and bacterial species. bioRxiv 2021&ndash;09 (2021)).</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Machine Learning Framework for High-Resolution Air Temperature Downscaling Using LiDAR-Derived Urban Morphological Features

<p>This dataset supports the study titled <em>"Machine Learning Framework for High-Resolution Air Temperature Downscaling Using LiDAR-Derived Urban Morphological Features"</em>, published in&nbsp;<em>Urban Climate</em> (<a href="https://doi.org/10.1016/j.uclim.2024.102102" target="_new" rel="noopener">DOI: 10.1016/j.uclim.2024.102102</a>).</p> <p>&nbsp;</p> <p><strong>Content Overview:</strong></p> <ul> <li> <p><strong>Building Label Data for Footprint Detection</strong>:</p> <ul> <li><em>Amsterdam_BDG_Label.rar</em></li> <li><em>MiamiDade_BDG_Label.rar</em></li> </ul> <p>These are the label datasets used for training the building detection segmentation models. They have been instrumental in accurately detecting building footprints in Amsterdam.</p> </li> <li> <p><strong>Amsterdam_3D_Buildings.rar</strong>: &nbsp;CityGML file of 3D building models for Amsterdam, derived from LiDAR data and U-Net3+ model.</p> </li> </ul> <ul> <li> <p><strong>Morphological Features.rar</strong>: Contains urban morphological features (in raster format) extracted from LiDAR data used in the study.</p> </li> <li> <p><strong>Training and Test Data for Air Temperature Estimation</strong>:</p> <ul> <li><em>Train_Test_AvgTemp_Amsterdam.rar</em></li> <li><em>Train_Test_MaxTemp_Amsterdam.rar</em></li> <li><em>Train_Test_MinTemp_Amsterdam.rar</em></li> </ul> <p>This dataset includes training and testing data for estimating air temperatures in three scenarios: average daily temperature, minimum daily temperature, and maximum daily temperature for the city of Amsterdam.</p> </li> </ul>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Supporting Information for "Machine Learning interpretation of the correlation between infrared emission features of interstellar polycyclic aromatic hydrocarbons"

<p>This is the supporting information for the article&nbsp;&quot;Machine Learning interpretation of the correlation between infrared emission features of interstellar polycyclic aromatic hydrocarbons&quot;. It contains:</p> <p>In&nbsp;Supporting_Information.pdf:</p> <p>1. A map of&nbsp;the correlation between descriptors.</p> <p>2. A spectral distribution.</p> <p>3. An explanation of the file&nbsp;example_code.zip.</p> <p>In&nbsp;example_code.zip:</p> <p>An example code of a machine learning model based on ECFP and corresponding input files.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Blood pressure monitoring during anesthesia induction using PPG morphology features and machine learning

<p>PPG-BP dataset of forty patients undergoing general anesthesia, as described in the corresponding journal publication at PLOS ONE (10.1371/journal.pone.0279419).</p> <p>When using this data, please cite the corresponding journal publication.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
dryad36/100

ECG and EEG stress features for: ECG and EEG based detection and multilevel classification of stress using machine learning for specified genders: A preliminary study

<p>Mental health, especially stress, plays a crucial role in the quality of life. During different phases (luteal and follicular phases) of the menstrual cycle, women may exhibit different responses to stress from men. This, therefore, may have an impact on stress detection and classification accuracy of machine learning models that genders are not taken into account. However, this has never been investigated before. In addition, only a handful of stress detection devices are scientifically validated. To this end, this work proposes stress detection and multilevel stress classification models for unspecified and specified genders through ECG and EEG signals. Models for stress detection are achieved through developing and evaluating multiple individual classifiers. On the other hand, stacking technique is employed to obtain models for multilevel stress classification. ECG and EEG features extracted from 40 subjects (21 females and 19 males) were used to train and validate the models. In the low&amp;high combined stress condition, RBF-SVM and kNN yielded the highest average classification accuracy for females (79.81%) and males (73.77%), respectively. Combining ECG and EEG, the average classification accuracy increased to at least 87.58% (male, high stress) and up to 92.70% (female, high stress). For multilevel stress classification from ECG and EEG, the accuracy for females was 62.60% and for males was 71.57%. This study shows that the difference in genders influences the classification performance for both the detection and multilevel classification of stress. The developed models can be used for both personal (through ECG) and clinical (through ECG and EEG) stress monitoring with and without taking genders into account.</p>

opencc-zeroMar 2023View details →
zenodo36/100

Supplemetary Data for the article: Machine-learning identified molecular fragments responsible for infrared emission features of polycyclic aromatic hydrocarbons

<p>This is a set of&nbsp;Supplementary materials for&nbsp;the article &#39;Machine-learning identified molecular fragments responsible for infrared emission features of polycyclic aromatic hydrocarbons&#39;, by Meng et al.</p> <p>Supplementary_Data_I.pdf&nbsp;contains an extensive table spanning 36 pages that lists the top-10 molecular fragments accountable for the spectral bands between 2.761 and 1172.745 &mu;m. To access this table, hyperlinks within the document can be used for navigation.</p> <p>Supplementary_Data_II.pdf comprises a large table that encompasses 10,691 pages, including the top-100 molecular fragments responsible for the spectral bands between 2.761 and 1172.745 &mu;m. Navigation through the hyperlinks enables access to this table.</p> <p>Supplementary_Data_III.csv&nbsp;encompasses the chemical formulas, number of unpaired valence electrons, spin multiplicities, xyz data, and SMILES strings of the PAHs carrying the additional spectra.</p> <p>Supplementary_data_IV.zip&nbsp;includes the input and output datasets along with the ML code. The code script is written in Python 3.7, and is supported by the following libraries: sklearn, json, numpy, and pandas.</p> <p>Supplementary_Information.pdf contains the evidence supporting the choice of the cutoff radius, as well as the figures of the count of the molecules in the dataset, the FI with changing datasets and hyperparameters, of cross-validation, and of UIE bands and emission features of four SH PAHs. Importance of three&nbsp;carbon skeleton fragments for&nbsp;bands in different&nbsp;intervals is also demonstrated.</p>

opencc-by-4.0Mar 2023View details →
ClinicalTrials.gov36/100

Predicting Pathological Complete Response in Esophageal Squamous Cell Carcinoma Using a Multimodal Model Integrating Clinical, Radiomics, and Deep Learning Features

ClinicalTrials.gov study NCT07181850. IPD Sharing: NO. Countries: 1. Publications: 4.

closedIPD-NOFeb 2026View details →
dryad36/100

Data from: Machine learning identification of microhabitat features associated with occupancy of artificial nestboxes by hazel dormice (Muscardinus avellanarius) in a UK woodland site

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad36/100

ECG and EEG stress features for: ECG and EEG based detection and multilevel classification of stress using machine learning for specified genders: A preliminary study

Open the record for dataset details and reuse information.

publicMar 2023View details →
dryad36/100

Machine learning feature data from EHR, labels, and estimates for next generation sequencing-based assay

Open the record for dataset details and reuse information.

publicNov 2024View details →
zenodo32/100

Prediction of cardiovascular diseases by integrating multi-modal features with machine learning methods

<p>Electrocardiogram (ECG) and Phonocardiogram (PCG) play important roles in early prevention and diagnosis of cardiovascular diseases. As the development of machine learning technique, detection of cardiovascular diseases from ECG and PCG has been attracted much attention.&nbsp;However, current available methods are mostly based on single data resource. It is desirable to develop efficient multi-modal machine learning methods to predict and diagnose cardiovascular diseases. In this study, we propose a novel multi-modal method for predicting cardiovascular diseases based on ECG and PCG features. By building up conventional neural networks, we extract ECG and PCG deep coding features respectively. The genetic algorithm is used to screen the combined features and obtain the best feature subset. Then support vector machine makes classification decision. Experimental results show that compared with using single-modal features ECG and PCG, the performance of this method reaches an AUC value of 0.936 when using multi-modal data resources.</p> <p>This&nbsp;dataset is&nbsp;developed&nbsp;from a real-world dataset which was assembled by PhysioNet/CinC Challenge&nbsp;in 2016. The original dataset can be downloaded from website (<a href="http://www.physionet.org/challenge/2016/">http://www.physionet.org/challenge/2016/</a>).</p>

opencc-by-4.0Nov 2020View details →
zenodo32/100

Prediction of femoral osteoporosis using machine-learning analysis with radiomics features and abdomen-pelvic CT: A retrospective single center preliminary study

<p>Dataset of prediction of osteoporosis using APCT, radiomics and machine learning analysis.</p>

opencc-by-4.0Jan 2021View details →
zenodo32/100

Multilingual bottle-neck feature learning from untranscribed speech for track 1 in zerospeech2017 (system 2 -- with VTLN)

<p>We investigate the extraction of bottle-neck features (BNFs) for multiple languages without access to manual transcription. Multilingual BNFs are derived from a multi-task learning deep neural network which is trained with unsupervised phoneme-like labels. The unsupervised phoneme-like labels are obtained from language-dependent Dirichlet process Gaussian mixture models separately trained on untranscribed speech of multiple languages.</p> <blockquote> <p>In this version, the input MFCC for DPGMM is processed with VTLN.</p> </blockquote> <p> </p>

opencc-by-4.0Jul 2017View details →
zenodo32/100

Best learned models : Land Cover Classification with Gaussian Processes using spatio-spectro-temporal features

<p>Best learned models (Gaussian Processes, Random Forest, Multilayer Perceptron and Lightweight Temporal Self-Attention models) for each region based on the classification data set DS-A with seed 0 (see description <a href="https://zenodo.org/deposit/7099785">here</a>).&nbsp;</p><p>Models: GP non spatial, GP spatial (sum), GP spatial (product),RF non spatial, RF spatial, MLP non spatial, MLP spatial, LTAE non spatial, LTAE spatial</p><p>For further details see section VI-C of the pre-print article "Land Cover Classification with Gaussian Processes using spatio-spectro-temporal features ". This article is available <a href="https://hal.archives-ouvertes.fr/hal-03781332">here</a>.</p><p>The implementation of the models is available in the <a href="https://gitlab.cesbio.omp.eu/belletv/land_cover_southfrance_gp">open source repository</a>.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Supplementary Information: CHAPTER 3 - Classification of genomic features of plant-associated bacteria using machine learning

<p>Appendix A-&nbsp; List of all bacterial genomes used in orthologous&nbsp; genes clustering in the feature extraction step and in the further steps to build and test classifiers&rsquo; models. The list includes the isolation source information and the&nbsp; related category for the genome classification and features selection purposes.</p> <p>Appendix B - Distribution of genomes by phylum, family, and genus&nbsp; among the categories defined according to bacteria lifestyle association.</p> <p>Appendix C - Enriched orthogroups by genus according to each enrichment test (Material and Methods). Values for each test are "Y" (enriched), "N" (not enriched), or "Untested" (clusters were untested when there was insufficient phylogenetic signal, they were too small or were found in all genomes).</p> <p>Appendix D - Classification performance of&nbsp; random forest and logistic regression techniques applied to genus-specific datasets of genomic features (orthogroups) using both matrices from gene count number and presence/absence values. Sensitivity is a measure of how well a test identifies true positives; Specificity: is a measure how well a test or model avoids false positives; Positive Predictive Value (Pos. Pred. Value): The probability that a positive prediction is correct; Negative Predictive Value (Neg. Pred. Value): The probability that a negative prediction is correct; Precision: The accuracy of positive predictions; Recall (Sensitivity): The ability to find all relevant cases; F1 Score: A combined measure of precision and recall; Prevalence: The proportion of positive cases in the total; Detection Rate: The proportion of true positive cases identified; Detection Prevalence: The proportion of positive predictions; Balanced Accuracy: An average of sensitivity and specificity;&nbsp; Area Under the Curve (AUC): The overall performance of the model in distinguishing between positive and negative cases.</p> <p>Appendix E - Orthogroups assigned with predicted COGs as an important feature for classifying plant-associated genomes. COG categories: A - RNA processing and modification; B - Chromatin structure and dynamics; C - Energy production and conversion; D - Cell cycle control, cell division, chromosome partitioning; E - Amino acid transport and metabolism; F - Nucleotide transport and metabolism; G - Carbohydrate transport and metabolism; H - Coenzyme transport and metabolism; I - Lipid transport and metabolism; J - Translation, ribosomal structure and biogenesis; K - Transcription; L - Replication, recombination and repair; M - Cell wall/membrane/envelope biogenesis; N - Cell motility; O - Posttranslational modification, protein turnover, chaperones; P - Inorganic ion transport and metabolism; Q - Secondary metabolites biosynthesis, transport and catabolism; R - General function prediction only; S - Function unknown; T - Signal transduction mechanisms; U - Intracellular trafficking, secretion, and vesicular transport; V - Defense mechanisms; W - Extracellular structures; X - Mobilome: prophages, transposons; Y - Nuclear structure; Z - Cytoskeleton.</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

dataset for "Feature distribution learning by passive exposure"

<p>Dataset for &quot;Feature distribution learning by passive exposure&quot;.</p> <p>The dataset contains .csv files for the two experiments described in the manuscript and an example R script to read the data and perform model comparison.</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

Machine Learning Constructs Color Features to Accelerate Development of Long-Term Continuous Water Quality Monitoring

<p>This is a machine learning method for predicting the concentration of colored pollutants based on RGB and kmeans methods. This dataset includes raw images of pollutants as well as characteristic data of pollutants, as well as code for the model. You can see the contents of the zip file for details.</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

Blinded Predictions and Post-hoc Analysis of the Second Solubility Challenge Data: Exploring Training Data and Feature Set Selection for Machine and Deep Learning Models

<p>Training and test datasets and scripts for training models.</p>

openmit-licenseSep 2022View details →
zenodo32/100

Gaussian approximation of dispersion potentials for efficient featurization and machine-learning predictions of metal–organic frameworks

<p>Scripts and data for the publication</p>

opencc-by-4.0May 2022View details →
zenodo32/100

AggMapNet: Enhanced and Explainable Low-Sample Omics Deep Learning with Feature-Aggregated Multi-Channel Networks

<p>This data contains the datasets used in the paper &quot;AggMapNet: Enhanced and Explainable Low-Sample Omics Deep Learning with Feature-Aggregated Multi-Channel Networks&quot;, each folder is named by the dataset name in the paper</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record