Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
345
datasets available to search
ShareScore release 0.9.0
Dataset results
345 results for “classification analysis”
Animal Recognition Using Methods Of Fine-Grained Visual Analysis - YOLOv5 Breed Classification Dataset (Tsinghua Dogs)
<p>Tsinghua Dogs Dataset with ground truth labels for breeds in YOLOv5 format.</p>
Datasets corresponding to "Real-time intelligent classification of COVID-19 and thrombosis via massive image-based analysis of platelet aggregates"
<p>Datasets corresponding to "Real-time intelligent classification of COVID-19 and thrombosis via massive image-based analysis of platelet aggregates"</p> <p> </p> <p>Please find below an explanation for the <strong>files </strong>in this repository:</p> <p><br> <br> <strong>DiseaseClassifPaper_Dataset_01.7z, DiseaseClassifPaper_Dataset_02.7z</strong></p> <p>Experimental data. To reproduce the analyses, unzip both files and put the content into a folder called "Dataset"</p> <p><strong>02_CNN_PhenotypeClassif.7z</strong></p> <p>CNN Phenotype classification. Model was trained using AIDeveloper. using manually labelled data. Labelled Data is contained in folder "03_GatedData". The AIDeveloper session file in "02_Model\M10_Nitta6l_32pix_8class_meta.xlsx" shows, which files correspond to which subpopulation. The final model "M10_Nitta6l_32pix_8class_448.model" and corresponding .pb files are also located in that folder.</p> <p><strong>03_ExampleMeasurement.zip</strong></p> <p>One measurement file and a corresponding scatterplot</p> <p><strong>04_Dataset_load.zip</strong></p> <p>The python script "03_ExtractFeatures.py" loads the list of available experiment files (01_Dataset_Table_v02.csv). The experiment files are contained in DiseaseClassifPaper_Dataset_01.7z, DiseaseClassifPaper_Dataset_02.7z. The scrip then evaluates each experiment file to obtain distribution parameters for Area and Solidity. These values are written to new "01_Dataset_Table_v03.csv".</p> <p><strong>05_RF_training</strong></p> <p>Scripts to train and evaluate the Random Forest model (using features contained in "01_Dataset_Table_v03.csv").</p> <p><strong>07_pytranskit</strong></p> <p>Scripts for training and evaluating CDT-PLDA classifier</p> <p> </p> <p> </p>
1H HRMAS NMR ERETIC-CPMG dataset for survival analysis and pathological classification of gliomas
<p>This repository contains the raw ERETIC-CPMG HRMAS NMR data to reproduce the results reported in the following preprint: PiDeeL: Pathway-informed deep learning model for survival analysis and pathological classification of gliomas</p> <p>The FID spectra can be found under the FID_Samples directory. The pathological classification and survival analysis labels can be found in the Dataset_Labels.xlsx file.</p>
Machine learning-based pulse wave analysis for classification of circle of Willis topology: an in silico study with 30,618 virtual subjects (database: Complete CoW)
<p>This repository contains the dataset for the complete CoW described in the article with the same name. MATLAB and Python codes for post-processing the dataset and the code for training and testing all machine learning models using the open-source library TensorFlow 2.12, the Keras application programming interface, and the Scikit-learn Python package can be found in here (<a href="https://zenodo.org/records/12519322" target="_blank" rel="noopener">https://zenodo.org/records/12519322</a>).</p>
Figure 2 in Revised classification design of the Anatolian species of Nannospalax (Rodentia: Spalacidae) using RFLP analysis
Figure 2. Principal coordinate analysis plot of the chromosomal races of Nannospalax.
Auditory Scene Analysis dataset (Multichannel universal sound separation & polyphonic audio classification)
<p>We constructed a new dataset for <strong>multichannel universal sound separation</strong> and <strong>polyphonic audio classification</strong> tasks.</p> <p>We constructed a new dataset for multichannel USS and polyphonic audio classification tasks. The proposed dataset is designed to reflect various conditions, including moving sources with temporal onsets and offsets. For foreground sound sources, signals from 13 audio classes were selected from open-source databases (Pixabay and FSD50K, Librispeech, MUSDB18, Vocalsound). These signals were resampled to 16 kHz and pre-processed by either padding zeros or cropping to 4 seconds. Each sound source has a 75% probability of being a moving source, with speeds ranging from 0 to 3 m/s. The dataset features between 2 to 4 foreground sound sources, along with one background noise from the diffused TAU-SNoise dataset with a signal-to-noise ratio (SNR) ranging from 6 to 30 dB. The simulations were conducted using gpuRIR. Room dimensions were set to a width and length between 5 and 8 meters, and a height between 3 and 4 meters, with reverberation times ranging from 0.2 to 0.6 seconds. These parameters were sampled from uniform distributions. We simulated spatialized sound sources using a 4-channel tetrahedral microphone array with a radius of 4.2 cm. The procedure for dataset generation and details about class configuration and durations of audio clips are provided in the paper. This dataset poses a significant challenge for separation tasks due to the inclusion of moving sources, onset and offset conditions, overlapped in-class sources, and noisy reverberant environments.</p> <p>The procedure for dataset generation and details about class configuration and durations of audio clips are provided in the paper. This dataset poses a significant challenge for separation tasks due to the inclusion of moving sources, onset and offset conditions, overlapped in-class sources, and noisy reverberant environments.</p>
DIVERSE: Deciphering Internet Views on the U.S. Military Through Video Comment Stance Analysis: A Novel Benchmark Dataset for Stance Classification
<p>Paper citation: Cruickshank, Iain J., and Lynnette Hui Xian Ng. "DIVERSE: Deciphering Internet Views on the US Military Through Video Comment Stance Analysis, A Novel Benchmark Dataset for Stance Classification." <em>arXiv preprint arXiv:2403.03334</em> (2024).</p> <p>Link to paper: https://arxiv.org/abs/2403.03334</p>
Additional files for Horvath et al., 2024. Detection and classification of long terminal repeat sequences in plant LTR-retrotransposons and their analysis using explainable machine learning.
<p>Additional data for Horvath et al., 2024 (source code freeze, models, data, supplementary figures, tables and files(.</p>
Data from: Integrating a UAV-derived DEM in object-based image analysis increases habitat classification accuracy on coral reefs
<p>Very shallow coral reefs (< 5 m deep) are naturally exposed to strong sea surface temperature variations, UV radiation and other stressors exacerbated by climate change, raising great concern over their future. As such, accurate and ecologically informative coral reef maps are fundamental for their management and conservation. Since traditional mapping and monitoring methods fall short in very shallow habitats, shallow reefs are increasingly mapped with Unmanned Aerial Vehicles (UAVs). UAV-imagery is commonly processed with Structure-from-Motion (SfM) to create orthomosaics and Digital Elevation Models (DEMs) spanning several hundred metres. Techniques to convert these SfM products to ecologically relevant habitat maps are still relatively underdeveloped. Here we demonstrate that incorporating geomorphometric variables (the DEM and its derivatives) in addition to spectral information (the orthomosaic) can greatly enhance the accuracy of automatic habitat classification. Therefore, we mapped three very shallow reef areas off KAUST on the Saudi Arabian Red Sea coast with an RTK-ready UAV. Imagery was processed with SfM, and classified through Object-Based Image Analysis (OBIA). Within our OBIA workflow, we observed overall accuracy increases of up to 11% when training a Random Forest classifier on both spectral and geomorphometric variables as opposed to traditional methods that only use spectral information. Our work highlights the potential of incorporating a UAV's DEM in OBIA for benthic habitat mapping, a promising but still scarcely exploited asset.</p>
Deep Neural Network for Stroke Patient Gait Analysis and Classification
ClinicalTrials.gov study NCT04968418. IPD Sharing: NO. Countries: 1. Publications: 3.
Data from: Integrating a UAV-derived DEM in object-based image analysis increases habitat classification accuracy on coral reefs
Open the record for dataset details and reuse information.
A revised classification of the assassin bugs (Hemiptera: Heteroptera: Reduviidae) based on combined analysis of phylogenomic and morphological data
Open the record for dataset details and reuse information.
Data from: Advancing Pyrus phylogeny: Deep genome skimming-based inference coupled with paralogy analysis yields a robust phylogenetic backbone and an updated infrageneric classification of the pear genus (Maleae, Rosaceae)
Open the record for dataset details and reuse information.
Datasets of the article "From Classification to Quantification in Tweet Sentiment Analysis"
<p>Datasets used for the following SNAM paper:<br> ---------------------------------------------------------------------------------------------------<br> Title: From Classification to Quantification in Tweet Sentiment Analysis<br> Authors: Wei Gao and Fabrizio Sebastiani<br> Organization: Qatar Computing Research Institute, Hamad Bin Khalifa University, Doha, Qatar<br> ---------------------------------------------------------------------------------------------------</p> <p>[Content]</p> <p>* SemEval2013, SemEval2014, SemEval2015 datasets:<br> - semeval.train.feature.txt: Training set for learning sentiment models at development stage<br> - semeval.dev.feature.txt: Held-out set for tuning parameters<br> - semeval.train+dev.feature.txt: Training set for learning the final sentiment model<br> - semeval13.test.feature.txt: SemEval2013 test set<br> - semeval14.test.feature.txt: SemEval2014 test set<br> - semeval15.test.feature.txt: SemEval2015 test set<br> <br> * Other datasets: semeval2016, sanders, sst, omd, hcr, gasp, wa, wb<br> - X.train.feature.txt: Training set for learning sentiment models at development stage<br> - X.dev.feature.txt: Held-out set for tuning parameters<br> - X.train+dev.feature.txt: Training set for learning the final sentiment model<br> - X.test.feature.txt (or X.dev-test.feature.txt for semeval2016 only): Test set<br> where X is one of semeval2016, sanders, sst, omd, hcr and gasp.</p> <p>* Training files are saved in ./data/train directory, and held-out and test files are in ./data/test directory</p> <p><br> For more details, please refer to the paper.</p> <p><br> [Citation]<br> You can cite the following paper when referring to the dataset:</p> <pre>@article{gao2016classification, title={From classification to quantification in tweet sentiment analysis}, author={Gao, Wei and Sebastiani, Fabrizio}, journal={Social Network Analysis and Mining}, volume={6}, number={1}, pages={19}, year={2016}, publisher={Springer} }</pre> <p> </p>
Data from: Survival analysis and classification methods for forest fire size
Factors affecting wildland-fire size distribution include weather, fuels, and fire suppression activities. We present a novel application of survival analysis to quantify the effects of these factors on a sample of sizes of lightning-caused fires from Alberta, Canada. Two events were observed for each fire: the size at initial assessment (by the first fire fighters to arrive at the scene) and the size at "being held" (a state when no further increase in size is expected). We developed a statistical classifier to try to predict cases where there will be a growth in fire size (i.e., the size at "being held" exceeds the size at initial assessment). Logistic regression was preferred over two alternative classifiers, with covariates consistent with similar past analyses. We conducted survival analysis on the group of fires exhibiting a size increase. A screening process selected three covariates: an index of fire weather at the day the fire started, the fuel type burning at initial assessment, and a factor for the type and capabilities of the method of initial attack. The Cox proportional hazards model performed better than three accelerated failure time alternatives. Both fire weather and fuel type were highly significant, with effects consistent with known fire behaviour. The effects of initial attack method were not statistically significant, but did suggest a reverse causality that could arise if fire management agencies were to dispatch resources based on a-priori assessment of fire growth potentials. We discuss how a more sophisticated analysis of larger data sets could produce unbiased estimates of fire suppression effect under such circumstances.
FIGURE 15 in Generic classification for the Gasteruptiinae (Hymenoptera: Gasteruptiidae) based on a cladistic analysis, with the description of two new Neotropical genera and the revalidation of Plutofoenus Kieffer
FIGURE 15. Geographic distribution of the species of Plutofoenus, Spinolafoenus and Trilobitofoenus.
FIGURE 9 in Generic classification for the Gasteruptiinae (Hymenoptera: Gasteruptiidae) based on a cladistic analysis, with the description of two new Neotropical genera and the revalidation of Plutofoenus Kieffer
FIGURE 9. First metasomal segment in ventral view. a: Gasteruption bispinosum; b: Plutofoenus edwardsi. S2: first metasomal sternum; T2: first metasomal tergum. Without scale. Numbers indicate characters and respective synapomorphic states (within brackets).
FIGURE 14 in Generic classification for the Gasteruptiinae (Hymenoptera: Gasteruptiidae) based on a cladistic analysis, with the description of two new Neotropical genera and the revalidation of Plutofoenus Kieffer
FIGURE 14. Propodeum and metacoxa in dorsal view. a: Gasteruption bispinosum; b–c: Plutofoenus edwardsi; d: Gasteruption sp. I; e: Spinolafoenus ruficornis; f: Trilobitofoenus plaumanni. Scale bar: 1.0 mm. Numbers indicate characters and respective synapomorphic states (within brackets).
FIGURE 10 in Generic classification for the Gasteruptiinae (Hymenoptera: Gasteruptiidae) based on a cladistic analysis, with the description of two new Neotropical genera and the revalidation of Plutofoenus Kieffer
FIGURE 10. Subgenital sternum in ventral view. a: notched, Y-shaped; b: notched, V-shaped; c: simple. Without scale. Numbers indicate characters and respective synapomorphic states (within brackets).
FIGURE 8 in Generic classification for the Gasteruptiinae (Hymenoptera: Gasteruptiidae) based on a cladistic analysis, with the description of two new Neotropical genera and the revalidation of Plutofoenus Kieffer
FIGURE 8. Propleuron in ventral view. a: Gasteruption bispinosum; b: Trilobitofoenus plaumanni. Without scale. Numbers indicate characters and respective synapomorphic states (within brackets).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.