Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
598
datasets available to search
ShareScore release 0.9.0
Dataset results
598 results for “classifier”
Deep learning based on convolutional neural networks to classify nanobiomechanical data
Open the record for dataset details and reuse information.
Cane Toad Acoustic Classifier Audio Training Data
<h3><strong>Cane Toad Audio Dataset for Machine Learning Classifier Development</strong></h3> <h3><strong>Description:</strong></h3> <p>This dataset was created as part of a study aimed at developing a machine learning classifier to detect the advertisement calls of the cane toad (<em>Rhinella marina</em>) using BirdNET. The dataset comprises 3-second audio snippets that capture a variety of sounds, including cane toad vocalizations, calls from spectrally overlapping species, environmental noises, and unidentified sounds. These labelled sound data were collected from various Australian Acoustic Observatory' recording sites, covering a broad range of geographic locations and environmental conditions in Australia.</p> <h3><strong>Sound Classes:</strong></h3> <ul> <li>Background</li> <li>Canis lupus dingo (Dingo)</li> <li>Centropus phasianinus (Pheasant Coucal)</li> <li>Cyclorana australis (Water Holding Frog)</li> <li>Cyclorana cryptotis (Hidden Ear Frog)</li> <li>Cyclorana novaehollandiae (New Holland Frog)</li> <li>Dacelo novaeguineae (Kookaburra)</li> <li>Ninox boobook (Southern Boobook)</li> <li>Notaden melanoscaphus (Northern Spadefoot Toad)</li> <li>Rhinella marina (Cane Toad) – Processed with a low-pass filter to remove frequencies above 1300 Hz, minimizing interference from co-occurring sounds.</li> <li>Unidentified Sounds</li> </ul> <h3><strong>Use and Applications:</strong></h3> <p>This dataset is valuable for training machine learning models focused on the acoustic detection of cane toads, could be useful for researchers and professionals working in bioacoustics, machine learning, ecological monitoring and invasive species management.</p>
Clusters of topic modelling and naive Bayes classifier - FR CH newspapers
<p>Clusters of articles based on annotations produced by topic modelling and naive Bayes classifier applied to French language newspapers of Switzerland, published between 1900 and 1944 and containing the characters "europ", extracted from the impresso app. </p>
Mapping of glacial lakes using Sentinel-1 and Sentinel-2 data and a random forest classifier: Strengths and challenges
<p>The water body detection and mapping algorithm named 'glakemap' that I designed was aimed at specifically mapping glacial lakes across alpine regions where their detection and mapping are challenged by many factors such as shadows, cloud cover, turbidity, and ice surface. The algorithm uses Copernicus Sentinel-1 and -2 satellites data and machine learning model (random forest) in an integrated manner to automatically classify glacial lakes from other surface features. In specific, the algorithm takes Sentinel-1 and -2 satellites data as the main inputs. It calculates radar backscatter and Normalised Difference Water Indices (NDWIs) using these datasets, respectively. The radar backscatter and NDWIs products (images) are segmented using a set of rules producing many polygons including lake polygons. Lake polygons are then automatically separated/retained using the random forest model which is trained using features relevant to lakes.</p> <p>The dataset is also available at https://www.mountcryo.org/</p>
Research data supporting: "A Data-Driven Dimensionality Reduction Approach to Compare and Classify Lipid Force Fields"
<p>This repository contains the data used in the paper of Capelli <em>et al. </em>"A Data-Driven Dimensionality Reduction Approach to Compare and Classify Lipid Force Fields", published on Journal of Physical Chemistry B (DOI: 0.1021/acs.jpcb.1c02503).<br> <br> The archive traj_processed.tar.gz contains the trajectories converted in xyz format with the dimensions of the box.</p> <p>The archive trajectories_xtc.tar.gz contains the raw trajectories (of the membranes without solvent) in gromacs xtc format with a .tpr binary file. <br> </p>
Supplementary datasets for the paper of "Multi-resBind: a residual network-based multi-label classifier for in vivo RNA binding prediction and preference visualization"
<p>There are two eCLIP datasets (cell lines of K562 and HepG2). The eCLIP datasets were then divided into five categories for each cell line: low, medium 1, medium 2, high 1 and high 2 with peaks of >1,000 but <2,000, >2,000 but <4,000, >4,000 but <7,000, >7,000 but <10,000 and >10,000, respectively. </p>
Photorealistic and classified urban point cloud dataset (Helsinki)
<p>The dataset contains a photorealistic and classified urban point cloud from Kalasatama region in Helsinki, Finland. Data was created using both terrestrial laser scanning (Leica RTC360) and UAV-based (DJI P4 Pro+) photogrammetry. 3D reconstruction was completed with RealityCapture and point cloud classification with TerraScan. The dataset is georeferenced in ETRS-TM35FIN (EPSG 3067) coordinate system. Work was done in Aalto University (The Research Institute of Measuring and Modeling for the Built Environment) with support from the City of Helsinki.</p>
World Atlas of Classifier Languages
<p>Cite the source of the dataset as:</p> <blockquote> <p>Her, One-Soon, Harald Hammarström and Marc Allassonnière-Tang. 2022. Defining numeral classifiers and identifying classifier languages of the world. Linguistics Vanguard. https://doi.org/10.1515/lingvan-2022-0006</p> </blockquote>
Image to Classify type of SSW in MERRA (1979-2005)
<p>Imágenes de altura geopotencial a 10 hPa durante los SSW en el hemisferio Norte de 1979-2005 de MERRA. Usadas para clasificar el tipo de evento (División o desplazamiento de vórtice).</p>
Grey Level Co-Occurrence Matrix and Learning Algorithms to Quantify and Classify Use-Wear on Experimental Flint Tools
<p>These are the images and the grey levels co-occurence matrix analysis results from my experimental dataset</p>
Multiplex imaging of breast cancer lymph node metastases identifies prognostic single-cell populations independent of clinical classifiers
<p>This repository contains the raw IMC data of ZTMA 26 as continuation of dataset <strong>10.5281/zenodo.7494413.</strong> The zip files starting with ZTMA contain the raw IMC measurements (mcd and txt) of those parts of the TMA. The TMA measurements are split up into parts in order to avoid huge files.</p> <p>Additionally, this repository contains the metadata of the patients analyzed in this study, the panel information and the single-cell data that was extracted from the multiplexed images together with the associated metadata in SingleCellExperiment format for analysis in R.</p> <p>The analysis.zip folder contains files that were written out during the analysis according to the scripts in https://github.com/BodenmillerGroup/BC_LN_metastses.</p> <p>The single-cell and other data outputs from CellProfiler can be found in the cpout.zip file.</p> <p>The IF_whole_sections.zip file contains the IF images of the primary breast cancer sections (czi files) and the extracted single-cell data.</p>
Multiplex imaging of breast cancer lymph node metastases identifies prognostic single-cell populations independent of clinical classifiers
<p>This repository contains the raw IMC data of ZTMA 21 and 25 of the matched primary breast cancer and lymph node metastasis study presented in Fischer and Jackson et al., 2023. The code that was used to process and analyze this data can be found at https://github.com/BodenmillerGroup/BC_LN_metastses.</p> <p>The zip files starting with ZTMA contain the raw IMC measurements (mcd and txt) of the respective parts of the TMA. The TMA measurements are split up into parts in order to avoid huge files.</p>
Online supplement to manuscript: "Ability of ChatGPT to generate competent radiology reports for distal radius fracture by use of RSNA template items and integrated AO classifier." Current problems in diagnostic radiology (2023).
<p>Online supplement to manuscript: </p> <p>Bosbach, Wolfram A., Jan F. Senge, Bence Nemeth, Siti H. Omar, Milena Mitrakovic, Claus Beisbart, András Horváth, Johannes Heverhagen, and Keivan Daneshvar. "Ability of ChatGPT to generate competent radiology reports for distal radius fracture by use of RSNA template items and integrated AO classifier." <em>Current problems in diagnostic radiology</em> (2023). <a href="https://doi.org/10.1067/j.cpradiol.2023.04.001">doi.org/10.1067/j.cpradiol.2023.04.001</a></p>
Training data for the Singapore Butterfly Classifier used in the XPRIZE Rainforest semifinals
<p>Training data based on Lepidoptera observation data from Singapore with images attached from the Global Biodiversity Index Facility (GBIF) (de Vries H, Lemmens M (2023). Observation.org, Nature data from around the World. Observation.org. Occurrence dataset https://doi.org/10.15468/5nilie accessed via GBIF.org on 2023-04-17; Nature data from around the World, Observation.org, https://doi.org/10.15468/5nilie accessed via GBIF.org on 2023-04-17; iNaturalist Research-grade Observations, iNaturalist.org, https://doi.org/10.15468/ab3s5x; Earth Guardians Weekly Feed, Occurrence dataset <a href="https://doi.org/10.15468/slqqt8">https://doi.org/10.15468/slqqt8</a>)</p> <p>Annotations were produced with a generic butterfly detector trained on annotations available at https://www.kaggle.com/datasets/mistag/arthropod-taxonomy-orders-object-detection-dataset, accessed on 2023-05-08). Annotations were produced with the generic butterfly detector while labels assigned using the specied ID assigned in the GBIF observation.</p> <p>Training data contains 230 species. These were the ones that contained species ID in relevant columns, contained 50+ images, and received predictions with the generic butterfly detector.</p> <p> </p> <p>Dataset: Observation.org, Nature data from around the World <br> Rights as supplied: http://creativecommons.org/licenses/by-nc/4.0/legalcode<br> Dataset: iNaturalist Research-grade Observations <br> Rights as supplied: http://creativecommons.org/licenses/by-nc/4.0/legalcode<br> Dataset: Earth Guardians Weekly Feed <br> Rights as supplied: http://creativecommons.org/licenses/by-nc/4.0/legalcode</p> <p> </p> <p> </p>
Data from "Benchmark Generation Framework with Customizable Distortions for Image Classifier Robustness"
<p>This repository contains the data from the paper, "Benchmark Generation Framework with Customizable Distortions for Image Classifier Robustness." </p> <p>Relevant URLs:</p> <p>https://hewlettpackard.github.io/trust-ml/</p> <p>https://github.com/HewlettPackard/trust-ml/</p> <p> </p> <p>Abstract:</p> <p>We present a novel framework for generating adversarial benchmarks to evaluate the robustness of image classification models. The RLAB framework allows users to customize the types of distortions to be optimally applied to images, which helps address the specific distortions relevant to their deployment. The benchmark can generate datasets at various distortion levels to assess the robustness of different image classifiers. Our results show that the adversarial samples generated by our framework with any of the image classification models, like ResNet-50, Inception-V3, and VGG-16, are effective and transferable to other models causing them to fail. These failures happen even when these models are adversarially retrained using state-of-the-art techniques, demonstrating the generalizability of our adversarial samples. Our framework also allows the creation of adversarial samples for non-ground truth classes at different levels of intensity, enabling tunable benchmarks for the evaluation of false positives. We achieve competitive performance in terms of net $L_2$ distortion compared to state-of-the-art benchmark techniques on CIFAR-10 and ImageNet; however, we demonstrate our framework achieves such results with simple distortions like Gaussian noise without introducing unnatural artifacts or color bleeds. This is made possible by a model-based reinforcement learning (RL) agent and a technique that reduces a deep tree search of the image for model sensitivity to perturbations, to a one-level analysis and action. The flexibility of choosing distortions and setting classification probability thresholds for multiple classes makes our framework suitable for algorithmic audits.</p>
Online classified adverts reflect the broader United Kingdom trade in turtles and tortoises rather than drive it.
<p>The accompanying file represents a flat dataset as used in the article: "Online classified adverts reflect the broader United Kingdom trade in turtles and tortoises rather than drive it."</p> <p>Any potentially personal or sensitive information was removed at the collection and processing stages and seller names or IDs (as listed on the original adverts) have been anonymised.<br> </p>
Dataset for: AutoSourceID-Classifier
<p>This is the dataset used in the paper AutoSourceID-Classifier for Training, Test, Validation and Calibration. </p> <p>This repository contains:</p> <ul> <li>the folders with Training, Test, Validation and Calibration images and labels used for the analysis</li> <li>the catalogue with information on all the individual sources used in this paper</li> </ul> <p>If you use the dataset, please cite the appropriate DOIs.</p> <p>Code available at <a href="https://github.com/FiorenSt/AutoSourceID-Classifier">https://github.com/FiorenSt/AutoSourceID-Classifier</a> </p>
Endmember spectra and classified mosaics derived from CRISM mapping data at the south pole of Mars
<p><strong>Overview</strong></p> <p>Multispectral mapping data from the Compact Reconnaissance Imaging Spectrometer for Mars (CRISM) provide a unique opportunity to characterize south polar ice deposits at higher spectral sampling, spatial resolution, or spatiotemporal coverage than previous work. This new perspective can help to constrain the nature and distribution of different mixtures of CO<sub>2</sub> ice, H<sub>2</sub>O ice, and dust that influence the formation, evolution, and preservation of Mars climate records. We processed 1103 CRISM observations spanning southern summer of six Mars Years through a combination of <em>k</em>-means clustering and random forest classification. Using a set of 12 spectral endmembers directly tied to previous work with high-resolution CRISM targeted data, we made a series of temporally restricted mosaics showing surface spectral variation over time. The endmember set and classified mosaics produced in this work can provide critical context for future studies of the dynamic processes that shape south polar ice deposits.</p> <p>This is follow-on work to a previous study that produced a series of classified maps and a spectral library from CRISM targeted data. That work can be accessed with the following links:</p> <ul> <li>Publication: <a href="https://doi.org/10.1029/2022JE007372">https://doi.org/10.1029/2022JE007372</a></li> <li>Repository: <a href="https://doi.org/10.5281/zenodo.6960943">https://doi.org/10.5281/zenodo.6960943</a></li> </ul> <p><strong>Contents</strong></p> <p>The repository contains two .zip files, which can be expanded to access the files described below:</p> <ul> <li><strong>classified_mosaics.zip</strong> <ul> <li>lookup_files <ul> <li><em>SP_CRISM_RF_ColorMap.clr</em> : An ESRI-formatted color map file that can be used to apply the endmember color scheme to random forest-classified maps in ArcGIS (see Symbology settings). Note that for continuity with <em>Cartwright (2022),</em> there are 21 colors in this file, though only endmembers associated with 12 of those colors are present in the classified mosaics.</li> <li><em>SP_CRISM_RF_ColorMap.txt</em> : A text file that can be loaded into Python as a numpy array and used to generate a matplotlib color ramp. Row indices correspond to endmember numbers (see below) and columns are red, green, and blue values scaled from 0 to 1. Note that for continuity with <em>Cartwright (2022),</em> there are 21 colors in this file, though only endmembers associated with 12 of those colors are present in the classified mosaics.</li> <li><em>SP_CRISM_RF_Mapping_Endmember_Lookup.csv</em> : A lookup table that can be used to cross-reference between endmember numbers (as stored in GeoTiffs or the indices of the colormaps above) and corresponding endmember names (C1m, Dc3m, etc.). Note that for continuity with <em>Cartwright (2022),</em> the endmember numbers are not sequential and instead correspond to the number of the original reference spectrum (e.g., C1m in this work has the same endmember number as C1 in <em>Cartwright (2022)</em>).</li> </ul> </li> <li>random_forest_classification <ul> <li>Contains GeoTiff mosaics of random forest classification results. Mosaics compile all observations falling in 10º bins of solar longitude (Ls) for a given year, provided that data was acquired in that range. Classified observations are stacked in ascending order of Ls. Mosaics of MSW data have a resolution of ~90 m/pixel while mosaics of MSP or combined MSP and MSW data have a resolution of ~180 m/pixel. Filenames indicate the observation type ("MSP", "MSW", or combined "MSP-MSW"), Mars Year ("MY28", "MY29", etc.) and Ls range ("Ls300-310" indicates all observations acquired between Ls 300º and Ls 310º). These images are not rendered to display colors consistent with the publication figures, but instead store the endmember classification for each pixel as a number from 0 to 21; use the contents of <em>lookup_files</em> to find the corresponding endmember name or render the image with the color scheme from the publication. To view rendered summary plots of each observation, see the contents of <em>summary_plots </em>below.</li> </ul> </li> <li>summary_plots <ul> <li>Contains summary plots of the endmember-classified mosaics provided in <em>random_forest_classification </em>above<em>.</em> Each plot presents the classified mosaic over a grey outline approximating the extent of high-albedo CO<sub>2</sub> ice in the south polar residual cap (SPRC). Note that the summary view focuses on the area in and around the SPRC, but the full mosaics may extend beyond these bounds. </li> </ul> </li> </ul> </li> <li><strong>spectral_library.zip</strong> <ul> <li>Contains spectral libraries with the median spectrum of each endmember as presented in Figure 3 in the publication. A basic text file listing the wavelength (WVL) and normalized reflectance values for each endmember is included (<em>SP_CRISM_MappingEndmemberMedians.txt</em>) as well as an ENVI-formatted spectral library (S<em>P_CRISM_MappingEndmemberMedians_ENVI.sli</em>). Note that a subset of the 55 wavelengths sampled in the source CRISM data were removed to avoid error-prone bands, leaving these spectral libraries with 51 bands. The spectra are normalized by the value at 1.330 µm.</li> </ul> </li> </ul>
Data files and taxonomic classifiers for Pseudomonas syringae classification and virulence factor prediction
<p>PSSC.tree : core-genome tree of 2,161 high quality <em>Pseudomonas syringae</em> genomes</p> <p>metadata.csv: A CSV file containing taxonomic data, type strain designations, phylogroups as assigned in this study, LIN clusters assigned for classification purposes, presence/absence of key virulence factors, and metadata found in each genome’s Biosample record for all genomes found in PSSC.tree</p> <p>CLASSIFIER_xxx: QIIME 2 classifier artifacts trained on amplicons generated from primer sets indicated in file name</p> <p>xxx_VFOC.JSON: HMMER results for T3SS and effectors and WHOP genes, structured with both genome and gene product accession numbers as primary key, depending on file</p> <p> </p> <p> </p>
Developing Deep Learning Approaches to Find and Classify Architectural Design Decisions in Issue Tracking Systems
<p>This upload contains three files:</p> <ol> <li>mongodump-JiraRepos_2023-03-07-16 00.archive: Archive containing the issue data pulled from the Jira API.</li> <li>mongodump-MiningDesignDecisions.archive: Archive containing the data of our deep learning models and the labelled issues.</li> <li>mongodump-MiningDesignDecisions-lite.archive: Similar to the archive above, except this one only contains the best trained model (BERT). Also, it does not contain any embeddings or other files.</li> </ol> <p>This archive contains the data of our deep learning models and the labelled issues (MiningDesignDecisions archive).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.