Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.7.1
Dataset results
6 results for “domain shift”
Crop classification dataset for testing domain adaptation or distributional shift methods
<p>In this upload we share processed crop type datasets from both France and Kenya. These datasets can be helpful for testing and comparing various domain adaptation methods. The datasets are processed, used, and described in this paper: <a href="https://doi.org/10.1016/j.rse.2021.112488">https://doi.org/10.1016/j.rse.2021.112488</a> (arXiv version: <a href="https://arxiv.org/pdf/2109.01246.pdf">https://arxiv.org/pdf/2109.01246.pdf</a>). </p> <p>In summary, each point in the uploaded datasets corresponds to a particular location. The label is the crop type grown at that location in 2017. The 70 processed features are based on Sentinel-2 satellite measurements at that location in 2017. The points in the France dataset come from 11 different departments (regions) in Occitanie, France, and the points in the Kenya dataset come from 3 different regions in Western Province, Kenya. Within each dataset there are notable shifts in the distribution of the labels and in the distribution of the features between regions. Therefore, these datasets can be helpful for testing for testing and comparing methods that are designed to address such distributional shifts.</p> <p>More details on the dataset and processing steps can be found in <a href="https://doi.org/10.1016/j.rse.2021.112488">Kluger et. al. (2021)</a>. Much of the processing steps were taken to deal with Sentinel-2 measurements that were corrupted by cloud cover. For users interested in the raw multi-spectral time series data and dealing with cloud cover issues on their own (rather than using the 70 processed features provided here), the raw dataset from Kenya can be found in <a href="https://openreview.net/forum?id=5HR3vCylqD">Yeh et. al. (2021)</a>, and the raw dataset from France can be made available upon request from the authors of this Zenodo upload.</p> <p>All of the data uploaded here can be found in "CropTypeDatasetProcessed.RData". We also post the dataframes and tables within that .RData file as separate .csv files for users who do not have R. The contents of each R object (or .csv file) is described in the file "Metadata.rtf".</p> <p><strong>Preferred Citation:</strong></p> <p>-Kluger, D.M., Wang, S., Lobell, D.B., 2021. Two shifts for crop mapping: Leveraging aggregate crop statistics to improve satellite-based maps in new regions. Remote Sens. Environ. 262, 112488. https://doi.org/10.1016/j.rse.2021.112488.</p> <p>-URL to this Zenodo post https://zenodo.org/record/6376160</p>
IEEE ICME 2024 Grand Challenge: Semi-supervised Acoustic Scene Classification under Domain Shift Evaluation Dataset
<p>The Chinese Acoustic Scene (CAS) 2023 dataset is a large-scale dataset that serves as a foundation for research related to environmental acoustic scenes. The dataset includes 10 common acoustic scenes, with a total duration of over 130 hours. Each audio clip is 10 seconds long with metadata about the recording location and timestamp. The dataset was collected by members of the <em>Joint Laboratory of Environmental Sound Sensing at the School of Marine Science and Technology, Northwestern Polytechnical University</em>. The data collection period spanned from April 2023 to September 2023, covering 22 different cities across China. The CAS 2023 dataset was collected using the XS-SN-2BE1 manufactured by <em>Xi'an Lianfeng Acoustic Technologies Co., Ltd</em> (https://www.lfxstek.com/). </p> <p>The ICME 2024 <em>Semi-supervised Acoustic Scene Classification under Domain Shift</em> challenge (https://2024.ieeeicme.org/grand-challenge-proposals/, https://ascchallenge.xshengyun.com/) dataset consists of development (https://zenodo.org/records/10616533) and evaluation datasets, all derived from the CAS 2023 dataset. The evaluation dataset includes 1,100 recordings, where data are selected from 12 cities, with 5 unseen cities specifically chosen to provide a more comprehensive evaluation of submissions under domain shift.</p> <p>Baseline: https://github.com/JishengBai/ICME2024ASC</p> <p>Acoustic scenes (10): Bus, Airport, Metro, Restaurant, Shopping mall, Public square, Urban park, Traffic street, Construction site, Bar</p>
ImageNet-Cartoon and ImageNet-Drawing: two domain shift datasets for ImageNet
<p>Benchmarking the robustness to distribution shifts traditionally relies on dataset collection which is typically laborious and expensive, in particular for datasets with a large number of classes like ImageNet. An exception to this procedure is ImageNet-C (Hendrycks & Dietterich, 2019), a dataset created by applying common real-world corruptions at different levels of intensity to the (clean) ImageNet images. Inspired by this work, we introduce ImageNet-Cartoon and ImageNet-Drawing, two datasets constructed by converting ImageNet images into cartoons and colored pencil drawings, using a GAN framework (Wang & Yu, 2020) and simple image processing (Lu et al., 2012), respectively.</p> <p>This repository contains ImageNet-Cartoon and ImageNet-Drawing. Checkout the <a href="https://github.com/oberman-lab/imagenet-shift">official GitHub Repo</a> for the code on how to reproduce the datasets.</p> <p>If you find this useful in your research, please consider citing:</p> <p> @inproceedings{imagenetshift,<br> title={ImageNet-Cartoon and ImageNet-Drawing: two domain shift datasets for ImageNet},<br> author={Tiago Salvador and Adam M. Oberman},<br> booktitle={ICML Workshop on Shift happens: Crowdsourcing metrics and test datasets beyond ImageNet.},<br> year={2022}<br> }</p>
GradDA – A novel dataset for investigating domain shifts in image classification
<p>A domain shift occurs when the testing data is drawn from a distribution different from that of the training dataset. This shift presents a significant challenge and may compromise the performance of machine learning models, which leads to poor generalization. Over the past years, various models have been developed and evaluated on benchmark datasets such as VisDA, Office-Home and DomainNet. These datasets consist of discrete domains with different object classes. However, a notable limitation when addressing the domain shift is the absence of data samples where the exact same object exists in both domains. </p> <p>We propose a new dataset designed to address this challenge. In particular, we introduce a domain shift from a purely synthetic style (grey object on white background) to a more realistic appearance (object with texture against a realistic background) with differential modifications, which enables the representation of the same object in both synthetic and real domains, consequently facilitating the analysis of a transition between the two domains. The dataset comprises five distinct classes (Airplane, Bicycle, Bus, Car, Train), with multiple objects per class. Additionally, each object is depicted from 20 different perspectives, resulting in a total of 101 images per perspective that captures the transition from pure synthetic to a more real-world-like domain. This dataset offers a unique opportunity to investigate the impact of domain shift on model performance in classification tasks, as it focuses solely on domain changes without other interfering effects. It is the objective of our work to trigger new discussions about the domain shift problem, and how it can be tackled with alternative data driven model designs.</p>
IMAD-DS: A Dataset for Industrial Multi-Sensor Anomaly Detection Under Domain Shift Conditions
<p>IMAD-DS is a dataset developed for multi-rate multi-sensor anomaly detection (AD) in industrial environments, that considers varying operational and environmental conditions known as domain shifts.</p> <p><strong>Dataset Overview:</strong></p> <p>This dataset includes data from two scaled industrial machines: a robotic arm and a brushless motor.</p> <p>It includes both normal and abnormal data recorded under various operating conditions to account for domain shifts. These shifts are categorized into:</p> <p>Robotic Arm: The robotic arm is a scaled version of a robotic arm used to move silicon wafers in a factory. Anomalies are created by removing bolts at the nodes of the arm, resulting in an imbalance in the machine.<br>Brushless Motor: The brushless motor is a scaled representation of an industrial brushless motor. Two anomalies are introduced: first, a magnet is moved closer to the motor load, causing oscillations by interacting with two symmetrical magnets on the load; second, a belt that rotates in unison with the motor shaft is tightened, creating mechanical stress.</p> <p>The following domain shifts are included in the dataset:</p> <p>Operational Domain Shifts: Variations caused by changes in machine conditions (e.g., load changes for the robotic arm and speed changes for the brushless motor).</p> <p>Environmental Domain Shifts: Variations due to changes in background noise levels.</p> <p>Combinations of operating and environmental conditions divide each machine's dataset into two subsets: the <em>source domain</em> and the <em>target domain</em>. The source domain has a large number of training examples. The target domain, instead, has limited training data. This discrepancy highlights a common issue in the industry where sufficient training data is often unavailable for the target domain, as machine data is collected under controlled environments that do not fully represent the deployment environments.</p> <p> </p> <p><strong>Data Collection and Processing:</strong></p> <p>Data is collected using the STEVAL-STWINBX1 IoT Sensor Industrial Node. The sensor used to record the dataset are the following.</p> <p>· Analog Microphone (16 kHz)</p> <p>· 3-axis Accelerometer (6.7 kHz)</p> <p>· 3-axis Gyroscope (6.7 kHz)</p> <p>Recordings are conducted in an anechoic chamber to control acoustic conditions precisely</p> <p><strong>Data Format:</strong><strong><br></strong>Files are already divided into train and test sets. Inside each folder, each sensor's data is stored in a separate '.parquet' file.</p> <p>Sensor files related to the <em>same</em> segment of machine data share a unique ID. The mapping of each machine data segment to the sensor files is given in .csv files inside the train and test folders. Those .csv files also contain metadata denoting the operational and environmental conditions of a specific segment.</p> <p> </p> <p> </p> <p> </p>
Maize Root Domain Shift Image Datasets, Segmentation Models, and RhizoVision Explorer Settings
<p>This .zip file contains the following: 1) a .csv file used for RhizoVision Explorer settings, 2) root segmentation .pkl models developed using RootPainter, and 3) folders of maize root images collected in the field and greenhouse experiment that were used to train models and that were manually annotated to generate ground-truth datasets.</p> <p>For clarification, the segmentation model file named "000053_1679678936_V7_2021_tiled.pkl" is the growth stage-specific V7 field model. The segmentation model file named "000040_1679498247_R2_2021_tiled.pkl" is the growth stage-specific R2 field model. The segmentation model file named "000022_1679880442_V7plusR2_2021_tiled.pkl" is the fine-tuned V7+R2 field model. The segmentation model file named "000059_1687007135_V7R2_Tiled_Combined_2021.pkl" is the combined V7+R2 field model. The segmentation model file named "000068_1687361521_V7extendedR2_tiled_2021.pkl" is the extended V7+R2 field model. The segmentation model file named "000027_1681159536_GH_R2_tiled.pkl" the R2 greenhouse model. The segmentation model file named "000059_1680366870_GH_V9_tiled.pkl" is the V9 greenhouse model.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.