Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
18
datasets available to search
ShareScore release 0.9.0
Dataset results
18 results for “domain classification”
Maximum likelihood classification of 2006 AISA hyperspectral imagery of the GCE domain for vegetation
Airborne Imaging Spectrometer for Applications (AISA) Eagle hyperspectral imagery were acquired on June 20-21, 2006, by the Center for Advanced Land Management Information Technologies (CALMIT). This included four flight lines flown for the examination of vegetation for the Duplin River salt marshes. Imagery was acquired for 63 bands from 400-980 nm at a 1 m spatial resolution. Imagery were classified using the maximum likelihood classifier (MLC) and a post-classification decision tree to achieve an overall classification accuracy of 90%. Classification training and validation data were obtained from the 2006 Hyperspectral ground survey. See Hladik (2012) and Hladik, Alber, and Schalles (2013) and Schalles, et. al. (2013) for additional details.
Crop classification dataset for testing domain adaptation or distributional shift methods
<p>In this upload we share processed crop type datasets from both France and Kenya. These datasets can be helpful for testing and comparing various domain adaptation methods. The datasets are processed, used, and described in this paper: <a href="https://doi.org/10.1016/j.rse.2021.112488">https://doi.org/10.1016/j.rse.2021.112488</a> (arXiv version: <a href="https://arxiv.org/pdf/2109.01246.pdf">https://arxiv.org/pdf/2109.01246.pdf</a>). </p> <p>In summary, each point in the uploaded datasets corresponds to a particular location. The label is the crop type grown at that location in 2017. The 70 processed features are based on Sentinel-2 satellite measurements at that location in 2017. The points in the France dataset come from 11 different departments (regions) in Occitanie, France, and the points in the Kenya dataset come from 3 different regions in Western Province, Kenya. Within each dataset there are notable shifts in the distribution of the labels and in the distribution of the features between regions. Therefore, these datasets can be helpful for testing for testing and comparing methods that are designed to address such distributional shifts.</p> <p>More details on the dataset and processing steps can be found in <a href="https://doi.org/10.1016/j.rse.2021.112488">Kluger et. al. (2021)</a>. Much of the processing steps were taken to deal with Sentinel-2 measurements that were corrupted by cloud cover. For users interested in the raw multi-spectral time series data and dealing with cloud cover issues on their own (rather than using the 70 processed features provided here), the raw dataset from Kenya can be found in <a href="https://openreview.net/forum?id=5HR3vCylqD">Yeh et. al. (2021)</a>, and the raw dataset from France can be made available upon request from the authors of this Zenodo upload.</p> <p>All of the data uploaded here can be found in "CropTypeDatasetProcessed.RData". We also post the dataframes and tables within that .RData file as separate .csv files for users who do not have R. The contents of each R object (or .csv file) is described in the file "Metadata.rtf".</p> <p><strong>Preferred Citation:</strong></p> <p>-Kluger, D.M., Wang, S., Lobell, D.B., 2021. Two shifts for crop mapping: Leveraging aggregate crop statistics to improve satellite-based maps in new regions. Remote Sens. Environ. 262, 112488. https://doi.org/10.1016/j.rse.2021.112488.</p> <p>-URL to this Zenodo post https://zenodo.org/record/6376160</p>
DeepAstroUDA: Semi-Supervised Universal Domain Adaptation for Cross-Survey Galaxy Morphology Classification and Anomaly Detection
<p>We present the data used in "DeepAstroUDA: Semi-Supervised Universal Domain Adaptation for Cross-Survey Galaxy Morphology Classification and Anomaly Detection". It was also used in the conference paper presented in Machine Learning and the Physical Sciences workshop at NeurIPS 2022: "Semi-Supervised Domain Adaptation for Cross-Survey Galaxy Morphology Classification and Anomaly Detection".</p> <p>A plethora of AI methods, has already shown huge promise in increasing quality and speed of work with astronomical datasets, but high complexity of AI methods leads to extraction of dataset-specific non-robust features, which leads to models that cannot work on multiple datasets at the same time. We develop a Universal Domain Adaptation method <em><strong>DeepAstroUDA</strong></em>, capable of performing <strong>semi-supervised domain adaptation, that can be applied to datasets with different data distributions and class overlap</strong>. Extra classes can be present in any of the two datasets, and the method can even be used in the presence of unknown classes. We apply our model to three examples of galaxy morphology classification tasks of different complexities (3-class and 10-class problems), with anomaly detection i.e. in all our experiments we have one extra class in the unlabeled target dataset, which represents our anomaly class.</p> <p> </p> <p><strong>DATA:</strong></p> <p><strong>1) DA across two different data releases of the same survey (LSST 1 and 10 years of observation):</strong> We use data from Ciprijanovic et al. 2022. which can also be found on Zenodoo: <a href="https://zenodo.org/record/5514180#.Y6SM7y-B2_w">https://zenodo.org/record/5514180</a> . Data contains three classes: spiral (0), elliptical (1) and merging galaxies (3, anomaly class).</p> <p><strong>2) DA across two surveys (SDSS and DeCALS): </strong>We create datasets using data and labels from the Galaxy Zoo project. Datasets contain 10 classes (9 known classes present in both SDSS and DeCALS data, and one unknown anomaly class present only in DeCALS data): disturbed (0), merging (1), round smooth (2), cigar shaped smooth (3), barred spiral (4), unbarred tight spiral (5), unbarred loose spiral (6), edge-on without bulge (7), edge-on with bulge (8), lenses (9, unknown anomaly class).</p> <p>SDSS (wide filed): datasets is split into two files - sdss_1.h5, sdss_2.h5</p> <p>DeCALS: decals.zip</p> <p><strong>3) DA between wide and deep observing fields of the same survey (SDSS):</strong> We create datasets using data and labels from the Galaxy Zoo project. Datasets contain same 10 classes as in 2), with the final lens anomaly class being only present in the SDSS deep field.</p> <p>SDSS (wide filed): the same data as in 2)</p> <p>SDSS (Strip 82 deep field): sdss_stripe82.zip</p> <p>All SDSS and DECaLS files contain full datasets (train, validation and test). Exact split that we performed (0.6 : 0.2 : 0.2) can be done using the code that accompanies this publication: <a href="https://github.com/deepskies/DeepAstroUDA">https://github.com/deepskies/DeepAstroUDA</a> .</p>
IEEE ICME 2024 Grand Challenge: Semi-supervised Acoustic Scene Classification under Domain Shift Evaluation Dataset
<p>The Chinese Acoustic Scene (CAS) 2023 dataset is a large-scale dataset that serves as a foundation for research related to environmental acoustic scenes. The dataset includes 10 common acoustic scenes, with a total duration of over 130 hours. Each audio clip is 10 seconds long with metadata about the recording location and timestamp. The dataset was collected by members of the <em>Joint Laboratory of Environmental Sound Sensing at the School of Marine Science and Technology, Northwestern Polytechnical University</em>. The data collection period spanned from April 2023 to September 2023, covering 22 different cities across China. The CAS 2023 dataset was collected using the XS-SN-2BE1 manufactured by <em>Xi'an Lianfeng Acoustic Technologies Co., Ltd</em> (https://www.lfxstek.com/). </p> <p>The ICME 2024 <em>Semi-supervised Acoustic Scene Classification under Domain Shift</em> challenge (https://2024.ieeeicme.org/grand-challenge-proposals/, https://ascchallenge.xshengyun.com/) dataset consists of development (https://zenodo.org/records/10616533) and evaluation datasets, all derived from the CAS 2023 dataset. The evaluation dataset includes 1,100 recordings, where data are selected from 12 cities, with 5 unseen cities specifically chosen to provide a more comprehensive evaluation of submissions under domain shift.</p> <p>Baseline: https://github.com/JishengBai/ICME2024ASC</p> <p>Acoustic scenes (10): Bus, Airport, Metro, Restaurant, Shopping mall, Public square, Urban park, Traffic street, Construction site, Bar</p>
Domain-Independent Reviews' Sentiment Polarity Classification using Shallow Word2Seq Convolutional Neural Network
<p>Reviews and comments are perceptions about specific services or products. They are embedded with hidden sentiments the reviewer has towards certain subjects. Business owners use customer reviews to understand customers’ perceptions about specific services or products. The ability to understand reviews’ sentiment from different domains or areas give decision makers and business owners the opportunity to make critical business decisions which can help them to increase profits of their businesses. Previous studies had focussed on classifying sentiment polarity by using traditional machine learning and deep learning methods. However, these suffered from low model generalization ability, causing the models to perform better only on single domain datasets rather than multiple domain datasets. The problem is the inability of the classification model to learn domain-restricted knowledge from multi-domain datasets. Aiming to improve the accuracy of the cross-domain classification, this paper proposes a method which uses Word2Seq Convolutional Neural Network (CNN) to classify reviews’ sentiment across multiple domain datasets (i.e. digital worker, movie, product, hotel and restaurant reviews). The evaluation showed that the proposed method had achieved the state-of-the-art performance. The high classification performance also promoted the reliability and effectiveness of implementing the Word2Seq CNN to classify reviews’ sentiment across different domains and learn domain restricted knowledge while improving the model generalization ability.</p> <p>The uploaded dataset is a sampled dataset with 5000 observations for both training and testing sets.</p>
Classification of domain names of scholarly email addresses
<p>This dataset contains the classification of email domain names used in the paper 'Analyzing the use of email addresses in scholarly publications' by Marc Luwel and Nees Jan van Eck. Email domain names are classified as institutional if they are linked to a scholarly organization or as non-institutional if they are linked to an email service provider.</p> <p>The file 'email_domain_name_classification.txt' contains for 11,608 email domain names the classification based on a rule-based approach and after manual validation.</p>
Complete classification of six-dimensional iso-edge domains
<p>This dataset contains the complete list of the 55.083.357 iso-edge domains in dimension 6. <br>The data is stored in a netCDF-4 data file. <br>For each iso-edge domain it contains 63 canonicalized 6-dimensional coordinates that represent the iso-edge domain.<br>Furthermore, for each iso-edge domain it contains the number of neighbouring iso-edge domains.</p> <p>The netCDF dimensions and variables are as follows:</p> <pre><code>netcdf ctype_dim6 { dimensions: number_ctype = UNLIMITED ; // (55083357 currently) n = 6 ; n_vect = 63 ; variables: int Ctype(number_ctype, n_vect, n) ; Ctype:long_name = "Ctype canonicalized coordinates" ; Ctype:units = "nondimensional" ; int nb_adjacent(number_ctype) ; nb_adjacent:long_name = "number of adjacent Ctypes" ; nb_adjacent:units = "nondimensional" ; }</code></pre> <p>The data can be extracted to a human readable format using for example the <em>ncdump</em> utility. </p>
Padding Ain't Enough: Assessing the Privacy Guarantees of Encrypted DNS – Subpage-Agnostic Domain Classification Firefox
<p>This dataset contains one part for the "Subpage-Agnostic Domain Classification" section of our FOCI 2020 paper "Padding Ain’t Enough: Assessing the Privacy Guarantees of Encrypted DNS".</p> <p><a href="https://www.usenix.org/conference/foci20/presentation/bushart">https://www.usenix.org/conference/foci20/presentation/bushart</a></p> <p>You can find the source code for this project on GitHub: <a href="https://github.com/jonasbb/padding-aint-enough">https://github.com/jonasbb/padding-aint-enough</a></p> <p>When using this software or our dataset, please cite our FOCI 20 paper.</p> <pre>@inproceedings {PaddingAintEnough, author = {Jonas Bushart and Christian Rossow}, booktitle = {10th {USENIX} Workshop on Free and Open Communications on the Internet ({FOCI} 20)}, month = aug, publisher = {{USENIX} Association}, title = {Padding Ain{\textquoteright}t Enough: Assessing the Privacy Guarantees of Encrypted {DNS}}, year = {2020}, }</pre>
Padding Ain't Enough: Assessing the Privacy Guarantees of Encrypted DNS – Subpage-Agnostic Domain Classification Tor Browser
<p>This dataset contains the second part of the "Subpage-Agnostic Domain Classification" section of our FOCI 2020 paper "Padding Ain’t Enough: Assessing the Privacy Guarantees of Encrypted DNS".</p> <p><a href="https://www.usenix.org/conference/foci20/presentation/bushart">https://www.usenix.org/conference/foci20/presentation/bushart</a></p> <p>You can find the source code for this project on GitHub: <a href="https://github.com/jonasbb/padding-aint-enough">https://github.com/jonasbb/padding-aint-enough</a></p> <p>When using this software or our dataset, please cite our FOCI 20 paper.</p> <pre>@inproceedings {PaddingAintEnough, author = {Jonas Bushart and Christian Rossow}, booktitle = {10th {USENIX} Workshop on Free and Open Communications on the Internet ({FOCI} 20)}, month = aug, publisher = {{USENIX} Association}, title = {Padding Ain{\textquoteright}t Enough: Assessing the Privacy Guarantees of Encrypted {DNS}}, year = {2020}, }</pre>
Image-based taxonomic classification of bulk biodiversity samples using deep learning and domain adaptation
<p>Complex bulk samples of insects from biodiversity surveys present a challenge for taxonomic identification, which could be overcome by high-throughput imaging combined with machine learning for rapid classification of specimens. These procedures require that taxonomic labels from an existing source data set are used for model training and prediction of an unknown target sample. However, such transfer learning may be problematic for the study of new samples not previously encountered in an image set, e.g. from unexplored ecosystems, and require methods of domain adaptation that reduce the differences in the feature distribution of the source and target domains (training and test sets). We assessed the efficiency of domain adaptation for family-level classification of bulk samples of Coleoptera, as a critical first step in the characterisation of biodiversity samples. Neural network models trained with images from a global database of Coleoptera were applied to a biodiversity sample from understudied forests in Cyprus as the target. Within-dataset classification accuracy reached 98% and depended on the number and quality of training images and on dataset complexity. The accuracy of between-datasets predictions (across disparate source-target pairs that do not share any species or genera) was at most 82% and depended greatly on the standardisation of the imaging procedure. Algorithms for domain adaptation significantly improved the prediction performance of models trained by non-standardised, low-quality images. Our findings demonstrate that existing databases can be used to train models and successfully classify images from unexplored biota, but the imaging conditions and classification algorithms need careful consideration.</p>
GradDA – A novel dataset for investigating domain shifts in image classification
<p>A domain shift occurs when the testing data is drawn from a distribution different from that of the training dataset. This shift presents a significant challenge and may compromise the performance of machine learning models, which leads to poor generalization. Over the past years, various models have been developed and evaluated on benchmark datasets such as VisDA, Office-Home and DomainNet. These datasets consist of discrete domains with different object classes. However, a notable limitation when addressing the domain shift is the absence of data samples where the exact same object exists in both domains. </p> <p>We propose a new dataset designed to address this challenge. In particular, we introduce a domain shift from a purely synthetic style (grey object on white background) to a more realistic appearance (object with texture against a realistic background) with differential modifications, which enables the representation of the same object in both synthetic and real domains, consequently facilitating the analysis of a transition between the two domains. The dataset comprises five distinct classes (Airplane, Bicycle, Bus, Car, Train), with multiple objects per class. Additionally, each object is depicted from 20 different perspectives, resulting in a total of 101 images per perspective that captures the transition from pure synthetic to a more real-world-like domain. This dataset offers a unique opportunity to investigate the impact of domain shift on model performance in classification tasks, as it focuses solely on domain changes without other interfering effects. It is the objective of our work to trigger new discussions about the domain shift problem, and how it can be tackled with alternative data driven model designs.</p>
From Galaxy Zoo DECaLS to BASS+MzLS: detailed galaxy morphological classification with unsupervised domain adaption
<p>This repository contains the data released in the paper "From Galaxy Zoo DECaLS to BASS+MzLS: Detailed Galaxy Morphological Classification with Unsupervised Domain Adaption".</p> <p>We release detailed galaxy morphological classification in DESI Legacy Imaging Surveys (LIS) DECaLS, BASS, MzLS for m_r<17.77 galaxies and z<0.15.</p> <p>-morphology_GZD.csv contains prediction for DESI LIS DECaLS footprint.</p> <p>-morphology_BMz.csv contains prediction for DESI LIS BASS+MzLS footprint.</p> <p>They include the information about: ra, dec, {question}_{answer}_alpha,{question}_{answer}_prob,{question}_{answer}_var</p> <p>Predictions of the Dirichlet parameter alpha for each galaxy on each feature of each problem, with all the original multiple MC Dropout results, are included to make it easier for you to know all the original predictions.</p>
Datasets for "A classification-based approach to override cross-domain data bias in materials discovery"
<p>This repository provides the featurized versions of the specialized datasets, SuperCon and ESTM, utilized in the study titled 'Classification-based detection and quantification of cross-domain data bias in<br>materials discovery'.</p>
Multi-channel auto-encoders for learning domain invariant representations enabling superior classification of histopathology images
<p>A partially synthetic histopathology dataset containing image patches of colon tissue from 3 staining and scanning conditions.</p> <p>This dataset can be used to develop novel histopathology image analysis algorithms that are better able to generalise to novel data domains.</p> <p>See repo for more information.</p>
DPAM Domain Classification of Human Proteins against ECOD Reference
<p>Domain definitions of AlphaFold classifications of the human proteome (v1) from the AlphaFold Database. Also included are classifications of <em>Danio rerio</em>, <em>Mus musculus</em>, <em>Pan paniscus</em>, <em>Drosophila melanogaster</em>, <em>Caenorhabditis elegans </em>used for comparative analysis to human. See README file for descriptions of file formats.</p>
Image-based taxonomic classification of bulk biodiversity samples using deep learning and domain adaptation
Open the record for dataset details and reuse information.
Evaluating the use of paralogous protein domains to increase data availability for missense variant classification [dataset]
Open the record for dataset details and reuse information.
Data associated with Automated Biomedical Text Classification with Research Domain Criteria
<p>Each file contains a set of abstracts for a given RDoC construct. Each file is named after one RDoC category. In each file, each line represents one PubMed abstract that belong to the RDoC category. Each line has a PubMeD ID and an abstract text separated by a tab. The dataset was created on August 2018. </p> <p>Please cite: M. Anani and I. Kahanda, Automated Biomedical Text Classification with Research Domain Criteria, International Conference on Bioinformatics and Computational Biology, Las Vegas, NV, 2018.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.