Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

34

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

34 results for “Domain Adaptation”

Learn how ShareScore rates datasets ↗
zenodo44/100

Crop classification dataset for testing domain adaptation or distributional shift methods

<p>In this upload we share processed crop type datasets from both France and Kenya. These datasets can be helpful for testing and comparing various domain adaptation methods. The datasets are processed,&nbsp;used, and described&nbsp;in this paper:&nbsp;<a href="https://doi.org/10.1016/j.rse.2021.112488">https://doi.org/10.1016/j.rse.2021.112488</a>&nbsp;(arXiv version: <a href="https://arxiv.org/pdf/2109.01246.pdf">https://arxiv.org/pdf/2109.01246.pdf</a>).&nbsp;</p> <p>In summary, each point in the uploaded datasets corresponds to a particular location. The label&nbsp;is the crop type grown at that location in 2017.&nbsp;The 70 processed features are based on&nbsp;Sentinel-2 satellite measurements at that location in 2017. The points in the France dataset come from 11 different departments (regions) in Occitanie, France, and the points in the Kenya dataset come from 3 different regions in Western Province, Kenya. Within each dataset there&nbsp;are&nbsp;notable shifts in the distribution of the labels and in the distribution of the features between regions. Therefore, these datasets can be helpful for testing&nbsp;for testing and comparing methods that are designed to address such distributional shifts.</p> <p>More details on the dataset and processing steps can be found in&nbsp;<a href="https://doi.org/10.1016/j.rse.2021.112488">Kluger et. al. (2021)</a>. Much of the&nbsp;processing steps were taken to deal with Sentinel-2 measurements that were corrupted by cloud cover. For users interested in the raw multi-spectral time series data and dealing with cloud cover issues on their own (rather than using the 70 processed features provided here), the raw dataset from Kenya can be found in <a href="https://openreview.net/forum?id=5HR3vCylqD">Yeh et. al. (2021)</a>, and the raw dataset from France can be made available upon request from the authors of this Zenodo upload.</p> <p>All of the data uploaded here can be found in &quot;CropTypeDatasetProcessed.RData&quot;. We also post the dataframes and tables within that .RData file&nbsp;as separate .csv&nbsp;files for users who do not have R. The contents of each R object (or&nbsp;.csv file) is described in the file &quot;Metadata.rtf&quot;.</p> <p><strong>Preferred Citation:</strong></p> <p>-Kluger, D.M., Wang, S., Lobell, D.B., 2021. Two shifts for crop mapping: Leveraging aggregate crop statistics to improve satellite-based maps in new regions. Remote Sens. Environ. 262, 112488. https://doi.org/10.1016/j.rse.2021.112488.</p> <p>-URL to this Zenodo post https://zenodo.org/record/6376160</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

DeepAstroUDA: Semi-Supervised Universal Domain Adaptation for Cross-Survey Galaxy Morphology Classification and Anomaly Detection

<p>We present the data used in &quot;DeepAstroUDA: Semi-Supervised Universal Domain Adaptation for Cross-Survey Galaxy Morphology Classification and Anomaly Detection&quot;. It was also used in the&nbsp;conference paper presented in&nbsp;Machine Learning and the Physical Sciences workshop at&nbsp;NeurIPS&nbsp;2022:&nbsp;&quot;Semi-Supervised Domain Adaptation for Cross-Survey Galaxy Morphology Classification and Anomaly Detection&quot;.</p> <p>A plethora of AI methods, has already shown huge promise&nbsp;in increasing quality and speed of work with astronomical&nbsp;datasets, but high complexity&nbsp;of AI methods leads to extraction of dataset-specific non-robust features, which&nbsp;leads to models that cannot work on multiple datasets at the same time. We develop a Universal Domain Adaptation method <em><strong>DeepAstroUDA</strong></em>,&nbsp;capable of performing&nbsp;<strong>semi-supervised domain adaptation, that can be applied&nbsp;to datasets with different data distributions and class overlap</strong>. Extra classes&nbsp;can be present in any of the two datasets, and the method can even be used&nbsp;in the presence of unknown classes. We&nbsp;apply our model to three examples&nbsp;of galaxy morphology classification tasks of different complexities (3-class and&nbsp;10-class&nbsp;problems), with anomaly detection i.e.&nbsp;in all our experiments we have one extra class in the unlabeled target dataset, which represents our anomaly class.</p> <p>&nbsp;</p> <p><strong>DATA:</strong></p> <p><strong>1) DA across two different data releases of the same survey (LSST 1&nbsp;and 10 years of observation):</strong> We use data from Ciprijanovic et al. 2022. which&nbsp;can also be found&nbsp;on Zenodoo:&nbsp;<a href="https://zenodo.org/record/5514180#.Y6SM7y-B2_w">https://zenodo.org/record/5514180</a>&nbsp;. Data contains three classes: spiral (0), elliptical (1)&nbsp;and merging galaxies (3, anomaly class).</p> <p><strong>2) DA across two surveys (SDSS and DeCALS): </strong>We create datasets using data and labels from the Galaxy Zoo project. Datasets contain&nbsp;10 classes (9 known classes present in both SDSS and DeCALS data, and one unknown anomaly class present only in DeCALS data):&nbsp;disturbed&nbsp;(0), merging (1), round smooth (2), cigar shaped&nbsp;smooth (3), barred spiral (4), unbarred tight spiral (5),&nbsp;unbarred loose spiral (6), edge-on without bulge (7),&nbsp;edge-on with bulge (8), lenses (9, unknown anomaly class).</p> <p>SDSS (wide filed): datasets is split into two files &nbsp;-&nbsp;sdss_1.h5, sdss_2.h5</p> <p>DeCALS:&nbsp; decals.zip</p> <p><strong>3) DA between wide and&nbsp;deep observing fields of the same survey (SDSS):</strong> We create&nbsp;datasets using data and labels from the Galaxy Zoo project. Datasets contain same 10 classes as in 2), with the final lens anomaly class being only present in the SDSS deep field.</p> <p>SDSS (wide filed):&nbsp;the same data as in 2)</p> <p>SDSS (Strip 82 deep field):&nbsp;sdss_stripe82.zip</p> <p>All SDSS and DECaLS files contain full datasets (train, validation and test). Exact split that we performed (0.6 : 0.2 : 0.2) can be done using the code that accompanies this publication:&nbsp;<a href="https://github.com/deepskies/DeepAstroUDA">https://github.com/deepskies/DeepAstroUDA</a>&nbsp;.</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

EBSD datasets for cross-sectioned structural steel hardness indentations - Adaptive Domain Misorientation

<p>Open access datasets for structural steel hardness indentations from the following publication: Ultramicroscopy 2021, Volume 222:&nbsp;<a href="https://doi.org/10.1016/j.ultramic.2021.113203">https://doi.org/10.1016/j.ultramic.2021.113203</a></p> <p>Files included:</p> <ul> <li>Adaptive domain misorientation calculated for Indentation 1 and 2 using misorientation thresholds (Delta theta) 0.5deg and 2deg, corresponding to dense dislocation walls and sub-grain boundaries</li> <li>Indentation 2: Raw dataset and associated mask file for excluding the edge of the data</li> </ul> <p>The methodology for analysing and plotting of the data is found at:&nbsp;<a href="https://doi.org/10.5281/zenodo.4430623">https://doi.org/10.5281/zenodo.4430623</a></p> <p>For further information visit:&nbsp;Aalto University Wiki -&nbsp;<a href="https://wiki.aalto.fi/display/EMDIDS">https://wiki.aalto.fi/display/EMDIDS</a></p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

AdaptFerm: Bioprocess Monitoring Using FTIR spectroscopy: Insights into Substrate Effects and Domain Adaptation

<h2>&nbsp;</h2> <h2><strong>1. Introduction</strong></h2> <p>The AdaptFerm dataset is designed to support the development of a monitoring framework for lactic acid production fermentation using Fourier Transform Infrared (FTIR) spectroscopy. Its primary goal is to facilitate the control strategies for continuous fermentation processes to maximize the lactic acid production. The AdaptFerm encompasses data from two distinct batch fermentation environments: one employing simple sugar (glucose) as the substrate and the other utilizing complex sugars derived from bio-waste. The study focuses on developing accurate predictive models for glucose and lactic acid concentrations, with an emphasis on applying classical machine learning techniques and enhancing domain generalization capabilities.</p> <h2><strong>2. Prediction Model for Different Substrate Environments</strong></h2> <p>The chemical composition of substrates are presented in <a title="Digital model of biochemical reactions in lactic acid bacterial fermentation of simple glucose and biowaste substrates" href="https://doi.org/10.1016/j.heliyon.2024.e38791" target="_blank" rel="noopener">Table 1 [1]</a>. The dataset is utilized to train and test models within the same substrate domain. For instance, data from a single fermentation environment (e.g., glucose substrate) is used for both training and testing phases. The applied machine learning models showed accurate prediction within the same domain <a title="Digital model of biochemical reactions in lactic acid bacterial fermentation of simple glucose and biowaste substrates" href="https://doi.org/10.1016/j.heliyon.2024.e38791" target="_blank" rel="noopener">[1]</a>. For more details on the methods applied, please refer to the following link: <a title="Digital model of biochemical reactions in lactic acid bacterial fermentation of simple glucose and biowaste substrates" href="https://doi.org/10.1016/j.heliyon.2024.e38791" target="_blank" rel="noopener">https://doi.org/10.1016/j.heliyon.2024.e38791</a>. In this study, the MIR results correspond to the AdaptFerm dataset. The&nbsp;spectra of the glucose and biowaste hydrolysate fermentation process are presented in <a title="Digital model of biochemical reactions in lactic acid bacterial fermentation of simple glucose and biowaste substrates" href="https://doi.org/10.1016/j.heliyon.2024.e38791" target="_blank" rel="noopener">Figure 3</a> and <a title="Digital model of biochemical reactions in lactic acid bacterial fermentation of simple glucose and biowaste substrates" href="https://doi.org/10.1016/j.heliyon.2024.e38791" target="_blank" rel="noopener">Figure 4</a>.</p> <h2><strong>3. Domain Adaptation</strong></h2> <p>The dataset was also used to address the challenge posed by shifts in FTIR data when substrates change. Transitioning from simple sugar (glucose) to complex sugar (bio-waste) causes significant variations in the FTIR spectra, making it difficult for models trained on glucose fermentation data to maintain prediction accuracy in the complex sugar fermentation environment. This results in reduced robustness and performance when applied to out-of-distribution data. To address these challenges, we explore methods that improve the generalization ability and robustness of models in such scenarios without using labels from complex sugar fermentation <a title="Domain-Invariant Monitoring for Lactic Acid Production: Transfer Learning from Glucose to Bio-Waste Using Machine Learning Interpretation" href="https://dx.doi.org/10.2139/ssrn.5012080" target="_blank" rel="noopener">[2]</a>. It shows the application of machine learning interpretation to find domain invariant features for glucose and lactic acid. For more details on the methods applied, please refer to the following link: <a title="Domain-Invariant Monitoring for Lactic Acid Production: Transfer Learning from Glucose to Bio-Waste Using Machine Learning Interpretation" href="https://dx.doi.org/10.2139/ssrn.5012080" target="_blank" rel="noopener">https://dx.doi.org/10.2139/ssrn.5012080</a>. The code is available at <a title="ShapFS" href="https://github.com/shl-shawn/ShapFS" target="_blank" rel="noopener">https://github.com/shl-shawn/ShapFS</a>.</p> <h2><strong>4. Real-World Use Cases</strong></h2> <h3><strong>4.1. Regression Task</strong></h3> <p>AdaptFerm serves as a benchmark for machine learning model applications in fermentation processes, specifically for predicting glucose and lactic acid concentrations, measured in g/L (grams per liter), while considering issues of out-of-distribution generalization.</p> <h3><strong>4.2. Domain Adaptation Regression Task</strong></h3> <p>The dataset is also suitable for evaluating different &nbsp;domain adaptation methods. In particular, the glucose substrate fermentation data can be used as the source domain, while the complex sugar fermentation data from bio-waste serves as the target domain. For semi-supervised domain adaptation approaches, it is recommended to use the initial data points (i.e., those collected at the beginning of the fermentation process) from the target domain, as the dataset is organized chronologically by collection day. These approaches aim to improve the robustness of models by transferring knowledge across domains and mitigating the effects of out-of-distribution data.</p> <h3><strong>4.3. Anomaly Detection</strong></h3> <p>The dataset can be used to train anomaly detection models to identify outliers or deviations from normal fermentation behavior. This could be valuable in industrial bioprocessing, where early detection of issues like contamination or process failure is crucial. Techniques like Isolation Forests, One-Class SVM, or Autoencoders could be applied to identify unusual patterns in FTIR spectra.</p> <h3><strong>4.4. Classification Task</strong></h3> <p>Although the main task is regression, the dataset could also be used in classification tasks by discretizing the concentrations of glucose and lactic acid into categories (e.g., low, medium, high). This would allow for the application of classification algorithms like Support Vector Machines (SVM), Random Forests, or Neural Networks for predicting the fermentation phase or identifying specific operational conditions.</p> <h3><strong>4.5. Transfer Learning</strong></h3> <p>Given the nature of the domain adaptation approach in this dataset, transfer learning models can be explored. Models pre-trained on glucose fermentation data can be fine-tuned on complex sugar fermentation data, enabling quicker model convergence and improved performance in data-scarce environments.</p> <h3><strong>4.6. Multi-Task Learning</strong></h3> <p>In a multi-task learning scenario, models could simultaneously predict both glucose and lactic acid concentrations from the same FTIR data. This could help in improving model accuracy by leveraging shared representations across the two tasks.</p> <h3><strong>4.7. Feature Selection</strong></h3> <p>The FTIR spectral data contains a large number of features (wavelengths), and feature selection techniques such as Recursive Feature Elimination (RFE), Lasso regression, or mutual information could be applied to identify the most relevant wavelengths for predicting glucose and lactic acid concentrations, improving model performance and interpretability.</p> <h2><strong>5. Dataset Structure and Meta Information</strong></h2> <p>The dataset is organized into four Excel files, corresponding to two main fermentation domains (different substrates) and two key process variables:</p> <p><strong>a) Simple Sugar Substrate</strong><br>This domain contains data for the fermentation process using glucose as the substrate to produce lactic acid. It includes two files&mdash;one for glucose concentrations and one for lactic acid concentrations. Both are measured in&nbsp;g/L.</p> <p><strong>b) Complex Sugar Substrate</strong><br>This doman contains data for the fermentation process using bio-waste as the substrate to produce lactic acid. Similar to the previous domain, it includes two files&mdash;one for glucose concentrations and one for lactic acid concentrations. Both are measured in g/L.</p> <p>Each file is structured as follows:</p> <ul> <li>The first column contains the<strong> </strong>sample ID, which serves as the timeline of measurements (Sample ID 1 represents the first measurement in the fermentation process).</li> <li>From the second column onwards, the FTIR data is provided, covering the spectral range from 549.6 cm-1 to 3999.6 cm-1 comprising 3,579 features.</li> <li>The final column contains the ground truth data, the chemical measurements of fermentation variables such as glucose and lactic acid concentrations, both measured in g/L.</li> </ul> <h2><strong>6. Conclusion</strong></h2> <p>The AdaptFerm features FTIR spectra data from two distinct fermentation environments: simple sugar (glucose) and complex sugar (bio-waste). The dataset is designed to be used in regression tasks, including domain adaptation, and can be applied in machine learning model development for fermentation process monitoring, with a focus on enhancing model robustness and handling out-of-distribution data.&nbsp;This dataset provides a valuable resource for exploring<strong> </strong>domain shift and improving the robustness of machine learning models in bioengineering and fermentation processes. It enables further research into domain generalization techniques and offers a wide range of possibilities for machine learning applications.</p> <h2>References</h2> <p>&nbsp;[1] Arman Arefi, Barbara Sturm, Majharulislam Babor, Michael Horf, Thomas Hoffmann, Marina H&ouml;hne, Kathleen Friedrich, Linda Schroedter, Joachim Venus, Agata Olszewska-Widdrat, Digital model of biochemical reactions in lactic acid bacterial fermentation of simple glucose and biowaste substrates, Heliyon, Volume 10, Issue 19, 2024, e38791, ISSN 2405-8440, DOI: 10.1016/j.heliyon.2024.e38791, <a href="https://doi.org/10.1016/j.heliyon.2024.e38791" target="_blank" rel="noopener">https://doi.org/10.1016/j.heliyon.2024.e38791</a>.</p> <p>[2] &nbsp;Majharulislam Babor, Shanghua Liu, Arman Arefi, Agata Olszewska-Widdrat, &nbsp;Barbara Sturm, Joachim Venus, and Marina M.-C. H&ouml;hne, Domain-Invariant Monitoring for Lactic Acid Production: Transfer Learning from Glucose to Bio-Waste Using Machine Learning Interpretation. Available at <a href="https://dx.doi.org/10.2139/ssrn.5012080" target="_blank" rel="noopener">http://dx.doi.org/10.2139/ssrn.5012080.</a></p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation

<p>The benchmark code is available at:&nbsp;<a href="https://github.com/Junjue-Wang/LoveDA">https://github.com/Junjue-Wang/LoveDA</a></p> <p><strong>Highlights:&nbsp;</strong></p> <ol> <li>5987 high spatial resolution (0.3 m) remote sensing images from Nanjing, Changzhou, and Wuhan</li> <li>Focus on different geographical environments between Urban and Rural</li> <li>Advance both semantic segmentation and domain adaptation tasks</li> <li>Three considerable challenges: multi-scale objects, complex background samples, and inconsistent class distributions</li> </ol> <p><strong>Reference:</strong></p> <pre><code>@inproceedings{wang2021loveda, title={Love{DA}: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation}, author={Junjue Wang and Zhuo Zheng and Ailong Ma and Xiaoyan Lu and Yanfei Zhong}, booktitle={Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks}, editor = {J. Vanschoren and S. Yeung}, year={2021}, volume = {1}, pages = {}, url={https://datasets-benchmarks proceedings.neurips.cc/paper/2021/file/4e732ced3463d06de0ca9a15b6153677-Paper-round2.pdf} }</code></pre> <p><strong>License:</strong></p> <p>The owners of the data and of the copyright on the data are RSIDEA, Wuhan University. Use of the Google Earth images must respect the &quot;Google Earth&quot; terms of use. All images and their associated annotations in LoveDA can be used for academic purposes only, <strong>but any commercial use is prohibited. (CC BY-NC-SA 4.0)</strong></p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Cross-Modality Domain Adaptation Challenge 2022 (crossMoDA)

<p>Official training and validation sets of crossMoDA 2022.</p> <p><strong>All data will be made available online with a permissive non-commercial copyright-license (<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">CC BY-NC-SA 4.0</a>), allowing for data to be shared, distributed and improved upon</strong>.</p> <p>&nbsp;</p> <p>If you use the data, please cite:</p> <p>1. Shapey, J., Kujawa, A., Dorent, R., Wang, G., Bisdas, S., Dimitriadis, A., Grishchuck, D., Paddick, I., Kitchen, N., Bradford, R., Saeed, S., Ourselin, S., &amp; Vercauteren, T. (2021). Segmentation of Vestibular Schwannoma from Magnetic Resonance Imaging: An Open Annotated Dataset and Baseline Algorithm [Data set]. The Cancer Imaging Archive. <a href="https://doi.org/10.7937/TCIA.9YTJ-5Q73">https://doi.org/10.7937/TCIA.9YTJ-5Q73</a>&nbsp;</p> <p>2. Dorent, R. et al (2022).&nbsp; CrossMoDA 2021 challenge: Benchmark of Cross-Modality Domain Adaptation techniques for Vestibular Schwannoma and Cochlea Segmentation.&nbsp; ArXiv <a href="https://arxiv.org/abs/2201.02831">https://arxiv.org/abs/2201.02831</a></p> <p>&nbsp;</p> <p>Acknowledgments:</p> <p>This challenge is supported by Wellcome Trust (203145Z/16/Z, 203148/Z/16/Z), EPSRC (NS/A000050/1,<br> NS/A000049/1) and ZonMw (project number: 10070012010006) funding. All the organizers will have access to the<br> test set if needed.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

SemEval-2021 Task 10: Source-Free Domain Adaptation for Semantic Processing

<p>Data sharing restrictions are common in NLP datasets. For example, Twitter policies do not allow sharing of tweet text, though tweet IDs may be shared. The situation is even more common in clinical NLP, where patient health information must be protected, and annotations over health text, when released at all, often require the signing of complex data use agreements. The SemEval-2021 Task 10 framework asks participants to develop semantic annotation systems in the face of data sharing constraints. A participant&#39;s goal is to develop an accurate system for a target domain when annotations exist for a related domain but cannot be distributed. Instead of annotated training data, participants are given a model trained on the annotations. Then, given unlabeled target domain data, they are asked to make predictions.</p> <p>Website: <a href="https://machine-learning-for-medical-language.github.io/source-free-domain-adaptation/">https://machine-learning-for-medical-language.github.io/source-free-domain-adaptation/</a></p> <p>CodaLab site: <a href="https://competitions.codalab.org/competitions/26152">https://competitions.codalab.org/competitions/26152</a></p> <p>Github repository: <a href="https://github.com/Machine-Learning-for-Medical-Language/source-free-domain-adaptation">https://github.com/Machine-Learning-for-Medical-Language/source-free-domain-adaptation</a></p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Domain-adaptive Data Synthesis for Large-scale Supermarket Product Recognition

<p><strong>Domain-Adaptive Data Synthesis for Large-Scale Supermarket Product Recognition</strong></p> <p>This repository contains the data synthesis pipeline and synthetic product recognition datasets proposed in [1].</p> <p><strong>Data Synthesis Pipeline:</strong></p> <p>We provide the Blender 3.1 project files and Python source code of our data synthesis pipeline <em>pipeline.zip,&nbsp;</em>accompanied by the<em>&nbsp;</em><a href="https://github.com/taesungp/contrastive-unpaired-translation">FastCUT</a> models used for synthetic-to-real domain translation<em>&nbsp;models.zip</em>. For the synthesis of new shelf images, a product assortment list and product images must be provided in the corresponding directories <em>products/assortment/</em> and <em>products/img/</em>. The pipeline expects product images to follow the naming convention <em>c</em>.png, with <em>c</em> corresponding to a GTIN or generic class label (e.g., 9120050882171.png). The assortment list, <em>assortment.csv</em>, is expected to use the sample format [<em>c, w, d, h</em>], with <em>c</em> being the class label and <em>w, d,</em> and <em>h</em> being the packaging dimensions of the given product in mm (e.g., [4004218143128, 140, 70, 160]). The assortment list to use and the number of images to generate can be specified in <em>generateImages.py </em>(see comments). The rendering process is initiated by either executing&nbsp;<em>load.py</em> from within Blender or within a command-line terminal as a background process.&nbsp;</p> <p><strong>Datasets:</strong></p> <ul> <li><strong>SG3k</strong> -&nbsp;Synthetic GroZi-3.2k (SG3k) dataset, consisting of 10,000 synthetic shelf images with 851,801 instances of 3,234&nbsp;GroZi-3.2k products.&nbsp;Instance-level bounding boxes and generic class labels are provided for all product instances.</li> <li><strong>SG3kt</strong>&nbsp;-&nbsp;Domain-translated version of&nbsp;SGI3k, utilizing GroZi-3.2k as the target domain.&nbsp;Instance-level bounding boxes and generic class labels are provided for all product instances.</li> <li><strong>SGI3k</strong> -&nbsp;Synthetic GroZi-3.2k (SG3k) dataset, consisting of 10,000 synthetic shelf images with 838,696&nbsp;instances of 1,063&nbsp;GroZi-3.2k products.&nbsp;Instance-level bounding boxes and&nbsp;generic class labels&nbsp;are provided for all product instances.</li> <li><strong>SGI3kt</strong>&nbsp;-&nbsp;Domain-translated version of&nbsp;SGI3k, utilizing GroZi-3.2k as the target domain.&nbsp;Instance-level bounding boxes and&nbsp;generic class labels are provided for all product instances.</li> <li><strong>SPS8k</strong> - Synthetic Product Shelves 8k (SPS8k) dataset, comprised&nbsp;of 16,224 synthetic shelf images with 1,981,967 instances of 8,112 supermarket products. Instance-level bounding boxes and GTIN class labels are provided for all product instances.</li> <li><strong>SPS8kt</strong>&nbsp;- Domain-translated version of&nbsp;SPS8k, utilizing&nbsp;SKU110k as the target domain.&nbsp;Instance-level bounding boxes and GTIN class labels for all product instances.</li> </ul> <p>Table 1: Dataset characteristics.&nbsp;</p> <table> <tbody> <tr> <td><strong>Dataset</strong></td> <td><strong>#images</strong></td> <td><strong>#products</strong></td> <td><strong>#instances</strong></td> <td>&nbsp;&nbsp;<strong>labels &nbsp; &nbsp; </strong></td> <td><strong>translation</strong></td> </tr> <tr> <td>SG3k</td> <td>10,000</td> <td>3,234</td> <td>851,801</td> <td>bounding box &amp; generic class&sup1;</td> <td>none</td> </tr> <tr> <td>SG3kt</td> <td>10,000</td> <td>3,234</td> <td>851,801</td> <td>bounding box &amp; generic class&sup1;</td> <td>GroZi-3.2k</td> </tr> <tr> <td>SGI3k</td> <td>10,000</td> <td>1,063</td> <td>838,696</td> <td>bounding box &amp; generic class&sup2;</td> <td>none</td> </tr> <tr> <td>SGI3kt</td> <td>10,000</td> <td>1,063</td> <td>838,696</td> <td>bounding box &amp; generic class&sup2;</td> <td>GroZi-3.2k</td> </tr> <tr> <td>SPS8k</td> <td>16,224</td> <td>8,112</td> <td>1,981,967</td> <td>bounding box &amp; GTIN</td> <td>none</td> </tr> <tr> <td>SPS8kt</td> <td>16,224</td> <td>8,112</td> <td>1,981,967</td> <td>bounding box &amp; GTIN</td> <td>SKU110k</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Sample Format</strong></p> <p>A sample consists of an RGB image (i.png) and an accompanying label file (i.txt), which contains the labels for all product instances present in the image. Labels use the YOLO format&nbsp;[c, x, y, w, h].</p> <p>&sup1;SG3k and&nbsp;SG3kt&nbsp;use generic pseudo-GTIN&nbsp;class labels, created&nbsp;by combining the&nbsp;GroZi-3.2k food product category number <em>i</em> (1-27) with the product image index <em>j </em>(j.jpg)<em>, </em>following the convention<em>&nbsp;i0000j </em>(e.g., 13000097).</p> <p>&sup2;SGI3k and&nbsp;SGI3kt&nbsp;use the generic&nbsp;GroZi-3.2k class labels from&nbsp;<a href="https://arxiv.org/abs/2003.06800">https://arxiv.org/abs/2003.06800</a>.</p> <p><strong>Download and Use</strong><br>This data may be used for non-commercial research purposes only.&nbsp;If you publish material based on this data, we request that you include a reference to our paper [1].</p> <p>[1] Strohmayer, Julian, and Martin Kampel. "Domain-Adaptive Data Synthesis for Large-Scale Supermarket Product Recognition."&nbsp;<em>International Conference on Computer Analysis of Images and Patterns</em>. Cham: Springer Nature Switzerland, 2023.</p> <p>BibTeX&nbsp;citation:</p> <pre>@inproceedings{strohmayer2023domain, title={Domain-Adaptive Data Synthesis for Large-Scale Supermarket Product Recognition}, author={Strohmayer, Julian and Kampel, Martin}, booktitle={International Conference on Computer Analysis of Images and Patterns}, pages={239--250}, year={2023}, organization={Springer} }</pre>

opencc-by-4.0Sep 2023View details →
dryad40/100

Data from: Leveraging data mining, active learning, and domain adaptation for efficient discovery of advanced oxygen evolution electrocatalysts

Open the record for dataset details and reuse information.

publicMar 2025View details →
zenodo36/100

Reference Architecture for Autonomy and Adaptivity in the Space Domain

<p>File 1 - <strong>Protocol SLR - Reference Architecture for autonomous and adaptive space systems.xlsx</strong>: This data extraction worksheet presents the relevant data from the identified studies and provides the selection process in this systematic literature review. The main information reported in this worksheet summarizes the followed protocol, containing an identifier (id) for each returned study. This catalog helped us in the data extraction and synthesis procedures and may be used by potentially interested, for example, for updating or replication.&nbsp;</p> <p>File 2 - <strong>Interview_Guide_Questionnaire.pdf</strong>: The questionnaire was used to guide the interviews and also to ensure consistency between responses. It covers key aspects of the proposed reference architecture, such as general applicability, autonomy and adaptability, scalability, and standards compliance. There is also space for interviewers to add comments in each section and at the end of the form.</p> <p>File 3 - <strong>Filled_Questionnaire_Interviewee_1.pdf</strong>: The questionnaire answered by the interviewee 1, with the scores and comments.</p> <p>FIle 4 - <strong>Expert 1 Transcription.docx</strong>: The transcripts of the interviewee 1. The original language was not English, so we used OpenAI to generate the English version.</p> <p>File 5 - <strong>Filled_Questionnaire_Interviewee_2.pdf</strong>: The questionnaire answered by the interviewee 2, with the scores and comments.</p> <p>FIle 6 - <strong>Expert 2 Transcription.docx</strong>: The transcripts of the interviewee 2. The original language was not English, so we used OpenAI to generate the English version.</p> <p>File 7 - <strong>Filled_Questionnaire_Interviewee_3.pdf</strong>: The questionnaire answered by the interviewee 3, with the scores and comments.</p> <p><strong>Expert 3</strong> did not authorize us to release the interviews, so we are not sending the transcripts.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Reference Architecture for Autonomy and Adaptivity in the Space Domain

<p>File 1 - <strong>Protocol SLR - Reference Architecture for autonomous and adaptive space systems.xlsx</strong>: This data extraction worksheet presents the relevant data from the identified studies and provides the selection process in this systematic literature review. The main information reported in this worksheet summarizes the followed protocol, containing an identifier (id) for each returned study. This catalog helped us in the data extraction and synthesis procedures and may be used by potentially interested, for example, for updating or replication.&nbsp;</p> <p>File 2 - <strong>Interview_Guide_Questionnaire.pdf</strong>: The questionnaire was used to guide the interviews and also to ensure consistency between responses. It covers key aspects of the proposed reference architecture, such as general applicability, autonomy and adaptability, scalability, and standards compliance. There is also space for interviewers to add comments in each section and at the end of the form.</p> <p>File 3 - <strong>Filled_Questionnaire_Interviewee_1.pdf</strong>: The questionnaire answered by the interviewee 1, with the scores and comments.</p> <p>FIle 4 - <strong>Expert 1 Transcription.docx</strong>: The transcripts of the interviewee 1. The original language was not English, so we used OpenAI to generate the English version.</p> <p>File 5 - <strong>Filled_Questionnaire_Interviewee_2.pdf</strong>: The questionnaire answered by the interviewee 2, with the scores and comments.</p> <p>FIle 6 - <strong>Expert 2 Transcription.docx</strong>: The transcripts of the interviewee 2. The original language was not English, so we used OpenAI to generate the English version.</p> <p>File 7 - <strong>Filled_Questionnaire_Interviewee_3.pdf</strong>: The questionnaire answered by the interviewee 3, with the scores and comments.</p> <p><strong>Expert 3</strong> did not authorize us to release the interviews, so we are not sending the transcripts.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

MultiCardioNER Corpus: Multilingual Adaptation of Clinical NER Systems to the Cardiology Domain

<h1><strong>MultiCardioNER</strong></h1> <p><strong>MultiCardioNER</strong> is a shared task about the adaptation of clinical NER systems to the cardiology domain. It uses a combination of two existing datasets (DisTEMIST for diseases and the newly-released DrugTEMIST for medications), as well as a new, smaller dataset of cardiology clinical cases annotated using the same guidelines.</p> <p>Participants are provided DisTEMIST and DrugTEMIST as training data to use as they see fit (1,000 documents, with the original partitions splitting them into 750 for training and 250 for testing). The cardiology clinical cases (cardioccc) are meant to be used as a development or validation set (258 documents), although participants are encourage to experiment with the documents and annotations as they see fit. The evaluation is done using a different collection of cardiology clinical cases (250).</p> <p>MultiCardioNER proposes two tracks:</p> <p>- Track 1: Spanish adaptation of disease recognition systems to the cardiology domain.<br>- Track 2: Multilingual (Spanish, English and Italian) adaptation of medication recognition systems to the cardiology domain.</p> <p>Please read the README file attached for more information on folder structure and file format.</p> <p><strong>MultiCardioNER</strong> was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of BioASQ 2024. For more information on the corpus, annotation scheme and task in general, please visit: <a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">https://temu.bsc.es/multicardioner</a>. This task is promoted by Spanish and European projects such as DataTools4Heart, AI4HF, BARITONE and AI4ProfHealth.</p> <p><strong>UPDATE MAY 28th 2024: </strong>The test set annotations are now out! We've also included the original background set files, as well as a file with the mappings from the masked filenames used during the evaluation phase to the original filenames. Please check the README for more information.</p> <h2><strong>Resources</strong></h2> <ul> <li><a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">MultiCardioNER website</a></li> <li><a href="http://bioasq.org/" target="_blank" rel="noopener">BioASQ website</a></li> <li><a href="../doi/10.5281/zenodo.6458078" target="_blank" rel="noopener">DisTEMIST Guidelines</a></li> <li><a href="../doi/10.5281/zenodo.11065432" target="_blank" rel="noopener">DrugTEMIST Guidelines</a></li> </ul> <h2><strong>License</strong></h2> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <h2><strong>Contact</strong></h2> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-L&oacute;pez (&lt;salvador [dot] limalopez [at] gmail [dot] com&gt;)<br>- Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p> <h2><strong>Additional resources and corpora</strong></h2> <p>If you are interested in MultiCardioNER, you might want to check out these corpora and resources:</p> <ul> <li><a href="../records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT)</li> <li><a href="../records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT)</li> <li><a href="../records/10635215">SympTEMIST</a> (Corpus of clinical findings and normalization to SNOMED CT)</li> <li><a href="../records/4270158">PharmaCoNER</a> (Corpus of medications, drugs, chemical substances, genes, proteins and vaccine mentions and normalization)</li> <li><a href="../records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization)</li> <li><a href="../records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization)</li> <li><a href="../records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI))</li> <li><a href="../records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization)</li> <li><a href="../records/3837305">CodiESP</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version)</li> <li><a href="../records/7684093">LivingNER</a> (Corpus of mentions of species, including human/family members, pathogens, food, etc.. and normalization to NCBI Taxonomy)</li> <li><a href="../records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags)</li> <li><a href="../records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries))</li> <li><a href="../records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags)</li> <li><a href="../records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts)</li> </ul>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Neural network prediction of strong lensing systems with domain adaptation and uncertainty quantification

<p>This project combines the emerging field of Domain Adaptation with Uncertainty Quantification, working towards applying machine learning to real scientific datasets with limited labelled data. For this project, simulated images of strong gravitational lenses are used as source and target dataset, and the Einstein radius &theta; E and its uncertainty are determined through regression.</p> <p>Applying machine learning in science domains such as astronomy is difficult. With models trained on simulated data being applied to real data, models frequently underperform - simulations cannot perfectlty capture the true complexity of real data. Enter domain adaptation (DA). The DA techniques used in this work use Maximum Mean Discrepancy (MMD) Loss to train a network to being embeddings of labelled "source" data gravitational lenses in line with unlabeled "target" gravitational lenses. With source and target datasets made similar, training on source datasets can be used with greater fidelity on target datasets.</p> <p>Scientific analysis requires an estimate of uncertainty on measurements. We adopt an approach known as mean-variance estimation, which seeks to estimate the variance and control regression by minimizing the beta negative log-likelihood loss. To our knowledge, this is the first time that domain adaptation and uncertainty quantification are being combined, especially for regression on an astrophysical dataset.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

DAGHAR: A Benchmark for Domain Adaptation and Generalization in Smartphone-Based Human Activity Recognition

<p>DAGHAR benchmark is a curated dataset collection designed for domain adaptation and domain generalization studies in HAR tasks, using inertial sensors such as accelerometers and gyroscopes, from "A benchmark for domain adaptation and generalization in smartphone-based human activity recognition" work.&nbsp;It features raw inertial sensor data sourced exclusively from smartphones. Six public datasets were selected and standardized in terms of accelerometer units of measurement, sampling rate, gravity component, activity labels, user partitioning, and time window size. This standardization process allows for creating a comprehensive benchmark for evaluating the generalization capabilities of HAR models in cross-dataset scenarios.</p> <p>The benchmark is based on the following datasets:</p> <ul> <li><strong>Ku-HAR</strong>, from "Sikder, N. and Nahid, A.A., 2021. KU-HAR: An open dataset for heterogeneous human activity recognition. Pattern Recognition Letters, 146, pp.46-54", avaliable at <a href="https://data.mendeley.com/datasets/45f952y38r/5">Mendeley</a>. Distributed under CC BY 4.0.</li> <li><strong>MotionSense</strong>, from "Malekzadeh, M., Clegg, R.G., Cavallaro, A. and Haddadi, H., 2019, April. Mobile sensor data anonymization. In Proceedings of the international conference on internet of things design and implementation (pp. 49-58)", available at <a href="https://www.kaggle.com/datasets/malekzadeh/motionsense-dataset" target="_blank" rel="noopener">Kaggle</a>. Distributed under Open Data Commons Open Database License (ODbL) v1.0.</li> <li><strong>RealWorld</strong>, from "Sztyler, T. and Stuckenschmidt, H., 2016, March. On-body localization of wearable devices: An investigation of position-aware activity recognition. In 2016 IEEE international conference on pervasive computing and communications (PerCom) (pp. 1-9). IEEE", available at <a href="https://www.uni-mannheim.de/dws/research/projects/activity-recognition/dataset/dataset-realworld/" target="_blank" rel="noopener">this link</a>. We obtained explicitly permission to distribute a copy of the preprocessed data from the original authors.</li> <li><strong>UCI-HAR</strong>, from "Reyes-Ortiz, J.L., Oneto, L., Sam&agrave;, A., Parra, X. and Anguita, D., 2016. Transition-aware human activity recognition using smartphones. Neurocomputing, 171, pp.754-767", available at <a href="https://archive.ics.uci.edu/dataset/240/human+activity+recognition+using+smartphones">UCI Repository</a>. Distributed under CC BY 4.0.</li> <li><strong>WISDM</strong>, from "Weiss, G.M., Yoneda, K. and Hayajneh, T., 2019. Smartphone and smartwatch-based biometrics using activities of daily living. Ieee Access, 7, pp.133190-133202", available at <a href="https://archive.ics.uci.edu/dataset/507/wisdm+smartphone+and+smartwatch+activity+and+biometrics+dataset">UCI repository</a>. Distributed under CC BY 4.0.</li> </ul>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Structural Re-weighting Improves Graph Domain Adaptation

<p>This contains the High Energy Physics (HEP) datasets included in the paper:&nbsp;Structural Re-weighting Improves Graph Domain Adaptation. The paper can be found&nbsp;<a href="https://arxiv.org/abs/2306.03221">https://arxiv.org/abs/2306.03221</a>.</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

DREDDA: Drug Repositioning through Expression Data Domain Adaptation

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
dryad32/100

Image-based taxonomic classification of bulk biodiversity samples using deep learning and domain adaptation

<p>Complex bulk samples of insects from biodiversity surveys present a challenge for taxonomic identification, which could be overcome by high-throughput imaging combined with machine learning for rapid classification of specimens. These procedures require that taxonomic labels from an existing source data set are used for model training and prediction of an unknown target sample. However, such transfer learning may be problematic for the study of new samples not previously encountered in an image set, e.g. from unexplored ecosystems, and require methods of domain adaptation that reduce the differences in the feature distribution of the source and target domains (training and test sets). We assessed the efficiency of domain adaptation for family-level classification of bulk samples of Coleoptera, as a critical first step in the characterisation of biodiversity samples. Neural network models trained with images from a global database of Coleoptera were applied to a biodiversity sample from understudied forests in Cyprus as the target. Within-dataset classification accuracy reached 98% and depended on the number and quality of training images and on dataset complexity. The accuracy of between-datasets predictions (across disparate source-target pairs that do not share any species or genera) was at most 82% and depended greatly on the standardisation of the imaging procedure. Algorithms for domain adaptation significantly improved the prediction performance of models trained by non-standardised, low-quality images. Our findings demonstrate that existing databases can be used to train models and successfully classify images from unexplored biota, but the imaging conditions and classification algorithms need careful consideration.</p>

opencc-zeroJan 2022View details →
zenodo32/100

The Development of Adaptation Aftereffects in the Vibrotactile Domain

<p>Dataset for the study &quot;The Development of Adaptation Aftereffects in the Vibrotactile Domain&quot;.</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

RIGA+ Dataset for Unsupervised Domain Adaptation in Medical Image Segmentation

<p>Different from the previous combined multi-domain dataset for unsupervised domain adaptation (UDA) in medical image segmentation, this multi-domain fundus image dataset contains annotations made by&nbsp;the same group of ophthalmologists. Hence the&nbsp;annotator bias&nbsp;among different datasets can be&nbsp;mitigated. Therefore, this dataset can provide a&nbsp;relatively fair benchmark for evaluating UDA methods in fundus image segmentation.</p> <p>This dataset is based on the RIGA[1] dataset and MESSIDOR[2] dataset. We&nbsp;appreciate their&nbsp;efforts&nbsp;devoted by the authors of [1] and [2].</p> <p>The six&nbsp;duplicated cases in the&nbsp;RIGA dataset are&nbsp;filtered out according to the&nbsp;<a href="https://www.adcis.net/en/third-party/messidor/">Errata</a>. We also remove the duplicated cases that exist in both the&nbsp;RIGA dataset and the&nbsp;MESSIDOR dataset by hash value matching.</p> <table align="center"> <caption>Details of the RIGA+ dataset</caption> <thead> <tr> <th scope="row">Domain</th> <th scope="col">Dataset</th> <th scope="col"> <p>Labeled Samples</p> <p>(Train+Test)</p> </th> <th scope="col"> <p>Unlabeled</p> <p>Samples</p> </th> </tr> </thead> <tbody> <tr> <th scope="row">Source</th> <td>BinRushed</td> <td>195 (195+0)</td> <td>0</td> </tr> <tr> <th scope="row">Source</th> <td>Magrabia</td> <td>95 (95+0)</td> <td>0</td> </tr> <tr> <th scope="row">Target</th> <td>MESSIDOR-BASE1</td> <td>173 (138+35)</td> <td>227</td> </tr> <tr> <th scope="row">Target</th> <td>MESSIDOR-BASE2</td> <td>148 (118+30)</td> <td>238</td> </tr> <tr> <th scope="row">Target</th> <td>MESSIDOR-BASE3</td> <td>133 (106+27)</td> <td>252</td> </tr> </tbody> </table> <p>[1]&nbsp;Almazroa A, Alodhayb S, Osman E, et al. Retinal fundus images for glaucoma analysis: the RIGA dataset[C]//Medical Imaging 2018: Imaging Informatics for Healthcare, Research, and Applications. International Society for Optics and Photonics, 2018, 10579: 105790B.</p> <p>[2]&nbsp;Decenci&egrave;re E, Zhang X, Cazuguel G, et al. Feedback on a publicly distributed image database: the Messidor database[J]. Image Analysis &amp; Stereology, 2014, 33(3): 231-234.</p> <p>If you find this dataset useful for your research, please consider citing the paper as follows:</p> <pre><code>@inproceedings{hu2022domain, title={Domain Specific Convolution and High Frequency Reconstruction based Unsupervised Domain Adaptation for Medical Image Segmentation}, author={Shishuai Hu and Zehui Liao and Yong Xia}, booktitle={International Conference on Medical Image Computing and Computer-Assisted Intervention}, year={2022}, organization={Springer} }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

WildCLIP: Scene and animal attribute retrieval from camera trap data with domain-adapted vision-language models

<p>#############</p> <h1>WildCLIP: Scene and animal attribute retrieval from camera trap data with domain-adapted vision-language models</h1> <p>#############</p> <p>Authors: Valentin Gabeff, Marc Russwurm, Devis Tuia &amp; Alexander Mathis</p> <p>Affiliation: EPFL</p> <p>Date: January, 2024</p> <p>Link to the article: <a href="https://link.springer.com/article/10.1007/s11263-024-02026-6">https://link.springer.com/article/10.1007/s11263-024-02026-6</a></p> <p>--------------------------------</p> <p>WildCLIP is a fine-tuned CLIP model that allows to retrieve camera-trap events with natural language from the Snapshot Serengeti dataset. This project intends to demonstrate how vision-language models may assist the annotation process of camera-trap datasets.</p> <p>Here we provide the processed Snapshot Serengeti data used to train and evaluate WildCLIP, along with two versions of WildCLIP (model weights).</p> <p>Details on how to run these models can be found in the project <a href="https://github.com/amathislab/wildclip">github repository</a>.</p> <h2>Provided data (images and attribute annotations):&nbsp;</h2> <p>The data consists of 380 x 380 image crops corresponding to the MegaDetector output of Snapshot Serengeti with a confidence threshold above 0.7. We considered only camera trap images containing single individuals.</p> <p>A description of the original data can be found on LILA <a href="https://lila.science/datasets/snapshot-serengeti">here</a>, released under the <a href="https://cdla.dev/permissive-1-0/" rel="nofollow">Community Data License Agreement (permissive variant)</a>.</p> <p>We warmly thank the authors of LILA for making the MegaDetector outputs publicly available, as well as for structuring the dataset and facilitating its access.</p> <h2>Adapted CLIP model (model weights):&nbsp;</h2> <p>WildCLIP models provided:</p> <ul> <li><strong>[New] WildCLIP_vitb16_t1.pth:&nbsp;</strong>CLIP model with the ViT-B/16 visual backbone trained on data with captions following template 1. Trained on both base and novel vocabulary (see paper for details).</li> <li><strong>[New] WildCLIP_vitb16_t1_lwf.pth:&nbsp;</strong>CLIP model with the ViT-B/16 visual backbone trained on data with captions following template 1, and with the additional VR-LwF loss. Trained on both base and novel vocabulary (see paper for details).</li> <li><strong>WildCLIP_vitb16_t1_base.pth:</strong> CLIP model with the ViT-B/16 visual backbone trained on data with captions following template 1. Model used for evaluation and trained on base vocabulary only. (previously named <em>WildCLIP_vitb16_t1.pth</em>)</li> <li><strong>WildCLIP_vitb16_t1t7_lwf_base.pth</strong>: CLIP model with the ViT-B/16 visual backbone trained on data with captions following templates 1 to 7, and with the additional VR-LwF loss. Model used for evaluation and trained on base vocabulary only.&nbsp;(previously named <em>WildCLIP_vitb16_t1t7_lwf.pth</em>)</li> </ul> <p>We also provide the CSV files containing the train / val / test splits. The train / test splits follow camera split from LILA (https://lila.science/datasets/snapshot-serengeti). The validation split is custom, and also at the camera level.</p> <ul> <li><strong>train_dataset_crops_single_animal_template_captions_T1T7_ID.csv</strong>: Train set with captions from templates 1 through 7 (column "all captions") or template 1 only (column "template 1")</li> <li><strong>val_dataset_crops_single_animal_template_captions_T1T7_ID.csv</strong>: Validation set with captions from templates 1 through 7 (column "all captions") or template 1 only (column "template 1")</li> <li><strong>test_dataset_crops_single_animal_template_captions_T1T8T10.csv</strong>: Test set with captions from templates 1, 8, 9 and 10 (columns "all captions")</li> </ul> <p>Details on how the models were trained can be found in the associated&nbsp;<a href="https://link.springer.com/article/10.1007/s11263-024-02026-6" target="_blank" rel="noopener">publication</a>.</p> <h2>References:&nbsp;</h2> <p>If you find our code, or weights, please cite:</p> <pre>@article{gabeff2024wildclip, title={WildCLIP: Scene and animal attribute retrieval from camera trap data with domain-adapted vision-language models}, author={Gabeff, Valentin and Ru{\ss}wurm, Marc and Tuia, Devis and Mathis, Alexander}, journal={International Journal of Computer Vision}, pages={1--17}, year={2024}, publisher={Springer} }</pre> <p>If you use the adapted Snapshot Serengeti data please also cite their article:</p> <pre>@article{swanson2015snapshot, title={Snapshot Serengeti, high-frequency annotated camera trap images of 40 mammalian species in an African savanna}, author={Swanson, Alexandra and Kosmala, Margaret and Lintott, Chris and Simpson, Robert and Smith, Arfon and Packer, Craig}, journal={Scientific data}, volume={2}, number={1}, pages={1--14}, year={2015}, publisher={Nature Publishing Group} }</pre>

opencdla-permissive-1.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record