Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

48

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

48 results for “contrastive learning”

Learn how ShareScore rates datasets ↗
zenodo44/100

[Data] Qualify-As-You-Go: Sensor Fusion of Optical and Acoustic Signatures with Contrastive Deep Learning for Multi-Material Composition Monitoring in Laser Powder Bed Fusion Process

<p><br>Growing demand for multi-material Laser Powder Bed Fusion (LPBF) faces process control and quality monitoring challenges, particularly in ensuring precise material composition. This study explores optical and acoustic emission signals during LPBF processes with multiple materials, addressing challenges in process control and ensuring accurate material composition. Experimental data from processing five powder compositions were collected using a custombuilt monitoring system in a commercial LPBF machine. The research categorised signals from LPBF processing various compositions, enhancing prediction accuracy by combining optical with acoustic data and training convolutional neural networks using contrastive learning. Latent spaces of trained models using two contrastive loss functions, clustered acoustic and optical<br>emissions based on similarities, aligning with five compositions. Contrastive learning and sensor fusion were found to be essential for monitoring LPBF processes involving multiple materials. This research advances the understanding of multi-material LPBF, highlighting sensor fusion strategies&rsquo; potential for improving quality control in additive manufacturing. Data set for this work is hosted here</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Multimodala Dataset for multimodal contrastive learning for crop classification

<p>We developed this dataset using an existing dataset name DENETHOR developed by TUM <a href="https://openreview.net/forum?id=uUa4jNMLjrL">https://openreview.net/forum?id=uUa4jNMLjrL</a> to conduct our multi-modal contrastive learning experiments.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

LocScale-EMmerNet deep learning models for contrast optimisation of cryo-EM maps

<p>EMmerNet deep learning models for local optimisation of cryo-EM map contrast using <a href="https://gitlab.tudelft.nl/aj-lab/locscale">LocScale</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Short and long term impacts of Covid-19 on Older childreN's healTh-Related behAviours, learning and wellbeing STudy (CONTRAST) dataset

<p>The CONTRAST study explored how the Covid-19 (lockdown) restrictions affected lives of older children in the UK, particularly how they have influenced&nbsp;learning, eating, physical and other activities and wellbeing.</p>

opencc-by-nc-4.0Nov 2023View details →
zenodo40/100

Data from: Coronary artery segmentation in non-contrast calcium scoring CT images using deep learning

<p><strong>Abstract</strong></p> <p>Precise segmentation of coronary arteries in non-contrast Computed Tomography (CT) scans plays an important role in the assessment of the coronary artery disease, where it is the key component for evaluating the Calcium Score (Agatston et al. 1990). In the paper by Bujny et al. (2024), a deep-learning approach for high-precision segmentation of coronary arteries in non-contrast CT was proposed along with a novel method for generating Ground Truth (GT) test data (<em>test-GT</em>) via manual registration of high-resolution coronary tree models obtained based on contrast CT with the non-contrast CT scans. In this dataset, we present the inferences of the neural network model together with the corresponding <em>test-GT</em> samples, based on 6 CT scans from the openly available OrCaScore dataset (Wolterink et al. 2016). The geometrical models included in the dataset can be used both for inspection of the proposed deep learning model and for testing of new non-contrast coronary vessel segmentation approaches, which is a unique opportunity since, to the best of our knowledge, manual generation of GT for non-contrast coronary artery segmentation was not addressed so far due to very challenging character of this particular segmentation task.</p> <p>&nbsp;</p> <p><strong>Methods</strong></p> <p><strong><em>Manual Generation of test-GT</em></strong></p> <p>The geometric models of coronary arteries used for the evaluation of the proposed neural network model were generated according to the manual mesh-to-image registration process as described by Bujny et al. (2024). In this approach, the high-resolution coronary artery masks obtained based on contrast CT scans are manually aligned with the corresponding non-contrast CT images using tools available in the open-source 3D computer graphics software, Blender (<a href="https://www.blender.org/">https://www.blender.org/</a>). To ease the manual alignment process, specialized add-ons for medical image processing such as Cardiac add-on for Blender of Graylight Imaging (<a href="https://graylight-imaging.com/3d-modelling/">https://graylight-imaging.com/3d-modelling/</a>) can be used, as well. The STL models in this dataset were manually generated by a medical expert with 4 years of experience.</p> <p><strong><em>Segmentation of Coronary Arteries using a Deep Learning Model</em></strong></p> <p>For each of the cases presented in this dataset, we run an inference of an nnU-Net (Isensee et al. 2021) model trained according to the process described in our paper (Bujny et al. 2024). Since we use a standard nnU-Net, which utilizes a sliding window approach for processing of the CT scan, the context information within a patch is limited, which can lead to some false-positive detections. To mitigate this problem, we additionally post-process the inferences by eliminating small vessel fragments of less than 50 [mm^3] volume and structures outside of pericardium, which we segment using another nnU-Net model, SegTHOR (Lambert et al. 2020). The resulting geometric models are stored using the STL format and presented as green masks in the HTML reports with an embedded viewer based on the K3D-jupyter library (<a href="https://k3d-jupyter.org/">https://k3d-jupyter.org/</a>).</p> <p>&nbsp;</p> <p><strong>Dataset organization</strong></p> <p>The root folder contains 6 folders whose names correspond to the CT scans from the OrCaScore dataset (Wolterink et al. 2016). In each of the folders, there are the following 4 files available:</p> <ul> <li><span>&lsquo;manualGT_rater1.stl&rsquo; &ndash; high-resolution STL model of coronary arteries obtained via manual alignment of the geometric model segmented in contrast CT with the corresponding non-contrast CT scan by the first rater.</span>&nbsp;A sample belonging to the <em>test-GT</em> set (Bujny et al. 2024).</li> <li>&lsquo;manualGT_rater2.stl&rsquo; &ndash; corresponding <em>test-GT</em> sample by the second rater.</li> <li>&lsquo;ML.stl&rsquo; &ndash; post-processed inference of the nnU-Net ML model in the STL format.</li> <li>&lsquo;report.html&rsquo; &ndash; interactive HTML report consisting of a manually-aligned <em>test-GT</em> sample (red mask), the ML segmentation based on the non-contrast CT scan (green mask), and selected slices of the non-contrast CT scan. The reports contain the relevant information related to the scanning device and present the main segmentation quality metrics for the ML model inference.</li> </ul>

opencc-by-4.0May 2023View details →
zenodo40/100

Baseline embeddings from the BBBC022 dataset used in "Semisupervised contrastive learning for bioactivity prediction using Cell Painting image data"

<p>3 Baseline embeddings aclculated from the BBBC022 dataset. A self-supervised contrastive learning-based model, DINO and CellProfiler were used&nbsp; to calculate the embeddings.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Raw data for: Use the 4S (Signal-Safe Speckle Subtraction): Explainable Machine Learning reveals the Giant Exoplanet AF Lep b in High-Contrast Imaging Data from 2011

<p>This collection of data contains all raw data needed to reproduce the results in the paper:</p> <p>Use the 4S (Signal-Safe Speckle Subtraction): Explainable Machine Learning reveals the Giant Exoplanet AF Lep b in High-Contrast Imaging Data from 2011</p> <p>It can also be used as a demonstration dataset for our Python package fours.</p> <p>More details can be found in the online documentation of our python package:<br><a href="https://fours.readthedocs.io/en/latest/">https://fours.readthedocs.io/en/latest/</a></p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Intermediate results for: Use the 4S (Signal-Safe Speckle Subtraction): Explainable Machine Learning reveals the Giant Exoplanet AF Lep b in High-Contrast Imaging Data from 2011

<p>This collection contains all intermediate results needed to reproduce the results in the paper:</p> <p>Use the 4S (Signal-Safe Speckle Subtraction): Explainable Machine Learning reveals the Giant Exoplanet AF Lep b in High-Contrast Imaging Data from 2011</p> <p>You can use these intermediate results to create all plots in our paper without the need to run all experiments on a large cluster.</p> <p>More details can be found in the online documentation of our python package:<br><a href="https://fours.readthedocs.io/en/latest/">https://fours.readthedocs.io/en/latest/</a></p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Automatic Choroid Vascularity Index Calculation in Optical Coherence Tomography Images with Low Contrast Sclerocho-roidal Junction Using Deep Learning

<p>This project aims to calculate Choroid Vascularity Index (CVI) in optical coherenece tomography (OCT) images, using loss modified U-Net. The method is detailed in &quot;Automatic Choroid Vascularity Index Calculation in Optical Coherence Tomography Images low contrast sclerochoroidal junction Using Deep Learning&quot;. The dataset consists of&nbsp;Enhanced-depth imaging optical coherence tomography images from two patient groups.</p> <p>&bull; First dataset is including Raster OCT B-scans from patients with diabetic retinopathy.</p> <p>&bull; Second dataset is including EDI-HD OCT B-scans from patients with pachychoroid spectrum.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Modeling surface pCO2 variability in two contrasting basins of North Indian Ocean using advanced machine learning algorithms

<p>The dataset contains surface ocean <em>p</em>CO2, uncertainty and air-sea CO2 flux for the North Indian Ocean region. The data is available from 1993 to 2020 on a monthly time scale. Each of these data has a spatial resolution of 1/12&ordm;. Air-sea CO2 flux is calculated using a bulk parameterization, which is a function of wind speed. A positive CO2 flux value signifies CO2 outgassing, while a negative value indicates atmospheric CO2 uptake.&nbsp;</p> <p><strong>**It is recommended to use the latest version V3. Previous versions are depricated.</strong></p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Data for contrastive learning framework

<p>Data for contrastive learning framework, containing data&nbsp;for training and evaluation in two settings:&nbsp;detection of functionally equivalent programs on the&nbsp;<br> POJ-104 dataset, and the plagiarism detection task on the dataset of solutions to competitive programming contests held on the Codeforces platform. In both tasks, the datasets contain pairs of programs, labeled whether they are clones or not.</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

CAKE: clustering scRNA-seq data via combining contrastive learning with knowledge distillation

<p>All datasets used in our paper &quot;<strong>CAKE: clustering scRNA-seq data via combining contrastive learning with knowledge distillation&quot;</strong></p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Videos For: One-shot recognition of any material anywhere using contrastive learning with physics-based renderin

<p>Demonstration Video:</p> <p>Image Recognition of Materials States And Types From Single Example Identifying the state of the material in the main video by matching it to a few reference images (top). Each image contains a specific material state or type. The best match image is marked green. Based on the net for identifying similarity between any material state and type. The code for the net and the trained model used for this is available at: https://github.com/sagieppel/Contrastive-learning-for-one-shot-materials-and-textures-similarity-recognition-from-images Paper : One-shot recognition of any material anywhere using contrastive learning with physics-based rendering. https://arxiv.org/abs/2212.00648</p>

opencc-by-4.0Sep 2023View details →
zenodo32/100

Large-scale semantic indexing of Spanish biomedical literature using contrastive transfer learning

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo32/100

spatiAlign: An Unsupervised Contrastive Learning Model for Data Integration of Spatially Resolved Transcriptomics

<p>Integrative analysis of spatially resolved transcriptomics datasets empowers a deeper understanding of complex biological systems. However, integrating multiple tissue sections presents challenges for batch effect removal, particularly when the sections are measured by various technologies or collected at different times. Here, we propose spatiAlign, an unsupervised contrastive learning model that employs the expression of all measured genes and the spatial location of cells, to integrate multiple tissue sections. It enables the joint downstream analysis of multiple datasets not only in low-dimensional embeddings but also in the reconstructed full expression space. In benchmarking analysis, spatiAlign outperforms state-of-the-art methods in learning joint and discriminative representations for tissue sections, each potentially characterized by complex batch effects or distinct biological characteristics. Furthermore, we demonstrate the benefits of spatiAlign for the integrative analysis of time-series brain sections, including spatial clustering, differential expression analysis, and particularly trajectory inference that requires a corrected gene expression matrix.</p>

opencc-zeroJan 2024View details →
zenodo32/100

Change-Aware Sampling and Contrastive Learning for Satellite Images

<p><strong>Instructions</strong></p> <div>Change-Aware Sampling and Contrastive Learning for Satellite Images</div> <div>The 1 million sized dataset in compressed format.</div> <div>This dataset is split in 4 parts due to Zenodo's size restictions.</div> <div>Each part can be downloaded using the following link.</div> <p><strong>Part 1: </strong><a href="../records/10913216">https://zenodo.org/records/10913216</a></p> <p><strong>Part 2: </strong><a href="../records/10914902">https://zenodo.org/records/10914902</a></p> <p><strong>Part 3: </strong><a href="10915715">https://zenodo.org/uploads/10915715</a></p> <p><strong>Part 4: </strong><a href="../records/10916979">https://zenodo.org/records/10916979</a></p> <p>&nbsp;</p> <div>Use the following commands to combine and extract the compressed file.</div> <blockquote> <div>cat clean_1m_geography_part* &gt; clean_1m_geography.tar.gz</div> <div>tar -xvf clean_1m_geography.tar.gz</div> </blockquote> <p>&nbsp;</p> <p><strong>Paper Abstract</strong></p> <p>Automatic remote sensing tools can help inform many large-scale challenges such as disaster management, climate change, etc. While a vast amount of spatio-temporal satellite image data is readily available, most of it remains unlabelled. Without labels, this data is not very useful for supervised learning algorithms. Self-supervised learning instead provides a way to learn effective representations for various downstream tasks without labels. In this work, we leverage characteristics unique to satellite images to learn better self-supervised features. Specifically, we use the temporal signal to contrast images with long-term and short-term differences, and we leverage the fact that satellite images do not change frequently. Using these characteristics, we formulate a new loss contrastive loss called Change-Aware Contrastive (CACo) Loss. Further, we also present a novel method of sampling different geographical regions. We show that leveraging these properties leads to better performance on diverse downstream tasks. For example, we see a 6.5% relative improvement for semantic segmentation and an 8.5% relative improvement for change detection over the best-performing baseline with our method.<br><br><br></p> <div> <div>&nbsp;</div> </div>

opencc-by-4.0Jun 2023View details →
zenodo32/100

Contrast MR Based Radiomics and Machine Learning Analysis to assess clinical outcomes following liver resec-tion in Colorectal Liver Metastases: a preliminary study

<p>We uploaded the images of the manuscript &quot;Contrast MR Based Radiomics and Machine Learning Analysis to assess clinical outcomes following liver resection in Colorectal Liver Metastases: a preliminary study&quot; accepted on Cancers.</p> <p>Vincenza Granata1*, Roberta Fusco2, Federica De Muzio3, Carmen Cutolo4, Sergio Venanzio Setola1, Federica dell&rsquo; Aversana5, Alessandro Ottaiano6, Antonio Avallone6, Guglielmo Nasti6, Francesca Grassi5, Vincenzo Pilone4, Vittorio Miele7-8, Luca Brunese3, Francesco Izzo9, Antonella Petrillo1</p> <p>1Division of Radiology, &ldquo;Istituto Nazionale Tumori IRCCS Fondazione Pascale &ndash; IRCCS di Napoli&rdquo;, Naples, Italy</p> <p>2Medical Oncology Division, Igea SpA, Napoli, Italy</p> <p>3Department of Medicine and Health Sciences &ldquo;V. Tiberio&rdquo;, University of Molise, 86100 Campobasso, Italy</p> <p>4Department of Medicine, Surgery and Dentistry, University of Salerno, Salerno, Italy</p> <p>5Division of Radiology, &ldquo;Universit&agrave; degli Studi della Campania Luigi Vanvitelli&rdquo;, Naples, Italy</p> <p>.6Division of Abdominal Oncology, &ldquo;Istituto Nazionale Tumori IRCCS Fondazione Pascale &ndash; IRCCS di Napoli&rdquo;, Naples, Italy</p> <p>7Division of Radiology, &ldquo;Azienda Ospedaliera Universitaria Careggi&rdquo;, Florence, Italy</p> <p>8Italian Society of Medical and Interventional Radiology (SIRM), SIRM Foundation, via della Signora 2, 20122 Milan, Italy</p> <p>9Division of Epatobiliary Surgical Oncology,&ldquo;Istituto Nazionale Tumori IRCCS Fondazione Pascale &ndash; IRCCS di Napoli&rdquo;, Naples, Italy</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Spatial domains identification in spatial transcriptomics by domain knowledge-aware and subspace-enhanced graph contrastive learning

<p>We propose a graph contrastive learning framework, GRAS4T, which combines contrastive learning and subspace module to accurately distinguish different spatial domains by capturing tissue microenvironment through self-expressiveness of spots within the same domain. To uncover the pertinent features for spatial domain identification, GRAS4T employs a graph augmentation based on histological images prior, preserving information crucial for the clustering task. Experimental results on 8 ST datasets from 5 different platforms show that GRAS4T outperforms five state-of-the-art competing methods in spatial domain identification. Significantly, GRAS4T excels at separating distinct tissue structures and unveiling more detailed spatial domains. GRAS4T combines the advantages of subspace analysis and graph representation learning with extensibility, making it an ideal framework for ST domain identification.</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Valentwin: Using Self-Supervised Contrastive Learning on Language Model for Schema Matching Datasets

<div>ValenTwin is a schema matching framework that uses self-supervised contrastive learning to train the model,&nbsp;uses the model to generate embeddings of table columns, then uses different similarity measures to match the column embeddings.</div> <div>&nbsp;</div> <div> <div>We provide two types of zip files for the datasets:<br>1. `data.zip` contains the raw data files, the ground truth files, the sampled data (n=[100, 200, 300, 400, 500] used in the experiments, as well as the contrastive data used to train the model.<br>2. `data-raw.zip` contains only the raw data files and the ground truth files. You can sample the data and generate the contrastive dataset yourself by following step 1 and 2 in the `How to Run` section. <br>Download and unzip one of the zip files to the `data` folder.</div> </div>

opencc-by-4.0May 2024View details →
zenodo32/100

Contrastive Learning for Fine-Grained Ship Classification in Remote Sensing Images

<p>Dataset for Contrastive Learning for Fine-Grained Ship Classification in Remote Sensing Images from https://github.com/WindVChen/Push-and-Pull-Network?tab=readme-ov-file</p>

opencc-by-4.0Jun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record