Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

36

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

36 results for “Self-supervised learning”

Learn how ShareScore rates datasets ↗
zenodo32/100

Self-Supervised Learning for Avian Diversity Monitoring SC22

<p><strong>Clusters&nbsp;</strong>The clusterization generated from the output of the pre-trained backbone</p> <p><strong>Morton Spectrograms</strong> The original spectrograms with which we trained the model to check the clusterization&nbsp;</p> <p><strong>Pretrained Model</strong> The pre-trained model</p> <p><strong>Single Image Attentional Maps</strong>&nbsp;The attentional maps and masks generated for a single image</p> <p><strong>features attentional&nbsp;maps and names</strong>&nbsp;Features, attentional maps and names of all the Spectrogram Images</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

Valentwin: Using Self-Supervised Contrastive Learning on Language Model for Schema Matching Datasets

<div>ValenTwin is a schema matching framework that uses self-supervised contrastive learning to train the model,&nbsp;uses the model to generate embeddings of table columns, then uses different similarity measures to match the column embeddings.</div> <div>&nbsp;</div> <div> <div>We provide two types of zip files for the datasets:<br>1. `data.zip` contains the raw data files, the ground truth files, the sampled data (n=[100, 200, 300, 400, 500] used in the experiments, as well as the contrastive data used to train the model.<br>2. `data-raw.zip` contains only the raw data files and the ground truth files. You can sample the data and generate the contrastive dataset yourself by following step 1 and 2 in the `How to Run` section. <br>Download and unzip one of the zip files to the `data` folder.</div> </div>

opencc-by-4.0May 2024View details →
zenodo32/100

images for self-supervised learning

<p>Real images with reduced resolution with associated masks generated from data obtained from High Frequency Receiver onboard Van Allen Probes.</p> <p>Sythetic images with associated masks generated based on statistics and radomlization.&nbsp;</p> <p>Inputs for the self-supervised contrastive learning software published at https://github.com/Yi-JiunSu/SSL-Contrastive</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Supplementary to "Generalizable biomarker prediction from cancer pathology slides with self-supervised deep learning - a retrospective multicentric study"

<p>High-resolution images of&nbsp;heatmaps. Top tiles and GradCam of figure 5</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Unlabeled Sentinel 2 time series dataset (training, T30TXT): Self-supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<strong> T30TXT unlabeled S2 dataset </strong></p> <p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30TXT</strong> are available. To download the full pretraining dataset, see : <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table> <p>&nbsp;</p>

openApr 2023View details →
zenodo32/100

Unlabeled Sentinel 2 time series dataset (training, T30TYS): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30TYS</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Structure-based self-supervised learning enables ultrafast prediction of stability changes upon mutation at the protein universe scale

<p>Pythia computed all single mutations of <em>E.coli</em> proteome, high quality high quality of Swiss-Prot structures and thermophilic proteins used in analysis.</p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

Improving Sub-pixel Accuracy in Ultrasound Localization Microscopy Using Supervised and Self-supervised Deep Learning

<p>These are&nbsp;the original PSFs used for generating the training, validation, and evaluation dataset for the paper &quot;Improving Sub-pixel Accuracy in Ultrasound Localization Microscopy Using Supervised and Self-supervised Deep Learning&quot;.</p>

opencc-by-4.0Aug 2023View details →
zenodo28/100

The impacts of active and self-supervised learning on efficient annotation of single-cell expression data - source data

<p>Source data used to create all figures in the manuscript.</p>

opencc-by-4.0Dec 2023View details →
zenodo28/100

Audio-visual self-supervised learning

<p>CVPR 2021 tutorial</p>

opencc-by-4.0Jul 2021View details →
geo24/100

Self-supervised learning for predicting transcriptomic groups on whole slides images in intrahepatic cholangiocarcinoma

GEO Series GSE244807. Homo sapiens. 246 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2023View details →
zenodo24/100

Self-supervised retinal thickness prediction enables deep learning from unlabeled data to boost classification of diabetic retinopathy

<p><strong>This data repository contains the OCT images and binary annotations for&nbsp;segmentation of retinal tissue using deep learning. To use, please refer to the Github repository&nbsp;</strong><a href="https://github.com/theislab/DeepRT">https://github.com/theislab/DeepRT</a>.</p> <p>&nbsp;</p> <p><strong>#######</strong></p> <p><strong>Access to large, annotated samples represents a considerable challenge for training accurate deep-learning models in medical imaging. While current leading-edge transfer learning from pre-trained models can help with cases lacking data, it limits design choices, and generally results in the use of unnecessarily large models. We propose a novel, self-supervised training scheme for obtaining high-quality, pre-trained networks from unlabeled, cross-modal medical imaging data, which will allow for creating accurate and efficient models. We demonstrate this by accurately predicting optical coherence tomography (OCT)-based retinal thickness measurements from simple infrared (IR) fundus images. Subsequently, learned representations outperformed advanced classifiers on a separate diabetic retinopathy classification task in a scenario of scarce training data. Our cross-modal, three-staged scheme effectively replaced 26,343 diabetic retinopathy annotations with 1,009 semantic segmentations on OCT and reached the same classification accuracy using only 25% of fundus images, without any drawbacks, since OCT is not required for predictions. We expect this concept will also apply to other multimodal clinical data-imaging, health records, and genomics data, and be applicable to corresponding sample-starved learning problems.</strong></p> <p><strong>#######</strong></p>

opencc-by-4.0Jan 2020View details →
zenodo24/100

Synthetic dataset for dual-perspective self-supervised learning

<p>The synthetic Ca datasets for training and testing,&nbsp;including training dataset with bidirectional collinear scan (N<sub>y</sub> = 2N<sub>x</sub>) for MP-SSL, training dataset with normal scan (N<sub>y</sub> = N<sub>x</sub>) for TP-SSL, and testing data (N<sub>y</sub> = N<sub>x</sub>)</p> <p>If you use these data simulated using our modified&nbsp;<a href="https://doi.org/10.1016/j.jneumeth.2021.109173">NAOMi</a> model, please cite the corresponding work:</p> <p><a href="https://doi.org/10.1186/s43074-023-00117-0"><strong><span>B. Shen</span></strong><span>, C. Luo, W. Pang, Y. Jiang, W. Wu, R. Hu, J. Qu, B. Gu, L. Liu. Surmounting photon limits and motion artifacts for biological dynamics imaging via dual-perspective self-supervised learning. PhotoniX 5, 1 (2024).&nbsp;</span></a></p>

openAug 2023View details →
zenodo24/100

Experimental dataset for dual-perspective self-supervised learning

<p>Experimental data for training and testing, including astrocyte data, rapid hemodynamic data (larger vessels), vascular data (smaller vessels), zebrafish cardiac data.</p> <p>If you use these data acquired using our imaging system, please cite the corresponding work:</p> <p><a href="https://doi.org/10.1186/s43074-023-00117-0"><span><span><span>&nbsp;</span></span></span><strong><span>B. Shen</span></strong><span>, C. Luo, W. Pang, Y. Jiang, W. Wu, R. Hu, J. Qu, B. Gu, L. Liu. Surmounting photon limits and motion artifacts for biological dynamics imaging via dual-perspective self-supervised learning. PhotoniX 5, 1 (2024). </span></a></p> <p>&nbsp;</p>

openAug 2023View details →
ClinicalTrials.gov20/100

AI-Based Self-Supervised Learning Model Using Non-Contrast Breast MRI for Early Screening and Clinical Utility Evaluation

ClinicalTrials.gov study NCT07205276. IPD Sharing: YES. Countries: 0. Publications: 0.

controlledIPD-YESFeb 2026View details →
zenodo16/100

Self-Supervised Learning Cell Image Dataset of Master Thesis "Enhancing Cell Instance Segmentation in 3D Microscopy using Self-Supervised ViTs"

<p>This is the self-supervised learning cell image dataset of master thesis "Enhancing Cell Instance Segmentation in 3D Microscopy using Self-Supervised ViTs". We gather images from datasets such as the LIVECell dataset, the EVICAN dataset, as well as datasets available on Image Data Resource (https://idr.openmicroscopy.org/) and the Broad Bioimage Benchmark Collection (https://bbbc.broadinstitute.org/). Only images with sizes larger than 512x512 are collected. For datasets containing more than 1000 images, we randomly select 1000 images. Otherwise, we retain all images in the dataset.&nbsp;</p> <p>&nbsp;</p> <p>After download, please put all compressed folders of subdatasets in the "image" folder under the root directory.</p>

restrictedcc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record