Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
36
datasets available to search
ShareScore release 0.9.0
Dataset results
36 results for “Self-supervised learning”
Self-Supervised Learning for Avian Diversity Monitoring SC22
<p><strong>Clusters </strong>The clusterization generated from the output of the pre-trained backbone</p> <p><strong>Morton Spectrograms</strong> The original spectrograms with which we trained the model to check the clusterization </p> <p><strong>Pretrained Model</strong> The pre-trained model</p> <p><strong>Single Image Attentional Maps</strong> The attentional maps and masks generated for a single image</p> <p><strong>features attentional maps and names</strong> Features, attentional maps and names of all the Spectrogram Images</p>
Valentwin: Using Self-Supervised Contrastive Learning on Language Model for Schema Matching Datasets
<div>ValenTwin is a schema matching framework that uses self-supervised contrastive learning to train the model, uses the model to generate embeddings of table columns, then uses different similarity measures to match the column embeddings.</div> <div> </div> <div> <div>We provide two types of zip files for the datasets:<br>1. `data.zip` contains the raw data files, the ground truth files, the sampled data (n=[100, 200, 300, 400, 500] used in the experiments, as well as the contrastive data used to train the model.<br>2. `data-raw.zip` contains only the raw data files and the ground truth files. You can sample the data and generate the contrastive dataset yourself by following step 1 and 2 in the `How to Run` section. <br>Download and unzip one of the zip files to the `data` folder.</div> </div>
images for self-supervised learning
<p>Real images with reduced resolution with associated masks generated from data obtained from High Frequency Receiver onboard Van Allen Probes.</p> <p>Sythetic images with associated masks generated based on statistics and radomlization. </p> <p>Inputs for the self-supervised contrastive learning software published at https://github.com/Yi-JiunSu/SSL-Contrastive</p>
Supplementary to "Generalizable biomarker prediction from cancer pathology slides with self-supervised deep learning - a retrospective multicentric study"
<p>High-resolution images of heatmaps. Top tiles and GradCam of figure 5</p>
Unlabeled Sentinel 2 time series dataset (training, T30TXT): Self-supervised Spatio-Temporal Representation Learning of Satellite Image Time Series
<p> <strong> T30TXT unlabeled S2 dataset </strong></p> <p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article "Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series" available <a href="https://hal.science/hal-04084839">here</a>. Each patch is constituted of the 10 bands [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks ['CLM_R1', 'EDG_R1', 'SAT_R1']. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30TXT</strong> are available. To download the full pretraining dataset, see : <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table> <p> </p>
Unlabeled Sentinel 2 time series dataset (training, T30TYS): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series
<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article "Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series" available <a href="https://hal.science/hal-04084839">here</a>. Each patch is constituted of the 10 bands [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks ['CLM_R1', 'EDG_R1', 'SAT_R1']. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30TYS</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>
Structure-based self-supervised learning enables ultrafast prediction of stability changes upon mutation at the protein universe scale
<p>Pythia computed all single mutations of <em>E.coli</em> proteome, high quality high quality of Swiss-Prot structures and thermophilic proteins used in analysis.</p>
Improving Sub-pixel Accuracy in Ultrasound Localization Microscopy Using Supervised and Self-supervised Deep Learning
<p>These are the original PSFs used for generating the training, validation, and evaluation dataset for the paper "Improving Sub-pixel Accuracy in Ultrasound Localization Microscopy Using Supervised and Self-supervised Deep Learning".</p>
The impacts of active and self-supervised learning on efficient annotation of single-cell expression data - source data
<p>Source data used to create all figures in the manuscript.</p>
Audio-visual self-supervised learning
<p>CVPR 2021 tutorial</p>
Self-supervised learning for predicting transcriptomic groups on whole slides images in intrahepatic cholangiocarcinoma
GEO Series GSE244807. Homo sapiens. 246 samples. Type: Expression profiling by high throughput sequencing.
Self-supervised retinal thickness prediction enables deep learning from unlabeled data to boost classification of diabetic retinopathy
<p><strong>This data repository contains the OCT images and binary annotations for segmentation of retinal tissue using deep learning. To use, please refer to the Github repository </strong><a href="https://github.com/theislab/DeepRT">https://github.com/theislab/DeepRT</a>.</p> <p> </p> <p><strong>#######</strong></p> <p><strong>Access to large, annotated samples represents a considerable challenge for training accurate deep-learning models in medical imaging. While current leading-edge transfer learning from pre-trained models can help with cases lacking data, it limits design choices, and generally results in the use of unnecessarily large models. We propose a novel, self-supervised training scheme for obtaining high-quality, pre-trained networks from unlabeled, cross-modal medical imaging data, which will allow for creating accurate and efficient models. We demonstrate this by accurately predicting optical coherence tomography (OCT)-based retinal thickness measurements from simple infrared (IR) fundus images. Subsequently, learned representations outperformed advanced classifiers on a separate diabetic retinopathy classification task in a scenario of scarce training data. Our cross-modal, three-staged scheme effectively replaced 26,343 diabetic retinopathy annotations with 1,009 semantic segmentations on OCT and reached the same classification accuracy using only 25% of fundus images, without any drawbacks, since OCT is not required for predictions. We expect this concept will also apply to other multimodal clinical data-imaging, health records, and genomics data, and be applicable to corresponding sample-starved learning problems.</strong></p> <p><strong>#######</strong></p>
Synthetic dataset for dual-perspective self-supervised learning
<p>The synthetic Ca datasets for training and testing, including training dataset with bidirectional collinear scan (N<sub>y</sub> = 2N<sub>x</sub>) for MP-SSL, training dataset with normal scan (N<sub>y</sub> = N<sub>x</sub>) for TP-SSL, and testing data (N<sub>y</sub> = N<sub>x</sub>)</p> <p>If you use these data simulated using our modified <a href="https://doi.org/10.1016/j.jneumeth.2021.109173">NAOMi</a> model, please cite the corresponding work:</p> <p><a href="https://doi.org/10.1186/s43074-023-00117-0"><strong><span>B. Shen</span></strong><span>, C. Luo, W. Pang, Y. Jiang, W. Wu, R. Hu, J. Qu, B. Gu, L. Liu. Surmounting photon limits and motion artifacts for biological dynamics imaging via dual-perspective self-supervised learning. PhotoniX 5, 1 (2024). </span></a></p>
Experimental dataset for dual-perspective self-supervised learning
<p>Experimental data for training and testing, including astrocyte data, rapid hemodynamic data (larger vessels), vascular data (smaller vessels), zebrafish cardiac data.</p> <p>If you use these data acquired using our imaging system, please cite the corresponding work:</p> <p><a href="https://doi.org/10.1186/s43074-023-00117-0"><span><span><span> </span></span></span><strong><span>B. Shen</span></strong><span>, C. Luo, W. Pang, Y. Jiang, W. Wu, R. Hu, J. Qu, B. Gu, L. Liu. Surmounting photon limits and motion artifacts for biological dynamics imaging via dual-perspective self-supervised learning. PhotoniX 5, 1 (2024). </span></a></p> <p> </p>
AI-Based Self-Supervised Learning Model Using Non-Contrast Breast MRI for Early Screening and Clinical Utility Evaluation
ClinicalTrials.gov study NCT07205276. IPD Sharing: YES. Countries: 0. Publications: 0.
Self-Supervised Learning Cell Image Dataset of Master Thesis "Enhancing Cell Instance Segmentation in 3D Microscopy using Self-Supervised ViTs"
<p>This is the self-supervised learning cell image dataset of master thesis "Enhancing Cell Instance Segmentation in 3D Microscopy using Self-Supervised ViTs". We gather images from datasets such as the LIVECell dataset, the EVICAN dataset, as well as datasets available on Image Data Resource (https://idr.openmicroscopy.org/) and the Broad Bioimage Benchmark Collection (https://bbbc.broadinstitute.org/). Only images with sizes larger than 512x512 are collected. For datasets containing more than 1000 images, we randomly select 1000 images. Otherwise, we retain all images in the dataset. </p> <p> </p> <p>After download, please put all compressed folders of subdatasets in the "image" folder under the root directory.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.