Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

36

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

36 results for “Self-supervised learning”

Learn how ShareScore rates datasets ↗
zenodo44/100

Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction - Datasets

<p>Datasets to NeurIPS 2021 accepted paper &quot;Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction&quot;.</p> <p>Datasets are pytorch files containing a dictionary with training, validation and test sets. Train, validation and test sets are custom dataset classes which inherit from the standard torch dataset class. Corresponding code an be found at https://github.com/HSG-AIML/NeurIPS_2021-Weight_Space_Learning.</p> <p>Datasets 41, 42, 43 and 44 are our dataset format wrapped around the zoos from Unterthiner et al, 2020 (https://github.com/google-research/google-research/tree/master/dnn_predict_accuracy)<br> <br> Abstract:<br> Self-Supervised Learning (SSL) has been shown to learn useful and information-preserving representations. Neural Networks (NNs) are widely applied, yet their weight space is still not fully understood. Therefore, we propose to use SSL to learn neural representations of the weights of populations of NNs. To that end, we introduce domain specific data augmentations and an adapted attention architecture. Our empirical evaluation demonstrates that self-supervised representation learning in this domain is able to recover diverse NN model characteristics. Further, we show that the proposed learned representations outperform prior work for predicting hyper-parameters, test accuracy, and generalization gap as well as transfer to out-of-distribution settings.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Data Set for 'Self-Supervised Machine Learning for Live Cell Imagery Segmentation'

<p><strong>Self-supervised machine learning code and data for segmenting live cell imagery (Matlab)</strong></p> <p><em>Running the Code</em></p> <p>SSL_Demo_2.m : main program for self-supervised machine learning segmentation</p> <p>SSL_Declumping_2.m : main program for declumping application (applied to output of SSL_Demo_2.m)</p> <p>This Matlab code is designed to be used with time-resolved live cell microscopy images (tiffs) for the automated segmentation of cells from background.</p> <p>It is recommended you first run this code with its accompanying demo data (included in this package), keeping the current directory structure.</p> <p>Simply open SSL_Demo_2.m or SSL_Declumping_2.m in Matlab and hit Run.</p> <p><em>Code Methodology</em></p> <p>The principle of self-supervised machine learning is that you simply load your images and Run - no parameter tuning needed, no training imagery required.</p> <p>Run from start to finish, the SSL_Demo_2.m code uses consecutive pairs of images to generate training data of &#39;cells&#39; and &#39;background&#39; via dynamic feature vectors based on optical flow (unsupervised). These self-labeled pixels are then used to generate static feature vectors (entropy, gradient), which in turn are used to train a classifier model. The training data is updated every image in order to automatically adapt to temporal changes in cell morphologies or background illumination.</p> <p>The code was tested for high fidelity segmentation using five different modes of light microscopy: transmitted light, DIC, phase contrast, fluorescence and interference reflection microscopy.</p> <p>Six different cell lines were imaged to cover a range of morphologies and phenotypic dynamics using three cameras of differing resolutions.</p> <p>The associated manuscript for this work can be found here (although the latest version is under peer review as of this writing):&nbsp;</p> <p><a href="https://www.biorxiv.org/content/10.1101/2021.01.07.425773v1">https://www.biorxiv.org/content/10.1101/2021.01.07.425773v1</a></p> <p>This code was tested on Matlab v2020a and v2021a using commercially available laptop computers running the Windows 10 operating system.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Self-supervised learning of seismological data reveals undocumented eruptive sequences at the Mayotte submarine volcano - Supplementary Materials

<p>The following files are shared:<br> &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; - The scripts used to train the model and generate the figures of the article<br> &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; - The input images used to train the model as well as the final outputs (embedding matrix and the associated filenames matrix)<br> &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; - The clusters organization with their associated images</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Small PASTIS training dataset config: Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>Files to run the small dataset experiments used in the preprint&nbsp; &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>. This .csv files enables to generate balanced small dataset from the <a href="https://zenodo.org/record/5012942#.ZFDfUJHP1H4">PASTIS dataset</a>. These files are required to run the experiment with a small training data-set, from the open source code <a href="https://src.koda.cnrs.fr/iris.dumeur/ssl_ubarn.git">ssl_ubarn</a>. In the .csv file name selected_patches_fold_{FOLD}_nb_{NSITS}_seed_{SEED}.csv :</p> <ul> <li>FOLD: id which corresponds to one of the 5 experiments run due to PASTIS K-fold.</li> <li>NSITS: Number of SITS selected to construct this training data-set</li> <li>SEED: the randomness used to create this small dataset</li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (validation): Self-supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only validation data</strong> are available. To download the full pretraining dataset, see <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UVU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T30TUVU): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30UVU</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T30TYQ): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30TYQ</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset : Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This repository list all the available repositories, to load the unlabeled Sentinel 2 (S2) L2A dataset used in the article<a href="https://ieeexplore.ieee.org/document/10414422/"> "Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series"</a>. This dataset is composed of patch time series acquired over France. For further details, see section IV.A of the pre-print article, available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks ['CLM_R1', 'EDG_R1', 'SAT_R1']. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <ul> <li>The validation dataset is available here : <a href="https://doi.org/10.5281/zenodo.7890452">10.5281/zenodo.7890452</a></li> <li>The training dataset is composed of 9 zenodo repositories, one for each S2 tiles. Here are the available repositories: <ul> <li>T31UEP<a href="http://https://doi.org/10.5281/zenodo.7899943"> 10.5281/zenodo.7899943</a></li> <li>T31TGJ <a href="https://doi.org/10.5281/zenodo.7899237">10.5281/zenodo.7899237</a></li> <li>T30TYS <a href="https://doi.org/10.5281/zenodo.7924193">10.5281/zenodo.7924193</a></li> <li>T31TFN <a href="https://doi.org/10.5281/zenodo.7896621">10.5281/zenodo.7896621</a></li> <li>T31TDL <a href="http://10.5281/zenodo.7896082">10.5281/zenodo.7896082</a></li> <li>T31TDJ <a href="https://doi.org/10.5281/zenodo.7895498">10.5281/zenodo.7895498</a></li> <li>T30UVU <a href="https://doi.org/10.5281/zenodo.7892410">10.5281/zenodo.7892410</a></li> <li>T30TYQ<a href="https://doi.org/10.5281/zenodo.7890542"> 10.5281/zenodo.7890542</a></li> <li>T30TXT <a href="https://doi.org/10.5281/zenodo.7875977">10.5281/zenodo.7875977</a></li> </ul> </li> </ul> <table> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T31TDJ): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T31TDJ</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T31TFN): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T31TFN</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T31TDL): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T31TDL</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T31TGJ): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T31TGJ</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T31UEP): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T31UEP</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0May 2023View details →
dryad40/100

Classifications of auroral phenomena in THEMIS All-Sky images obtained via self-supervised learning

Open the record for dataset details and reuse information.

publicNov 2024View details →
zenodo36/100

[Data] Real-time monitoring and quality assurance for laser-based directed energy deposition: integrating co-axial imaging and self-supervised deep learning framework

<p>The experimental setup utilized a co-axial color Charged Couple Device (CCD) camera, integrated into the laser deposition head. This camera operates at a frame rate of 30 frames per second and captures the morphology of the process area. The captured images consist of three RGB channels with a 640&thinsp;&times;&thinsp;480 pixels resolution. To enable the camera to capture the radiation from the process zone, a beam splitter is installed on Precitec's laser applicator head. An optical notch filter within the 650&ndash;675 nm range also blocks the laser wavelengths.</p> <p>The dataset consists of four categories that covers the process map of DED process [.rar file].<br>The dataset consist of around 48,000 images that are labelled into 4 categories [P1-P2-P3-P4]. The images correspond to DED process zone captured co-axially<br>The categories are function of linear laser energy deposited. The folder is already split into Train and Test.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

RS3L: A jet tagging dataset for self-supervised learning based on re-simulation

<p>Jet tagging dataset used for "Re-Simulation-based Self-Supervised Learning" (RS3L, <a href="https://arxiv.org/abs/2403.07066">arXiv:2403.07066</a>). Simulated partons are re-showered with various parton shower configurations and reconstructed with the Delphes3 detector software.&nbsp;<br><br>The <code>singletons</code> array contains information about the jet and has dimensions <code>N_examples x N_augmentations x N_singletons</code>. The augmenations are ordered by: <br><br>1. nominal scenario: jet showered with Pythia8<br>2. changing the numerical seed in Pythia8<br>3. changing the scale controlling the probability for final state radiation by 1/sqrt(2)<br>4. changing the scale controlling the probability for final state radiation by sqrt(2)<br>5. using Herwig7 as parton shower</p> <p>The variables stored in the <code>singletons</code> array are:</p> <p><code>singletons = ['jettype','parton1_pt','parton1_eta','parton1_phi','parton1_e','parton2_pt','parton2_eta','parton2_phi','parton2_e', 'jet_pt','jet_eta','jet_phi','jet_e','jet_msd','jet_n2']</code><br><br></p> <p>For gluon-initiated jets,&nbsp;<code>parton1</code> and <code>parton2</code> are the two gluon daughters (see below). For any other jet type, parton1 is the main particle (W, H, Z, or single quark), and parton2 is filled with 0s.</p> <p>The <code>jet_pflow_cands</code> contains the particles clustered into the jet. The array has a shape <code>N_examples x N_augmentations x N_features x N_particles</code>.<br><br>The features per particle are:</p> <p><code>features = ['pt','relpt','eta','phi','dr','e','rele','charge','pdgid','d0','dz']</code><br><br>The <code>jettype</code> gives the parton at the origin of the jet:&nbsp;<br><br><code>jettype: <br>1: q<br>2: c<br>3: b<br>4: H-&gt;bb<br>5: g-&gt;qq<br>6: g-&gt;cc<br>7: g-&gt;bb<br>8: g-&gt;gg<br>9: W-&gt;two quarks<br>10: Z-&gt;qq<br>11: Z-&gt;bb</code></p> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Three-dimensional Unfolding and Unfaulting Dataset for Structural Interpretation Using Self-Supervised Learning

<p>This dataset accompanies the study "Three-dimensional unfolding and unfaulting for structural interpretation using self-supervised learning." It provides synthetic data used to develop and validate a novel 3-D lightweight neural network framework for restoring deformed geological structures to their flattened state. The data include highly deformed examples with complex faulting and folding, designed to test the model's ability to compute precise shifts that realign stratigraphic layers.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

scPretrain: Multi-task self-supervised learning for cell type classification

<p>The dataset and code for paper, scPretrain: Multi-task self-supervised learning for cell type classification.</p>

opencc-by-4.0Nov 2020View details →
zenodo36/100

Multi-task self-supervised learning for wearables - human activity recognition

<p>Datasets used to train and evaluated the self-supervised-learning model</p>

opencc-by-4.0May 2022View details →
zenodo32/100

[Data] Self-Supervised Bayesian Representation Learning of Acoustic Emissions from Laser Powder Bed Fusion Process for In-situ Monitoring

<div> <div> <div> <p>Different Laser Powder Bed Fusion (LPBF) process spaces were deliberately introduced by employing two distinct 316L stainless steel powder distributions (with particle sizes &gt;45 &mu;m and &lt; 45 &mu;m) and processing them with two sets of laser parameters, resulting in the creation of four datasets [D1, D2, D3, and D4]. These datasets encompass LoF pores, conduction mode, and keyhole formations, each associated with three LPBF regimes denoted as D1, D2, D3, and D4. The experiments utilized a Sisma MYSINT 100 commercial LPBF printer and an airborne AE sensor system with a flat frequency response ranging from 0 to 150 kHz.&nbsp;Validation of the ground truths for the three laser regimes across the four datasets, representing distinct process spaces, was accomplished through the confirmation of cross-sectional images. In the course of fabricating a cube using a powder bed and laser, data acquisition from an AE sensor was triggered when the optical intensity reached a threshold of 0.5 V for each scan length. The photodiode trigger gain was adjusted to saturate at 5 V, and the ensuing continuous-time window, where the optical signal remained at 5 V for 12.5 ms, was calculated and segmented to generate the dataset.&nbsp;Irrespective of the specific regime (Lack of Fusion, Conduction, and Keyhole) or the cube being fabricated (with two powder distributions), the signals obtained during this process were then segmented into a 12.5 ms window comprising 5000 data points. To eliminate any noise, an offline application of a low-pass Butterworth filter with a 150 kHz cut-off frequency was employed, aligned with the frequency response specification of the AE sensor. Each dataset has two files against it [raw/groundtruth label].</p> </div> </div> </div>

opencc-by-4.0Nov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record