Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

69

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

69 results for “Representation learning”

Learn how ShareScore rates datasets ↗
zenodo48/100

Transfer of sensorimotor learning reveals phoneme representations in preliterate children - Dataset

<p>This file provides formants values in each speaker and for each trial of the experiment described in the article : Transfer of sensorimotor learning reveals phoneme representations in preliterate children.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction - Datasets

<p>Datasets to NeurIPS 2021 accepted paper &quot;Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction&quot;.</p> <p>Datasets are pytorch files containing a dictionary with training, validation and test sets. Train, validation and test sets are custom dataset classes which inherit from the standard torch dataset class. Corresponding code an be found at https://github.com/HSG-AIML/NeurIPS_2021-Weight_Space_Learning.</p> <p>Datasets 41, 42, 43 and 44 are our dataset format wrapped around the zoos from Unterthiner et al, 2020 (https://github.com/google-research/google-research/tree/master/dnn_predict_accuracy)<br> <br> Abstract:<br> Self-Supervised Learning (SSL) has been shown to learn useful and information-preserving representations. Neural Networks (NNs) are widely applied, yet their weight space is still not fully understood. Therefore, we propose to use SSL to learn neural representations of the weights of populations of NNs. To that end, we introduce domain specific data augmentations and an adapted attention architecture. Our empirical evaluation demonstrates that self-supervised representation learning in this domain is able to recover diverse NN model characteristics. Further, we show that the proposed learned representations outperform prior work for predicting hyper-parameters, test accuracy, and generalization gap as well as transfer to out-of-distribution settings.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

CITRIS - Causal Representation Learning Datasets

<p>This repository contains the datasets from the paper &quot;CITRIS: Causal Identifiability from Temporal Intervened Sequences&quot; (<a href="https://arxiv.org/abs/2202.03169">link</a>) by Phillip Lippe,&nbsp;Sara Magliacane,&nbsp;Sindy L&ouml;we,&nbsp;Yuki M. Asano,&nbsp;Taco Cohen,&nbsp;Efstratios Gavves.</p> <p><strong>Temporal Causal3DIdent</strong>&nbsp;-&nbsp;The Temporal Causal3DIdent dataset is a collection of 3D object shapes, which are observed under varying positions, rotations, lightning, and colors. Overall, we this dataset contains&nbsp;7 (multidimensional) causal factors. The 7 shapes used are&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Armadillo</a>,&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Bunny</a>,&nbsp;<a href="https://www.cs.cmu.edu/~kmcrane/Projects/ModelRepository/#spot">Cow</a>,&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Dragon</a>,&nbsp;<a href="https://gfx.cs.princeton.edu/proj/sugcon/models/">Head</a>,&nbsp;<a href="https://www.cc.gatech.edu/projects/large_models/horse.html">Horse</a>, <a href="https://github.com/brendel-group/cl-ica">Teapot</a>.&nbsp;We provide two versions of the dataset: one that only contains images of the Teapot, and one that uses all 7 shapes. For more details on the dataset, see&nbsp;<a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p> <p><strong>Interventional Pong</strong>&nbsp;- The Interventional Pong environment is inspired by the game dynamics of Pong, where both paddles follow the policy of moving towards the ball, and the ball has slightly random movements. This dataset considers the 5 causal variables paddle left, paddle right, the ball position, the ball velocity, and the score. For more details on the dataset, see&nbsp;<a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p> <p><strong>Ball-in-Boxes</strong> -&nbsp;The Ball-in-Boxes is a simple dataset for showcasing the concept of the minimal causal variables. The system consists of a ball which randomly moves within a box, but only under an intervention can swap between the two boxes. Thereby, the intervention does not affect the x-position in the box. Thus, one can only discover the box assignment as a causal variable, and not whether the inner x-position also belongs to it.&nbsp;For more details on the dataset, see&nbsp;<a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

iCITRIS - Causal Representation Learning Datasets

<p>This repository contains the datasets from the paper &quot;iCITRIS: Causal Representation Learning for Instantaneous Temporal Effects&quot;&nbsp;(<a href="http://arxiv.org/abs/2206.06169">link</a>) by Phillip Lippe,&nbsp;Sara Magliacane,&nbsp;Sindy L&ouml;we,&nbsp;Yuki M. Asano,&nbsp;Taco Cohen,&nbsp;Efstratios Gavves.&nbsp;</p> <p><strong>Instantaneous Temporal Causal3Ident&nbsp;</strong>-&nbsp;The Temporal Causal3DIdent dataset is a collection of 3D object shapes, which are observed under varying positions, rotations, lightning, and colors. Overall, we this dataset contains&nbsp;7 (multidimensional) causal factors with instantaneous and temporal causal relations between them. The 7 shapes used are&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Armadillo</a>,&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Bunny</a>,&nbsp;<a href="https://www.cs.cmu.edu/~kmcrane/Projects/ModelRepository/#spot">Cow</a>,&nbsp;<a href="http://graphics.stanford.edu/data/3Dscanrep/">Dragon</a>,&nbsp;<a href="https://gfx.cs.princeton.edu/proj/sugcon/models/">Head</a>,&nbsp;<a href="https://www.cc.gatech.edu/projects/large_models/horse.html">Horse</a>,&nbsp;<a href="https://github.com/brendel-group/cl-ica">Teapot</a>.&nbsp;For more details on the dataset, see&nbsp;<a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p> <p><strong>Causal Pinball </strong>- The Causal Pinball environment implements the simplified, real-world game dynamics of Pinball. This dataset considers 5 causal variables with instantaneous effects: the paddle position left, the paddle position right, the ball (velocity and position), the state of all bumpers, and the score.&nbsp;For more details on the dataset as well as the code to generate this dataset, see&nbsp;<a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Supplementary run files for the paper "Learning Effective Representations for Retrieval using Self-Distillation with Adaptive Relevance Margins"

<p>TREC-Format run files of all trained models as supplementary material for the paper "Learning Effective Representations for Retrieval using Self-Distillation with Adaptive Relevance Margins".</p> <p>File naming follows the schema:&nbsp;<code>{model}-{loss variant}-{in-batch usage}-{dataset}.txt.gz</code></p>

opencc-by-4.0May 2024View details →
zenodo44/100

CNN Wild Park - Graph Neural Networks for Learning Equivariant Representations of Neural Networks

<p>This repository contains the <strong>CNN Wild Park</strong> dataset from the paper:</p> <blockquote> <p><strong>Graph Neural Networks for Learning Equivariant Representations of Neural Networks</strong><br><a href="https://mkofinas.github.io/">Miltiadis Kofinas</a>*,&nbsp;<a href="https://bknyaz.github.io/">Boris Knyazev</a>, <a href="https://www.cyanogenoid.com/">Yan Zhang</a>,&nbsp;<a href="https://yunlu-chen.github.io/">Yunlu Chen</a>,&nbsp;<a href="https://gertjanburghouts.github.io/">Gertjan J. Burghouts</a>,&nbsp;<a href="https://egavves.com/">Efstratios Gavves</a>,&nbsp;<a href="https://www.ceessnoek.info/">Cees G. M. Snoek</a>,&nbsp;<a href="https://davzha.netlify.app/">David W. Zhang</a>*<br><em>ICLR 2024</em> (oral)<br><a href="https://arxiv.org/abs/2403.12143">https://arxiv.org/abs/2403.12143</a><br><a href="https://github.com/mkofinas/neural-graphs">https://github.com/mkofinas/neural-graphs</a><br>*Joint first and last authors</p> </blockquote> <p>We introduce a new dataset of CNNs, which we term <em>CNN Wild Park</em>.<br>The dataset consists of 117,241 checkpoints from 2,800 CNNs, trained for up to 1,000 epochs on CIFAR10.<br>The CNNs vary in the number of layers, kernel sizes, activation functions, and residual connections between arbitrary layers.</p> <p>More specifically, we construct the CNN Wild Park dataset by training 2,800 small CNNs with different architectures for 200 to 1,000 epochs on CIFAR10. We retain a checkpoint of its parameters every 10 steps and also record the test accuracy. The CNNs vary by:</p> <ul> <li>Number of layers L in [2, 3, 4, 5] (note that this does not count the input layer).</li> <li>Number of channels per layer c_l in [4, 8, 16, 32].</li> <li>Kernel size of each convolution k_l in [3, 5, 7].</li> <li>Activation functions at each layer are one of ReLU, GeLU, tanh, sigmoid, leaky ReLU, or the identity function.</li> <li>Skip connections between two layers with at least one layer in between. Each layer can have at most one incoming skip connection. We allow for skip connections even in the case when the number of channels differ, to increase the variety of architectures and ensure independence between different architectural choices. We enable this by adding the skip connection only to the min(c_n, c_m) nodes.</li> </ul> <p>We divide the dataset into train/val/test splits such that checkpoints from the same run are <strong>not</strong> contained in both the train and test splits.&nbsp;</p> <div>&nbsp;</div> <div>&nbsp;</div>

opencc-by-4.0May 2024View details →
zenodo44/100

Genetic Variants Representation Learning (GV-Rep)

<p>This dataset is used for Genetic Variants (GV) representation learning.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Variational Inference for Learning Representations of Natural Language Edits

<p>Performance Evaluation of Edit Representations (PEER), the dataset we use in the paper <a href="https://arxiv.org/abs/2004.09143">&quot;Variational Inference for Learning Representations of Natural Language Edits&quot;</a>.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Learning Rich Representation of Keyphrases from Text

<p>In this work, we explore how to learn task-specific language models aimed towards learning rich representation of keyphrases from text documents. We experiment with different masking strategies for pre-training transformer language models (LMs) in discriminative as well as generative settings. In the discriminative setting, we introduce a new pre-training objective - Keyphrase Boundary Infilling with Replacement (KBIR), showing large gains in performance (up to 9.26 points in F1) over SOTA, when LM pre-trained using KBIR is fine-tuned for the task of keyphrase extraction. In the generative setting, we introduce a new pre-training setup for BART - KeyBART, that reproduces the keyphrases related to the input text in the CatSeq format, instead of the denoised original input. This also led to gains in performance (up to 4.33 points in F1@M) over SOTA for keyphrase generation. Additionally, we also fine-tune the pre-trained language models on named entity recognition (NER), question answering (QA), relation extraction (RE), abstractive summarization and achieve comparable performance with that of the SOTA, showing that learning rich representation of keyphrases is indeed beneficial for many other fundamental NLP tasks.</p> <p>As a part of this zip file we release the KBIR model which is continually pre-trained on RoBERTa-Large and also the KeyBART model which is continually pre-trained on BART-Large. Both these models can be used in place of a RoBERTa-Large or BART-Large model in PyTorch codebases and also with HuggingFace.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

PGB: A PubMed Graph Benchmark for Heterogeneous Network Representation Learning

<p>PubMed Graph Benchmark (PGB)&nbsp;aggregates&nbsp;the metadata associated with the biomedical articles from PubMed into a unified source.&nbsp;The benchmark contains metadata including&nbsp;title, abstract, authors, in/out citations, MeSH terms, MeSH hierarchy, venue, publication type, and chemicals.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Deep Representation Learning of Physical Activity and Sleep Patterns During Pregnancy Identifies post-hoc Inferences Associated with Prematurity

<p><strong>Running title</strong>: series2signal gestational age &quot;clock&quot; for pregnancy monitoring</p> <p><strong>Summary</strong>:&nbsp;</p> <p>Preterm birth (PTB) is the leading cause of infant mortality globally. While research has focused on the development of predictive models for PTB, cost-effective interventions have remained understudied. Physical activity and&nbsp;sleep present unique opportunities for interventions in low- and middle-income populations.&nbsp;However, objective&nbsp;measurement of physical activity and sleep remains challenging and self-reported metrics suffer from low-resolution and accuracy that decays over time. In this study, we use physical activity data collected using a wearable device&nbsp;comprising over 181,&nbsp;944 hours of data across&nbsp;N&nbsp;= 1,&nbsp;083 patients. Using a new state-of-the art deep learning time-series classification architecture, we first develop a &rdquo;clock&rdquo; of healthy dynamics in physical activity patterns during pregnancy by using gestational age (GA) as a surrogate for progression of pregnancy. We also developed a novel interpretability algorithm that integrates unsupervised clustering, model error analysis, feature attribution, and automated actigraphy analysis, allowing for model interpretation with respect to sleep, activity, and static clinical variables. Our model performs significantly better than 7 other machine learning and AI methods for modeling the progression of pregnancy based on measures of physical activity and sleep.</p> <p>Importantly, we found that deviations from this normal &rdquo;clock&rdquo; of physical activity and sleep changes during&nbsp;pregnancy are strongly associated with pregnancy outcomes. When our model underestimates GA, there are 0.52&nbsp;fewer preterm births than expected (P&nbsp;= 1.01e&nbsp;&minus;&nbsp;67) and when our model overestimates GA, there are 1.44 times&nbsp;(P&nbsp;= 2.82e&nbsp;&minus;&nbsp;39) more preterm births than expected. Model error is negatively correlated with interdaily stability&nbsp;(P&nbsp;= 0.043), indicating that our model assigns a more advanced GA when an individual&rsquo;s daily rhythms are less&nbsp;precise. Supporting this, our model attributes higher importance to sleep periods in predicting higher-than-actual&nbsp;GA, relative to lower-than-actual GA (P&nbsp;= 1.01e&nbsp;&minus;&nbsp;21).&nbsp;Combining prediction with interpretability allows us&nbsp;to robustly signal when activity behaviors increase or decrease the likelihood of preterm birth and advocates for the future development of clinical decision support through passive monitoring and suggestions around exercise&nbsp;habits and sleep patterns, which are easily implemented in low- and middle-income countries (LMICs).&nbsp;Beyond&nbsp;this particular application, the presented pipeline can be used to analyze high-fidelity time-series data in other translational studies utilizing wearable devices.</p> <p>&nbsp;</p> <p><strong>Data description (brief)</strong>: the raw wearables data is available as .mtn files with the GA encoded in the filename after the underscore. The processed data with sleep annotations can be loaded using the pickle module for serialized objects in python. See https://github.com/nealgravindra/wearables for examples.</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Small PASTIS training dataset config: Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>Files to run the small dataset experiments used in the preprint&nbsp; &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>. This .csv files enables to generate balanced small dataset from the <a href="https://zenodo.org/record/5012942#.ZFDfUJHP1H4">PASTIS dataset</a>. These files are required to run the experiment with a small training data-set, from the open source code <a href="https://src.koda.cnrs.fr/iris.dumeur/ssl_ubarn.git">ssl_ubarn</a>. In the .csv file name selected_patches_fold_{FOLD}_nb_{NSITS}_seed_{SEED}.csv :</p> <ul> <li>FOLD: id which corresponds to one of the 5 experiments run due to PASTIS K-fold.</li> <li>NSITS: Number of SITS selected to construct this training data-set</li> <li>SEED: the randomness used to create this small dataset</li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (validation): Self-supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only validation data</strong> are available. To download the full pretraining dataset, see <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UVU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T30TUVU): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30UVU</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T30TYQ): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30TYQ</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset : Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This repository list all the available repositories, to load the unlabeled Sentinel 2 (S2) L2A dataset used in the article<a href="https://ieeexplore.ieee.org/document/10414422/"> "Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series"</a>. This dataset is composed of patch time series acquired over France. For further details, see section IV.A of the pre-print article, available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks ['CLM_R1', 'EDG_R1', 'SAT_R1']. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <ul> <li>The validation dataset is available here : <a href="https://doi.org/10.5281/zenodo.7890452">10.5281/zenodo.7890452</a></li> <li>The training dataset is composed of 9 zenodo repositories, one for each S2 tiles. Here are the available repositories: <ul> <li>T31UEP<a href="http://https://doi.org/10.5281/zenodo.7899943"> 10.5281/zenodo.7899943</a></li> <li>T31TGJ <a href="https://doi.org/10.5281/zenodo.7899237">10.5281/zenodo.7899237</a></li> <li>T30TYS <a href="https://doi.org/10.5281/zenodo.7924193">10.5281/zenodo.7924193</a></li> <li>T31TFN <a href="https://doi.org/10.5281/zenodo.7896621">10.5281/zenodo.7896621</a></li> <li>T31TDL <a href="http://10.5281/zenodo.7896082">10.5281/zenodo.7896082</a></li> <li>T31TDJ <a href="https://doi.org/10.5281/zenodo.7895498">10.5281/zenodo.7895498</a></li> <li>T30UVU <a href="https://doi.org/10.5281/zenodo.7892410">10.5281/zenodo.7892410</a></li> <li>T30TYQ<a href="https://doi.org/10.5281/zenodo.7890542"> 10.5281/zenodo.7890542</a></li> <li>T30TXT <a href="https://doi.org/10.5281/zenodo.7875977">10.5281/zenodo.7875977</a></li> </ul> </li> </ul> <table> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T31TDJ): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T31TDJ</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T31TFN): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T31TFN</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T31TDL): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T31TDL</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0May 2023View details →
zenodo40/100

Unlabeled Sentinel 2 time series dataset (training, T31TGJ): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T31TGJ</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record