Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
101
datasets available to search
ShareScore release 0.9.0
Dataset results
101 results for “Supervised learning”
Datasets for evaluating scalable supervised learning for synthesize-on-demand chemical libraries
<p>This repository contains datasets for the manuscript "Evaluating scalable supervised learning for synthesize-on-demand chemical libraries":</p> <ul> <li><strong>ams_all_preds.csv.gz</strong>: The AMS dataset predictions when using an RF or baseline model trained on the training dataset. Includes the predicted score and rank from each model for each compound. We started with 8,434,707 AMS compounds and detected that 247,025 were in the LC or MLPCN training data. These were removed from the AMS list, leaving 8,187,682 compounds to score. The compound matching was done on the SMILES that we canonicalized in rdkit.</li> <li><strong>ams_order_results.csv.gz</strong>: Information about the 1,024 compounds purchased from the AMS library. Excludes the 4 AMS compounds that were incompletely dissolved. Includes the chemical feature representation, information from the vendor, RF and baseline model predictions, screening results, and clustering results.</li> <li><strong>baseline_weight.npy</strong>: The saved Similarity Baseline model, which consists of the active compounds in the training data. This model was used to score the AMS library. See the <a href="https://github.com/gitter-lab/pria-ams-enamine">GitHub repository</a> for code to load the model and make predictions on new compounds.</li> <li><strong>cdd_training_data.tar.gz</strong>: The LC1234 and MLPCN PriA-SSB screening data exported from CDD.</li> <li><strong>enamine_costs_clustered_v3_with_nneighbor.csv.gz</strong>: Contains 5,620 Enamine compounds that were selected based on the RF prediction score and availability. This file also contains the Taylor-Butina cluster ID when clustering the training compounds, 1,024 tested AMS compounds, and top-ranked Enamine compounds at a 0.4 threshold. The nearest neighbor compounds in the training and AMS sets are also included along with compound information from Enamine, RF model scores, and chemical feature representations.</li> <li><strong>enamine_dose_response_curve_plots.xlsx</strong>: Images of the dose response curves from all three runs on the 68 Enamine compounds. If a compound was tested multiple times, multiple curves are shown in the same plot. The compound structure images and SMILES are exported from CDD, not generated with RDKit.</li> <li><strong>enamine_dose_response_curves.tsv</strong>: The dose response curve summaries from all three runs on the 68 Enamine compounds. If a compound was tested multiple times, only the highest-quality dose response curve was used.</li> <li><strong>enamine_final_list.csv.gz</strong>: The final 100 filtered compounds from <code>enamine_top_10000.csv.gz</code>. Contains compound information from Enamine as well as RF model scores, chemical feature representations, and clustering results.</li> <li><strong>enamine_PriA-SSB_dose_response_data.tar.gz</strong>: The dose response screening data from all three runs on the 68 Enamine compounds. The 2021-06-16 run was originally screened on 2020-08-24. 2021-06-16 is the date the compound identities were corrected. This run contains two 1,536 well plates.</li> <li><strong>enamine_top_10000.csv.gz</strong>: Top 10,000 predictions from the Enamine REAL dataset using the selected RF model. Contains compound information from Enamine as well as RF model scores, chemical feature representations, and clustering results.</li> <li><strong>master_df.csv.gz</strong>: The output of preprocessing the files in <code>cdd_training_data.tar.gz</code>. Contains 441,900 rows.</li> <li><strong>random_forest_classification_139.pkl</strong>: The saved RF classification model with hyperparameter ID 139. This model was used to score the AMS and Enamine REAL libraries. See the <a href="https://github.com/gitter-lab/pria-ams-enamine">GitHub repository</a> directory for code to load the model and make predictions on new compounds.</li> <li><strong>train_ams_real_cluster.csv.gz</strong>: Contains cluster IDs for Taylor-Butina clustering at a 0.4 threshold applied to the training compounds, 1,024 tested AMS compounds, and top-ranked compounds from Enamine. Includes the chemical features, dataset to which the compound belongs, leader compound for each cluster, and whether the compound is a known hit.</li> <li><strong>training_df_single_fold.csv.gz</strong>: This is all ten folds in <code>training_folds.tar.gz</code> merged for convenience. Contains 427,300 compounds.</li> <li><strong>training_df_single_fold_with_ams_clustering.csv.gz</strong>: Contains cluster IDs for Taylor-Butina clustering applied to the 427,300 training compounds and the 1,024 tested AMS compounds. Different clustering results are shown at the 0.2, 0.3, and 0.4 thresholds. Includes the leader compound for each cluster. Although the training and AMS compounds were clustered jointly, only the training compounds' clusters are shown. The AMS compounds' clusters are in <code>ams_order_results.csv.gz</code>.</li> <li><strong>training_folds.tar.gz</strong>: The LC1234 and MLPCN training data split into ten folds. This dataset with 427,300 compounds was used for cross validation and model selection. This dataset is derived from <code>master_df.csv.gz.</code></li> </ul> <p>If you use these datasets in a publication, please cite:</p> <p>Moayad Alnammi, Shengchao Liu, Spencer S. Ericksen, Gene E. Ananiev, Andrew F. Voter, Song Guo, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter. <a href="https://doi.org/10.1021/acs.jcim.3c00912">Evaluating scalable supervised learning for synthesize-on-demand chemical libraries</a>. <em>Journal of Chemical Information and Modeling</em> 2023.</p> <p>See PubChem AID <a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1272365">1272365</a>, AID <a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1918986">1918986</a>, and the associated publications for details about the PriA-SSB screening data. The screening datasets were compiled from three separate sources that should all be cited if the training dataset is used in a publication:</p> <ul> <li>Moayad Alnammi, Shengchao Liu, Spencer S. Ericksen, Gene E. Ananiev, Andrew F. Voter, Song Guo, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter. <a href="https://doi.org/10.1021/acs.jcim.3c00912">Evaluating scalable supervised learning for synthesize-on-demand chemical libraries</a>. <em>Journal of Chemical Information and Modeling</em> 2023.</li> <li>Shengchao Liu<sup>+</sup>, Moayad Alnammi<sup>+</sup>, Spencer S. Ericksen, Andrew F. Voter, Gene E. Ananiev, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter. <a href="https://doi.org/10.1021/acs.jcim.8b00363">Practical model selection for prospective virtual screening</a>. <em>Journal of Chemical Information and Modeling</em> 2018.</li> <li>Andrew F. Voter<sup>+</sup>, Michael P. Killoran<sup>+</sup>, Gene E. Ananiev, Scott A. Wildman, F. Michael Hoffmann, James L. Keck. <a href="https://doi.org/10.1177/2472555217712001">A high-throughput screening strategy to identify inhibitors of SSB protein–protein interactions in an academic screening facility</a>. <em>SLAS Discovery</em> 2018.</li> </ul> <ul> </ul>
Datasets for "Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Classification Algorithms in Reiner Gamma and Mare Ingenii"
<p>Final surface reflectance data at 2.6 m/pixel resolution with floating point values are available as GeoTiff and ASCII text files. Definition files for the K-Means and MLC algorithms in classifying swirl units are also available as ASCII text files. See README file for further details.</p> <p>Data used in the research article:</p> <p>Chuang, F.C., M.D. Richardson, J.R. Weirich, A.A. Sickafoose, and D.L. Domingue, 2022. Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Image Classification Algorithms in Reiner Gamma and Mare Ingenii. The Planetary Science Journal, 3:231. doi://10.3847/PSJ/ac8f43</p> <p> </p>
Datasets for Supervised Learning Model Predicts Protein Adsorption to Carbon Nanotubes
<p>All used Datasets to pair with "Supervised Learning Model Predicts Protein Adsorption to Carbon Nanotubes" by Nicholas Ouassil*, Rebecca L. Pinals*, Jackson Travis Del Bonis-O'Donnell, Jeffrey W. Wang, and Markita P. Landry</p> <p>*Co-authors</p>
Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction - Datasets
<p>Datasets to NeurIPS 2021 accepted paper "Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction".</p> <p>Datasets are pytorch files containing a dictionary with training, validation and test sets. Train, validation and test sets are custom dataset classes which inherit from the standard torch dataset class. Corresponding code an be found at https://github.com/HSG-AIML/NeurIPS_2021-Weight_Space_Learning.</p> <p>Datasets 41, 42, 43 and 44 are our dataset format wrapped around the zoos from Unterthiner et al, 2020 (https://github.com/google-research/google-research/tree/master/dnn_predict_accuracy)<br> <br> Abstract:<br> Self-Supervised Learning (SSL) has been shown to learn useful and information-preserving representations. Neural Networks (NNs) are widely applied, yet their weight space is still not fully understood. Therefore, we propose to use SSL to learn neural representations of the weights of populations of NNs. To that end, we introduce domain specific data augmentations and an adapted attention architecture. Our empirical evaluation demonstrates that self-supervised representation learning in this domain is able to recover diverse NN model characteristics. Further, we show that the proposed learned representations outperform prior work for predicting hyper-parameters, test accuracy, and generalization gap as well as transfer to out-of-distribution settings.</p>
Weakly Supervised Learning for Industrial Optical Inspection
<p><strong>Abstract</strong></p> <p>In the following, we present a synthetic benchmark corpus for detect detection on statistically textured surfaces.We hope that it facilitates to further develop and benchmark classification algorithms for applications of industrial optical inspection. All data is publicly available and can be downloaded from this page.</p> <p><strong>Competition at DAGM 2007 symposium</strong></p> <p>The <a href="https://www.dagm.de/">DAGM (Deutsche Arbeitsgemeinschaft für Mustererkennung e.V., German chapter of the IAPR (International Association for Pattern Recognition))</a> and the <a href="http://www.gnns.de/">GNSS (German Chapter of the European Neural Network Society)</a> offered an open competition on <em>Weakly Supervised Learning for Industrial Optical Inspection</em> held as part of the DAGM symposium in 2007.<br><br>The competition was inspired by the fact that automated optical inspection allows to reduce the cost of industrial quality control significantly. The competitors had to design a classification algorithm which:</p> <ul> <li>detects miscellaneous defects on various statistically textured backgrounds.</li> <li>learns to discern defects automatically from a weakly labelled training data.</li> <li>works on data whose exact characteristics are unknown at development time.</li> <li>adapts all parameters automatically and does not require any human intervention.</li> <li>has a moderate running time (in this competition 24 hours for training and 12 hours for the test phase).</li> <li>takes into account asymmetric costs for false positive and false negative decisions (1:20 was used for the competition).</li> </ul> <p><strong>Data description</strong></p> <p>Preview Image: <a href="../api/iiif/record:12750201:examples_small.jpg/full/!800,800/0/default.jpg" target="_blank" rel="noopener">https://zenodo.org/api/iiif/record:12750201:examples_small.jpg/full/!800,800/0/default.jpg</a></p> <p>The data is artificially generated, but similar to real world problems. The first six out of ten datasets, denoted as development datasets, are supposed to be used for algorithm development. The remaining four datasets, which are referred to as competition datasets, can be used to evaluate the performance. Researchers should consider not using or analyzing the competition datasets before the development is completed as a code of honour.<br>In the following we provide some details about the datasets:</p> <ul> <li>Each development (competition) dataset consists of 1000 (2000) 'non-defective' and of 150 (300) 'defective' images saved in grayscale 8-bit PNG format.</li> <li>Each dataset is generated by a different texture model and defect model.</li> <li>'Non-defective' images show the background texture without defects, 'defective' images have exactly one labelled defect on the background texture.</li> <li>All datasets has been randomly split into a training and testing sub-dataset of equal size.</li> <li>Weak labels are provided as ellipses roughly indicating the defective area. Technically, defective images are augmented with a separate grayscale 8-bit image in the PNG format located in a folder 'Label'. The values 0 and 255 denote background and defective area, respectively.</li> </ul> <p>All meta-data is subsumed in a separate ASCII textfile called 'Labels.txt' which is located in the 'Label' folder. The structure is as follows:<br>1 \n<br>[id of item no. 1] \t [0 if non-defective, 1 if defective] \t [filename of raw image no. 1] \t 0 \t [filename of label image no. 1 if defective, 0 otherwise] \n<br>...<br>[id of item no. N] \t [0 if non-defective, 1 if defective] \t [filename of raw image no. N] \t 0 \t [filename of label image no. N if defective, 0 otherwise] \n</p>
Data Set for 'Self-Supervised Machine Learning for Live Cell Imagery Segmentation'
<p><strong>Self-supervised machine learning code and data for segmenting live cell imagery (Matlab)</strong></p> <p><em>Running the Code</em></p> <p>SSL_Demo_2.m : main program for self-supervised machine learning segmentation</p> <p>SSL_Declumping_2.m : main program for declumping application (applied to output of SSL_Demo_2.m)</p> <p>This Matlab code is designed to be used with time-resolved live cell microscopy images (tiffs) for the automated segmentation of cells from background.</p> <p>It is recommended you first run this code with its accompanying demo data (included in this package), keeping the current directory structure.</p> <p>Simply open SSL_Demo_2.m or SSL_Declumping_2.m in Matlab and hit Run.</p> <p><em>Code Methodology</em></p> <p>The principle of self-supervised machine learning is that you simply load your images and Run - no parameter tuning needed, no training imagery required.</p> <p>Run from start to finish, the SSL_Demo_2.m code uses consecutive pairs of images to generate training data of 'cells' and 'background' via dynamic feature vectors based on optical flow (unsupervised). These self-labeled pixels are then used to generate static feature vectors (entropy, gradient), which in turn are used to train a classifier model. The training data is updated every image in order to automatically adapt to temporal changes in cell morphologies or background illumination.</p> <p>The code was tested for high fidelity segmentation using five different modes of light microscopy: transmitted light, DIC, phase contrast, fluorescence and interference reflection microscopy.</p> <p>Six different cell lines were imaged to cover a range of morphologies and phenotypic dynamics using three cameras of differing resolutions.</p> <p>The associated manuscript for this work can be found here (although the latest version is under peer review as of this writing): </p> <p><a href="https://www.biorxiv.org/content/10.1101/2021.01.07.425773v1">https://www.biorxiv.org/content/10.1101/2021.01.07.425773v1</a></p> <p>This code was tested on Matlab v2020a and v2021a using commercially available laptop computers running the Windows 10 operating system.</p>
Phononic crystals dataset for supervised training of surrogate deep learning model
<p>The dataset contains shapes of unit cells of phononic crystals (inputs) in the form of images and corresponding dispersion diagrams (outputs). The dataset is used for deep learning (DL) model training.<br> Outputs are in the form of .mat files which contain vectors of reduced wavevector and corresponding frequencies, and also displacements u, v, w which can be used for polarization calculation.</p> <p>The dataset contains 11000 cases.</p> <p>Note: Ignore names "labels" as these are actually inputs to the DL model, not labels.</p>
Self-supervised learning of seismological data reveals undocumented eruptive sequences at the Mayotte submarine volcano - Supplementary Materials
<p>The following files are shared:<br> - The scripts used to train the model and generate the figures of the article<br> - The input images used to train the model as well as the final outputs (embedding matrix and the associated filenames matrix)<br> - The clusters organization with their associated images</p>
A didactical dataset to learn supervised classification with candy
<h2>A didactical dataset to learn supervised classification</h2><p>It was obtained from university level students measuring candy that was mixed and distributed in bowls to them. The goal of this dataset creation was to expose the students to the data taking process. Further, the dataset is meant for classification.</p><h3>Dataset Structure</h3><p>The dataset consists of 6 csv files:</p><ul><li><strong>peanuts.csv</strong> represents the entire dataset (a concatenation of all group?.csv files) omitting the sample column</li><li><strong>peanuts_all.csv</strong> represents the entire dataset (a concatenation of all group?.csv files)</li><li>files matching <strong>group[1-5].csv </strong>represent the measurements of each group</li></ul><h3>Data Representation</h3><p>Each file contains 5 columns. </p><ul><li>color, int values, 0: white, 1: black, 2: brown, 3: other</li><li>shape, int values, 0: irregular, 1 round, 2: lens-like</li><li>height, float values, in millimeter</li><li>width, float values, in millimeter</li><li>label, category, peanut/nopeanut</li></ul><p>For more information on the didactical background, see the <a href="https://proceedings.mlr.press/v141/huppenkothen21a.html">original publication</a> that presented the concept for this activity.</p>
Database used for the supervised learning of few dirty bosons with variable particle number
<p>This dataset includes files with points of different disordered potentials and their corresponding ground state energies depending on the number of particles of the system (N=1,2,3,4). It has been used to train a neural network. </p>
Datasets for CASSL: A cell-type annotation method for single cell transcriptomics data using semi-supervised learning
<p>This repository contains datasets used in the project CASSL: A cell-type annotation method for single cell transcriptomics data using semi-supervised learning. This project aims at learning cell annotations for missing cell labels via NMF and recursive k-Means clustering.</p>
Automated metabolic assignment: Semi-supervised learning in metabolic analysis employing two dimensional Nuclear Magnetic Resonance (NMR)
<p>This dataset is related to the paper <strong>“Automated metabolic assignment: Semi-supervised learning in metabolic analysis employing two dimensional Nuclear Magnetic Resonance (NMR)”.</strong></p> <p>https://www.sciencedirect.com/science/article/pii/S2001037021003792?via%3Dihub</p> <p>The dataset comprises horizontal and vertical frequencies of 2D NMR TOCSY of breast cancer-tissue sample. 2D TOCSY was acquired by employing a broadband high resolution 600.13 MHz (B0 = 14.1 T) NMR Bruker spectrometer (AVANCE III 600 with the Bruker magnet ASCEND 600) supported with the room temperature probe (BBO model-Bruker) and Magic Angle Spinning (MAS) probehead. 1D and 2D NMR spectra acquisition and processing were achieved by using the TopSpin software package 3.6.</p> <p>There are two files:</p> <p><strong>BreastCancerMetabolites.csv:</strong></p> <p>First column: numerical labels of the metabolites. Each number represent a metabolite. In total, there are 27 metabolites with multiple multiplets per metabolite.</p> <p>Second and third column: Horizontal and vertical frequencies for each metabolite.</p> <p><strong>Labels.csv:</strong></p> <p>The corresponding metabolites names.</p> <p> </p>
Assessing predictive performance of supervised machine learning algorithms for a diamond pricing model
<p>The diamond is 58 times harder than any other mineral in the world, and its elegance as a jewel has long been appreciated. Forecasting diamond prices is challenging due to nonlinearity in important features such as carat, cut, clarity, table, and depth. Against this backdrop, the study conducted a comparative analysis of the performance of multiple supervised machine learning models (regressors and classifiers) in predicting diamond prices. Eight supervised machine learning algorithms were evaluated in this work including Multiple Linear Regression, Linear Discriminant Analysis, eXtreme Gradient Boosting, Random Forest, k-Nearest Neighbors, Support Vector Machines, Boosted Regression and Classification Trees, and Multi-Layer Perceptron. The analysis is based on data preprocessing, exploratory data analysis (EDA), training the aforementioned models, assessing their accuracy, and interpreting their results. Based on the performance metrics values and analysis, it was discovered that eXtreme Gradient Boosting was the most optimal algorithm in both classification and regression, with a R<sup>2</sup> score of 97.45% and an Accuracy value of 74.28%. As a result, eXtreme Gradient Boosting was recommended as the optimal regressor and classifier for forecasting the price of a diamond specimen.</p>
Toward a semi-supervised learning approach to phylogenetic estimation
<p>Models have always been central to inferring molecular evolution and to reconstructing phylogenetic trees. Their use typically involves the development of a mechanistic framework reflecting our understanding of the underlying biological processes, such as nucleotide substitutions, and the estimation of model parameters by maximum likelihood or Bayesian inference. However, deriving and optimizing the likelihood of the data is not always possible under complex evolutionary scenarios or even tractable for large datasets, often leading to unrealistic simplifying assumptions in the fitted models. To overcome this issue, we coupled stochastic simulations of genome evolution with a new supervised deep learning model to infer key parameters of molecular evolution. Our model is designed to directly analyze multiple sequence alignments and estimate per-site evolutionary rates and divergence, without requiring a known phylogenetic tree. The accuracy of our predictions matched that of likelihood-based phylogenetic inference, when rate heterogeneity followed a simple gamma distribution, but it strongly exceeded it under more complex patterns of rate variation, such as codon models. Our approach is highly scalable and can be efficiently applied to genomic data, as we showed on a dataset of 26 million nucleotides from the clownfish clade. Our simulations also showed that the integration of per-site rates obtained by deep learning within a Bayesian framework led to significantly more accurate phylogenetic inference, particularly with respect to the estimated branch lengths. We thus propose that future advancements in phylogenetic analysis will benefit from a semi-supervised learning approach that combines deep-learning estimation of substitution rates, which allows for more flexible models of rate variation, and probabilistic inference of the phylogenetic tree, which guarantees interpretability and a rigorous assessment of statistical support.</p>
A Supervised Machine-Learning Prediction of Textile's Antimicrobial Capacity Coated with Nanomaterials
<p>The dataset contains P-Chem properties of NMs and experimental conditions for assessing the antimicrobial properties of inorganic and organic NMs using machine learning tools.</p>
Small PASTIS training dataset config: Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series
<p>Files to run the small dataset experiments used in the preprint "Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series" available <a href="https://hal.science/hal-04084839">here</a>. This .csv files enables to generate balanced small dataset from the <a href="https://zenodo.org/record/5012942#.ZFDfUJHP1H4">PASTIS dataset</a>. These files are required to run the experiment with a small training data-set, from the open source code <a href="https://src.koda.cnrs.fr/iris.dumeur/ssl_ubarn.git">ssl_ubarn</a>. In the .csv file name selected_patches_fold_{FOLD}_nb_{NSITS}_seed_{SEED}.csv :</p> <ul> <li>FOLD: id which corresponds to one of the 5 experiments run due to PASTIS K-fold.</li> <li>NSITS: Number of SITS selected to construct this training data-set</li> <li>SEED: the randomness used to create this small dataset</li> </ul> <p> </p>
Unlabeled Sentinel 2 time series dataset (validation): Self-supervised Spatio-Temporal Representation Learning of Satellite Image Time Series
<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article "Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series" available <a href="https://hal.science/hal-04084839">here</a>. Each patch is constituted of the 10 bands [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks ['CLM_R1', 'EDG_R1', 'SAT_R1']. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only validation data</strong> are available. To download the full pretraining dataset, see <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UVU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table> <p> </p>
Unlabeled Sentinel 2 time series dataset (training, T30TUVU): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series
<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article "Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series" available <a href="https://hal.science/hal-04084839">here</a>. Each patch is constituted of the 10 bands [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks ['CLM_R1', 'EDG_R1', 'SAT_R1']. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30UVU</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>
Unlabeled Sentinel 2 time series dataset (training, T30TYQ): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series
<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article "Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series" available <a href="https://hal.science/hal-04084839">here</a>. Each patch is constituted of the 10 bands [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks ['CLM_R1', 'EDG_R1', 'SAT_R1']. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30TYQ</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>
Unlabeled Sentinel 2 time series dataset : Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series
<p>This repository list all the available repositories, to load the unlabeled Sentinel 2 (S2) L2A dataset used in the article<a href="https://ieeexplore.ieee.org/document/10414422/"> "Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series"</a>. This dataset is composed of patch time series acquired over France. For further details, see section IV.A of the pre-print article, available <a href="https://hal.science/hal-04084839">here</a>. Each patch is constituted of the 10 bands [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks ['CLM_R1', 'EDG_R1', 'SAT_R1']. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <ul> <li>The validation dataset is available here : <a href="https://doi.org/10.5281/zenodo.7890452">10.5281/zenodo.7890452</a></li> <li>The training dataset is composed of 9 zenodo repositories, one for each S2 tiles. Here are the available repositories: <ul> <li>T31UEP<a href="http://https://doi.org/10.5281/zenodo.7899943"> 10.5281/zenodo.7899943</a></li> <li>T31TGJ <a href="https://doi.org/10.5281/zenodo.7899237">10.5281/zenodo.7899237</a></li> <li>T30TYS <a href="https://doi.org/10.5281/zenodo.7924193">10.5281/zenodo.7924193</a></li> <li>T31TFN <a href="https://doi.org/10.5281/zenodo.7896621">10.5281/zenodo.7896621</a></li> <li>T31TDL <a href="http://10.5281/zenodo.7896082">10.5281/zenodo.7896082</a></li> <li>T31TDJ <a href="https://doi.org/10.5281/zenodo.7895498">10.5281/zenodo.7895498</a></li> <li>T30UVU <a href="https://doi.org/10.5281/zenodo.7892410">10.5281/zenodo.7892410</a></li> <li>T30TYQ<a href="https://doi.org/10.5281/zenodo.7890542"> 10.5281/zenodo.7890542</a></li> <li>T30TXT <a href="https://doi.org/10.5281/zenodo.7875977">10.5281/zenodo.7875977</a></li> </ul> </li> </ul> <table> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.