Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

53

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

53 results for “autoencoder”

Learn how ShareScore rates datasets ↗
zenodo52/100

Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease subtyping

<h3>TCGA pan-cancer mRNA and DNA data augmented with artificial confounders utilised in "Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease&nbsp;subtyping" by Zuqi Li and Sonja Katz (manuscript in preparation).</h3> <p>The following data curation steps were carried out:&nbsp;</p> <ul> <li><strong>Step 1. Download data from TCGA</strong> <ul> <li>R package `TCGAbiolinks`</li> <li>2547 patients (after step 2) with 6 cancer types: <ul> <li>BRCA (731)</li> <li>THCA (408)</li> <li>BLCA (387)</li> <li>LUSC (297)</li> <li>HNSC (412)</li> <li>KIRC (312)</li> </ul> </li> <li>mRNA expression profiles</li> <li>DNAm expression profiles</li> <li>Clinical data: <ul> <li>tumor stage: i, ia, ib, ii, iia, iib, iii, iiia, iiib, iiic, iv, iva, ivb, ivc, x</li> <li>age at diagnosis</li> <li>race: 'white', 'black or african amarican', 'asian', 'american indian or alaska native'</li> <li>gender<br><br></li> </ul> </li> </ul> </li> <li><strong>Step 2. Removal criteria</strong> <ul> <li>Patients with <ul> <li>NA or 'not reported' clinical data</li> <li>race 'american indian or alaska native'</li> <li>tumor stage x</li> </ul> </li> <li>mRNA and DNAm probes with <ul> <li>0 variance across all included patients</li> <li>not shared across all cancer types</li> <li>with missing values<br><br></li> </ul> </li> </ul> </li> <li>&nbsp;<strong>Step 3. Encode clinical vairables and save datasets</strong> <ul> <li>mRNA dataset: 2547 patients x 58,456 mRNAs</li> <li>DNAm dataset: 2547 patients x 232,088 DNAm</li> <li>clinic dataset: 2547 patients x 6 variables<br>&nbsp; &nbsp; 1. patient ID<br>&nbsp; &nbsp; 2. tumor stage: 1, 1, 1, 2, 2, 2, 3, 3, 3, 3, 4, 4, 4, 4<br>&nbsp; &nbsp; 3. age at diagnosis<br>&nbsp; &nbsp; 4. race: asian(1), black or african amarican(2), white(3)<br>&nbsp; &nbsp; 5. gender: female(0), male(1)<br>&nbsp; &nbsp; 6. cancer type: BRCA(1), THCA(2), BLCA(3), LUSC(4), HNSC(5), KIRC(6)<br>&nbsp; &nbsp;&nbsp;</li> </ul> </li> <li><strong>&nbsp;Step 4. Pre-process the datasets</strong> <ul> <li>mRNA dataset: '<em>TCGA_mRNAs_processed.csv'</em><br> <ul> <li>Take the 2000 mRNAs with highest variance</li> <li>Rescale every feature to [0,1]</li> <li>--&gt; 2547 patients x 2000 mRNAs</li> </ul> </li> <li>DNAm dataset: <em>'TCGA_DNAm_processed.csv'</em><br> <ul> <li>Take the 2000 DNAm with highest variance</li> <li>Rescale every feature to [0,1]</li> <li>--&gt; 2547 patients x 2000 DNAm</li> </ul> </li> <li>clinic dataset:<em> 'TCGA_clinic.csv'<br><br></em></li> </ul> </li> <li><strong>Step 5. Simulate confounders (instructions can be found in Methods section of manuscript)</strong> <ul> <li>Linear confounder: <ul> <li><em>'TCGA_confounder_linear.csv' -</em> linear confounding classes<em><br></em></li> <li><em>'TCGA_DNAm_confounded_linear.csv' </em>- linearly confounded DNAm data<em><br></em></li> <li><em>'TCGA_mRNA2_confounded_linear.csv'&nbsp;</em> - linearly confounded mRNA data<em><br></em></li> </ul> </li> <li>Squared confounder <ul> <li><em>'TCGA_confounder.csv' -</em> squared confounding classes<em><br></em></li> <li><em>'TCGA_DNAm_confounded.csv' </em>- squared confounded DNAm data<em><br></em></li> <li><em>'TCGA_mRNA2_confounded.csv'&nbsp;</em> - squared confounded mRNA data</li> </ul> </li> <li>Categorical confounder&nbsp; <ul> <li><em>'TCGA_confounder_categ2.csv' -</em> categorical confounding classes<em><br></em></li> <li><em>'TCGA_DNAm_confounded_categ2.csv' </em>- categorically confounded DNAm data<em><br></em></li> <li><em>'TCGA_mRNA2_confounded_categ2.csv'&nbsp;</em> - categorically&nbsp; confounded mRNA data</li> </ul> </li> <li>Multiple confounders - combined effect (linear + squared + categorical)<br> <ul> <li><em>'TCGA_confounder_multi.csv' -</em> confounding classes for combined effect<em><br></em></li> <li><em>'TCGA_DNAm_confounded_multi.csv' </em>- DNAm data with combined effect<em><br></em></li> <li><em>'TCGA_mRNA2_confounded_multi.csv'&nbsp;</em> - mRNA data&nbsp;with combined effect</li> </ul> </li> </ul> </li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

ArrayCGH microarray images for 'Autoencoder and NCA based neural network model to estimate survival prognosis in multiple myeloma using arrayCGH data'

<p>ArrayCGH microarray images for &#39;Autoencoder and NCA based neural network model to estimate survival prognosis in multiple myeloma using arrayCGH data&#39;</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Dragon_Pi: IoT Side-Channel Power Data Intrusion Detection Dataset and Unsupervised Convolutional Autoencoder for Intrusion Detection

<h2><strong>Dragon_Pi</strong></h2> <div> <div>For a more in depth description of the Dragon_Pi dataset, please consult the journal article of the same name:</div> <div>Lightbody <em>et al.</em>, Future Internet, 2024, <a href="https://doi.org/10.3390/fi16030088">https://doi.org/10.3390/fi16030088</a> - specifically Section 3.2: Dataset Overview.</div> <div>&nbsp;</div> </div> <p>Dragon_Pi is an intrusion detection dataset for IoT devices. In the field of IoT security there are few datasets, and those which do exist tend to focus solely on network traffic. The Dragon_Pi dataset seeks to provide not only more data for the field of IoT security, but also, data of a somewhat under-published type: linear time series power consumption data.</p> <p>Dragon_Pi is a fully labelled Intrusion Detection dataset for IoT devices. It is composed of both normal and under-attack power consumption data obtained from two separate testbeds - one using a DragonBoard 410c and the other a Raspberry Pi Model 3 - Hence the moniker&nbsp;<em>Dragon_Pi</em>.&nbsp;</p> <p>These testbeds were set up with predefined normal behavour as described in the attached publications. The normal linear time series power consumption&nbsp; was sampled from the testbed under these normal conditions. Both testbeds were then attacked using some common attacks on IoT - the linear time series power consumption captured under these condtions as well.&nbsp;</p> <p>Specifically, the testbeds were subjected to the Port Scan (using Nmap), SSH Brute Force (using Hydra) and SYNFlood Denial of Service (using Hping3) attacks. These attacks were repeated to gain insight to what their signatures looked like and also how varying the tool settings effected the resultant signature.&nbsp; A fourth type of scenario was also conducted on the testbeds - the "Capture the Flag" scenarios. In these files multiple attack types were used with a more specific target - to exfiltrate a hidden file from the testbeds.</p> <p>Each file has three hierarchical levels of annotation for <strong>each sample</strong> within:</p> <ol> <li>A simple "Normal or Anomaly" label for the specific sample</li> <li>A specifc attack type label e.g. "SSH Bruteforce", for the specific sample</li> <li>A specific tool setting for that attack e.g. "Hydra_T16", for the specific sample</li> </ol> <p>Users can decide for themselves what level of annotation they require for their specific task.&nbsp;</p> <p>Each file in the Dragon_Pi dataset is accompanied by its own legend file. This file explains the contents of the specific .csv file and the specific indexes of the events within.</p> <p>The Dragon_Pi dataset consists of approximately 67 files, as shown in Table 1. Compressed, the datset totals approximately 13GB. Completely decompressed the dataset is approximately 80GB ( 30GB Pi data, 50 GB Dragon data).&nbsp;</p> <div>&nbsp;</div> <div> <table> <tbody> <tr> <td>Label Type</td> <td>Specific Label&nbsp;</td> <td>Number of Files DragonBoard 410c</td> <td>Number of Files Raspberry Pi</td> </tr> <tr> <td>Normal&nbsp;</td> <td>Normal&nbsp;</td> <td>3&nbsp;</td> <td>2</td> </tr> <tr> <td>Port Scan Attack&nbsp;</td> <td>Nmap_T5</td> <td>2</td> <td>1</td> </tr> <tr> <td>&nbsp;</td> <td>Nmap_T4</td> <td>1</td> <td>1</td> </tr> <tr> <td>&nbsp;</td> <td>Nmap_T3</td> <td>1</td> <td>1</td> </tr> <tr> <td>&nbsp;</td> <td>Nmap_T2</td> <td>1</td> <td>1</td> </tr> <tr> <td>SSH Brute Force</td> <td>Hydra_T32</td> <td>4</td> <td>2</td> </tr> <tr> <td>&nbsp;</td> <td>Hydra_T16</td> <td>16</td> <td>2</td> </tr> <tr> <td>&nbsp;</td> <td>Hydra_T3</td> <td>8</td> <td>2</td> </tr> <tr> <td>&nbsp;</td> <td>Hydra_T1</td> <td>5</td> <td>2</td> </tr> <tr> <td>SYNFlood DOS</td> <td>SYNFlood DOS</td> <td>1</td> <td>1</td> </tr> <tr> <td>Capture the Flag</td> <td>Misc Attacks</td> <td>3</td> <td>5</td> </tr> </tbody> </table> </div> <div>Table 1. Enumeration of the in the Dragon_Pi dataset.</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>For a more in depth description of the Dragon_Pi dataset, please consult the journal article of the same name:</div> <div>Lightbody <em>et al.</em>, Future Internet, 2024, <a href="https://doi.org/10.3390/fi16030088">https://doi.org/10.3390/fi16030088</a> - specifically Section 3.2: Dataset Overview.</div> <div>&nbsp;</div> <div>&nbsp;</div> <div><strong>Publication of this dataset:</strong></div> <div>&nbsp;</div> <div>This dataset was published in Lightbody&nbsp;<em>et al.</em>, Future Internet, 2024, <a href="https://doi.org/10.3390/fi16030088">https://doi.org/10.3390/fi16030088</a>. Consult and cite this article for a more in depth dataset description, as well as an in depth review of first AI Intrusion Detection model trained on this dataset.&nbsp;</div> <div>&nbsp;</div> <div>See article Lightbody <em>et al.</em>, Future Internet, 2023, <a href="https://doi.org/10.3390/fi15050187">https://doi.org/10.3390/fi15050187</a> for a detailed investigation on&nbsp; the attack signatures discovered while creating this dataset. This work was an inital investigation of the dataset and can serve as a part 1 to the Dragon_Pi paper.</div> <div>&nbsp;</div> <div>&nbsp;</div> <div><strong>How to cite this dataset in your work:&nbsp;</strong></div> <div>&nbsp;</div> <div>Please cite these two DOIs when publishing using this dataset:</div> <div> <ol> <li>Dragon_Pi release publication: <a href="https://doi.org/10.3390/fi16030088">https://doi.org/10.3390/fi16030088</a> (most important)</li> <li>Zenodo Dataset DOI: https://doi.org/10.5281/zenodo.10784947</li> </ol> </div> <div> <div>&nbsp;</div> </div> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Particle-based Fast Jet Simulation at the LHC with Variational Autoencoders: generator-level and reconstruction-level jets dataset

<p>Jets at generator and reconstruction level saved in .npy format.</p> <p>Each jet is represented as an array of jet constituents characterized by their particle momentum in Cartesian coordinates, i.e., (px, py, pz). For both generator-level and reconstruction-level jets, jet constituents are ordered by decreasing pT.</p> <p>The shape of the datasets is [N, 50, 3], where N is the total number of jets,&nbsp;50 is the number of particles per jet, and 3 is the number of particle features (in order): [px, py, pz].<br> About 1.7M jets split into training, validation and testing sets at 60%, 20% and 20% respectively.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

A cooperative deep learning model for stock market prediction using deep autoencoder and sentiment analysis

<p>This data is used for Stock Market Prediction.&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Dataset for "Solving deep-learning density-functional theory via variational autoencoder"

<p>The dataset contains the ground state energies, the ground state density profiles, and the external potentials of a 3D single particle system with a Gaussian-like external potential.<br>The number of grid points for each dimension is \(N_g=18\), the linear length of the box is \(L=a_0\) with \(a_0\) the unit of length. The unit of energy is \(E_0=\frac{ \hbar^2}{(m a_0^2)}\).</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>The dataset is zip file of a Python npz file with the following keys:</p> <p>- "density" that corresponds to the ground state density profile.<br>- "potential" is the external potential.<br>-"energy" is the ground state energy.</p> <p><br>The number of instances is 36000.&nbsp;</p> <p>-3D_gaussian.zip -&gt; 3D_gaussian.npz</p> <p>&nbsp; &nbsp; a dictionary with three keys -density, potential, energy-.<br>&nbsp; &nbsp; The dimension of both potential and density is \([N_d,N_g,N_g,N_g]\).<br>&nbsp; &nbsp; The shape of energy is \([N_d]\).<br>&nbsp; &nbsp; \(N_d=36000\)</p> <p>-3D_gaussian_transfer_test_1.npz</p> <p>&nbsp; &nbsp; a dictionary with three keys -density, potential, energy-.<br>&nbsp; &nbsp; The dimension of both potential and density is \([N_d,N_g,N_g,N_g]\).<br>&nbsp; &nbsp; The shape of energy is \([N_d]\).<br>&nbsp; &nbsp; \(N_d=500\)</p> <p>-3D_gaussian_transfer_test_2.npz</p> <p>&nbsp; &nbsp; a dictionary with three keys -density, potential, energy-.<br>&nbsp; &nbsp; The dimension of both potential and density is \([N_d,N_g,N_g,N_g]\).<br>&nbsp; &nbsp; The shape of energy is \([N_d]\).<br>&nbsp; &nbsp; \(N_d=500\)</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Elucidating the Hierarchical Nature of Behavior with Masked Autoencoders

<p>#########</p> <p>Elucidating the Hierarchical Nature of Behavior with Masked Autoencoders</p> <p>#########</p> <p>Authors: Lucas Stoffl, Andy Bonnetto, St&eacute;phane D'Ascoli &amp; Alexander Mathis</p> <p>Affiliation: Ecole Polytechnique de Lausanne (EPFL)</p> <p>Date: 25/09/2024</p> <p>Link to the BiorXiv article : https://doi.org/10.1101/2024.08.06.606796</p> <p>-----------------</p> <h2>Provided data (hBehaveMAE checkpoints)</h2> <p>We provide a collection of pre-trained models that were reported in our paper, allowing you to reproduce our results for MABe22, hBABEL and Shot7M2 datasets.</p> <p>Note that you can <a href="https://huggingface.co/datasets/amathislab/SHOT7M2">download Shot7M2</a> on HuggingFace and <a href="https://github.com/amathislab/BehaveMAE/tree/main/hBABEL">generate hBABEL</a> by following the instructions on the <a href="https://github.com/amathislab/BehaveMAE">github page.</a></p> <ul> <li><strong>hBehaveMAE_hBABEL.pth </strong>: checkpoint for the hBehaveMAE pre-trained on the hBABEL dataset</li> <li><strong>hBehaveMAE_Shot7M2.pth</strong> : checkpoint for the hBehaveMAE pre-trained on the Shot7M2 dataset</li> <li><strong>hBehaveMAE_MABe22.pth</strong>: checkpoint for the hBehaveMAE pre-trained on the MABe22 dataset</li> </ul> <h2>References</h2> <p>If you find our code, weights or ideas useful, please cite:</p> <table> <tbody> <tr> <td>@article {Stoffl2024hBehaveMAE,<br>&nbsp; &nbsp; author = {Stoffl, Lucas and Bonnetto, Andy and d{\textquoteright}Ascoli, St{\'e}phane and Mathis, Alexander},<br>&nbsp; &nbsp; title = {Elucidating the Hierarchical Nature of Behavior with Masked Autoencoders},<br>&nbsp; &nbsp; elocation-id = {2024.08.06.606796},<br>&nbsp; &nbsp; year = {2024},<br>&nbsp; &nbsp; doi = {10.1101/2024.08.06.606796},<br>&nbsp; &nbsp; publisher = {Cold Spring Harbor Laboratory},<br>&nbsp; &nbsp; URL = {https://www.biorxiv.org/content/early/2024/08/08/2024.08.06.606796},<br>&nbsp; &nbsp; eprint = {https://www.biorxiv.org/content/early/2024/08/08/2024.08.06.606796.full.pdf},<br>&nbsp; &nbsp; journal = {bioRxiv}<br>}</td> </tr> </tbody> </table>

openapache2.0Aug 2024View details →
zenodo40/100

Source-Agnostic Gravitational-Wave Detection with Recurrent Autoencoders: H1 detector

<p>Gravitational Wave signals and random noise datasets, generated with the&nbsp;PyCBC library (https://pycbc.org).&nbsp;</p> <p>Data consists of 8 sec time sequences for a single detector (the H1 detector, mimicking LIGO Hanford), sampled at&nbsp;2048 Hz and&nbsp;&nbsp;represented as a one-dimensional array with 16,384 entries.&nbsp;</p> <p>Signal samples are provided, corresponding to Binary Black Hole and Binary Neutron Star mergers overlapped&nbsp;to noise.&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Source-Agnostic Gravitational-Wave Detection with Recurrent Autoencoders: L1 detector

<p>Gravitational Wave signals and random noise datasets, generated with the&nbsp;PyCBC library (https://pycbc.org).&nbsp;</p> <p>Data consists of 8 sec time sequences for a single detector (the L1 detector, mimicking LIGO Livingston), sampled at&nbsp;2048 Hz and&nbsp;&nbsp;represented as a one-dimensional array with 16,384 entries.&nbsp;</p> <p>Signal samples are provided, corresponding to Binary Black Hole and Binary Neutron Star mergers overlapped&nbsp;to noise.&nbsp;</p>

opencc-by-4.0Jul 2021View details →
dryad40/100

Towards a more informative representation of the fetal-neonatal brain connectome using Variational Autoencoder

<p>Recent advances in functional magnetic resonance imaging (fMRI) have helped elucidate previously inaccessible trajectories of early-life prenatal and neonatal brain development. To date, the interpretation of fetal-neonatal fMRI data has relied on linear analytic models, akin to adult neuroimaging data. However, unlike the adult brain, the fetal and newborn brain develops extraordinarily rapidly, far outpacing any other brain development period across the lifespan. Consequently, conventional linear computational models may not adequately capture these accelerated and complex neurodevelopmental trajectories during this critical period of brain development along the prenatal-neonatal continuum. To obtain a nuanced understanding of fetal-neonatal brain development, including non-linear growth, for the first time, we developed quantitative, systems-wide representations of brain activity in a large sample (&gt;500) of fetuses, preterm, and full-term neonates using an unsupervised deep generative model called Variational Autoencoder (VAE), a model previously shown to be superior to linear models in representing complex resting state data in healthy adults. Here, we demonstrated that non-linear brain features, i.e., latent variables, derived with the VAE pretrained on rsfMRI of human adults, carried important individual neural signatures, leading to improved representation of prenatal-neonatal brain maturational patterns and more accurate and stable age prediction in the neonate cohort compared to linear models. Using the VAE decoder, we also revealed distinct functional brain networks spanning the sensory and default mode networks. Using the VAE, we are able to reliably capture and quantify complex, non-linear fetal-neonatal functional neural connectivity. This will lay the critical foundation for detailed mapping of healthy and aberrant functional brain signatures that have their origins in fetal life.</p>

opencc-zeroMay 2023View details →
zenodo40/100

STGMVA: clustering, imputation, and integration for spatial resolved transcriptomics using spatiotemporal gaussian mixture variational autoencoder

<p>&nbsp;In this study, we present STGMVA, a comprehensive analysis toolkit employs a spatiotemporal gaussian mixture variational autoencoder to tackle these tasks effectively. STGMVA consists of two stages: pretraining the gene expression and spatial location using a gaussian mixture model, and learning the embedding vectors through a variational graph autoencoder. Results demonstrate STGMVA surpasses state-of-the-art approaches on various spatial transcriptomics datasets, exhibiting superior performance across different scales and resolutions. Notably, STGMVA achieves the highest clustering accuracy in human brain, mouse hippocampus, and mouse olfactory bulb tissues. Furthermore, STGMVA enhances and denoises gene expression patterns for gene imputation task. Additionally, STGMVA has the capability to correct batch effects and achieve joint analysis when integrating multiple tissue slices.</p>

opencc-by-4.0Jul 2023View details →
dryad40/100

Towards a more informative representation of the fetal-neonatal brain connectome using Variational Autoencoder

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad40/100

Data from: Bayesian estimation of muscle mechanisms and therapeutic targets using variational autoencoders

Open the record for dataset details and reuse information.

publicMar 2025View details →
zenodo36/100

Dataset of "Denoising Image-based Experimental Data without Clean Targets based on Deep Autoencoders"

<p>Dataset of the paper "Denoising Image-based Experimental Data without Clean Targets based on Deep Autoencoders", published in Experimental Thermal and Fluid Science (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.expthermflusci.2024.111195" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.expthermflusci.2024.111195</a>)</p> <p>The project received funding from: the European Research Council (ERC) under the European Union&rsquo;s Horizon 2020 research and innovation program (grant agreement No 949085); the National Natural Science Foundation of China (NSFC No 12227803 and No 12372276).</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Training and test data, plus saved models for the paper "Top-down effects in an early visual cortex inspired hierarchical Variational Autoencoder" submitted to the SVRHM 2022 Workshop @ NeurIPS

<p>Each .pkl&nbsp;file contains a training or test dataset&nbsp;in the form of a Python dictionary (generated with Python 3.8.5) with the following fields:</p><ul><li>'train_images': 640,000 float32 images&nbsp;used&nbsp;for model training. 20px images contain 400 pixel intensities, 40px images contain 1600 pixel intensities each.</li><li>'train_labels': float32 labels for each image in&nbsp;'train_images'. All natural images are&nbsp;labeled&nbsp;with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0,&nbsp;according to their texture family.</li><li>'test_images': 64,000 float32 images&nbsp;used&nbsp;for model testing.&nbsp;20px images contain 400 pixel intensities, 40px images contain 1600 pixel intensities each.</li><li>'test_labels': float32 labels for each image in&nbsp;'test_images'. All natural images are&nbsp;labeled&nbsp;with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0,&nbsp;according to their texture family.</li></ul><p>Each .zip file contains a saved model.&nbsp;Details on these are coming soon.</p><p>For more details, see the paper&nbsp;"Top-down effects in an early visual cortex inspired hierarchical Variational Autoencoder" published at the SVRHM 2022 Workshop @ NeurIPS&nbsp;(<a href="https://openreview.net/forum?id=8dfboOQfYt3">link</a>).</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Trained models and code accompanying 'Downscaling using Deep Convolutional Autoencoders, a case study for South East Asia'

<p>Trained models and code accompanying &#39;Downscaling using Deep Convolutional Autoencoders, a case study for South East Asia&#39;.</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Datasets accompanying 'Downscaling using Deep Convolutional Autoencoders, a case study for South East Asia'

<p>Datasets required to replicate the experiment described in the paper: &#39;Downscaling using Deep Convolutional Autoencoders, a case study for South East Asia&#39;</p> <p>This includes NetCDF files used to generate training data and processed climate data for model training, trained models and outputs from model predictions. It also contains the datasets used for comparisons (CORDEX and CMIP6)</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Polytopic autoencoders for very low-dimensional parametrizations of fluid flow models [source code]

<p>J. Heiland &amp; Y. Kim, 'Polytopic autoencoders for very low-dimensional parametrizations of fluid flow models', GAMM 2024</p>

opencc-by-4.0May 2024View details →
zenodo36/100

TB-CARE: A Novel Convolutional Autoencoder-based Tuberculosis Classification System with Enhanced EfficientNet

<p><span>The novel framework for TB classification using Convolutional AutoencodeR with EfficientNet (TB-CARE) involves the utilization of a convolutional autoencoder for feature extraction, an affinity propagation clustering method for selecting templates, and an enhanced EfficientNet (EEffNet) for classification. Extensive tests are performed on datasets that are freely accessible. The results of our methodology surpassed those of previous approaches, demonstrating its practicality for real-world applications. By leveraging deep learning models within the ensemble method, TB classification achieves a notable area under the receiver operating characteristic of up to 0.99, outperforming other tested classifiers and setting a new benchmark. EEffNet exhibits outstanding performance with an accuracy of 99.8%, sensitivity of 99.8%, and specificity of 99.7% on the NIH chest X-ray dataset and accuracy of 99.6%, sensitivity of 99.9%, and specificity of 99.4% on TBX11 k Dataset. These results indicate that employing features extracted from various image sources can significantly enhance the detection rate.</span></p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Unmixing Autoencoder for Image Reconstruction from Hyperspectral Data

<p>The NIR handwriting imaging data and the noise simulated data of five <span>hydroxyl compounds: methanol, ethanol, 2-phenylethanol, 1-propanol, and 2-chloroethanol.</span></p>

opencc-by-4.0Jul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record