Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
53
datasets available to search
ShareScore release 0.9.0
Dataset results
53 results for “autoencoder”
Benchmarking the Autoencoder Design for Imputing Single-Cell RNA Sequencing Data
<p>This repository contains the real and synthetic datasets used in the paper "Benchmarking the Autoencoder Design for Imputing Single-Cell RNA Sequencing Data". The zip file includes three folders:</p> <p>1. overall imputation accuracy: the 12 real scRNA-seq datasets used in the evaluation of overall imputation accuracy.</p> <p>2. cell clustering: the 20 real scRNA-seq datasets with cell type labels used in the evaluation of cell clustering.</p> <p>3. DE gene: the 20 scRNA-seq syntehtic datasets with ground-truth DE genes used in the evaluation of DE gene analysis. These datasets are simulated by simulator scDesign and 20 real datasets. </p> <p> </p>
Denoising Autoencoders for Phenotype Stratification (DAPS) Sample Trained Simulated Patient Data
<p>DAPS Trained data</p>
Denoising Autoencoders for Phenotype Stratification (DAPS) Sample Trained Patient Data
<p>Data for: https://github.com/greenelab/DAPS</p>
Source-Agnostic Gravitational-Wave Detection with Recurrent Autoencoders: BNS Dataset
<p>Source-Agnostic Gravitational-Wave Detection with Recurrent Autoencoders: BNS Dataset</p>
Source-Agnostic Gravitational-Wave Detection with Recurrent Autoencoders: BBH Dataset
<p>Source-Agnostic Gravitational-Wave Detection with Recurrent Autoencoders: BBH Dataset</p>
SpatialCVGAE: Spatial Domain Identification via Consensus Clustering Integrated Variational Graph Autoencoder
Open the record for dataset details and reuse information.
Dataset of "Towards Automatic Feature Extraction and Sample Generation of Grain Structure by Variational Autoencoder"
<p><em>No description.</em></p>
An interpretable and adaptive autoencoder for efficient tissue deconvolution
GEO Series GSE297720. Homo sapiens. 23 samples. Type: Expression profiling by high throughput sequencing.
Data from: Mirrored STDP implements autoencoder learning in a network of spiking neurons
The autoencoder algorithm is a simple but powerful unsupervised method for training neural networks. Autoencoder networks can learn sparse distributed codes similar to those seen in cortical sensory areas such as visual area V1, but they can also be stacked to learn increasingly abstract representations. Several computational neuroscience models of sensory areas, including Olshausen & Field's Sparse Coding algorithm, can be seen as autoencoder variants, and autoencoders have seen extensive use in the machine learning community. Despite their power and versatility, autoencoders have been difficult to implement in a biologically realistic fashion. The challenges include their need to calculate differences between two neuronal activities and their requirement for learning rules which lead to identical changes at feedforward and feedback connections. Here, we study a biologically realistic network of integrate-and-fire neurons with anatomical connectivity and synaptic plasticity that closely matches that observed in cortical sensory areas. Our choice of synaptic plasticity rules is inspired by recent experimental and theoretical results suggesting that learning at feedback connections may have a different form from learning at feedforward connections, and our results depend critically on this novel choice of plasticity rules. Specifically, we propose that plasticity rules at feedforward versus feedback connections are temporally opposed versions of spike-timing dependent plasticity (STDP), leading to a symmetric combined rule we call Mirrored STDP (mSTDP). We show that with mSTDP, our network follows a learning rule that approximately minimizes an autoencoder loss function. When trained with whitened natural image patches, the learned synaptic weights resemble the receptive fields seen in V1. Our results use realistic synaptic plasticity rules to show that the powerful autoencoder learning algorithm could be within the reach of real biological networks.
PILOT-GM-VAE: Patient-Level Analysis of Single Cell Disease Atlas with Optimal Transport of Gaussian Mixtures Variational Autoencoders
<p><strong>Datasets for PILOT-GM-VAE.<br>This part of the data includes Breast Cancer, Kidney_KPMP(only kidney tissue and all samples), Myocardial Infarction(2), and Kidney Cancer.<br>Please for more details read the data sets section of the paper.</strong></p>
SpaMask: Dual Masking Graph Autoencoder with Contrastive Learning for Spatial Transcriptomics
<p>Understanding the spatial locations of cell within tissues is crucial for unraveling the organization of cellular diversity. Recent advancements in spatial resolved transcriptomics (SRT) have enabled the analysis of gene expression while preserving the spatial context within tissues. Spatial domain characterization is a critical first step in SRT data analysis, providing the foundation for subsequent analyses and insights into biological implications. Graph neural networks (GNNs) have emerged as a common tool for addressing this challenge due to the structural nature of SRT data. However, current graph-based deep learning approaches often overlook the instability caused by the high sparsity of SRT data. <strong>Masking mechanisms</strong>, as an effective self-supervised learning strategy, can enhance the robustness of these models. To this end, we propose <strong>SpaMask, dual masking graph autoencoder with contrastive learning for SRT analysis</strong>. Unlike previous GNNs, SpaMask masks a portion of spot nodes and spot-to-spot edges to enhance its performance and robustness. SpaMask combines <strong>Masked Graph Autoencoders (MGAE) and Masked Graph Contrastive Learning (MGCL)</strong> modules, with MGAE using node masking to leverage spatial neighbors for improved clustering accuracy, while MGCL applies edge masking to create a contrastive loss framework that tightens embeddings of adjacent nodes based on spatial proximity and feature similarity. We conducted a comprehensive evaluation of SpaMask on <strong>eight datasets from five different platforms</strong>. Compared to existing methods, SpaMask achieves superior clustering accuracy and effective batch correction.</p>
Data from: Mirrored STDP implements autoencoder learning in a network of spiking neurons
Open the record for dataset details and reuse information.
Dataset used in COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations
<p>This dataset consists of two hdf5 files that contain pre-computed log-mel spectrograms that have been used to to train audio embedding models. The dataset is split into a training set and a validation set containing respectively 170793 and 19103 spectrogram patches with their accompanying multi-hot encoded tags from a vocabulary of 1000 tags provided by <a href="https://freesound.org/">Freesound</a> users.</p> <p>More details can be found in "COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations" by X. Favory, <a href="https://kdrossos.net">K. Drossos</a>, <a href="https://tutcris.tut.fi/portal/en/persons/tuomas-virtanen(210e58bb-c224-40a9-bf6c-5b786297e841).html">T. Virtanen</a>, and X. Serra. The code is available at this <a href="https://github.com/xavierfav/coala">GitHub repository</a>.</p> <p> </p> <p>License:</p> <p>This dataset is derived from content from the Freesound collection. All sounds are released under Creative Commons (CC) licenses from either <a href="https://creativecommons.org/publicdomain/zero/1.0/">CC0</a>, <a href="https://creativecommons.org/licenses/by/3.0/">CC-BY,</a> <a href="https://creativecommons.org/licenses/sampling+/1.0/">CC-S+</a>, or <a href="https://creativecommons.org/licenses/by-nc/3.0/">CC-BY-NC</a>. We attribute authors of all the sounds used in the dataset and provide their corresponding licenses in the attributions.txt file.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.