Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

69

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

69 results for “Representation learning”

Learn how ShareScore rates datasets ↗
zenodo24/100

Database For Finite Volume Features, Global Geometry Representations, and Residual Training for Deep Learning-based CFD Simulation

Open the record for dataset details and reuse information.

restrictedcc-by-4.0May 2024View details →
zenodo24/100

Pancancer survival prediction using a deep learning architecture with multimodal representation and integration

<p>The clinical data, gene expression (mRNA) data, and microRNA expression (miRNA) data were downloaded from the PanCanAtlas TCGA project (https://gdc.cancer.gov/about-data/publications/pancanatlas). The&nbsp;gene copy number variation (CNV) data was downloaded from UCSC Xena (http://xena.ucsc.edu/public/).</p> <p>Files prefixed with &#39;Pre&#39; represent the preprocessed datasets. The preprocessing details can be found in &#39;Datasets and data preprocessing&#39; subsection.</p>

opencc-by-4.0Jan 2023View details →
zenodo24/100

Learning-induced reorganization of number neurons and emergence of numerical representations in a biologically-inspired neural network

<p>Data published along with paper: <em>Learning-induced reorganization of number neurons and emergence of numerical&nbsp;representations in a biologically-inspired neural network</em></p> <p>In this data is shared: models (before &amp; after training), stimuli used to train and test, and activity of the model in test.</p>

opencc-by-4.0Jun 2023View details →
zenodo20/100

Dataset used in COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations

<p>This dataset consists of two hdf5 files that contain pre-computed log-mel spectrograms that have been used to to train audio embedding models. The dataset is split into a training set and a validation set containing respectively 170793 and&nbsp;19103 spectrogram patches with their accompanying multi-hot encoded tags from a vocabulary of 1000 tags provided by <a href="https://freesound.org/">Freesound</a> users.</p> <p>More details can be found in &quot;COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations&quot; by X. Favory, <a href="https://kdrossos.net">K. Drossos</a>, <a href="https://tutcris.tut.fi/portal/en/persons/tuomas-virtanen(210e58bb-c224-40a9-bf6c-5b786297e841).html">T. Virtanen</a>, and X. Serra. The code is available at this <a href="https://github.com/xavierfav/coala">GitHub repository</a>.</p> <p>&nbsp;</p> <p>License:</p> <p>This dataset is derived from content from the Freesound collection. All sounds are released under Creative Commons (CC) licenses from either <a href="https://creativecommons.org/publicdomain/zero/1.0/">CC0</a>, <a href="https://creativecommons.org/licenses/by/3.0/">CC-BY,</a> <a href="https://creativecommons.org/licenses/sampling+/1.0/">CC-S+</a>, or <a href="https://creativecommons.org/licenses/by-nc/3.0/">CC-BY-NC</a>. We attribute authors of all the sounds used in the dataset and provide their corresponding licenses in the attributions.txt file.</p> <p>&nbsp;</p>

openother-atJun 2020View details →
zenodo20/100

Deep Learning-Ready Voxel Representation of Protein-Ligand Complexes from an Enhanced PBDbind v.2020 Dataset

<p>A critical aspect of successful deep learning (DL) modelling in computer-aided drug discovery (CADD) is the representation of biomolecular data. Voxel grid representations have emerged as a straightforward method for depicting 3D molecular structures of protein-ligand complexes. Proper structural preparation of these complexes is also crucial, particularly in models where the orientation of hydrogen atoms and the accurate assignment of protonation/tautomeric states are vital. The PDBbind, a widely used dataset, can be improved in this regard. This work presents an enhanced version of the PDBbind v.2020 refined set concerning structural preparation, a voxel representation of these structures suitable for DL model training and a diverse set of docking-generated poses that could be used to develop new scoring functions for pose prediction. With this dataset, we aim to provide the CADD community with high-quality, accessible resources to facilitate the development of DL models for drug discovery.</p>

restrictedcc-by-4.0Dec 2023View details →
zenodo20/100

Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program Repair

<p>Dataset for experiment 2</p>

opencc-by-4.0Sep 2020View details →
zenodo16/100

MR2-Net: Retinal OCTA Image Stitching via Multi-Scale Representation Learning and Dynamic Location Guidance

<p>These datasets are released for academic research use only.</p>

restrictedcc-by-4.0May 2024View details →
zenodo12/100

Learning to read transforms phonological into phonographic representations: Evidence from a Mismatch Negativity study

<p>The EGI EEG dataset of the research project - Learning to read transforms phonological into phonographic representations: Evidence from a Mismatch Negativity study.</p><p>- MMNEGI_scripts.zip: MATLAB and R Scripts for data analysis.</p><p>- sub-*.zip: Raw EEG data organized in Brain Imaging Data Structure (BIDS).</p>

restrictedNov 2023View details →
zenodo12/100

Machine Learning Potential for Electrochemical Interfaces with Hybrid Representation of Dielectric Response

Open the record for dataset details and reuse information.

restrictedcc-by-4.0Jul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record