Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
69
datasets available to search
ShareScore release 0.9.0
Dataset results
69 results for “Representation learning”
Database For Finite Volume Features, Global Geometry Representations, and Residual Training for Deep Learning-based CFD Simulation
Open the record for dataset details and reuse information.
Pancancer survival prediction using a deep learning architecture with multimodal representation and integration
<p>The clinical data, gene expression (mRNA) data, and microRNA expression (miRNA) data were downloaded from the PanCanAtlas TCGA project (https://gdc.cancer.gov/about-data/publications/pancanatlas). The gene copy number variation (CNV) data was downloaded from UCSC Xena (http://xena.ucsc.edu/public/).</p> <p>Files prefixed with 'Pre' represent the preprocessed datasets. The preprocessing details can be found in 'Datasets and data preprocessing' subsection.</p>
Learning-induced reorganization of number neurons and emergence of numerical representations in a biologically-inspired neural network
<p>Data published along with paper: <em>Learning-induced reorganization of number neurons and emergence of numerical representations in a biologically-inspired neural network</em></p> <p>In this data is shared: models (before & after training), stimuli used to train and test, and activity of the model in test.</p>
Dataset used in COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations
<p>This dataset consists of two hdf5 files that contain pre-computed log-mel spectrograms that have been used to to train audio embedding models. The dataset is split into a training set and a validation set containing respectively 170793 and 19103 spectrogram patches with their accompanying multi-hot encoded tags from a vocabulary of 1000 tags provided by <a href="https://freesound.org/">Freesound</a> users.</p> <p>More details can be found in "COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations" by X. Favory, <a href="https://kdrossos.net">K. Drossos</a>, <a href="https://tutcris.tut.fi/portal/en/persons/tuomas-virtanen(210e58bb-c224-40a9-bf6c-5b786297e841).html">T. Virtanen</a>, and X. Serra. The code is available at this <a href="https://github.com/xavierfav/coala">GitHub repository</a>.</p> <p> </p> <p>License:</p> <p>This dataset is derived from content from the Freesound collection. All sounds are released under Creative Commons (CC) licenses from either <a href="https://creativecommons.org/publicdomain/zero/1.0/">CC0</a>, <a href="https://creativecommons.org/licenses/by/3.0/">CC-BY,</a> <a href="https://creativecommons.org/licenses/sampling+/1.0/">CC-S+</a>, or <a href="https://creativecommons.org/licenses/by-nc/3.0/">CC-BY-NC</a>. We attribute authors of all the sounds used in the dataset and provide their corresponding licenses in the attributions.txt file.</p> <p> </p>
Deep Learning-Ready Voxel Representation of Protein-Ligand Complexes from an Enhanced PBDbind v.2020 Dataset
<p>A critical aspect of successful deep learning (DL) modelling in computer-aided drug discovery (CADD) is the representation of biomolecular data. Voxel grid representations have emerged as a straightforward method for depicting 3D molecular structures of protein-ligand complexes. Proper structural preparation of these complexes is also crucial, particularly in models where the orientation of hydrogen atoms and the accurate assignment of protonation/tautomeric states are vital. The PDBbind, a widely used dataset, can be improved in this regard. This work presents an enhanced version of the PDBbind v.2020 refined set concerning structural preparation, a voxel representation of these structures suitable for DL model training and a diverse set of docking-generated poses that could be used to develop new scoring functions for pose prediction. With this dataset, we aim to provide the CADD community with high-quality, accessible resources to facilitate the development of DL models for drug discovery.</p>
Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program Repair
<p>Dataset for experiment 2</p>
MR2-Net: Retinal OCTA Image Stitching via Multi-Scale Representation Learning and Dynamic Location Guidance
<p>These datasets are released for academic research use only.</p>
Learning to read transforms phonological into phonographic representations: Evidence from a Mismatch Negativity study
<p>The EGI EEG dataset of the research project - Learning to read transforms phonological into phonographic representations: Evidence from a Mismatch Negativity study.</p><p>- MMNEGI_scripts.zip: MATLAB and R Scripts for data analysis.</p><p>- sub-*.zip: Raw EEG data organized in Brain Imaging Data Structure (BIDS).</p>
Machine Learning Potential for Electrochemical Interfaces with Hybrid Representation of Dielectric Response
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.