Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
848
datasets available to search
ShareScore release 0.9.0
Dataset results
848 results for “representation”
Multiple interactive memory representations underlie the induction of false memory
Open the record for dataset details and reuse information.
Dynamic representation of the subjective value of information
Open the record for dataset details and reuse information.
Adaptive memory distortions are predicted by feature representations in parietal cortex
Open the record for dataset details and reuse information.
Elevation Models for Reproducible Evaluation of Terrain Representation - Inventory of Renderings
<p>This is an inventory of 155 renderings from 78 publications on terrain visualization techniques. The renderings guided the selection of landform types in the elevation models that are proposed in the following article:</p> <p><em>Kennelly, P. J., Patterson, T., Jenny, B., Huffman, D. P., Marston, B. E., Bell, S. and Tait, A. M. (2021). Elevation models for reproducible evaluation of terrain representation. Cartography and Geographic Information Science, 48:1, 63–77. DOI: <a href="http://doi.org/10.1080/15230406.2020.1830856">10.1080/15230406.2020.1830856</a></em></p> <p>Visualization techniques include colored aspect, contour lines, hypsometric tints, plan oblique relief, relief shading, rock and scree representation, and spot heights. The inventory contains information about display scale of the sample renderings, landform types, cell size, and geographic location of the digital elevation models. Also inventoried are how authors evaluated their renderings, how scale was indicated on the renderings, and whether the cell size and source of elevation data was included.</p> <p>The inventory is formatted as a single table in CSV UTF-8 and MS Excel .xlsx formats. Papers are grouped by visualization type; attributes of each rendering are stored on a single row.</p>
Audio clips of Orca (Orcinus orca) and non-orca sounds for the exploration of multiple acoustic representations
<p>Data and code associated with "Comparing acoustic representations for deep learning-based classification of underwater acoustic signals: a case study on orca (Orcinus orca) vocalizations."</p> <p>A collection of 9600 audio clips recorded by a hydrophone off San Juan Island, WA, USA. The clips are 3 seconds in duration with a sampling rate of 64KHz, and contain a variety of orca vocalizations (in the srkw folder), as well as non-orca sounds, both humpbacks (hb folder) and unspecified sounds typical of the location (neg folder). </p> <p>The code for each of the representations used in this study is also included.</p>
Dataset for "Development of allocentric representations using self-motion information"
<p>The present .csv file contains the raw data from a 2017 data collection. Specifically, children from six to 11 years old were tested with a non-visual spatial orientation task, in which they were required to i) observe animal-shaped landmarks located in each of the 4 angles of the experimental room; ii) be guided by the experimenter along a two-legged segment; iii) indicate the location of the four landmarks.</p>
KDE Representations of the Gravitational Wave Background Free Spectra Present in the NANOGrav 15-Year Dataset
<p><i><strong>OVERVIEW</strong></i></p><p><i><strong>----------------</strong></i></p><p>This is a downloadable file of probability densities from KDEs (Kernel Density Estimator) of free spectrum analyses of the NANOGrav 15yr Dataset (DOI <a href="https://doi.org/10.5281/zenodo.7967584">10.5281/zenodo.7967584</a>) that can be used with the <a href="https://github.com/astrolamb/ceffyl">Ceffyl</a> and <a href="https://github.com/andrea-mitridate/PTArcade">PTArcade</a> packages. Please see the GitHub page for Ceffyl/PTArcade installation details and usage information.<br><br>Details on how this data product was produced can be found in <a href="https://journals.aps.org/prd/abstract/10.1103/PhysRevD.108.103019"><i>Lamb, Taylor & van Haasteren 2023 (DOI 10.1103/PhysRevD.108.103019).</i></a></p><p><i><strong>DIRECTORY AND FILE STRUCTURE</strong></i></p><p><i><strong>----------------------------------------------------</strong></i></p><p>Each directory contains a file with log-pdfs representing their corresponding free spectra (`density.npy`), the frequencies at which they were analysed (`freqs.npy`), the grid-points at which each log-pdf was computed (`log10rhogrid.npy`), frequency labels (`log10rholabels.txt`), analysis label (`pulsar_list.txt`), an array of bandwidths computed using the Sheather-Jones method (see Lamb et al. 2023; 'bandwidths.npy'), and a file with some metadata about the data (`log.txt`).</p><p>./30f_fs{cp}_ceffyl </p><ul><li>A representation of a 30 frequency CURN free spectrum.</li></ul><p>./30f_fs{hd}_ceffyl</p><ul><li>A representation of a 30 frequency HD-correlated free spectrum</li></ul><p>./30f_fs{hd+mp+dp}_ceffyl_hd-only</p><ul><li>A representation of an analysis that simultaneously modeled a HD-correlated free spectrum, a MP-correlated free spectrum, and a DP-correlated free spectrum. Only the HD component is represented here.</li></ul><p>./30f_fs{hd+mp+dp+cp}_ceffyl_hd-only</p><ul><li>A representation of an analysis that simultaneously modeled a HD-correlated free spectrum, a MP-correlated free spectrum, a DP-correlated free spectrum, and a CURN free spectrum. Only the HD component is represented here.</li></ul><p>./README</p><ul><li>this is a readme</li></ul><p><i><strong>SOFTWARE</strong></i></p><p><i><strong>------------------</strong></i></p><p>This data should ideally be used with the latest versions of:</p><ul><li><i><strong>ceffyl</strong></i> (https://github.com/astrolamb/ceffyl)</li><li><i><strong>PTArcade </strong></i>(https://github.com/andrea-mitridate/PTArcade)</li></ul><p><i><strong>PLANNED REVISIONS</strong></i></p><p><i><strong>---------------------------------</strong></i></p><p>None</p><p><i><strong>CHANGE LOG</strong></i></p><p><i><strong>----------------------</strong></i></p><p>10/12/2023 - updated KDE representations</p><p>A bug was found that produced a poor reflection at the lower prior boundary. Hence, data was being represented well at the lower prior boundary. This has now been corrected.</p>
Proposal of a domain model for 3D representation of buildings for the 3D cadastre in Ecuador
<p><span>The accelerated urban sprawl of cities around the world presents major challenges for urban planning and land resource management. In this context, it is crucial to have a detailed 3D representation of buildings enriched with accurate alphanumeric information. A distinctive aspect of this proposal is its specific focus on the spatial unit corresponding to buildings. In order to propose a domain model for the 3D representation of buildings, the national standard of Ecuador and the international standard (ISO 19152) were considered. The proposal includes a detailed specification of attributes, both for the general subclass of buildings and for their infrastructure. The application of the domain model proposal was crucial in a study area located in the Riobamba canton, due to the characteristics of the buildings in that area. For this purpose, a geodatabase was created in pgAdmin4 with official information, taking into account the structure of the proposed model and linking it with geospatial data for an adequate management and 3D representation of the buildings in an open-source Geographic Information System. This application improves cadastral management in the study region and has wider implications. This model is intended to serve as a benchmark for other countries facing similar challenges in cadastral management and 3D representation of buildings, promote efficient urban development and contribute to global sustainable development.</span></p>
Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction - Datasets
<p>Datasets to NeurIPS 2021 accepted paper "Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction".</p> <p>Datasets are pytorch files containing a dictionary with training, validation and test sets. Train, validation and test sets are custom dataset classes which inherit from the standard torch dataset class. Corresponding code an be found at https://github.com/HSG-AIML/NeurIPS_2021-Weight_Space_Learning.</p> <p>Datasets 41, 42, 43 and 44 are our dataset format wrapped around the zoos from Unterthiner et al, 2020 (https://github.com/google-research/google-research/tree/master/dnn_predict_accuracy)<br> <br> Abstract:<br> Self-Supervised Learning (SSL) has been shown to learn useful and information-preserving representations. Neural Networks (NNs) are widely applied, yet their weight space is still not fully understood. Therefore, we propose to use SSL to learn neural representations of the weights of populations of NNs. To that end, we introduce domain specific data augmentations and an adapted attention architecture. Our empirical evaluation demonstrates that self-supervised representation learning in this domain is able to recover diverse NN model characteristics. Further, we show that the proposed learned representations outperform prior work for predicting hyper-parameters, test accuracy, and generalization gap as well as transfer to out-of-distribution settings.</p>
Diagrammatic representation of ZCM, TAT, and PIM modes
<p>This representation is a vector adaptation of Figure 1 present in the Fekedulegn et al. article. If you use it, don't forget to cite the authors below (you don't need to cite me).</p> <p>Fekedulegn, D., Andrew, M. E., Shi, M., Violanti, J. M., Knox, S., & Innes, K. E. (2020). Actigraphy-based assessment of sleep parameters. <em>Annals of Work Exposures and Health</em>, <em>64</em>(4), 350-367. <a href="https://doi.org/10.1093/annweh/wxaa007/">https://doi.org/10.1093/annweh/wxaa007/</a>.</p>
epiGBS2: Improvements and evaluation of highly multiplexed, epiGBS-based reduced representation bisulfite sequencing
<p>We present epiGBS2, a laboratory protocol based on epiGBS (Gurp et al., 2016) with a revised and user-friendly bioinformatics pipeline for a wide range of species with or without reference genome. Performance of several critical steps in epiGBS2 was evaluated against baseline data sets from <em>Arabidopsis thaliana</em> and Great tit (<em>Parus major</em>), which confirmed overall good performance of epiGBS2. We provide here the raw bisulfite sequencing data of the epiGBS2 run for <em>Arabidopsis thaliana.</em></p> <p>A detailed description of the laboratory protocol and an extensive manual of the bioinformatics pipeline are publicly accessible on github (<a href="https://github.com/nioo-knaw/epiGBS2">https://github.com/nioo-knaw/epiGBS2)</a> and zenodo (https://doi.org/10.5281/zenodo.4764652).</p> <p>Demultiplexed data were deposited on NCBI under the BioProject ID PRJNA764918</p>
Sheep Creek Flowlines Generalized for 1:100,000 scale Representation
<p>This dataset comprises original and simplified versions of 28 hydrographic flowline features in North Dakota, USA. The original data were derived from National Hydrography Dataset (NHD) data for the 1:24,000 Sheep Creek Dam topographic map quadrangle. Flowline features were selected for 1:100,000 scale (100k) representation using the NHD VisibilityFilter attribute and further filtered and merged to form a contiguous network. Features were then simplified for 100k representation using a combination of automated and manual operations. The dataset was created for the 24<sup>th</sup> ICA Workshop on Map Generalisation and Multiple Representation; further details can be found in the workshop abstract. The dataset is intended to serve as a benchmark for cartographic generalization algorithms used to simplify and smooth hydrographic flowline features. </p> <p><strong>Statistics:</strong></p> <p>Original vertices: 9620</p> <p>Simplified vertices: 1485</p> <p>Vertex reduction: 84.6%</p> <p> </p> <table> <tbody> <tr> <td> <p> </p> </td> <td> <p><strong>Producer's Modified Hausdorff Distance (m)</strong></p> </td> <td> <p><strong>Producer's Average Distance (Vertex Influence Method)</strong></p> </td> <td> <p><strong>Sinuosity Reduction (%)</strong></p> </td> <td> <p><strong>Average Angular Deflection (degrees)</strong></p> </td> <td> <p><strong>Max Angular Deflection (degrees)</strong></p> </td> </tr> <tr> <td> <p><strong>AVG</strong></p> </td> <td> <p><strong>24.08</strong></p> </td> <td> <p><strong>4.75</strong></p> </td> <td> <p><strong>5.23</strong></p> </td> <td> <p><strong>29.61</strong></p> </td> <td> <p><strong>46.55</strong></p> </td> </tr> <tr> <td> <p><strong>MIN</strong></p> </td> <td> <p><strong>0.00</strong></p> </td> <td> <p><strong>0.00</strong></p> </td> <td> <p><strong>-0.49</strong></p> </td> <td> <p><strong>5.23</strong></p> </td> <td> <p><strong>5.23</strong></p> </td> </tr> <tr> <td> <p><strong>MAX</strong></p> </td> <td> <p><strong>38.88</strong></p> </td> <td> <p><strong>14.56</strong></p> </td> <td> <p><strong>30.40</strong></p> </td> <td> <p><strong>33.97</strong></p> </td> <td> <p><strong>53.80</strong></p> </td> </tr> </tbody> </table>
CITRIS - Causal Representation Learning Datasets
<p>This repository contains the datasets from the paper "CITRIS: Causal Identifiability from Temporal Intervened Sequences" (<a href="https://arxiv.org/abs/2202.03169">link</a>) by Phillip Lippe, Sara Magliacane, Sindy Löwe, Yuki M. Asano, Taco Cohen, Efstratios Gavves.</p> <p><strong>Temporal Causal3DIdent</strong> - The Temporal Causal3DIdent dataset is a collection of 3D object shapes, which are observed under varying positions, rotations, lightning, and colors. Overall, we this dataset contains 7 (multidimensional) causal factors. The 7 shapes used are <a href="http://graphics.stanford.edu/data/3Dscanrep/">Armadillo</a>, <a href="http://graphics.stanford.edu/data/3Dscanrep/">Bunny</a>, <a href="https://www.cs.cmu.edu/~kmcrane/Projects/ModelRepository/#spot">Cow</a>, <a href="http://graphics.stanford.edu/data/3Dscanrep/">Dragon</a>, <a href="https://gfx.cs.princeton.edu/proj/sugcon/models/">Head</a>, <a href="https://www.cc.gatech.edu/projects/large_models/horse.html">Horse</a>, <a href="https://github.com/brendel-group/cl-ica">Teapot</a>. We provide two versions of the dataset: one that only contains images of the Teapot, and one that uses all 7 shapes. For more details on the dataset, see <a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p> <p><strong>Interventional Pong</strong> - The Interventional Pong environment is inspired by the game dynamics of Pong, where both paddles follow the policy of moving towards the ball, and the ball has slightly random movements. This dataset considers the 5 causal variables paddle left, paddle right, the ball position, the ball velocity, and the score. For more details on the dataset, see <a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p> <p><strong>Ball-in-Boxes</strong> - The Ball-in-Boxes is a simple dataset for showcasing the concept of the minimal causal variables. The system consists of a ball which randomly moves within a box, but only under an intervention can swap between the two boxes. Thereby, the intervention does not affect the x-position in the box. Thus, one can only discover the box assignment as a causal variable, and not whether the inner x-position also belongs to it. For more details on the dataset, see <a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p>
iCITRIS - Causal Representation Learning Datasets
<p>This repository contains the datasets from the paper "iCITRIS: Causal Representation Learning for Instantaneous Temporal Effects" (<a href="http://arxiv.org/abs/2206.06169">link</a>) by Phillip Lippe, Sara Magliacane, Sindy Löwe, Yuki M. Asano, Taco Cohen, Efstratios Gavves. </p> <p><strong>Instantaneous Temporal Causal3Ident </strong>- The Temporal Causal3DIdent dataset is a collection of 3D object shapes, which are observed under varying positions, rotations, lightning, and colors. Overall, we this dataset contains 7 (multidimensional) causal factors with instantaneous and temporal causal relations between them. The 7 shapes used are <a href="http://graphics.stanford.edu/data/3Dscanrep/">Armadillo</a>, <a href="http://graphics.stanford.edu/data/3Dscanrep/">Bunny</a>, <a href="https://www.cs.cmu.edu/~kmcrane/Projects/ModelRepository/#spot">Cow</a>, <a href="http://graphics.stanford.edu/data/3Dscanrep/">Dragon</a>, <a href="https://gfx.cs.princeton.edu/proj/sugcon/models/">Head</a>, <a href="https://www.cc.gatech.edu/projects/large_models/horse.html">Horse</a>, <a href="https://github.com/brendel-group/cl-ica">Teapot</a>. For more details on the dataset, see <a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p> <p><strong>Causal Pinball </strong>- The Causal Pinball environment implements the simplified, real-world game dynamics of Pinball. This dataset considers 5 causal variables with instantaneous effects: the paddle position left, the paddle position right, the ball (velocity and position), the state of all bumpers, and the score. For more details on the dataset as well as the code to generate this dataset, see <a href="https://github.com/phlippe/CITRIS">our GitHub repository</a>.</p>
Effects of audio-motor training on spatial representations in long-term late blindness
<p>Datasets for behavioural data:</p> <p>-<em>Auditory horizontal localization task</em></p> <p>- <em>Auditory vertical localization task</em></p> <p>- <em>Position matching task</em></p> <p>- <em>Proprioceptive midline task</em></p> <p>Dataset for EEG data:</p> <p>- Spatial bisection task: mean ERP amplitude for 50-90ms timew window for each trial, separately for condition, session and roi</p> <p> </p>
Polidoc.net CODEBOOK: National and Regional Manifestos and other Political Documents Collected for the Research Projects "Representation in Europe: Congruence between Preferences of Elites and Voters" (REPCONG) and "The Impact of EU Cohesion Policy on European Identification" (COHESIFY)
<p>The Political Documents Archive http://www.polidoc.net/ contains election manifestos, coalition agreements, government declarations and various other documents of political actors from developed democracies. Currently, the archive builds on a stock of more than 3000 political documents from 20 European countries. The aim of the repository is to provide political texts in order to facilitate scholarly research in different areas of comparative politics such as party competition, coalition politics, legislative decision-making or electoral behavior.</p> <p>National electoral manifestos have been collected in the course of the REPCONG project ("Representation in Europe: Policy Congruence between Citizens and Elites"), and the archive includes party manifestos for regional elections in several European democracies. Because the process of European integration resulted in a strengthening of regions in EU member states and in countries that want to join the European Union, the relevance of the regional level for political decision-making has increased during the last decades. Therefore, also the policy profiles of regional parties are required to get a full picture of democratic responsiveness in European states across all levels of the political system. The collection of regional manifestos was supported by the COHESIFY project (www.cohesify.eu), funded under the Horizon 2020 Framework Programme for Research and Innovation. The aim of COHESIFY is to study whether the European Structural and Investment Funds affect people’s support for and identification with the European project.</p> <p>The archive is freely accessible (after a simple registration) and meant to foster rigorous research in these areas by enabling scholars to produce valid and reliable findings from empirical studies of textual data rather than unnecessarily struggling to obtain and process texts.</p>
Supplementary run files for the paper "Learning Effective Representations for Retrieval using Self-Distillation with Adaptive Relevance Margins"
<p>TREC-Format run files of all trained models as supplementary material for the paper "Learning Effective Representations for Retrieval using Self-Distillation with Adaptive Relevance Margins".</p> <p>File naming follows the schema: <code>{model}-{loss variant}-{in-batch usage}-{dataset}.txt.gz</code></p>
Physical inconsistencies in the representation of the ocean heat-carbon nexus in simple climate models (Ocean and Climate variables from Simple Climate Models)
<p>This dataset provide global ocean and climate variables from 8 Simple Climate Models used in the study "Physical inconsistencies in the representation of the ocean heat-carbon nexus in simple climate models"</p>
Symbol Representation of the Three Gluon Form Factor in N=4 Planar Super Yang-Mills Theory
<p>Datasets describing the symbol of the three-gluon form factor in N=4 planar super Yang-Mills theory, generated using the amplitude bootstrap approach. The file "EZ_symb_new_norm" contains the symbol form of this quantity at 1 through 5 loops of precision, while the file "EZ6_symb_new_norm" contains the symbol at 6 loops. The file "EZ_symb_quad_new_norm" contains the symbol at 1 through 6 loops in compressed "quad" form, where the final-entry conditions described in (https://arxiv.org/pdf/2204.11901) are used to dramatically reduce the total number of terms in the symbol. The file "EZ7_symb_quad_new_norm" contains the symbol at 7 loops in the "quad" form. </p> <p>The tag “new_norm” refers to the fact that in the symbols given here, the letters a,b,c are defined by a = sqrt(u/(v*w)), b = sqrt(v/(w*u)), c = sqrt(w/(u*v)), as in arXiv:2405.06107, in order to make all coefficients integers. In contrast, in arXiv:2204.11901, the letters a,b,c were defined by a = u/(v*w), b = v/(w*u), c = w/(u*v).</p> <p>In addition to the funding sources listed, MW was supported by research grant 00025445 from Villum Fonden.</p>
CNN Wild Park - Graph Neural Networks for Learning Equivariant Representations of Neural Networks
<p>This repository contains the <strong>CNN Wild Park</strong> dataset from the paper:</p> <blockquote> <p><strong>Graph Neural Networks for Learning Equivariant Representations of Neural Networks</strong><br><a href="https://mkofinas.github.io/">Miltiadis Kofinas</a>*, <a href="https://bknyaz.github.io/">Boris Knyazev</a>, <a href="https://www.cyanogenoid.com/">Yan Zhang</a>, <a href="https://yunlu-chen.github.io/">Yunlu Chen</a>, <a href="https://gertjanburghouts.github.io/">Gertjan J. Burghouts</a>, <a href="https://egavves.com/">Efstratios Gavves</a>, <a href="https://www.ceessnoek.info/">Cees G. M. Snoek</a>, <a href="https://davzha.netlify.app/">David W. Zhang</a>*<br><em>ICLR 2024</em> (oral)<br><a href="https://arxiv.org/abs/2403.12143">https://arxiv.org/abs/2403.12143</a><br><a href="https://github.com/mkofinas/neural-graphs">https://github.com/mkofinas/neural-graphs</a><br>*Joint first and last authors</p> </blockquote> <p>We introduce a new dataset of CNNs, which we term <em>CNN Wild Park</em>.<br>The dataset consists of 117,241 checkpoints from 2,800 CNNs, trained for up to 1,000 epochs on CIFAR10.<br>The CNNs vary in the number of layers, kernel sizes, activation functions, and residual connections between arbitrary layers.</p> <p>More specifically, we construct the CNN Wild Park dataset by training 2,800 small CNNs with different architectures for 200 to 1,000 epochs on CIFAR10. We retain a checkpoint of its parameters every 10 steps and also record the test accuracy. The CNNs vary by:</p> <ul> <li>Number of layers L in [2, 3, 4, 5] (note that this does not count the input layer).</li> <li>Number of channels per layer c_l in [4, 8, 16, 32].</li> <li>Kernel size of each convolution k_l in [3, 5, 7].</li> <li>Activation functions at each layer are one of ReLU, GeLU, tanh, sigmoid, leaky ReLU, or the identity function.</li> <li>Skip connections between two layers with at least one layer in between. Each layer can have at most one incoming skip connection. We allow for skip connections even in the case when the number of channels differ, to increase the variety of architectures and ensure independence between different architectural choices. We enable this by adding the skip connection only to the min(c_n, c_m) nodes.</li> </ul> <p>We divide the dataset into train/val/test splits such that checkpoints from the same run are <strong>not</strong> contained in both the train and test splits. </p> <div> </div> <div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.