Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3
datasets available to search
ShareScore release 0.9.0
Dataset results
3 results for “NoSQL”
EvoBench: Benchmarking Schema Evolution in NoSQL
<p>Docker containers for reproducing the proof of concept measurements with our NoSQL Schema Evolution Benchmark.</p>
Data for: Improving Kieker's Scalability by Employing Linked Read-Optimized and Write-Optimized NoSQL Storage
<p>We show, how polyglot persistence increase Kieker's scalability by employing separate read-optimized and write-optimized noSQL storage. For this purpose we extended Kieker to store its monitoring output in Apache Cassandra , which is a write-optimized wide-column noSQL database. For the analysis of the generated monitoring output we are using ElasticSearch, a read-optimized document store noSQL storage. We are interlinking read-optimized and write-optimized noSQL storage within our Regression Benchmarking Execution Environment (RBEE). To ensure scalability we are employing a container infrastructure. The mentioned noSQL storages and the linker are operating within one single Docker container which scales horizontally.</p> <p>For generating reference values, we instrumented a Java SE application with Kieker's file system writer and measured throughput and method's execution times. Consecutively, we instrumented the same Java SE application with our Apache Cassandra writer and measured throughput and method's execution time. Finally, we compared measurement results of Kieker's file system writer with the measurement results of our Apache Cassandra writer.</p>
PAX: Partition-Aware Autoscaling for the Cassandra NoSQL Database
<p>Apache Cassandra has emerged as one of the most widely adopted NoSQL databases. However, there is still a limited understanding on how to optimally operate Cassandra in the cloud using autoscaling methods, by which resources can be scaled up or down to reduce operational costs and meet service-level objectives (SLOs).</p> <p>To address this limitation, we present PAX, a partition-aware elastic resource management system for Apache Cassandra. PAX uses low-overhead query sampling and knowledge of the data-partitioning across the nodes to automatically adapt capacity in Cassandra clusters. Differently from existing autoscaling methods for Cassandra, which incur large acquisition times for new nodes, PAX exploits Cassandra's hinted handoff mechanism and a shared hints storage to minimize the time needed to acquire a node into the cluster.</p> <p>We propose a reactive and a proactive implementation of PAX and compare their performance against different workloads with varying intensities and item popularity distributions, finding that the proactive version significantly reduces SLO violations. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.