Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,549

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,549 results for “benchmarks”

Learn how ShareScore rates datasets ↗
zenodo40/100

Products and Models for "A benchmark JWST near-infrared spectrum for the exoplanet WASP-39 b"

<p>Publication Here: https://www.nature.com/articles/s41550-024-02292-x<br><br>Observing exoplanets through transmission spectroscopy supplies detailed information on their atmospheric composition, physics, and chemistry. Prior to <em>JWST,</em> these observations were limited to a narrow wavelength range across the near-ultraviolet to near-infrared, alongside broadband photometry at longer wavelengths. To understand more complex properties of exoplanet atmospheres, improved wavelength coverage and resolution are necessary to robustly quantify the influence of a broader range of absorbing molecular species. Here we show a combined analysis of <em>JWST</em> transmission spectroscopy across four different instrumental modes spanning 0.5&ndash;5.2 micron using Early Release Science observations of the Saturn-mass exoplanet WASP-39b. Our uniform analysis constrains the orbital and stellar parameters within sub-percent precision, including matching the precision obtained by the most precise asteroseismology measurements of stellar density to-date. Leveraging the advantages of a uniform light curve analysis, we improve the agreement between the transmission spectra of all modes, except for the NIRSpec PRISM, which is affected by partial saturation of the detector.&nbsp; Together, these collected data constitute the most comprehensive transmission spectrum of an exoplanet to date, providing unparalleled access to atmospheric absorbers including Na, K, H2O, CO, CO2, and SO2.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Scorpio Gene-Taxa Benchmark Dataset

<div> <div> <div> <div>&nbsp;</div> </div> </div> </div> <div> <div> <div> <div> <div> <div> <p>We used the Woltka pipeline to compile the complete Basic genome dataset, consisting of 4634 genomes, with each genus represented by a single genome. After downloading all coding sequences (CDS) from the NCBI database, we extracted 8 million distinct CDS, focusing on bacteria and archaea and excluding viruses and fungi due to inadequate gene information.</p> <p>To maintain accuracy, we excluded hypothetical proteins, uncharacterized proteins, and sequences without gene labels. We addressed issues with gene name inconsistencies in NCBI by keeping only genes with more than 1000 samples and ensuring each phylum had at least 350 sequences. This resulted in a curated dataset of 800,318 gene sequences from 497 genes across 2046 genera.</p> <p>We created four datasets to evaluate our model: a training set (Train_set), a test set (Test_set) with different samples but the same genus and gene as the training set, a Taxa_out_set excluding 18 phyla present in the training set but from different phyla, and a Gene_out_set excluding 60 genes from the training set but from the same phyla. We ensured each CDS had only one representation per genome, removing genes with multiple representations within the same species.</p> </div> </div> </div> </div> </div> </div>

opencc-by-4.0Jun 2024View details →
zenodo40/100

MedMNIST-C: Comprehensive benchmark and improved classifier robustness by simulating realistic image corruptions

<p><strong>Abstract: </strong>The integration of neural-network-based systems into clinical practice is limited by challenges related to domain generalization and robustness. The computer vision community established benchmarks such as ImageNet-C as a fundamental prerequisite to measure progress towards those challenges. Similar datasets are largely absent in the medical imaging community which lacks a comprehensive benchmark that spans across imaging modalities and applications. To address this gap, we create and open-source MedMNIST-C, a benchmark dataset based on the MedMNIST+ collection, covering 12 datasets and 9 imaging modalities. We simulate task and modality-specific image corruptions of varying severity to comprehensively evaluate the robustness of established algorithms against real-world artifacts and distribution shifts. We further provide quantitative evidence that our simple-to-use artificial corruptions allow for highly performant, lightweight data augmentation to enhance model robustness. Unlike traditional, generic augmentation strategies, our approach leverages domain knowledge, exhibiting significantly higher robustness when compared to widely adopted methods. By introducing MedMNIST-C and open-sourcing the corresponding library allowing for targeted data augmentations, we contribute to the development of increasingly robust methods tailored to the challenges of medical imaging. The code is available at <a href="https://github.com/francescodisalvo05/medmnistc-api">github.com/francescodisalvo05/medmnistc-api</a>.</p> <blockquote> <p>This work has been accepted at the Workshop on Advancing Data Solutions in Medical Imaging AI @ MICCAI 2024 [<a href="https://arxiv.org/pdf/2406.17536" target="_blank" rel="noopener">preprint</a>].</p> </blockquote> <p><strong>Note: </strong>Due to space constraints, we have uploaded all datasets except TissueMNIST-C. However, it can be reproduced via our APIs.&nbsp;</p> <p><strong>Usage:&nbsp;</strong>We recommend using the demo code and tutorials available on our GitHub&nbsp;<a href="https://github.com/francescodisalvo05/medmnistc-api">repository</a>.</p> <p><strong>Citation: </strong>If you find this work useful, please consider citing us:</p> <blockquote> <pre>@article{disalvo2024medmnist, title={MedMNIST-C: Comprehensive benchmark and improved classifier robustness by simulating realistic image corruptions}, author={Di Salvo, Francesco and Doerrich, Sebastian and Ledig, Christian}, journal={arXiv preprint arXiv:2406.17536}, year={2024} }</pre> </blockquote> <p><strong>Disclaimer: </strong>This repository is inspired by MedMNIST APIs and the ImageNet-C repository. Thus, please also consider citing <a href="https://www.nature.com/articles/s41597-022-01721-8" rel="nofollow">MedMNIST</a>, the respective source datasets (described&nbsp;<a href="https://medmnist.com/" rel="nofollow">here</a>), and&nbsp;<a href="https://arxiv.org/abs/1903.12261" rel="nofollow">ImageNet-C</a>.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Data for "A replicable and modular benchmark for long-read transcript quantification methods"

<p>This archive contains the input necessary to run the inital (TranSigner-protocol and IsoQuant-protocol) benchmarks associated with the <a href="https://github.com/COMBINE-lab/lr_quant_benchmarks" target="_blank" rel="noopener"><code>lr_quant_benchmarks repository</code></a>.&nbsp; The archive can be decompressed with <code>tar</code>&nbsp;and <code>zstd</code>&nbsp;using the command&nbsp;<code>tar --use-compress-program=zstd -xf input.tar.zstd</code>.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

AQL queries and benchmark results from PhD thesis "ANNIS: A graph-based query system for deeply annotated text corpora"

<p>These are the queries, the benchmark results and the evaluation scripts of the thesis &quot;ANNIS: A graph-based query system for deeply annotated text corpora&quot; (Thomas Krause 2018, Humboldt-Universit&auml;t zu Berlin)</p> <p><strong>diss_2018-01-12_v0.5.0.csv </strong><br> Results of all configurations of executed benchmarks for graphANNIS and also the baseline times of relANNIS.</p> <p><strong>queries.zip</strong><br> Contains folders for each corpus containing all queries used for the benchmark. Each file-name begins with the ID of the query. The extension denotes the type, which can be one of the following:</p> <ul> <li><em>&quot;.</em>aql&quot; contains the original AQL (ANNIS query language) query which was collected</li> <li>&quot;.json&quot; is JSON representation of the parsed AQL query</li> <li>&quot;.count&quot; is the number of matches a query should have</li> <li>&quot;.time&quot; is the average time in milliseconds that was needed to execute the query in relANNIS on the benchmark system</li> <li>&quot;.corpora&quot; contains the name of the corpus the query belongs to (should be only one corpus and the same as the folder name in the selection of queries in this data set)</li> <li>&quot;.relplan&quot; contains the PostgreSQL plan for the query</li> <li>&quot;.graphplan&quot; contains the graphANNIS plan for the query</li> </ul> <p><strong>evaluation-scripts.py/evaluation-scripts.ipynb</strong><br> Python scripts to perform the evaluation and generate the images. This are both a Python-file and the original notebook file that can be used with the Jupyter Notebook application.</p> <p><strong>relannis_benchmark_scripts.zip </strong><br> The files in this zip-file can be used to execute the benchmarks in the relANNIS system by piping the into the &quot;annis.sh&quot; command line tool of relANNIS</p>

opencc-by-4.0Jan 2018View details →
zenodo40/100

Efficient NAS Benchmark Kernels with C++ Parallel Programming Frameworks for Multi-Cores

<p>Benchmarking is a way to study the performance of new architectures and parallel programming frameworks. Well-established benchmark suites such as the NAS Parallel Benchmarks (NPB) comprise legacy codes that still lack portability to C++ language. As consequence, a set of high-level and easy-to-use C++ parallel programming frameworks cannot be tested in NPB. Our goal is to describe a C++ porting of the NPB kernels and to analyze the performance achieved by different parallel implementations written using the Intel TBB, OpenMP and FastFlow frameworks for Multi-Cores. The experiments show an efficient code porting from Fortran to C++ and a good parallel efficiency on average.</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

Datasets for "Precision and accuracy of single-molecule FRET measurements – a multi-laboratory benchmark study"

<p>Supplementary material (raw data) for Fig. 2 in &quot;<strong>Precision and accuracy of single-molecule FRET measurements &ndash; a multi-laboratory benchmark study</strong>&quot; to be published with Nature Methods</p> <p>The confocal data is given in ht3 and hdf5 format.</p> <p>For the TIRF data the original TIFF-stacks are uploaded including the calibration files.</p>

opencc-by-4.0May 2018View details →
zenodo40/100

Scalasca analysis report for SPEC MPI.2007 benchmark 132.zeump2 on 512 processes in virtual-node mode on Blue Gene/P

<p>A Cube3 performance analysis report written by the Scalasca 1.x parallel analyzer of a measurement of the SPEC MPI 2007 benchmark 132.zeusmp2 (ZeusMP/2) executed in virtual-node mode on 512 processes of the IBM Blue Gene/P system JUQUEEN, operated by Forschungszentrum J&uuml;lich GmbH.</p>

opencc-by-4.0Apr 2018View details →
zenodo40/100

Scalasca analysis report of the ASCI Sweep3D benchmark on 294,912 processes in virtual-node mode on IBM Blue Gene/P with manually annotated iterations

<p>A Cube3 performance analysis report written by the Scalasca parallel analyzer of a measurement of the ASCI benchmark Sweep3D, executed in virtual-node mode on 294,912 processes of the IBM Blue Gene/P system JUQUEEN, operated by Forschungszentrum J&uuml;lich GmbH. The measurement includes system topology information and manual annotations of the twelve iterations.</p>

opencc-by-4.0Aug 2018View details →
zenodo40/100

Open-Source Shared Memory implementation of the HPCG benchmark: analysis, improvements and evaluation on Cavium ThunderX2

<p>This archive contains the output files used to gerenate tables and plots in the paper &quot;Open-Source Shared Memory implementation of the HPCG benchmark: analysis, improvements and evaluation on Cavium ThunderX2&quot; submitted to the&nbsp;9th IEEE International Workshop on &quot;Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems&quot; (PMBS18) held as part of ACM/IEEE Supercomputing 2018 (SC18), Dallas, TX, USA.</p>

opencc-by-nc-4.0Sep 2018View details →
zenodo40/100

BenchPS: A Benchmark Dataset for Phrase Simplification

<p>BenchPS is a dataset built for the training and evaluation of phrase simplification systems. Each instance is composed of a sentence, target complex phrase, and a set of candidate simplifications ranked by simplicity. Each instance was annotated by humans through multiple annotations steps to ensure the reliability of the data.</p>

opencc-by-4.0Jan 2019View details →
zenodo40/100

Data for figures in "Reproducibility in Benchmarking Parallel Fast Fourier Transform based Applications"

<p>FFT benchmark data and Python plotting programs</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Code4Bench: A Multidimensional Benchmark of Codeforces Data for Different Program Analysis Techniques

<p>Reproducible research relies on well-designed benchmarks. However, evaluation on a single benchmark increases the risk of overfitting; that is, an optimization to reach a certain performance. In recent years several well-designed benchmarks have been constructed for different subfields of program analysis. However, they often involve real-world industrial projects in few languages such as C or Java. We provide Code4Bench, a benchmark comprising 3,421,357 programs totaling of 306,053,105 lines of code in 41 versions of 28 programming languages such as C/C++, Java, Python, and Kotlin. We have constructed this benchmark from Codeforces, a famous programming competition website, which is widely used by international programmers. Code4Bench advances the state-of-the-art in conducting reproducible and comparative experiments. It helps mitigate the bias and increase the generality and conclusiveness of the results. We present our methodology in construction of Code4Bench and give various descriptive statistics. We have also conducted an online survey on the users of Codeforces&rsquo; website whose code is included in the benchmark. The survey is concerned about the user&rsquo;s demographic information and programming habits, whose results are also provided in the benchmark. Finally, we leveraged an automatic process by which we localized faults within the faulty versions and categorize them according to a coarse-grained classification. In addition to its usage in empirical studies, Code4Bench can be used to teach programming and evolve algorithmic problems. We release Code4Bench in database format to allow researchers to extract other data of the benchmark by arbitrary queries.</p> <p>Code4Bench version 1.0.0 is publicly available at <a href="https://zenodo.org/record/2582968">https://zenodo.org/record/2582968</a>, with DOI 10.5281/zenodo.2582968, thereby providing long-term storage and versioning. It is released under the terms of Creative Commons Attribution 4.0 International license. Code4Bench is also publicly available at: <a href="https://github.com/code4bench/Code4Bench">https://github.com/code4bench/Code4Bench</a>, in which we have provided some additional information and script examples.</p>

opencc-by-4.0Mar 2019View details →
zenodo40/100

Supporting data for "Benchmarking Gate Fidelities in a Si/SiGe Two-Qubit Device" (arXiv:1811.04002)

<p>Supporting data and analysis scripts for FIG. 2&nbsp;- FIG. 6&nbsp;of &quot;Benchmarking Gate Fidelities in a Si/SiGe Two-Qubit Device&quot; (arXiv:1811.04002, Physical Review X in press)</p> <p>Data for FIG. 2 and FIG. 3 and analysis script&nbsp;are in the file &#39;single_two_qubit_RB.zip&#39;.</p> <p>Data for FIG. 4&nbsp;and analysis script&nbsp;are in the file &#39;CRB.zip&#39;.</p> <p>Data for FIG. 5&nbsp;and FIG. 6&nbsp;and analysis script&nbsp;are in the file &#39;Append.zip&#39;.</p> <p>To run the scripts, QCoDeS needs to be installed. Please refer to:</p> <p>http://qcodes.github.io/Qcodes/</p> <p>for instructions about installation and Dataset. If there&#39;s a problem, try to use an earlier version 0.1.7.</p> <p>For more information, contact the author: Xiao&nbsp;Xue, x.xue-1@tudelft.nl&nbsp;</p>

opencc-by-4.0Mar 2019View details →
zenodo40/100

Concrete mixed mode fracture test CARPIUC Benchmark « Predefined loadings » NP2 & P4

<p>The following data can be used for <strong>benchmarking numerical simulations of crack propagation test on quasi-brittle materials</strong>.</p> <p>Two crack propagation tests are proposed here, close to the well-known Nooru-Mohamed tests, but with modern instrumentation so that they provide trustworthy data. They present <strong>initiation and propagation</strong>. The goal is to compare your simulation results with the measured <strong>crack paths</strong> and <strong>force-displacement curves</strong>.</p> <p>The input data consists in specimen geometry, experimentally determined material properties (Young modulus, tensile strength, compressive strength and fracture energy) and the <strong>measured boundary conditions</strong>.</p>

opencc-by-4.0Dec 2018View details →
zenodo40/100

Concrete mixed mode fracture test CARPIUC Benchmark « Interactive loadings » IT1, IT2 & IT3

<p>The following data can be used for <strong>benchmarking numerical simulations of crack propagation test on quasi-brittle materials</strong>.</p> <p>Three crack propagation tests are proposed, inspired to some extent by the well-known Nooru-Mohamed tests. They present <strong>initiation, propagation, reorientation, link-up and branching</strong>. The goal is to compare your simulation results with the measured <strong>crack paths</strong> and <strong>force-displacement curves</strong>.</p> <p>The input data consists in specimen geometry, experimentally determined material properties (Young modulus, tensile strength, compressive strength and fracture energy) and the <strong>measured boundary conditions</strong>.</p>

opencc-by-4.0Dec 2018View details →
zenodo40/100

A Multi-Challenge Clustering Benchmark Dataset Embedding Large Differences in Spatial Extent

<p>This artificial clustering benchmark dataset was designed manually and draws its inspiration from structural aspects that can be seen in principal component plots of hyperspectral image data. Distance-separated, density-separated, gradient-separated as well as connected clusters have been placed into the dataset. Following the notion that clusters may vary significantly with respect to their spatial extent the respective separability problems are scaled at different levels and only become visible by magnifying certain parts of the dataset. Another special aspect of this dataset is that cluster borders have been kept rather ambiguous which, in our opinion, better resembles the situation in spectroscopic data.</p>

opencc-by-4.0Mar 2019View details →
zenodo40/100

Task graphs for benchmarking schedulers

<p><strong>Workflow Task Graph Dataset</strong></p> <p>This dataset contains three sets of task graphs representing different types of task workflows:</p> <ul> <li><em>Elementary</em> - contains trivial graph shapes, such as tasks with no dependencies or simple fork-join graphs. This set should test how the scheduler heuristics react to basic graph scenarios that frequently form parts of larger workflows.</li> <li><em>IRW</em> - is inspired by real-world workflows, such as machine learning cross-validation or map-reduce.</li> <li>&nbsp;<em>Pegasus</em> - is derived from graphs created by Pegasus Synthetic Workflow Generators (https://github.com/pegasus-isi/WorkflowGenerator)</li> </ul> <p>All of the provided task graphs are generated and compatible with ESTEE (https://github.com/It4innovations/estee) that allows to simulate their execution on a distributed system using various scheduling heuristics and environment conditions.</p> <p><strong>Data Format</strong></p> <p>Task graphs are stored in {elementary, irw, pegasus}.zip files that contain JSON representation of respective task graphs with the following fields:</p> <ul> <li>`graph_name` - Task graph name</li> <li>`graph_id` - Unique task graph identifier</li> <li>&nbsp;`graph` - Task graph representation - list of tasks where each task is represented as a dictionary with the following keys:</li> <li>&nbsp;`d`: Actual task duration in seconds (float value)</li> <li>&nbsp;`e_d`: User estimated task duration in seconds (float value)</li> <li>&nbsp;`cpus`: Task CPU core requirements (integer value)</li> <li>&nbsp;`outputs`: List of task outputs (list of integers indicating sizes of task outputs in MiB)</li> <li>&nbsp;`inputs`: List of task inputs in format of list [task\_id, output\_index]}. Output index is zero-based.</li> </ul> <p>For example this task graph:</p> <p>[{&#39;d&#39;: 200, &#39;e_d&#39;: 180, &#39;cpus&#39;: 1, &#39;outputs&#39;: [100], &#39;inputs&#39;: []},</p> <p>{&#39;d&#39;: 50, &#39;e_d&#39;: 60, &#39;cpus&#39;: 2, &#39;outputs&#39;: [], &#39;inputs&#39;: [[0, 0]]}]</p> <p>contains two tasks. One requiring no input, single CPU core with estimated duration 180s, actual duration 200s and producing a single output of 100 MiB. And another one requiring as an input task0&#39;s 0-th output, requiring 2 CPU cores, producing no output with estimated duration 60s and actual duration 50s.</p> <p>&nbsp;</p> <p><strong>Parsing the data</strong></p> <p>In Python, to load the elementary task graph set run the following snippet:</p> <pre><code class="language-python">import pandas as pd graphs = pd.read_json("./elementary.zip")</code></pre> <p>&nbsp;</p> <p>If you have Estee installed, you can use its provided `json_deserialize`</p> <p>function to parse the JSON encoded graphs into Estee TaskGraph data structure.</p> <p>&nbsp;</p> <pre><code class="language-python">from estee.serialization.dask_json import json_deserialize graph_json = graphs.loc[0, "graph"] graph = json_deserialize(graph)</code></pre> <p>&nbsp;</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

#nowplaying-RS: A New Benchmark Dataset for Building Context-Aware Music Recommender Systems

<p>Music recommender systems can offer users personalized and contextualized recommendation and are therefore important for music information retrieval. An increasing number of datasets have been compiled to facilitate research on different topics, such as content-based, context-based or next-song recommendation. However, these topics are usually addressed separately using different datasets, due to the lack of a unified dataset that contains a large variety of feature types such as item features, user contexts, and timestamps. To address this issue, we propose a large-scale benchmark dataset called #nowplaying-RS, which contains 11.6 million music listening events (LEs) of 139K users and 346K tracks collected from Twitter. The dataset comes with a rich set of item content features and user context features, and the timestamps of the LEs. Moreover, some of the user context features imply the cultural origin of the users, and some others&mdash;like hashtags&mdash;give clues to the emotional state of a user underlying an LE. In this paper, we provide some statistics to give insight into the dataset, and some directions in which the dataset can be used for making music recommendation. We also provide standardized training and test sets for experimentation, and some baseline results obtained by using factorization machines.</p> <p>The dataset contains three files:</p> <ul> <li>user_track_hashtag_timestamp.csv contains basic information about each listening event. For each listening event, we provide an id, the user_id, track_id, hashtag, created_at&nbsp;</li> <li>context_content_features.csv: contains all context and content features. For each listening event, we provide the id of the event, user_id, track_id, artist_id, content features regarding the track mentioned in the event (instrumentalness, liveness, speechiness, danceability, valence, loudness, tempo, acousticness, energy, mode, key) and context features regarding the listening event (coordinates (as geoJSON), place (as geoJSON), geo (as geoJSON), tweet_language, created_at, user_lang, time_zone, entities contained in the tweet).</li> <li>sentiment_values.csv contains sentiment information for hashtags. It contains the hashtag itself and the sentiment values gathered via four different sentiment dictionaries: AFINN, Opinion Lexicon, Sentistrength Lexicon and vader. For each of these dictionaries we list the minimum, maximum, sum and average of all&nbsp;sentiments of the tokens of the hashtag (if available, else we list empty values). However, as most hashtags only consist of a single token, these&nbsp;values are equal in most cases. Please note that the lexica are rather diverse and therefore, are able to resolve very different terms against a score. Hence,&nbsp;the resulting csv is rather sparse. The file contains the following comma-separated values: &lt;hashtag, vader_min, vader_max, vader_sum,vader_avg, &nbsp;afinn_min, afinn_max,&nbsp;afinn_sum, afinn_avg, ol_min, ol_max, ol_sum, ol_avg, ss_min, ss_max, ss_sum, ss_avg &gt;, where we abbreviate all scores gathered over the Opinion Lexicon with the&nbsp;prefix &#39;ol&#39;. Similarly, &#39;ss&#39; stands for SentiStrength.&nbsp;</li> </ul> <p>Please also find the training and test-splits for the dataset in this repo. Also, prototypical implementations of a context-aware recommender system based on the dataset can be found at&nbsp; <a href="https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM">https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM</a>.</p> <p>If you make use of this dataset, please cite the following paper where we describe and experiment with the dataset:</p> <p>@inproceedings{smc18,<br> title = {#nowplaying-RS: A New Benchmark Dataset for Building Context-Aware Music Recommender Systems},<br> author = {Asmita Poddar and Eva Zangerle and Yi-Hsuan Yang},<br> url = {http://mac.citi.sinica.edu.tw/~yang/pub/poddar18smc.pdf},<br> year = {2018},<br> date = {2018-07-04},<br> booktitle = {Proceedings of the 15th Sound &amp; Music Computing Conference},<br> address = {Limassol, Cyprus},<br> note = {code at https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM},<br> tppubtype = {inproceedings}<br> }</p>

opencc-by-4.0Jul 2018View details →
zenodo40/100

Benchmarking Smartphone Fluorescence-Based Microscopy with DNA Origami Nanobeads: Reducing the Gap toward Single-Molecule Sensitivity

<p>Smartphone-based fluorescence microscopy has been rapidly developing over the last few years, enabling point-of-need detection of cells, bacteria, viruses, and biomarkers. These mobile microscopy devices are cost-effective, field-portable, and easy to use, and benefit from economies of scale. Recent developments in smartphone camera technology have improved their performance, getting closer to that of lab microscopes. Here, we report the use of DNA origami nanobeads with predefined numbers of fluorophores to quantify the sensitivity of a smartphone-based fluorescence microscope in terms of the minimum number of detectable molecules per diffraction-limited spot. With the brightness of a single dye molecule as a reference, we compare the performance of color and monochrome sensors embedded in state-of-the-art smartphones. Our results show that the monochrome sensor of a smartphone can achieve better sensitivity, with a detection limit of &sim;10 fluorophores per spot. The use of DNA origami nanobeads to quantify the minimum number of detectable molecules of a sensor is broadly applicable to evaluate the sensitivity of various optical instruments.</p>

opencc-by-4.0Jan 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record