Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

40

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

40 results for “manual annotation”

Learn how ShareScore rates datasets ↗
zenodo36/100

Manual 4D annotations of Micro X-ray CT time-series (4D dataset)

<p>The 4D (3D+time) manual annotations of https://doi.org/10.5281/zenodo.4293394. For the annotation the SuRVoS workbench was used (https://doi.org/10.5281/10.5281/zenodo.247547) and our proposed hidden Markov model (HMM-T, https://doi.org/10.5281/zenodo.4416013 ) designed to refine 4D semantic segmentations made by a 3D semantic segmentation CNN after its applied on 4D data. Only slices 740-742 and 747-749 (refining to the first axis) are partially annotated. We acknowledge Diamond Light Source for the time on I13-2 under proposal mt9396.</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Manually Annotated Event Log of Users Prompts in LLMs for Conceptual Modeling

<p>This repository contains the supplementary material for our paper at ER 2024 Conference.&nbsp;</p> <p>The data contains the results of an empirical study with 76 undergraduate information systems students. The students submitted the course assignments in 39 groups (of one or two students). The assignment used for the study required use case modeling with UML use case diagrams and domain modeling with UML class diagrams. The groups were first expected to interact with an LLM and then, if needed, to manually improve their models. Groups were randomly assigned to interact with either GPT 4.0 or Code Llama 34B Instruct in one of three application domains.&nbsp;</p> <p>The participants were instructed to engage with the LLM until they were satisfied with the results or opted to skip further refinement. The interaction log contains the following fields: User ID, Input (the user prompt), Response (the modeling artifacts), and the Prompt Number (within user ID).</p>

opencc-by-4.0Aug 2024View details →
dryad36/100

Eastern Canada Flocks: Images and manually annotated bird positions

Open the record for dataset details and reuse information.

publicJan 2022View details →
zenodo32/100

Manual Topic Annotation of German Novels and Parlament Protocols by multiple Annotators

<p>This dataset was created in the research project hermA and contains topic annotations for 960 sentences, half of which were taken from transcripts of the German Bundestag and half from recent German novels.</p> <p>For each sentence, 30 different annotators evaluated whether illness is addressed, how central the topic is, if so, and how certain they are in the annotation.</p> <p>The dataset contains the following columns:</p> <ul> <li><strong>item_id </strong>(for each sentence)</li> <li><strong>worker_id </strong>(for each annotator)</li> <li><strong>worker_group </strong>(either &quot;crowdworker&quot;&nbsp;or &quot;student&quot;)</li> <li><strong>corpus </strong>(either &quot;protocol_corpus&quot;=transcript corpus&nbsp;or &quot;novel_corpus&quot;=fiction corpus)</li> <li><strong>text_source</strong></li> <li><strong>previous_sentences </strong>(in the text_source)</li> <li><strong>target_sentence</strong></li> <li><strong>following_sentences </strong>(in the text_source)</li> <li><strong>semantic_field_token </strong>(if existing)</li> <li><strong>semantic_field_status </strong>(either &quot;True&quot;&nbsp;or &quot;False&quot;)</li> <li><strong>1_wird_im_fett_gedruckten_satz_krankheit_thematisiert </strong>(topic annotations: either &quot;ja&quot;&nbsp;or &quot;nein&quot;)</li> <li><strong>1b_wie_zentral_ist_das_thema_krankheit_im_fettgedruckten_satz </strong>(topic centrality: &quot;NaN&quot;, &quot;krankheit_kommt_eher_am_rande_des_satzes_vor&quot;&nbsp;or &quot;krankheit_ist_das_zentrale_thema_des_satzes&quot;)</li> <li><strong>2_wie_sicher_bist_du_dir_bei_der_antwort_zu_frage_1_ </strong>(annotation certainty: &quot;sehr_sicher&quot;, &quot;eher_sicher&quot;, &quot;eher_unsicher&quot;&nbsp;or &quot;sehr_unsicher&quot;)</li> </ul> <p>&nbsp;</p> <p>We use the annotations to model ambiguity in:</p> <p>Andresen, Melanie; Vauth, Michael &amp; Zinsmeister, Heike. 2020. Modeling Ambiguity with Many Annotators and Self-Assessments of Annotator Certainty. <em>Proceedings of 14th Linguistic Annotation Workshop</em>.</p>

opencc-by-4.0Oct 2020View details →
zenodo32/100

Manually Curated Library of Transposable Elements (TEs) and TE Annotations for Drosophila amaguana

<p>This data collection provides a manually curated library of transposable elements (TEs) for <em>Drosophila amaguana</em>, including consensus sequences, genome-wide TE annotations, and individual TE copy sequences. The library was built using <em>de novo</em> generated by EDTA (Extensive <em>de novo</em> TE Annotator) and RepeatModeler, curated with MCHelper, and further processed for genome annotation using RepeatMasker and OneCodeToFindThemAll.&nbsp;</p> <p>Below is a description of the included files:</p> <ul> <li><strong><code>Dama_curated_TE_library.fasta</code>:</strong>&nbsp;Contains 737 consensus TE sequences manually curated for&nbsp;<em>D. amaguana</em>. Sequence identifiers include classification and origin (e.g., new families or similarity to known elements).&nbsp;The identifier for each sequence in the FASTA file includes information about its classification:<br><br> <ul> <li><strong>For sequences corresponding to potentially new families</strong>: The identifier consists of a three-letter abbreviation for <em>D. amaguana</em> (Dama), followed by the new family identifier and the superfamily name, all separated by underscores. <em>Example: </em>Dama_NF_BELPAO_1.</li> <li><strong>For consensus sequences that show similarity to TE sequences previously reported in other species</strong>: The identifier includes the abbreviation&nbsp;Dama, followed by the superfamily name and an abbreviation for the species in which the TE was previously reported, all separated by underscores. Example:&nbsp;Dama_Helitron-1_DVir.<br><br></li> </ul> </li> <li><strong><code>Dama_TE_annotations.out</code>:</strong> Genome-wide annotation file of TE insertions produced with RepeatMasker and post-processed using OneCodeToFindThemAll to merge fragmented elements.</li> <li><strong><code>Dama_TE_copies.fasta</code>: </strong>FASTA file containing the extracted sequences of all annotated TE copies from the <em>D. amaguana</em> genome.</li> <li><strong><code>TEcopies_sequences.sh</code>:</strong> Shell script used to extract TE copy sequences from the genome using the annotation coordinates.</li> <li><strong><code>Dynamics_Dama.ipynb</code>: </strong>Jupyter Notebook for the analysis of transposable element (TE) dynamics in<strong>&nbsp;</strong><em>D. amaguana.</em></li> </ul>

opencc-by-4.0Nov 2024View details →
dryad32/100

Automated bird sound classifications of long-duration recordings produce occupancy model outputs similar to manually annotated data

<p>Occupancy modeling is used to evaluate avian distributions and habitat associations, yet it typically requires extensive survey effort because a minimum of three repeat samples are required for accurate parameter estimation. Autonomous recording units (ARUs) can reduce the need for surveyors on site, yet ARUs utility were limited by hardware costs and the time required to manually annotate recordings. Software that identifies bird vocalizations may reduce expert time needed, if classification is sufficiently accurate. We assessed the performance of BirdNET – an automated classifier capable of identifying vocalizations from &gt;900 North American and European bird species – by comparing automated to manual annotations of recordings of 13 breeding bird species collected in northwestern California. We compared the parameter estimates of occupancy models evaluating habitat associations supplied with manually annotated data (9 min recording segments) to output from models supplied with BirdNET detections. We used three sets of BirdNET output to evaluate the duration of automatic annotation needed to approach manually annotated model parameter estimates: 9-min, 87-min, and 87-min of high-confidence detections. We incorporated 100 3-sec manually validated BirdNET detections per species to estimate true and false positive rates within an occupancy model. BirdNET correctly identified 90% and 65% of the bird species a human detected when data were restricted to detections exceeding a low or high confidence score threshold, respectively. Occupancy estimates, including habitat associations, were similar regardless of method. Precision (proportion of true positives to all detections) was &gt;0.70 for 9 of 13 species, and a low of 0.29. However, processing of longer recordings was needed to rival manually annotated data. We conclude that BirdNET is suitable for annotating multispecies recordings for occupancy modeling when extended recording durations are used. Together, ARUs and BirdNET may benefit monitoring and, ultimately, conservation of bird populations by greatly increasing monitoring opportunities.   </p>

opencc-zeroFeb 2022View details →
zenodo32/100

Manual acoustic signal annotation for species from sound libraries Jacques Vielliard Neotropical Sound Library

<h3>Description</h3> <p>We present a multispecies dataset of active acoustic recordings of anuran amphibians deposited in the Jacques Vielliard Neotropical Sound Library from the original AnuraSet localities. Some audio files were donated by professional acousticians, due to the small sample size of some species. The dataset consists of 270 audio files with a total of 305,211 minutes of annotated acoustic activity from 12 prioritised species.&nbsp;</p> <p>The annotation quality was classified as &lsquo;Clear&rsquo; by 46%, &lsquo;Medium&rsquo; by 43% and &lsquo;Far&rsquo; by 7%. &nbsp;The species with the highest recorded acoustic activity were Boana faber, Boana albomarginata, and Physalaemus cuvieri represented by 51, 50, and 50 recordings respectively, while the species with the lowest acoustic activity was Dendropsophus nahdereri represented by four recordings.&nbsp;</p> <p>To access the file you can request it at the following link:&nbsp;</p> <p>https://www2.ib.unicamp.br/fnjv/&nbsp;</p> <h3>Reference</h3> <p>Paper: https://www.nature.com/articles/s41597-023-02666-2</p> <p>Repository: https://github.com/soundclim/anuraset</p> <p>Webpage: https://soundclim.github.io/anuraweb/<br><br></p> <h3>Related projects</h3> <p>https://research.google/blog/in-search-of-a-generalizable-method-for-source-free-domain-adaptation/</p> <h3>Citation</h3> <p>Arcila-P&eacute;rez LF, Ulloa JS &amp; Soundclim Network (2024). Manual acoustic signal annotation for species from sound libraries Jacques Vielliard Neotropical Sound Library. Zenodo. https://doi.org/10.5281/zenodo.13786714</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

On the risk of manual annotations in 3D confocal microscopy image segmentation

<p>This dataset contains different annotated masks of human induced pluripotent stem cell nuclei from the dataset published with https://doi.org/10.1038/s41586-022-05563-7 and DL models trained using these masks. Napari-GT and Slicer-GT were manually annotated using the Napari and 3D Slicer software considering only the DNA channel, while for bioGT the Lamin B1 channel was annotated using the seeded watershed algorithm to obtain a reproducible and biologically plausible nucleus annotation. This dataset is provided to reproduce the results in the manuscript &quot;On the risk of manual annotations in 3D confocal microscopy image segmentation&quot;, more details can be found there.</p>

opencc-by-4.0Aug 2023View details →
ClinicalTrials.gov32/100

MANual vs. automatIC Local Activation Time Annotation for Guiding Premature Ventricular Complex Ablation

ClinicalTrials.gov study NCT03340922. IPD Sharing: Not stated. Countries: 1. Publications: 7.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad32/100

Automated bird sound classifications of long-duration recordings produce occupancy model outputs similar to manually annotated data

Open the record for dataset details and reuse information.

publicFeb 2022View details →
dryad28/100

Data from: Repertoire-wide gene structure analyses: a case study comparing automatically predicted and manually annotated gene models

The location and modular structure of eukaryotic protein-coding genes in genomic sequences can be automatically predicted by gene annotation algorithms. These predictions are often used for comparative studies on gene structure, gene repertoires, and genome evolution. However, automatic annotation algorithms do not yet correctly identify all genes within a genome, and manual annotation is often necessary to obtain accurate gene models and gene sets. As manual annotation is time-consuming, only a fraction of the gene models in a genome is typically manually annotated, and this fraction often differs between species. To assess the impact of manual annotation efforts on genome-wide analyses of gene structural properties, we compared the structural properties of protein-coding genes in seven diverse insect species sequenced by the i5k initiative. Our results show that the subset of genes chosen for manual annotation by a research community (3.5-7% of gene models) may have structural properties (e.g., lengths and exon counts) that are not necessarily representative for a species' gene set as a whole. Nonetheless, the structural properties of automatically generated gene models are only altered marginally (if at all) through manual annotation. Major correlative trends, for example a negative correlation between genome size and exonic proportion, can be inferred from either the automatically predicted or manually annotated gene models alike. Vice versa, some previously reported trends did not appear in either the automatic or manually annotated gene sets, pointing towards insect-specific gene structural peculiarities. In our analysis of gene structural properties, automatically predicted gene models proved to be sufficiently reliable to recover the same gene-repertoire-wide correlative trends that we found when focusing on manually annotated gene models only. We acknowledge that analyses on the individual gene level clearly benefit from manual curation. However, as genome sequencing and annotation projects often differ in the extent of their manual annotation and curation efforts, our results indicate that comparative studies analyzing gene structural properties in these genomes can nonetheless be justifiable and informative.

opencc-zeroAug 2020View details →
zenodo28/100

Kludt et al. Cell Reports Medicine - Dataset (manual annotations)

<p>Dataset of patches created from manually annotated whole-slide images of non-small cell lung cancer from the publication:</p> <p>"Next generation lung cancer pathology: development and validation of diagnostic and prognostic algorithms"</p> <p>in Cell Reports Medicine 2024</p> <p>The pixel-level ground truth information is included (classes.txt).</p> <p><br>The dataset can be used for academic research purposes only.</p> <p>&nbsp;</p> <p>(c) Yuri Tolkach &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;</p>

openJul 2024View details →
dryad28/100

Data from: Repertoire-wide gene structure analyses: a case study comparing automatically predicted and manually annotated gene models

Open the record for dataset details and reuse information.

publicAug 2020View details →
zenodo24/100

Stroke - Tasks Walking Total With Manual Annotation Bracelet Dataset

<p>The tasks_walking_total_with_manual_annotations.csv file contains accelerometer data and task labels of the pilot stroke patients that are given in folder datasets/bracelet/stroke and include only walking activities. It includes 11 patients that are manually annotated and consists of 7 columns. Those columns are:</p> <ol> <li> <p>x, y, z, which represent the accelerometer values of the bracelets sensors used on either left or right wrist of the patients</p> </li> <li> <p>T and time, which represent the timestamp of the activity (time) and the period (T).</p> </li> <li> <p>Patient ID column, which is the number id of the patients.</p> </li> <li> <p>Task column, which represents the task performed by the patient.</p> </li> </ol> <p>This file contains the walking tasks that are described below. Inside of each parenthesis, is given the name of each task, based on the annotations that were provided on datasets/annotations/raw_medical_tracking/stroke/<strong>intense-monitoring_with_manual_annotations_v2.xlsx</strong>&nbsp;file and on the accelerometer data that were available on stroke pilot folder mentioned above. Those tasks are:</p> <ol> <li> <ol> <li> <p>normal_walk (Normal walking)</p> </li> <li> <p>tandem_walk (Tandem walking)</p> </li> <li> <p>bicycle_walk (Bycicle walking)</p> </li> <li> <p>walk_with_knees_raised (Walking with the knees raised)</p> </li> </ol> </li> </ol> <p>&nbsp;</p> <p>The features used to recognize walking activities in stroke patients are x,y,z and Task.</p>

opencc-by-4.0Mar 2024View details →
geo20/100

Manual curation for improved genome annotation of the functionally extinct northern white rhinoceros (Ceratotherium simum cottoni)

GEO Series GSE300824. Ceratotherium simum simum. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2026View details →
geo20/100

High-throughput manual-quality annotation of full-length long noncoding RNAs with Capture Long-Read Sequencing (CLS)

GEO Series GSE93848. Mus musculus; Homo sapiens. 36 samples. Type: Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenJan 2017View details →
zenodo20/100

TRACES Telegram and Twitter Dataset with Bulgarian Journalists Manual Annotations of True/Untrue and Disinformation/Not and Automatic Annotations for Markers of Lies

<p>TRACES dataset of 4083 Twitter and Telegram posts automatically annotated for markers of lies and manually by Bulgarian journalists for containing true/untrue information and disinformation or not.</p> <p>Each message has been annotated by usually 3 (in under 10 cases by 2 annotators). The annotators came from different media, in order to obtain various views. They were asked to not get biased and were assured that their identities will not be revealed.&nbsp;</p> <p><strong>The dataset is a subset of these other datasets:</strong></p> <p>https://zenodo.org/record/7614247</p> <p>https://zenodo.org/record/7614318</p> <p>https://zenodo.org/record/7614357</p> <p>https://zenodo.org/record/7614294</p> <p>&nbsp;</p> <p><strong>It has been annotated following these Annotation Guidelines:</strong></p> <p>https://zenodo.org/record/7706743</p>

restrictedMar 2023View details →
zenodo16/100

VETO: Vessel topology manual annotation dataset

<p>We have labeled the vessel topology on four public-available retinal datasets:</p> <p>INSPIRE [1]<br>VICAVR [2]<br>IOSTAR [3]<br>DRIVE [4]<br>Two experts were asked to manually label the topological information of the retinal vascular structure by using a graph editing software we developed for this task. Expert one and two independently labeled each vessel segment or centerline for all datasets, based on the types of available manual annotations of the vessel structure. Then the consensus between them was released for public use.</p> <p>It is worth noting that the vessel segments or vessel centerlines were used for topology estimation were extracted either by human grader or automatic vessel segmentation method, i.e. the DRIVE and IOSTAR datasets include the manual annotations of retinal vessel for each image, so the topology reconstruction were performed at manual annotated vessel patterns; for VICAVR datasets, the topology estimation were performed at the automatic segmented vessels by using the automated segmentation method [5]; for INSPIRE dataset, the human expert graded the topology at the vessel centerline which were provided by [6].</p> <p>[1] M. Niemeijer, X. Xu, A. Dumitrescu, B. van Ginneken, J. Folk, and M. Abr&agrave;moff, &ldquo;Automated measurement of the arteriolar-to-venular width ratio in digital color fundus photographs,&rdquo; IEEE Trans. Med.Imaging, vol. 30, no. 11, pp. 1941&ndash;1950, 2011.</p> <p>[2] S. G. V&aacute;zquez, B. Cancela, N. Barreira, G. C. de Tuero, M. A. Barcel&oacute;, and M. Saez, &ldquo;Improving retinal artery and vein classification by means of a minimal path approach,&rdquo; Mach. Vis. Appl., vol. 24, no. 5, pp. 919&ndash; 930, 2013.</p> <p>[3] J. Zhang, B. Dashtbozorg, E. J. Bekkers, J. P. W. Pluim, R. Duits, and B. M. ter Haar Romeny, &ldquo;Robust retinal vessel segmentation via locally adaptive derivative frames in orientation scores,&rdquo; IEEE Trans. Med. Imaging, vol. 35, pp. 2631&ndash;2644, 2016.</p> <p>[4] J. Staal, M. D. Abr&agrave;moff, M. Niemeijer, M. A. Viergever, and B. van Ginneken, &ldquo;Ridge-based vessel segmentation in color images of the retina,&rdquo; IEEE Transactions on Medical Imaging, vol. 23, pp. 501&ndash;509, 2004.</p> <p>[5] Y. Zhao, L. Rada, K. Chen, , and Y. Zheng, &ldquo;Automated vessel segmentation using infinite perimeter active contour model with hybrid region information with application to retinal images,&rdquo; IEEE Trans. Med. Imaging, vol. 34, no. 9, pp. 1797&ndash;1807, 2015.</p> <p>[6] R. Estrada, C. Tomasi, S. Schmidler, and S. Farsiu, &ldquo;Tree topology estimation,&rdquo; IEEE Trans. Pattern Anal. Mach. Intell., vol. 37, no. 8, pp. 1688&ndash;1701, 2015.</p>

restrictedcc-by-4.0Jul 2020View details →
zenodo16/100

Single-cell RNA sequencing of 76,535 capillary PBMCs from 3 donors across 28 samples, freshly isolated or on ice for 24h with manual 4-layer annotation hierarchy

<p><strong>Sample Procurement</strong></p> <p>Samples were collected directly from participants as part of ImYoo&#39;s &quot;Single-cell immune profiling from self-collected capillary blood&quot; study, approved by Advarra IRB (Protocol #Pro00057361).</p> <p><strong>Sample Processing</strong></p> <p>Whole capillary blood samples were self-collected from participants using the <a href="https://yourbiohealth.com/en-us/virtually-painless-blood-collection-devices-for-clinical-trials-and-wellness-testing">TAP II</a> device. Samples were either processed shortly after collection, or if being stored for longer than 4 hours, were kept in a styrofoam cooler with ice packs. Cells were isolated using <a href="https://www.stemcell.com/products/easysep-direct-human-pbmc-isolation-kit.html">EasySep Direct Human PBMC Isolation Kit</a>&nbsp;(STEMCELL Technologies Catalog #19654) and cryopreserved using <a href="https://www.stemcell.com/cryostor-cs10.html">CryoStor CS10</a> (STEMCELL Technologies Catalog #07930). Upon thawing, samples were labeled in accordance with the MULTI-seq&nbsp;protocol (<a href="https://www.nature.com/articles/s41592-019-0433-8">https://www.nature.com/articles/s41592-019-0433-8</a>) and then processed on a 10X Genomics Chromium, using either the Chromium Next GEM Single Cell 3&#39; Kit v3.1 (10X Genomics Product Code&nbsp;1000269) or&nbsp;Chromium Next GEM Single Cell 3&rsquo; HT Kit v3.1 (10X Genomics Product Code 1000370). DNA libraries were sequenced on either a NovaSeq 6000 or NextSeq 550.</p> <p><strong>Data Processing</strong></p> <p>Transcriptomic sequencing data was processed using Cell Ranger v7.0.1 with default parameters. Multiplexing oligo sequencing data was processed through a custom python script that counts the number of occurrences of each sample barcode sequence&nbsp;and assigns it to the corresponding cell barcode. Samples were demultiplexed using a custom algorithm that estimates the background sample barcode counts, and assigns each cell a probability of belonging to each sample. Cell typing was done as part of a larger dataset and consisted of iterative manual assignments of clusters to cell types. For each cell subtype detected, a new model was trained on just the cells of that type, and the process was repeated.</p> <p><strong>Metadata Fields</strong></p> <ul> <li><strong>barcode:</strong> Original chromium cell barcode</li> <li><strong>Sample IDs:</strong> Unique ID for the experimental sample that was processed with 10x Chromium, could have come from the same biological sample (identified by <strong>original_sample_id</strong>)</li> <li><strong>Participant IDs:</strong> Unique ID for participant (here there are three participants: 2, 3 and 51)</li> <li><strong>Cell Barcoding Runs:</strong> Unique ID for the 10x Chromium cell barcoding run in which that sample was processed. Multiple samples can be processed in a cell barcoding run.</li> <li><strong>Lane:</strong> ID of which Chromium chip lane the cell came from</li> <li><strong>extraction_protocol:</strong>&nbsp;How the PBMCs were isolated from whole blood. In this dataset all samples were processed with the TAP device.</li> <li><strong>sample_processing_delay_seconds:</strong> The amount of time (in seconds) between when the blood was extracted from the participant and when PBMC isolation + cryopreservation was performed</li> <li><strong>cell_barcoding_delay_days:</strong> How long PBMC samples were stored in liquid nitrogen&nbsp;prior to being thawed and processed on 10x</li> <li><strong>cell_barcoding_protocol</strong>: Which single cell RNA sequencing experimental protocol was used. Here all samples were processed with 10x v3.1 chemistry.</li> <li><strong>run_lane_batch:</strong> Concatenation of columns <strong>Cell Barcoding Runs</strong> and <strong>Lane</strong> to provide a unique ID for experimental processing batch (i.e. the DNA library)</li> <li><strong>cell_type_level_1:</strong> Level 1 of a 4-tier PBMC ontology that does not provide a label for low quality cells - those were left as NaNs.</li> <li><strong>cell_type_level_2:</strong> Level 2 of a 4-tier PBMC ontology that does not provide a label for low quality cells - those were left as NaNs.</li> <li><strong>cell_type_level_3:</strong> Level 3 of a 4-tier PBMC ontology that does not provide a label for low quality cells - those were left as NaNs.</li> <li><strong>cell_type_level_4:</strong> Level 4 of a 4-tier PBMC ontology that does not provide a label for low quality cells - those were left as NaNs.</li> <li><strong>c1:</strong> Level 1 of a 4-tier PBMC ontology that also provides label for low quality cells, such as Debris, Doublets, experimental artifacts and others. These additional labels can be used for creating a junk detector.</li> <li><strong>c2:</strong> Level 1 of a 4-tier PBMC ontology that also provides label for low quality cells, such as Debris, Doublets, experimental artifacts and others. These additional labels can be used for creating a junk detector.</li> <li><strong>c3:</strong> Level 1 of a 4-tier PBMC ontology that also provides label for low quality cells, such as Debris, Doublets, experimental artifacts and others. These additional labels can be used for creating a junk detector.</li> <li><strong>c4:</strong> Level 1 of a 4-tier PBMC ontology that also provides label for low quality cells, such as Debris, Doublets, experimental artifacts and others. These additional labels can be used for creating a junk detector.</li> <li><strong>original_sample_id:</strong> Some samples were derived from the same originating whole blood sample. This field specifies the source of the whole blood sample.</li> </ul>

restrictedJun 2023View details →
zenodo12/100

Manually annotated video frames of bumblebees in flying arena

<p>Bounding boxes of</p> <ul> <li>Blue flowers</li> <li>Yellow flowers</li> <li>Bumblebees</li> <li>Bumblebees on flowers</li> </ul> <p>Annotated using CVAT.</p>

restrictedJan 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record