Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
105
datasets available to search
ShareScore release 0.9.0
Dataset results
105 results for “Ground truth”
segmentation results and ground truth for ComSeg benchmark
<p>This folder displays segmentation results and ground truth cell masks used in the benchmark of the paper: A point cloud segmentation framework for image-based spatial transcriptomics, Defard et al. </p> <p><br>The mouse iluem dataset raw images can be downloaded from the study : Petukhov, V., Xu, R.J., Soldatov, R.A. et al. Cell segmentation in imaging-based spatial transcriptomics. Nat Biotechnol 40, 345–354 (2022). https://doi.org/10.1038/s41587-021-01044-w</p> <p> </p> <p>the Vizgen MERFISH breast cancer dataset was sampled from MERSCOPE FFPE Human Immuno-oncology, https://vizgen.com/data-release-program/ : Breast cancer</p> <p> </p> <p> </p>
GT4HistOCR: Ground Truth for training OCR engines on historical documents in German Fraktur and Early Modern Latin
<p><strong>GT4HistOCR</strong> contains ground truth for research in Optical Character Recognition (OCR) technology applied to historical printings in German Fraktur and Early Modern Latin.</p> <p>The ground truth comes in pairs of images of single printed lines as they appear in book pages (*.png) and their corresponding diplomatic transcriptions (*.gt.txt), which are UTF-8 strings preserving the character forms (glyphs) as much as possible within the UNICODE standard. These pairs of line images and their transcriptions can be directly used to train recognition models with, e.g., the open source OCR engines <em>OCRopy</em> or <em>Tesseract</em>. A total of 313,173 ground truth lines are provided.</p> <p><strong>Please note that the subcorpora making up this collection used different transcription guidelines, so it is a bad idea to train a recognition model on the total collection! Rather train individual models for each subcorpus.</strong> Fur further information about the subcorpora, please see the README file and the accompanying publication.</p> <p>If these data are useful for you, please cite the accompanying publication:</p> <pre>@article{<a href="http://springmann.net/publications.html#springmann2018gt4hist">springmann2018gt4hist</a>, author = {Uwe Springmann and Christian Reul and Stefanie Dipper and Johannes Baiter}, title = {{Ground Truth for training {OCR} engines on historical documents in German Fraktur and Early Modern Latin}}, journal = {J. Lang. Technol. Comput. Linguistics}, volume = {33}, number = {1}, pages = {97--114}, year = {2018}, url = {https://jlcl.org/content/2-allissues/1-heft1-2018/jlcl_2018-1_5.pdf} }</pre>
Ground truth matches between maps with high disparity
<p>File describing the correct matches between the region of sketch maps and their model map, or between partial sensor built maps.</p> <p>The region were built using MAORIS segmentation algorithm</p>
TRIC Crawler Internal Localization on Pipe Segment Verification by External Ground Truth System
<p>Magnetic Crawling Robot on horizontal pipe segment</p> <p>External Trackikng with Optitrack System</p> <p>Internal Tracking with Odometry and Gravitation sensors</p> <p>Images, Plots, Movie, Raw-Data</p>
Ground truth for Neue Zürcher Zeitung black letter period
<p>The Neue Zürcher Zeitung (NZZ) has been publishing in black letter from its very first issue in 1780 until 1947. From this time period, we randomly sampled one frontpage per year, resulting in a total of 167 pages. We chose frontpages because they typically contain highly relevant material and because we want to make sure not to sample pages containing exclusively advertisements or stock information. During certain periods, the NZZ was published several times a day, and there were supplements, too. Due to incomplete metadata, the sampling included frontpages from supplements.</p> <p>We then manually corrected the pages, so it can be used as a ground truth to improve the OCR of black letter in historical newspapers.</p>
Bead tracking experimental ground truth for studying size segregation in bedload sediment transport
<p>Video sequences to study size segregation in bedload transport were recorded. Experiments consisted in mixtures of two-size spherical glass beads entrained by a turbulent supercritical free surface water flow over a mobile bed. The aim is to track all beads over time to obtain trajectories, particle velocities and concentrations, for studying bedload granular rheology, size segregation and associated morphology.</p> <p>This upload consists in :</p> <ul> <li>a 1000-frame experimental image sequence recorded at 130 fps with approximately 400 beads per frame (about 300 coarse and 100 small beads). The image resolution is 1280x320;</li> <li>the ground truth in the directory \result . It was obtained based on a tracking algorithm with subsequent expert modification. The tracking algorithm was developed by H. Lafaye de Micheaux et al. The code implementing the tracking algorithm is available on <a href="https://github.com/hugolafaye/BeadTracking">https://github.com/hugolafaye/BeadTracking</a>. The ground truth is a '.mat' file containing in particular the variable 'trackData' being a cell array of tracking matrices. There is one tracking matrix for each image of the sequence. Complete information on data format is given in the file readme.txt in the github BeadTracking package.</li> <li>In addition it contains three files allowing the user to run the BeadTracking package specifically on the experimental sequence : <ul> <li>sequence_param.txt : parameter file</li> <li>sequence_base_mask.tif : to remove the base</li> <li>template_transparent_bead_rOut10_rIn6.mat : a template for bead detection</li> </ul> </li> </ul>
UML Diagram Dataset from the paper Creating and Validating a Ground Truth Dataset of UML Diagrams Using Deep Learning Techniques
<p>Dataset of six UML diagram classes, comprising a total of 2,626 images (426 activity diagrams, 636 class diagrams, 352 component diagrams, 357 deployment diagrams, 435 sequence diagrams, and 420 use case diagrams). Importantly, unlike other existing datasets, ours contains no duplicate elements and all diagrams are correctly classified.</p>
Patrologia Graeca (OCR ground truth)
<p>Ground truth manually produced within the scope of the CGPG project (Calfa GREgORI Patrologia Graeca), led by Jean-Marie Auwers (UCLouvain), that aim to OCRize the remaining non-digital versions of the Patrologia Graeca volumes. This dataset compiles annotations from 2021 to 2022, and has been used for the Programming Historian online lesson "Transcription automatisée de graphies non latines" (link to come).</p> <p>The dataset contains a set of 100 images from the Patrologia Graeca, with their corresponding pageXML files. Annotation has been performed with the <a href="https://vision.calfa.fr">Calfa Vision platform</a>.</p> <p>Different level of annotations are proposed to overcome two tasks: (task1) detection of text regions in Greek (annotations at the region level only) and (task 2) recognition of ancient polytonic Greek (annotations at the line level).</p> <p><strong>Task 1:</strong><br> col_greek: 52<br> col_lat: 54<br> footnotes: 27<br> titles: 9</p> <p><strong>Task 2:</strong><br> lines: 2.579</p> <p>Final model of layout analysis is freely usable on the <a href="https://vision.calfa.fr">Calfa Vision platform</a>, by choosing the "Greek printed (Patrologia Graeca)" type of project.</p> <p>The project is sponsored by the ASBL <em>Byzantion</em>, the Fondation <em>Sedes Sapientiae</em>, the Institut <em>Religions, Spiritualités, Cultures, Sociétés</em> (RSCS, UCLouvain) and the <em>Centre d'études orientales</em> (CIOL, UCLouvain) and by a generous donor who wishes to remain anonymous. Other sponsors have recently expressed their willingness to support the project.</p>
Simulated industrial CT dataset for deep learning with dual-energy tomograms and ground truth material maps for copper and iron
<p>We use this dataset for training and evaluation of a deep learning model to discriminate multi-material systems with X-ray CT.</p> <p>The dataset consists of:</p> <ul> <li>inputs: dual-energy tomograms as binary files without a header (tensor <strong>shape for numpy: 2x128x128 @float32</strong>) <ul> <li>simulated spectra are 250kVp and 450kVp both prefiltered using 2mmCuSn</li> </ul> </li> <li>outputs: the material maps a.k.a. ground truths for the training (same shape as inputs) <ul> <li>sampled with a delaunay algorithm and randomly filled with iron and copper fractions</li> </ul> </li> </ul> <p>The <strong>dataset is normalized to [0, 1]</strong>, so you have to multiply by the mass densities of copper and iron to obtain effective fractions in g/cm^3.</p>
Klosterneuburg, Stiftsbibl., Cod. 48 - Ground Truth: Initial Release
<p>This is ground truth for the vast collection of sermons of Nikolaus von Dinkelsbühl (ca. 1360 to 17th March 1433), translated and reorganised by a German redactor, from the 15th century has never been edited until now. It consists of 361 folios of parchment and paper. The text speaks about various topics such as fasting and other religious practices. Being one of the leading intellectuals of his time, Nikolaus von Dinkelsbühl also contributed to the development of the University of Vienna. The manuscript was probably produced in the vicinity of Klosterneuburg in Austria and is still kept there today (Shelfmark: Cod. 48).</p> <p>Data collection and ground truth creation:</p> <p>The edition at hand was produced by an international team of researchers from various fields in the context of the Vienna HTR Winter School 2022 with the help of Transkribus Expert Client.</p> <p>We uploaded the images of the manuscript into the Transkribus platform, applied the line recognition tool and manually copied the transcribed text lines into the recognised line boxes. Various models were trained with the ground truth (20% of the entire codex) created by the team.</p> <p>Images of the Klosterneuburg, Augustiner-Chorherrenstift, Cod. 48 are available at: <a href="https://manuscripta.at/diglit/AT5000-48/0001">https://manuscripta.at/diglit/AT5000-48/0001</a></p>
A Semantically Annotated 15-Class Ground Truth Dataset for Substation Equipment
<p>This dataset contains 1660 images of electric substations with 50705 annotated objects. The images were obtained using different cameras, including cameras mounted on Autonomous Guided Vehicles (AGVs), fixed location cameras and those captured by humans using a variety of cameras. A total of 15 classes of objects were identified in this dataset, and the number of instances for each class is provided in the following table:</p> <table align="center"> <caption>Object classes and how many times they appear in the dataset.</caption> <thead> <tr> <th scope="col">Class</th> <th scope="col">Instances</th> </tr> </thead> <tbody> <tr> <td>Open blade disconnect</td> <td>310</td> </tr> <tr> <td>Closed blade disconnect switch</td> <td>5243</td> </tr> <tr> <td>Open tandem disconnect switch</td> <td>1599</td> </tr> <tr> <td>Closed tandem disconnect switch</td> <td>966</td> </tr> <tr> <td>Breaker</td> <td>980</td> </tr> <tr> <td>Fuse disconnect switch</td> <td>355</td> </tr> <tr> <td>Glass disc insulator</td> <td>3185</td> </tr> <tr> <td>Porcelain pin insulator</td> <td>26499</td> </tr> <tr> <td>Muffle</td> <td>1354</td> </tr> <tr> <td>Lightning arrester</td> <td>1976</td> </tr> <tr> <td>Recloser</td> <td>2331</td> </tr> <tr> <td>Power transformer</td> <td>768</td> </tr> <tr> <td>Current transformer</td> <td>2136</td> </tr> <tr> <td>Potential transformer</td> <td>654</td> </tr> <tr> <td>Tripolar disconnect switch</td> <td>2349</td> </tr> </tbody> </table> <p>All images in this dataset were collected from a single electrical distribution substation in Brazil over a period of two years. The images were captured at various times of the day and under different weather and seasonal conditions, ensuring a diverse range of lighting conditions for the depicted objects. A team of experts in Electrical Engineering curated all the images to ensure that the angles and distances depicted in the images are suitable for automating inspections in an electrical substation.</p> <p>The file structure of this dataset contains the following directories and files:</p> <p> images: This directory contains 1660 electrical substation images in JPEG format.</p> <p>images: This directory contains 1660 electrical substation images in JPEG format.</p> <ul> <li><strong>labels_json: </strong>This directory contains JSON files annotated in the VOC-style polygonal format. Each file shares the same filename as its respective image in the images directory.</li> <li><strong>15_masks:</strong> This directory contains PNG segmentation masks for all 15 classes, including the porcelain pin insulator class. Each file shares the same name as its corresponding image in the images directory.</li> <li><strong>14_masks:</strong> This directory contains PNG segmentation masks for all classes except the porcelain pin insulator. Each file shares the same name as its corresponding image in the images directory.</li> <li><strong>porcelain_masks:</strong> This directory contains PNG segmentation masks for the porcelain pin insulator class. Each file shares the same name as its corresponding image in the images directory.</li> <li><strong>classes.txt:</strong> This text file lists the 15 classes plus the background class used in LabelMe.</li> <li><strong>json2png.py:</strong> This Python script can be used to generate segmentation masks using the VOC-style polygonal JSON annotations.</li> </ul> <p>The dataset aims to support the development of computer vision techniques and deep learning algorithms for automating the inspection process of electrical substations. The dataset is expected to be useful for researchers, practitioners, and engineers interested in developing and testing object detection and segmentation models for automating inspection and maintenance activities in electrical substations.</p> <p>The authors would like to thank UTFPR for the support and infrastructure made available for the development of this research and COPEL-DIS for the support through project PD-2866-0528/2020—Development of a Methodology for Automatic Analysis of Thermal Images. We also would like to express our deepest appreciation to the team of annotators who worked diligently to produce the semantic labels for our dataset. Their hard work, dedication and attention to detail were critical to the success of this project.</p>
AimSeg ground truth and classifiers
<p>The data is divided into two distinct datasets: one dedicated to mice undergoing remyelination, known as the validation dataset, and the other focusing on a healthy control specimen. These datasets consist of transmission electron microscopy (TEM) images of the corpus callosum (CC) in adult mice. Both datasets are enriched with annotated ground truth information for two key tasks: instance segmentation (identifying individual myelinated axons) and semantic segmentation (discerning axons, inner cytoplasmic tongue, and myelin). Furthermore, we have incorporated ilastik pixel and object classifiers, specifically trained on remyelinating data, into the repository. To streamline usage, the training dataset has been integrated into the corresponding project file.</p>
H01 synapse ground truth
<p>Synapse ground truth used for the H01 study.</p> <p>Ground truth used to train a network for synapse identification (see associated paper for details) ares contained in the six Synapse_GT_6_subvolumes.zip</p> <p>Ground truth used for excitatory vs inhibitory classification of identified synapses (see associated paper for details) is contained in the files ei_gt_from_2945_synapses_from_104_proofread_neurons.csv and ei_gt_from_2367_verified_connections.csv</p>
Ground truth of binary disassembly
<p>Dataset of binary disassembly, including x86/x64, arm32/aarch64, and mips32/mips64 binaries. Ground truth of binary disassembly includes instruction starts, function entries and jump tables.</p>
Ground truth data used to train the synapse classifier used in Lillvis et al., 2022 for ExLLSM circuit reconstruction
<p class="MsoNormal">Brain function is mediated by the physiological coordination of a vast, intricately connected network of molecular and cellular components. The physiological properties of network components can be quantified with high throughput; the ability to assess many animals per study has been key to relating physiological properties to behavior. Conversely, detailed anatomical properties (e.g., the synaptic connectivity of molecularly-defined cell types across an entire circuit) are presently quantifiable only with low throughput; thus we know very little about how network structure, and structural variation, influences behavior. For neuroanatomical reconstruction there is a methodological gulf between electron-microscopic (EM) methods, which yield dense connectomes (but at great expense and low throughput) and light-microscopic methods, which provide molecular and cell-type specificity with high throughput (but without synaptic resolution). We developed a high-throughput analysis pipeline and imaging protocol using tissue expansion and light sheet microscopy (ExLLSM) to rapidly reconstruct selected circuits across many animals with single-synapse resolution and molecular contrast. Using <em>Drosophila </em>to validate this approach, we demonstrate that it yields synaptic counts similar to those obtained by EM, enables synaptic connectivity to be compared across sex and experience, and can be used to correlate structural connectivity, functional connectivity, and behavior. This approach fills a critical methodological gap in studying variability in the structure and function of neural circuits across individuals within and between species.</p> <p class="MsoNormal">Here, we share the data used to train the synapse classifier that was utilized in the analysis pipeline. All additional software, code, and usage examples to train and run the classifier can be found at Github: <a href="https://github.com/JaneliaSciComp/exllsm-circuit-reconstruction">https://github.com/JaneliaSciComp/exllsm-circuit-reconstruction</a></p>
SYNTHETIC VOLUMES AND 3D GROUND TRUTH ANNOTATIONS OF NUCLEI IN 3D MICROSCOPY VOLUMES
<p>Manual ground truth annotations of subvolumes of eight microscopy volumes are provided.</p> <p>Synthetic subvolumes, generated with an SpCycelGAN, of the same volumes are also provided. </p> <p> </p>
Dataset for the paper "Ground-truth Free Evaluation of HTR on Old French and Latin Medieval Literary Manuscripts"
<p>This dataset was used in the context of the article <em>Ground-truth Free Evaluation of HTR on Old French and Latin Medieval Literary Manuscripts</em>, at the Computational Humanities Research 2022 conference.</p> <p>Predictions.zip contains the XML ALTO for the HTR and Segmentation prediction of around 10 pages of 1900 manuscripts from the Bibliothèque nationale de France.</p> <p>Varying ground truth contains the original training material for training a classifier for CER classification (see the article).</p> <p>The ManuscriptsIIIF.csv contains metadata about the manuscripts.</p> <p>This work was funded by the DIM MAP under the CREMMALab funding.</p>
Ground Truthing Survey Data of Crop Type in Pakistan (Rabi 2022‒Kharif 2023)
<p>The dataset comprises ground truthing survey data collected during the winter (Rabi) season of 2022–23 and the summer (Kharif) season of 2023 in Pakistan. These surveys were conducted as part of the Asian Development Bank's (ADB) initiative to support Pakistan's Ministry of National Food Security and Research (MNFSR) and provincial Crop Reporting Service (CRS) departments in adopting technology-based data collection practices. There were 43,892 data points collected during the winter (Rabi) season and 92,951 during the summer (Kharif) season. The data collected is available in the below-mentioned format.</p> <div> <table> <tbody> <tr> <td> <p><strong>Variable Name</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Data Type</strong></p> </td> <td> <p><strong>Example Values</strong></p> </td> </tr> <tr> <td> <p>ID</p> </td> <td> <p>Unique identifier for each data point</p> </td> <td> <p>Text</p> </td> <td> <p>3-324-20-2-19082023-1-1</p> </td> </tr> <tr> <td> <p>Season</p> </td> <td> <p>Season in which data was collected</p> </td> <td> <p>Text</p> </td> <td> <p>Rabi</p> </td> </tr> <tr> <td> <p>Province</p> </td> <td> <p>Name of the province where data was collected</p> </td> <td> <p>Categorical</p> </td> <td> <p>Khyber Pakhtunkhwa</p> </td> </tr> <tr> <td> <p>District</p> </td> <td> <p>Name of the district where data was collected</p> </td> <td> <p>Categorical</p> </td> <td> <p>Malakand</p> </td> </tr> <tr> <td> <p>Date</p> </td> <td> <p>Date showing when the data was collected</p> </td> <td> <p>Date</p> </td> <td> <p>19/08/2023</p> </td> </tr> <tr> <td> <p>Latitude</p> </td> <td> <p>Latitude coordinate of the data point</p> </td> <td> <p>Float</p> </td> <td> <p>34.449521</p> </td> </tr> <tr> <td> <p>Longitude</p> </td> <td> <p>Longitude coordinate of the data point</p> </td> <td> <p>Float</p> </td> <td> <p>71.907877</p> </td> </tr> <tr> <td> <p>Code</p> </td> <td> <p>Numeric code representing specific crop (e.g. Wheat is given code 1)</p> </td> <td> <p>Integer</p> </td> <td> <p>14</p> </td> </tr> <tr> <td> <p>Land</p> </td> <td> <p>Type of land</p> </td> <td> <p>Categorical</p> </td> <td> <p>Rice, Intercropping</p> </td> </tr> <tr> <td> <p>Description</p> </td> <td> <p>Detail of land type</p> </td> <td> <p>Categorical</p> </td> <td> <p>Orchard (Apple)</p> </td> </tr> <tr> <td> <p>Stage</p> </td> <td> <p>Stage of crop at the time of data collection</p> </td> <td> <p>Categorical</p> </td> <td> <p>Reproductive</p> </td> </tr> </tbody> </table> </div> <p> </p> <p> </p>
Artificially-generated Lecture Video Fragmentation Dataset and Ground Truth
<p>We provide a large-scale lecture video dataset consisting of artificially-generated lectures, and the corresponding ground-truth fragmentation, for the purpose of evaluating lecture video fragmentation techniques.</p> <p>For creating this dataset, 1498 speech transcript files (generated automatically by ASR software) were used from the world's biggest academic online video repository, the VideoLectures.NET. These transcripts correspond to lectures from various fields of science, such as Computer science, Mathematics, Medicine, Politics etc. In order to create the synthetic video lectures, all transcripts were randomly split in fragments, the duration of which ranges between 4 and 8 minutes. Each synthetic lecture was then assembled by combining (stitching) exactly 20 randomly selected fragments. 300 such artificially-generated lectures are included in the released dataset. Each such lecture file has a mean duration of about 120 minutes, thus the dataset contains altogether about 600 hours of artificially-generated lectures. Every pair of consecutive fragments in these lectures originally comes from different videos, consequently the point in time where such two fragments are joined is a known ground-truth fragment boundary. All these boundaries form the dataset's ground truth. We should stress that we do not generate the corresponding video files for the artificially-generated lectures (only the transcripts), and one should not try to reverse-engineer the dataset creation process so as to use in some way the visual modality for detecting the fragments in this dataset.</p> <p><strong>File format</strong></p> <p>After you download the provided .zip and unpack it, the extracted folder will contain two sub-folders:</p> <pre><code>1. ALV_srt 2. ALV_srt_GT </code></pre> <p>Each of them contains 300 files.</p> <p>The <strong>ALV_srt</strong> folder contains the transcripts of every artificially-generated lecture, in the standard SRT format:</p> <pre><code>1. A numeric counter identifying each sequential subtitle 2. The time that the subtitle should appear on the screen, followed by --> and the time it should disappear 3. Subtitle's text itself on one or more lines 4. A blank line containing no text </code></pre> <p>The <strong>ALV_srt_GT</strong> folder contains the ground truth (GT) fragments corresponding to the lectures (transcripts) of the <strong>ALV_srt</strong> folder. Each GT file consists of 3 tab-separated columns and 20 rows, in the following format:</p> <pre><code><Fragment_ID_1> <StartTime_1> <EndTime_1> <Fragment_ID_2> <StartTime_2> <EndTime_2> <Fragment_ID_3> <StartTime_3> <EndTime_3> . . . <Fragment_ID_20> <StartTime_20> <EndTime_20> </code></pre> <p>Each row indicates a fragment. The first column indicates the ID of a fragment while the second and the third column indicate the start and the end time of the fragment respectively.</p> <p><strong>License and Citation</strong></p> <p>This dataset is provided for academic, non-commercial use only. If you find this dataset useful in your work, please cite the following publication where the dataset is introduced:</p> <p><em>D. Galanopoulos, V. Mezaris, “Temporal Lecture Video Fragmentation using Word Embeddings”, Proc. 25th Int. Conf. on Multimedia Modeling (MMM2019), Thessaloniki, Greece, Jan. 2019.</em></p> <p><strong>Acknowledgements</strong></p> <p>This work was supported by the EU’s Horizon 2020 research and innovation programme under grant agreement No 693092 MOVING. We are grateful to JSI/VideoLectures.NET for providing the lectures’ transcripts.</p>
Historical Newspapers Ground Truth
<p>This dataset contains 50 pages of ground truth data for digitized historical newspapers from the Berlin State Library for training and evaluation of OCR/OLR systems as produced in the context of the EU ICT-PSP project Europeana Newspapers (http://www.europeana-newspapers.eu/).</p> <p>The dataset comprises of the following resources:</p> <ul> <li><strong>gt_page.zip </strong>Ground Truth files in PAGE-XML format (cf. https://github.com/PRImA-Research-Lab/PAGE-XML)</li> <li><strong>img_full.zip </strong>Full resolution scanned images in TIF format</li> <li><strong>img_bin.zip</strong> Binarized (using the Gatos method) images in TIF format</li> <li><strong>ocr_full.zip</strong> OCR (FineReaderEngine11) results for full resolution images</li> <li><strong>ocr_bin.zip</strong> OCR (FineReaderEngine11) results for binarized images</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.