Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
zenodo40/100

Fig. 1 in Relationships Of The Heteronchocleidids (Heteronchocleidus, Eutrianchoratus And Trianchoratus) As Inferred From Ribosomal Dna Nucleotide Sequence Data

Fig. 1. Neighbour joining (NJ) tree constructed by PAUP* using partial 28S rDNA sequences (D1 domain) with Diplectanum spp. and Gyrodactylus spp. as outgroups. Percentages of the bootstrap values for neighbour joining (NJ)/maximum parsimony (MP) (NJ & MP=1,000 replicates) are shown along the branches. Bootstrap values lower than 50 are given as dashes (-).

opencc-by-4.0Aug 2011View details →
zenodo40/100

Solar flare forecasting based on magnetogram sequences learning with MViT and data augmentation

<p><strong>Source codes and dataset of the research "Solar flare forecasting based on magnetogram sequences learning with MViT and data augmentation".</strong></p><p>Our work employed PyTorch, a framework for training Deep Learning models with GPU support and automatic back-propagation, to load the MViTv2 s models with Kinetics-400 weights. To simplify the code implementation, eliminating the need for an explicit loop to train and the automation of some hyperparameters, we use the PyTorch Lightning module. The inputs were batches of 10 samples with 16 sequenced images in 3-channel resized to 224 × 224 pixels and normalized from 0 to 1.</p><p>Most of the papers in our literature survey split the original dataset chronologically. Some authors also apply k-fold cross-validation to emphasize the evaluation of the model stability. However, we adopt a hybrid split taking the first 50,000 to apply the 5-fold cross-validation between the training and validation sets (known data), with 40,000 samples for training and 10,000 for validation. Thus, we can evaluate performance and stability by analyzing the mean and standard deviation of all trained models in the test set, composed of the last 9,834 samples, preserving the chronological order (simulating unknown data).</p><p>We develop three distinct models to evaluate the impact of oversampling magnetogram sequences through the dataset. The first model, Solar Flare MViT (SF MViT), has trained only with the original data from our base dataset without using oversampling. In the second model, Solar Flare MViT over Train (SF MViT oT), we only apply oversampling on training data, maintaining the original validation dataset. In the third model, Solar Flare MViT over Train and Validation (SF MViT oTV), we apply oversampling in both training and validation sets.</p><p>We also trained a model oversampling the entire dataset. We called it the "SF_MViT_oTV Test" to verify how resampling or adopting a test set with unreal data may bias the results positively.</p><p><strong>GitHub version</strong></p><p>The .zip hosted here contains all files from the project, including the checkpoint and the output files generated by the codes. We have a clean version hosted on GitHub (<a href="https://github.com/lfgrim/SFF_MagSeq_MViTs">https://github.com/lfgrim/SFF_MagSeq_MViTs</a>), without the magnetogram_jpg folder (which can be downloaded directly on <a href="https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip">https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip)</a> and the output and checkpoint files. Most code files hosted here also contain comments on the Portuguese language, which are being updated to English in the GitHub version.</p><p><strong>Folders Structure</strong></p><p>In the Root directory of the project, we have two folders:&nbsp;</p><ul><li>magnetogram_jpg: holds the source images provided by Space Environment Artificial Intelligence Early Warning Innovation Workshop through the link <a href="https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip">https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip. </a>It comprises 73,810 samples of high-quality magnetograms captured by HMI/SDO from 2010 May 4 to 2019 January 26. The HMI instrument provides these data (stored in hmi.sharp_720s dataset), making new samples available every 12 minutes. However, the images from this dataset were collected every 96 minutes. Each image has an associated magnetogram comprising a ready-made snippet of one or most solar ARs. It is essential to notice that the magnetograms cropped by SHARP can contain one or more solar ARs classified by the National Oceanic and Atmospheric Administration (NOAA).</li><li>Seq_Magnetogram: contains the references for source images with the corresponding labels in the next 24 h. and 48 h. in the respectively M24 and M48 sub-folders.<ul><li>M24/M48: both present the following sub-folders structure:<ul><li>Seqs16;</li><li>SF_MViT;</li><li>SF_MViT_oT;</li><li>SF_MViT_oTV;</li><li>SF_MViT_oTV_Test.</li></ul></li></ul></li></ul><p>There are also two files in root:</p><ul><li>inst_packages.sh: install the packages and dependencies to run the models.</li><li>download_MViTS.py: download the pre-trained MViTv2_S from PyTorch and store it in the cache.</li></ul><p>M24 and M48 folders hold reference text files&nbsp;(flare_Mclass...) linking the images in the magnetogram_jpg folders or the sequences (Seq16_flare_Mclass...)&nbsp; in the Seqs16 folders with their respective labels. They also hold "cria_seqs.py" which was responsible for creating the sequences and "test_pandas.py" to verify head info and check the number of samples categorized by the label of the text files. All the text files with the prefix "Seq16" and inside the Seqs16 folder were created by "criaseqs.py" code based on the correspondent "flare_Mclass" prefixed text files.</p><p>Seqs16 folder holds reference text files, in which each file contains a sequence of images that was pointed to the magnetogram_jpg folders.</p><p>All SF_MViT... folders hold the model training codes itself (SF_MViT...py) and the corresponding job submission (jobMViT...), temporary input (Seq16_flare...),&nbsp;output (saida_MVIT... and MViT_S...), error (err_MViT...) and checkpoint files (sample-FLARE...ckpt). Executed model training codes generate output, error, and checkpoint files. There is also a folder called "lightning_logs" that stores logs of trained models.</p><p><strong>Naming pattern for the files:</strong></p><ul><li>magnetogram_jpg: follows the format<i> </i>"hmi.sharp_720s.&lt;SHARP-ID&gt;.&lt;date&gt;.magnetogram.fits.jpg" and</li><li>Seqs16: follows the format "hmi.sharp_720s.<i>&lt;</i>SHARP-ID<i>&gt;</i>.&lt;init-date&gt;.to.&lt;end-date&gt;", where:<ul><li>hmi: is the instrument that captured the image</li><li>sharp_720s: is the database source of SDO/HMI.</li><li>&lt;SHARP-ID&gt;: is the identification of SHARP region, and can contain one or more solar ARs classified by the (NOAA).</li><li>&lt;date&gt;: is the date-time the instrument captured the image in the format yyyymmdd_hhnnss_TAI (y:year, m:month, d:day, h:hours, n:minutes, s:seconds).</li><li>&lt;init-date&gt;: is the date-time when the sequence starts, and follow the same format of &lt;date&gt;.</li><li>&lt;end-date&gt;: is the date-time when the sequence ends, and follow the same format of &lt;date&gt;.</li></ul></li><li>Reference text files in M24 and M48 or inside SF_MViT... folders follows the format "&lt;prefix&gt;flare_Mclass_&lt;forecasting-horizon&gt;_&lt;dataset&gt;.txt&lt;over&gt;", where:<ul><li>&lt;prefix&gt;: is Seq16 if refers to a sequence, or void if refers direct to images.</li><li>&lt;forecasting-horizon&gt;: "24h" or "48h".</li><li>&lt;dataset&gt;: is "TrainVal&lt;n&gt;" or "Test". The &lt;n&gt; refers to the split of Train/Val.</li><li>&lt;over&gt;: void or "_over" after the extension (...txt_over): means temporary input reference that was over-sampled by a training model.</li></ul></li><li>All SF_MViT...folders:<ul><li>Model training codes: "SF_MViT_&lt;oversampling-type&gt;_M+_&lt;forecasting-horizon&gt;_&lt;split-type&gt;&lt;gpu-type&gt;", where:<ul><li>&lt;oversampling -type&gt;: void or "oT" (over Train) or "oTV" (over Train and Val) or "oTV_Test" (over Train, Val and Test);</li><li>&lt;forecasting-horizon&gt;: "24h" or "48h";</li><li>&lt;split-type&gt;: "oneSplit" for a specific split or "allSplits" if run all splits.</li><li>&lt;gpu-type&gt;: void is default to run 1 GPU or "2gpu" to run into 2 gpus systems;</li></ul></li><li>Job submission files: "jobMViT_&lt;queue&gt;", where:<ul><li>&lt;queue&gt;: point the queue in Lovelace environment hosted on CENAPAD-SP (<a href="https://www.cenapad.unicamp.br/parque/jobsLovelace">https://www.cenapad.unicamp.br/parque/jobsLovelace</a>)</li></ul></li><li>Temporary inputs: "Seq16_flare_Mclass_&lt;forecasting-horizon&gt;_&lt;dataset&gt;.txt&lt;over&gt;:<ul><li>&lt;dataset&gt;: train or val;</li><li>&lt;over&gt;: void or "_over" after the extension (...txt_over): means temporary input reference that was over-sampled by a training model.</li></ul></li><li>Outputs: "saida_MViT_Adam_10-7&lt;split&gt;", where:<ul><li>&lt;split&gt;: k0 to k4, means the correlated split of the output, or void if the output is from all splits.</li></ul></li><li>Error files: "err_MViT_Adam_10-7&lt;split&gt;", where:<ul><li>&lt;split&gt;: k0 to k4, means the correlated split of the error log file, or void if the error file is from all splits.</li></ul></li><li>Checkpoint files: "sample-FLARE_MViT_S_10-7-epoch=&lt;n-epoch&gt;-valid_loss=&lt;loss-value&gt;-Wloss_k=&lt;n-split&gt;.ckpt", where:<ul><li>&lt;n-opoch&gt;: epoch number of the checkpoint;</li><li>&lt;loss-value&gt;: corresponding valid loss;</li><li>&lt;n-split&gt;: 0 to 4.</li></ul></li></ul></li></ul>

opencc-by-4.0Nov 2023View details →
zenodo40/100

FIG. 1 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations

FIG. 1.—Map of the Hawaiian Islands with collection sitesfor Hawaiian hoary bat tissues used inthis study. Sites with n&gt; 1 are denoted with an asterisk.

opencc-by-4.0Aug 2020View details →
zenodo40/100

FIG. 2 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations

FIG. 2.—PCA result plot showing clustering of individual bats from four Hawaiian Islands using 21,808,031 SNPs. Sample information included in supplementary table S4, Supplementary Material online.

opencc-by-4.0Aug 2020View details →
zenodo40/100

FIG. 4 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations

FIG. 4.—SNAPP-based phylogenetic tree inference. (A) The maximum clade credibility or consensus tree, showing approximate divergence of hoary bats across the Hawaiian archipelago. The axis on the bottom of the figure corresponds to million years before present (Ma), using the emergence of Hawai'i (~0.43 Ma) as a calibration point (95% confidence intervals were given in square brackets). (B) The drawing of all sampled trees showing all ingroup nodes were supported by maximum posterior probabilities (1.00).

opencc-by-4.0Aug 2020View details →
dryad40/100

Code and sequence data pertaining to: A phylogenomic perspective on interspecific competition

<p>Evolutionary processes may have substantial impacts on community assembly, but evidence for phylogenetic relatedness as a determinant of interspecific interaction strength remains mixed. In this perspective, we consider a possible role for discordance between gene trees and species trees in the interpretation of phylogenetic signal in studies of community ecology. Modern genomic data show that the evolutionary histories of many taxa are better described by a patchwork of histories that vary along the genome rather than a single species tree. If a subset of genomic loci harbor trait-related genetic variation, then the phylogeny at these loci may be more informative of interspecific trait differences than the genome background. We develop a simple method to detect loci harboring phylogenetic signal and demonstrate its application through a proof of principle analysis of Penicillium genomes and pairwise interaction strength. Our results show that phylogenetic signal that may be masked genome-wide could be detectable using phylogenomic techniques and may provide a window into the genetic basis for interspecific interactions.</p>

opencc-zeroDec 2023View details →
zenodo40/100

Supplementary material for "Exploring Conceptual Data Modeling Processes: Insights from Clustering and Visualizing Modeling Sequences"

<p>This material supplements the following conference publication:</p> <p>Winkler, Rosenthal, Strecker (2024). "Exploring Conceptual Data Modeling Processes: Insights from Clustering and Visualizing Modeling Sequences". Modellierung 2024.</p>

opencc-by-4.0Dec 2023View details →
dryad40/100

Data from: A cost-effective blood DNA methylation-based age estimation method in domestic cats, Tsushima leopard cats (Prionailurus bengalensis euptilurus), and Panthera species, using targeted bisulfite sequencing and machine learning models

<p><span>Knowledge of individual age can help both in-situ and ex-situ conservation programs to design more efficient and suitable management plans for targeted wildlife species. DNA methylation is one of the epigenetic aging markers that has emerged as a promising tool that can estimate age with high accuracy using only a tiny amount of biological material, which can be collected in a minimally invasive way. Here, we sequenced five targeted genetic regions and used </span><span>8–23</span><span> selected CpG sites to build age estimation models with machine learning methods </span><span>with about only $3–7 per sample</span><span>, using blood samples of seven Felidae species—ranging from small to big, and domestic to endangered species: domestic cats (<em>Felis catus</em>, 139 samples), Tsushima leopard cats (<em>Prionailurus bengalensis euptilurus</em>, 84 samples), and five<em> Panthera </em>species (96 samples). </span><span>The models built achieved satisfactory accuracy—the mean absolute error of the best models was 1.966, 1.348, and 1.552 years in domestic cats, Tsushima leopard cats, and <em>Panthera</em> spp., respectively.</span><span> Our models in domestic cats and Tsushima leopard cats were applicable to individuals regardless of health conditions, indicating the high applicability of our models to samples collected from diverse situations, e.g., rescued individuals in the context of conservation. We also showed the possibility of developing universal age estimation models for the five<em> Panthera</em> spp. using two of the five genetic regions, suggesting an even lower cost to use our models for future applications.</span></p>

opencc-zeroJan 2024View details →
zenodo40/100

Station Data and Earthquake Catalogs - Distinct yet adjacent earthquake sequences near the Mendocino Triple Junction: 20 December 2021 Mw 6.1 and 6.0 Petrolia, and 20 December 2022 Mw 6.4 Ferndale

<p>Supplemental Material for publication from The Seismic Record (TSR):</p> <div> <div> <div> <p>Yoon, C. E. and D. R. Shelly (2024). Distinct Yet Adjacent Earthquake Sequences near the Mendocino Triple Junction: 20 December 2021 Mw 6.1 and 6.0 Petrolia, and 20 December 2022 Mw 6.4 Ferndale, The Seismic Record. 4(1), 81&ndash;92, doi: 10.1785/0320230053.</p> <p>Data Sets S0-S4 with station data and earthquake catalogs in text format</p> <p>See README_Data_Supplement.pdf for more details about contents of each data file.&nbsp; Please refer to the accompanying publication and its supplement for figures, tables, and equations.</p> </div> </div> </div> <p>&nbsp;</p>

opencc-by-4.0Feb 2024View details →
zenodo40/100

16S rRNA Sequencing Data of Fecal Microbiota in an Italian Cohort of Patients with CDKL5 Deficiency Disorder

<h3>Summary of the study&nbsp;</h3> <p>CDKL5 deficiency disorder (CDD) is a neurodevelopmental condition characterized by global developmental delay, early-onset seizures, intellectual disability, visual and motor impairments, distinct from Rett Syndrome (RTT) due to the absence of a clear regression period. Gastrointestinal (GI) disturbances and signs of subclinical immune dysregulation are common in CDD patients, yet the underlying causes are unknown. Recent studies hint at a possible link between neurological disorders and gut microbiota, an unexplored area in CDD. In this groundbreaking study, we examined fecal microbiota in CDD patients and their healthy relatives, revealing differences in bacterial diversity and composition. We further investigated microbiota changes based on various factors, including the severity of GI issues, seizure frequency, sleep disorders, food intake type, neuro-behavioral features (assessed through the RTT Behaviour Questionnaire &ndash; RSBQ), and ambulation capacity.&nbsp;</p> <p>Our findings suggest a potential connection between CDD, microbiota, and symptom severity. This study represents the first exploration of the gut-microbiota-brain axis in CDD patients, contributing to the growing body of research on the role of gut microbiota in neurodevelopmental disorders. It opens doors to potential interventions targeting intestinal microbes to enhance the well-being of individuals with CDD.</p> <h3>Mehods</h3> <p>The Dataset represent the raw data (.fastq) obtained from the sequencing of the fecal samples from 17 Italian Patients with CDD, and 17 Healthy Relatives (i.e. siblings or mother), collected at a single time-point.</p> <p>Samples from Patients affected by CDD are called CDD, samples from Healthy Relatives are called HC-CDD (i.e. healthy controls of patients affected by CDD). For details about the sample names see the &ldquo;Explanation Table&rdquo;.</p> <p>Bacterial DNA was extracted from fecal samples using the QIAmp Powerfexal DNA Kit (Qiagen, Germany) following the manufacturer's protocol. The 16S rRNA sequencing and analysis was performed by a service offered by Zymo Research (Germany).</p> <p><em>Targeted Library Preparation</em>: The DNA samples were prepared for targeted sequencing with the Quick-16S&trade; NGS Library Prep Kit (Zymo Research). The primer sets used were Quick-16S&trade; Primer Set V3-V4 (Zymo Research). The sequencing library was prepared using an innovative library preparation process in which PCR reactions were performed in real-time PCR machines to control cycles and therefore limit PCR chimera formation. The final PCR products were quantified with qPCR fluorescence readings and pooled together based on equal molarity. The final pooled library was cleaned up with the Select-a-Size DNA Clean &amp; Concentrator&trade;, then quantified with TapeStation&reg; (Agilent Technologies, Santa Clara, CA) and Qubit&reg; (Thermo Fisher Scientific, Waltham, WA).&nbsp;&nbsp;</p> <p><em>Sequencing:</em> The final library was sequenced on Illumina&reg; MiSeq&trade; with a v3 reagent kit (600 cycles).&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

Data from: "Rare earth elements sediment analysis tracing anthropogenic activities in the stratigraphic sequence of Alagankulam (India)"

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo40/100

Single-cell sequencing data of human umbilical cord and placental mesenchymal stem cells

<p>Expression matrix of umbilical cord and placenta single-cell sequencing data from the same donor.Table1 is the umbilical cord and Table2 is the placenta.</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Experimental data on "Sediment storage and fluvial sediment transport linkages across an experimental flood sequence"

<p>The repository contains data used in manuscript "Sediment storage and fluvial sediment transport linkages across an experimental flood sequence" by Hassan, Pierce, Chartrand.&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Figure 7 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies

Figure 7. Summary of Random Forests classifications for each empirical comparison. Each row shows results from the stratum with the smallest fraction of individuals correctly classified, with comparisons labeled by their taxonomic codes as listed in Table 2. Colors identify comparison type as species (blue), subspecies (green), and populations (red). Points show the fraction of individuals correctly classified with probabilities&gt; 50% (PD50, circles), and&gt; 95% (PD95, triangles). Thin colored lines show 95% confidence intervals (CI) around PD50 estimates. Gray bars show range of a priori random classification rates based on individual size (left) to maximum possible classification rates based on shared haplotypes (right).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Figure 6 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies

Figure 6. Frequency distributions of the change in observed diagnosability (x-axis) in the simulated data for increasing levels of the probability of misstratification (vertical panels). Figures on the left and right columns are censored by data sets for original diagnosability ≤50% and&gt;50%, respectively.

opencc-by-4.0Jun 2017View details →
zenodo40/100

Figure 5 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies

Figure 5. Two-dimensional GAM fits of theta (Ɵ), number of migrants (Nem), and divergence time in generations (T) from Model 2 simulated data. From left to right, columns show results from models without migration (m = 0), with migration and Nem &lt;1, and Nem ≥ 1. Colors indicate model prediction of percent correctly classified.

opencc-by-4.0Jun 2017View details →
zenodo40/100

Figure 4 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies

Figure 4. GAM fit of number of migrants (Nem) from Model 2 parameters. Solid line shows median value of predicted percent correctly classified, and shaded area shows 95% CI. The switch from bimodal distribution to a normal distribution occurs at Nem = 1 (log10Nem = 0).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Figure 3 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies

Figure 3. Two-dimensional GAM fits of effective population size (Ne), divergence time in generations (T), and mutation rate (µ) from Model 1 simulated data. Results from models without migration to the left and those with migration to the right. Colors indicate model prediction of percent correctly classified.

opencc-by-4.0Jun 2017View details →
zenodo40/100

Figure 1 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies

Figure 1. (A) Distribution of a hypothetical character for two putative subspecies (red and blue) demonstrating minimum overlap necessary to satisfy 75% rule of Amadon (1949). Character is continuous on the x-axis. Dashed lines indicate the point at which 75% of each distribution is outside of 99%+ of the other. Solid line indicates point of overlap where 97% of both distributions are outside one another. (B) Probability of membership to subspecies for specimens having values along the character axis. Probability is based on the ratio of the distribution frequencies at each point along the x-axis, with a 50:50 probability occurring at the threshold point.

opencc-by-4.0Jun 2017View details →
dryad40/100

Green turtle ddRAD raw sequencing data

<p>The occasional westward transport of warm water of the Agulhas Current, 'Agulhas leakage', around southern Africa has been suggested to facilitate tropical marine connectivity between the Atlantic and Indian oceans, but the 'Agulhas leakage' hypothesis doesn't explain the signatures of eastward gene flow observed in many tropical marine fauna. We investigated an alternative hypothesis: the establishment of a warm-water corridor during comparatively warm interglacial periods. The 'warm-water corridor' hypothesis was investigated by studying the population genomic structure of Atlantic and Southwest Indian Ocean green turtles (<i>N </i><span>= 27) </span>using 12,035 genome-wide single nucleotide polymorphisms (SNPs) obtained via ddRAD sequencing. Model-based and multivariate clustering suggested a hierarchical population structure with two main Atlantic and Southwest Indian Ocean clusters, and a Caribbean and East Atlantic sub-cluster nested within the Atlantic cluster. Coalescent-based model selection supported a model where Southwest Indian Ocean and Caribbean populations diverged from the East Atlantic population during the transition from the last interglacial period (130 – 115 thousand years ago; kya) to the last glacial period (115 – 90 kya). The onset of the last glaciation appeared to isolate Atlantic and Southwest Indian Ocean green turtles into three refugia, which subsequently came into secondary contact in the Caribbean and Southwest Indian Ocean when global temperatures increased after the Last Glacial Maximum. Our findings support the establishment of a warm-water corridor facilitating tropical marine connectivity between the Atlantic and Southwest Indian Ocean during warm interglacials.</p>

opencc-zeroDec 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record