Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
68
datasets available to search
ShareScore release 0.9.0
Dataset results
68 results for “De novo genome assembly”
Introduction to Ancient Metagenomics Textbook (Edition 2025): de novo Genome Assembly
<p>Data and conda software environment file for the chapter '<em>de novo</em> Genome Assembly' of the SPAAM Community's textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>
De novo genome assembly of the meadow brown butterfly, Maniola jurtina
<p>1. Whole-genome GFF file (raw and filtered for min. gene length) [<em>Maniola.jurtina.gff3</em>, <em>Maniola_jurtina_filtered.gff3</em>]</p> <p>2. List of <em>M. jurtina</em> proteins [<em>Mjurtina_proteins.fa</em>].</p> <p>3. Results of spot pattern genes BLAST against <em>M. jurtina</em> proteome [<em>Lepidoptera_MJ_protein_matches.xlsx</em>]. </p> <p>4. Annotations [blast2go_export.txt]</p>
Draft de novo genome assemblies of a male and female Amphibolurus muricatus (jacky dragon)
<p>Four de novo nuclear genome assemblies of <em>Amphibolurus muricatus</em></p> <p><strong>Assembly 1.0: A 10x Genomics linked-read sequencing assembly</strong><br> • AmpMurF_1.0.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_1.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 1.1: Further scaffolding of assembly 1.0 using RNA-seq data</strong><br> • AmpMurF_1.1.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_1.1.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 2.0: Further scaffolding of assembly 1.0 using SLR-superscaffolder</strong><br> • AmpMurF_2.0.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_2.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 3.0: An stLFR linked-read sequencing assembly</strong><br> • AmpMurF_3.0.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_3.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Methods<br> Assembly 1.0: A 10x Genomics linked-read sequencing assembly</strong><br> Male and female <em>A. muricatus</em> genome sequencing libraries were constructed on the Chromium system (10x Genomics, Pleasanton, CA, USA) by the Ramaciotti Centre for Genomics (Sydney, Australia). The Chromium instrument enables unique barcoding of long stretches of DNA on gel beads. The barcodes allow later reconstruction of long DNA fragments from a series of short DNA fragments with the same barcode (i.e., linked-reads). After barcoding, DNA was sheared into smaller fragments and sequenced on the NovaSeq 6000 platform (Illumina, CA, USA) to generate 151 bp paired-end (PE) reads. A total of 904.9 M raw 10x Genomics Chromium linked-reads were generated. Raw 10x data were assembled with Supernova v2.1.1 (Weisenfeld et al., 2017) and a FASTA file was generated using the ‘pseudohap style’ option in Supernova mkoutput. All female (~450 M) and male (~550 M) read pairs were utilised (female sequencing depth ca 50.3×; male, ca 47.8×). The resulting assemblies was further scaffolded with ARKS v1.0.3 (Coombe et al., 2018), reusing the 10x reads, and the companion LINKS program (v1.8.7) (Warren et al., 2015). ARKS employs a <em>k</em>-mer approach to map linked barcodes to the contigs in the initial Supernova assembly to generate a scaffold graph with estimated distances for LINKS input. These assemblies were denoted AmpMurF_1.0 (female) and AmpMurM_1.0 (male). We used GapCloser v1.12 (part of SOAPdenovo2) (Luo et al., 2012) to fill gaps in the assembly. GapCloser was run using the parameter -l 150) and clean 10x Genomics reads PE reads. </p> <p><strong>Assembly 1.1: Further scaffolding using RNA-seq data</strong><br> We attempted to improve the v1.0 genome assemblies’ contiguity using RNA-sequencing reads. RNA-seq reads (from brain, ovary, and testis; see below) were filtered (i.e., cleaned) to remove adapters and low-quality reads using Flexbar v3.4.0 and used to further re-scaffold the v1.0 assemblies (FASTA files before gapclosing) with P_RNA_scaffolder (Zhu et al., 2018). The default Flexbar settings discards all reads with any uncalled bases. A final round of scaffolding was performed on the resulting assemblies using L_RNA_scaffolder (Xue et al., 2013). These assemblies were denoted AmpMurF_1.1 (female) and AmpMurM_1.1 (male). As before, GapCloser and clean 10x Genomics reads were used to fill gaps. </p> <p><strong>Assembly 2.0: Further scaffolding using SLR-superscaffolder</strong><br> As an alternative approach, we attempted to improve the v1.0 genome assemblies’ contiguity using SLR-superscaffolder (Guo et al., 2021). Briefly, SLR-superscaffolder employs single tube long fragment read (stLFR) sequencing (Wang et al., 2019) reads (see section below) to generate hybrid genome assemblies. The software was run with default parameters except for PE_SEED_MIN=300 (minimum contig size to fill; default 1000). These assemblies were denoted AmpMurF_2.0 (female) and AmpMurM_2.0 (male). GapCloser and clean stLFR reads (with the barcode removed using https://github.com/BGI-Qingdao/stLFR_barcode_split) were used to fill gaps. </p> <p><strong>Assembly 3.0: An stLFR linked-read sequencing and supernova assembly</strong><br> We also generated independent assemblies for the individuals sequenced on the 10x Genomics Chromium system using single tube long fragment read (stLFR) sequencing (Wang et al., 2019). BGI (Brisbane, Australia) generated ~100×-coverage 100-bp paired-end reads (plus a 42-bp stLFR barcode on the right/_2 read). Low-quality reads, PCR duplicates, and adaptors were removed using SOAPnuke v1.5 (Chen et al. 2018). The stLFRdenovo pipeline (<a href="https://github.com/BGI-biotools/stLFRdenovo">https://github.com/BGI-biotools/stLFRdenovo</a>), which is based on Supernova and customized for stLFR data, was used to generate a <em>de novo</em> genome assembly. The stLFRdenovo tool ‘FillGaps’ was used to fill gaps.</p> <p><strong>References</strong><br> Chen, Y., Chen, Y., Shi, C., Huang, Z., Zhang, Y., Li, S., Li, Y., Ye, J., Yu, C., Li, Z., et al. (2018). SOAPnuke: a MapReduce acceleration-supported software for integrated quality control and preprocessing of high-throughput sequencing data. Gigascience 7, 1-6.<br> Coombe, L., Zhang, J., Vandervalk, B.P., Chu, J., Jackman, S.D., Birol, I., and Warren, R.L. (2018). ARKS: chromosome-scale scaffolding of human genome drafts with linked read kmers. BMC Bioinformatics 19, 234.<br> Guo, L., Xu, M., Wang, W., Gu, S., Zhao, X., Chen, F., Wang, O., Xu, X., Seim, I., Fan, G., et al. (2021). SLR-superscaffolder: a de novo scaffolding tool for synthetic long reads using a top-to-bottom scheme. BMC Bioinformatics 22, 158.<br> Luo, R., Liu, B., Xie, Y., Li, Z., Huang, W., Yuan, J., He, G., Chen, Y., Pan, Q., Liu, Y., et al. (2012). SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience 1, 18.<br> Wang, O., Chin, R., Cheng, X., Wu, M.K.Y., Mao, Q., Tang, J., Sun, Y., Anderson, E., Lam, H.K., Chen, D., et al. (2019). Efficient and unique cobarcoding of second-generation sequencing reads from long DNA molecules enabling cost-effective and accurate sequencing, haplotyping, and de novo assembly. Genome Res 29, 798-808.<br> Warren, R.L., Yang, C., Vandervalk, B.P., Behsaz, B., Lagman, A., Jones, S.J., and Birol, I. (2015). LINKS: Scalable, alignment-free scaffolding of draft genomes with long reads. Gigascience 4, 35.<br> Weisenfeld, N.I., Kumar, V., Shah, P., Church, D.M., and Jaffe, D.B. (2017). Direct determination of diploid genome sequences. Genome Res 27, 757-767.<br> Xue, W., Li, J.T., Zhu, Y.P., Hou, G.Y., Kong, X.F., Kuang, Y.Y., and Sun, X.W. (2013). L_RNA_scaffolder: scaffolding genomes with transcripts. BMC Genomics 14, 604.<br> Zhu, B.H., Xiao, J., Xue, W., Xu, G.C., Sun, M.Y., and Li, J.T. (2018). P_RNA_scaffolder: a fast and accurate genome scaffolder using paired-end RNA-sequencing reads. BMC Genomics 19, 175.</p>
Dataset for "Whole-genome de novo assemblies reveal structural variations and organelle-to-nucleus DNA transfers in Asian and African rice""
<p>DXCWR_O.rufipogon_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. rufipogon</em> DXCWR.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. glaberrima</em> IRGC104165.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. barthii</em> W1411.</p> <p>W1411_O.barthii_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. nivara</em> W2014.</p> <p>W2014_O.nivara_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly W2014_O.nivara_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly W2014_O.nivara_scaffolded_anchored.fa.</p>
Draft de novo genome assembly of the elusive jaguarundi, Puma yagouaroundi
<p>The Puma lineage within the family Felidae consists of three species that last shared a common ancestor around 4.9 million years ago. Whole-genome sequences of two species from the lineage were previously reported: the cheetah (<em>Acinonyx jubatus</em>) and the mountain lion (<em>Puma concolor</em>). The present report describes a whole-genome assembly of the remaining species, the jaguarundi (<em>Puma yagouaroundi</em>). We sequenced the genome of a male jaguarundi with 10X Genomics linked reads and assembled the whole-genome sequence. The assembled genome contains a series of scaffolds that reach the length of chromosome arms and is similar in scaffold contiguity to the genome assemblies of cheetah and puma, with a contig N50 = 100.2 kbp and a scaffold N50 = 49.27 Mbp. We assessed the assembled sequence of the jaguarundi genome using BUSCO, aligned reads of the sequenced individual and another published female jaguarundi to the assembled genome, annotated protein-coding genes, repeats, genomic variants and their effects with respect to the protein-coding genes, and analyzed differences of the two jaguarundis from the reference mitochondrial genome. The jaguarundi genome assembly and its annotation were compared in quality, variants and features to the previously reported genome assemblies of puma and cheetah. Computational analyzes used in the study were implemented in transparent and reproducible way to allow their further reuse and modification.</p>
PacBio HiFi de-novo assembled genome and mitochondrial genome for Orbicella faveolata
<p>Final assembly using Funannotate of <i>Orbicella faveolata</i> from PacBio HiFi reads. For full methods please see the publication. </p>
De novo whole genome assembly of the giant tiger prawn (Penaeus monodon) from Vietnam
<p>Basecalled Nanopore FastQ files for the Vietnamese giant tiger prawn and its genome assemblies.</p> <p>FastQ files (LSK109 sample prep sequenced on a MinION device for 48 hours). Read stats are in *_stat.txt:</p> <p>TP_A.fastq : gDNA was extracted using Zymo quick DNA minikit from ethanol-preserved muscle tissue</p> <p>TP_B.fastq : Same as TP_A.fastq</p> <p>TP_C.fastq : gDNA was extracted using conventional salting out method (longer read length but reduced yield)</p> <p>Assemblies:</p> <p>v1_MaSuRCA.fasta: Assembly using poly-G trimmed Illumina reads</p> <p>v2_NanoporeScaf.fasta : Scaffolding with Nanopore long reads</p> <p>v3_RNA_NanoporeScaf.fasta: Scaffolding of v2 with RNA reads</p> <p>v4_NCBI_Filt.fasta: post NCBI contaminant and carry-over adapter removal (final version)</p> <p>Annotation:</p> <p>Braker2_annotation.gff3.gz: Inititial Braker2 gff3 output</p> <p>Braker2_CDS.fna.gz: Initial Braker2 predicted genes</p> <p>Braker2_prot.faa.gz: Protein translation of Braker2_CDS.fna</p> <p>CAZy.tar.gz: CAZy annotation for four crustacean species</p> <p>Filtered_Gene.tar.gz: List of genes with functional annotation and/or orthologs</p> <p>Interproscan_result.tsv: Raw InterProScan output</p> <p>OrthoFinder2.tar.gz: OrthoFinder2 output. Proteins used to infer orthologs are included as ".faa".</p>
Data from: A de novo chromosome-level genome assembly of Coregonus sp. "Balchen": one representative of the Swiss Alpine whitefish radiation
<p>Salmonids are of particular interest to evolutionary biologists due to their incredible diversity of life-history strategies and the speed at which many salmonid species have diversified. In Switzerland alone, over 30 species of Alpine whitefish from the subfamily Coregoninae have evolved since the last glacial maximum, with species exhibiting a diverse range of morphological and behavioural phenotypes. This, combined with the whole genome duplication which occurred in the ancestor of all salmonids, makes the Alpine whitefish radiation a particularly interesting system in which to study the genetic basis of adaptation and speciation and the impacts of ploidy changes and subsequent rediploidization on genome evolution. Although well curated genome assemblies exist for many species within Salmonidae, genomic resources for the subfamily Coregoninae are lacking. To assemble a whitefish reference genome, we carried out PacBio sequencing from one wild-caught <i>Coregonus sp. "Balchen" </i>from Lake Thun to ~90x coverage. PacBio reads were assembled independently using three different assemblers, Falcon, Canu and wtdbg2 and subsequently scaffolded with additional Hi-C data. All three assemblies were highly contiguous, had strong synteny to a previously published <i>Coregonus</i>linkage map, and when mapping additional short-read data to each of the assemblies, coverage was fairly even across most chromosome-scale scaffolds. Here, we present the first <i>de novo</i>genome assembly for the Salmonid subfamily Coregoninae. The final 2.2 Gb wtdbg2 assembly included 40 scaffolds, an N50 of 51.9 Mb, and was 93.3% complete for BUSCOs. The assembly consisted of ~52% TEs and contained 44,525 genes.</p>
Baseline assemblies for "ntLink: a toolkit for de novo genome assembly scaffolding and mapping using long reads" protocol
<p>ntLink is a flexible <em>de novo</em> genome scaffolding toolkit which can be run in various modes depending on the desired user output, with multiple new functionalities recently introduced. Here, we provide the baseline assembly datasets used in the ntLink protocol paper "ntLink: a toolkit for <em>de novo </em>genome assembly scaffolding and mapping using long reads". The provided assemblies are ABySS (short-read) and Flye (long-read) assemblies of <em>Caenorhabditis elegans </em>genome sequencing data. The ABySS (v2.1.4) assembly utilized paired-end short reads (accession DRR008444), and was run with the following parameters: k=64 l=40 s=1000 q=15 B=10G j=8 kc=3 H=4 S=1000-10000 N=9.The <em>C. elegans</em> Flye (v2.5) assembly was run using Oxford Nanopore long reads (accession SRR10028109) and the following parameters: --nano-raw SRR10028109.fastq -g100m -t48.</p>
De novo assembly of a long-read Amblyomma americanum genome
<p>Genome assemblies of Amblyomma americanum generated from PacBio HiFi sequencing of 50 individual female ticks. This repository contains the unphased diploid assembly generated by the Flye assembler (Arcadia_Amblyomma_americanum_asm001.fasta). In addition, there are two associated fasta files containing sequences generated by submitting the unphased diploid assembly to separation by the Purge_Dups pipeline (purged pseudo-haploid assembly and haplotig assembly).</p> <p>Flye assembler: https://github.com/fenderglass/Flye</p> <p>Purge_Dups pipeline: https://github.com/dfguan/purge_dups</p> <p>NCBI Bioproject: PRJNA932813</p>
De novo assembly of a long-read Amblyomma americanum genome (NCBI/Genbank deposited genome)
<p>Genome assembly of Amblyomma americanum generated from PacBio HiFi sequencing of 50 individual female ticks. This repository contains the phased pseudo-haploid tick genome generated after assembly using Flye, phasing using Purge_Dups, and clean-up using custom python scripts generated in-house. </p> <p>NCBI Bioproject: PRJNA932813</p>
Assemblies for "Linear time complexity de novo long read genome assembly with GoldRush"
<p>GoldRush is a <em>de novo</em> genome assembly algorithm with linear time complexity in the number of input long sequencing reads. We tested GoldRush on Oxford Nanopore Technologies datasets with different base error profiles describing the genomes of three human cell lines (NA24385, HG01243 and HG02055), Oryza sativa (rice), and Solanum lycopersicum (tomato). Here, we provide the assemblies for the GoldRush, Flye, Redbean and Shasta assemblies of these long read datasets.</p>
Draft de novo genome assembly of the elusive jaguarundi, Puma yagouaroundi
Open the record for dataset details and reuse information.
Data from: A de novo chromosome-level genome assembly of Coregonus sp. “Balchen”: one representative of the Swiss Alpine whitefish radiation
Open the record for dataset details and reuse information.
De novo genome assembly of the Tobacco Hornworm moth (Manduca sexta)
<p><strong>We present the new reference genome for M sexta, JHU_Msex_v1.0, applying a combination of modern technologies in a de novo assembly to increase continuity, accuracy, and completeness. The assembly is 470 Mb and is ~25x more continuous than the original assembly, with scaffold N50 >14 Mb. We annotated the assembly by lifting over existing annotations and supplementing with additional supporting RNA-based data for a total of 25,256 genes. The new reference assembly is accessible in annotated form for public use.</strong></p>
de novo genome assembly of the LNCaP human prostate cancer cell line
<p>Whole-genome sequencing reads from the LNCaP human prostate cancer cell line were used to generate a <em>de novo </em>assembly with SGA v0.10.15. Please see https://github.com/sciseim/PCaWGS for associated scripts. Library preparation was performed using a TruSeq Nano DNA kit (Illumina) with a target insert size of 350bp. Paired-end libraries (150bp) were sequenced using a HiSeqX sequencer (Illumina).</p>
de novo genome assembly of the PC3 human prostate cancer cell line
<p>Whole-genome sequencing reads from the PC3 human prostate cancer cell line were used to generate a <em>de novo </em>assembly with SGA v0.10.15. Please see https://github.com/sciseim/PCaWGS for associated scripts. Library preparation was performed using a TruSeq Nano DNA kit (Illumina) with a target insert size of 350bp. Paired-end libraries (150bp) were sequenced using a HiSeqX sequencer (Illumina).</p> <p> </p> <p> </p> <p> </p>
De novo genome assembly of rice varieties using Nanopore long reads
<p>Genome sequences for Sugimura et al. (2024) of the rice (O. sativa) varieties 'Hitomebore' and 'Arroz da Terra.'</p> <p>Yusaku Sugimura, Kaori Oikawa, Yu Sugihara, Hiroe Utsushi, Eiko Kanzaki, Kazue Ito, Yumiko Ogasawara, Tomoaki Fujioka, Hiroki Takagi, Motoki Shimizu, Hiroyuki Shimono, Ryohei Terauchi, Akira Abe. Impact of rice GENERAL REGULATORY FACTOR14h (GF14h) on low-temperature seed germination and its application to breeding. PLoS Genet 20(8): e1011369. https://doi.org/10.1371/journal.pgen.1011369</p> <p>bioRxiv doi: https://doi.org/10.1101/2024.02.16.580620</p>
Chromosome-scale genome assembly and de novo annotation of Alopecurus aequalis.
<p><em>Alopecurus aequalis</em> is a winter annual or short-lived perennial bunchgrass which has in recent years emerged as the dominant agricultural weed of barley and wheat in certain regions of China and Japan, causing significant yield losses. Its robust tillering capacity and high fecundity, combined with the development of both target and non-target-site resistance to herbicides means it is a formidable challenge to food security. Here we report on a chromosome-scale assembly of <em>A. aequalis</em> with a genome size of 2.83 Gb. The genome contained 33,758 high-confidence protein-coding genes with functional annotation. Comparative genomics revealed that the genome structure of <em>A. aequalis</em> is more similar to <em>Hordeum vulgare </em>rather than the more closely related <em>Alopecurus myosuroides</em>. The datasets provided here are the assembly FASTA file (lpAloAequ1.1.prim.cur.20230912.fasta.gz), the high-confidence protein-coding genes (Alaeq_EIv0.2.release_HC_genes.gff3.gz) and the full annotation which includes both low and high confidence features of all biotypes (Alaeq_EIv0.2.release.gff3.gz) </p>
Data From: Oatk - a de novo assembly tool for complex plant organelle genomes
<p>This reposity hosts the data for 195 plant organelle genome assemblies generated in the manuscript "Oatk: a de novo assembly tool for complex plant organelle genomes". The sequence data were produced by the Tree of Life programme at the Sanger Institute, mostly from the Darwin Tree of Life (DToL) project, including 24 monocots, 154 eudicots, 16 mosses and one liverwort. See SAMPLE_LIST file for descriptions of these species.</p> <p>In each species subfolder, below files are included.</p> <ol> <li><code>PLTD.fasta</code> Plastome assembly file in FASTA format</li> <li><code>PLTD.annot.bed</code> Plastome assembly annotation file in BED format</li> <li><code>MITO.fasta</code> Mitogenome assembly file in FASTA format</li> <li><code>MITO.annot.bed</code> Mitogenome assembly annotation file in BED format</li> <li><code>MBG.gfa</code> Genome assembly file in GFA format generated with MBG</li> <li><code>PMAT.gfa</code> Genome assembly file in GFA format generated with OATK</li> <li><code>OATK.gfa</code> Genome assembly file in GFA format generated with PMAT (may not exist)</li> </ol> <p> </p> <p>Updates in the New Version:</p> <p>In the previous version, our raw PacBio HiFi read pre-processing pipeline had screened out some reads that it erroneously thought contained HiFi adapter sequence, which led to the gaps in the Hibiscus plastomes. We now fixed this and have rerun all the assemblies that led to any linear organelle components (37 species). All plastomes remain unchanged except for the three Hibiscuses, which are now also circular. Thirteen mitogenomes changed, with six of them now becoming circular.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.