Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
190
datasets available to search
ShareScore release 0.9.0
Dataset results
190 results for “Targeted enrichment”
Whole-genome capture and sequencing of Mycobacterium tuberculosis directly from clinical samples - Design of RNA oligonucleotide baits for Agilent Technologies' SureSelect target enrichment
<p>This dataset comprises the sequence of <strong>44 278 RNA oligonucleotide "baits" (120 bp each) </strong>designed to perform <strong>whole-genome capture and sequencing of <em>Mycobacterium tuberculosis</em> directly from clinical samples</strong> (DNA) using Agilent Technologies’ SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol. </p> <p>RNA oligonucleotide “baits” were designed to span the ∼4.5 Mb of the <em>M. tuberculosis</em> genome. In brief, the reference genome sequence of the MTBC H37Rv strain (Genbank #AL123456) was <em>in silico</em> fragmented into 120 bp sequences twice, to ensure an overlap of 60 bp between sequences. Due to their rich GC content, which could interfere with DNA capture, all MTBC genes of the PE, PPE and PE-PGRS family were also independently fragmented into 120 bp sequences, in order to increase capture sensitivity. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 42,278 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies. During synthesis, the 2198 sequences complementary to the PE, PPE and PE-PGRS family were unbalanced 8:1 to potentiate capture.</p> <p>More details can be found in the following publication:</p> <p>- Macedo, R., Isidro, J., Ferreira, R., Pinto, M., Borges, V., Duarte, S., Vieira, L., & Gomes, J. P. (2023). Molecular Capture of <em>Mycobacterium tuberculosis</em> Genomes Directly from Clinical Samples: A Potential Backup Approach for Epidemiological and Drug Susceptibility Inferences. <em>International journal of molecular sciences</em>, <em>24</em>(3), 2912. https://doi.org/10.3390/ijms24032912</p>
RESICE - Reusability-targeted Enriched Sea Ice Core Database - Part A
<div> <div>RESICE is described in detail in the article <em>Reusability-targeted enrichment of sea ice core data</em> published on 2025-03-20 in Scientific Data (DOI: <a href="https://doi.org/10.1038/s41597-025-04665-x" target="_blank" rel="noopener">10.1038/s41597-025-04665-x</a>). This is Part A of RESICE. RESICE_PartA.csv contains all data including profile data (several rows per core), RESICE_PartA_cores.csv provides all data excuding profile data (one row per core) and sources_PartA.csv provides a list of all sources. The database including its <a title="RE-SICE Part B" href="https://www.doi.org/10.5281/zenodo.14744942" target="_blank" rel="noopener">Part B</a> is described in the <a title="RE-SICE General Information" href="https://www.doi.org/10.5281/zenodo.14744912" target="_blank" rel="noopener">general information. </a>Part A and Part B had to be separated due to different licenses of the orginal data sources.</div> </div>
RESICE - Reusability-targeted Enriched Sea Ice Core Database - General Information
<div> <div>RESICE is described in detail in the article <em>Reusability-targeted enrichment of sea ice core data</em> published on 2025-03-20 in Scientific Data (DOI: <a href="https://doi.org/10.1038/s41597-025-04665-x" target="_blank" rel="noopener">10.1038/s41597-025-04665-x</a>).</div> <div> </div> <div>A large number of sea ice core data sets are available that have been acquired by research groups around the world and published in different data repositories. The structure of sea ice core data differs substantially across repositories and entries regarding combinations of content, level and quality of description, label names, formats, units, etc. Here, we have compiled sea ice core data and metadata available in data sets (DS) into a tabular database. Additionally, we have added data and metadata from articles (A) and expedition reports (ER). We have enriched the database with metadata from instrument manuals (IM) and controlled terminologies (CT) such as the <a title="SIN" href="https://library.wmo.int/idurl/4/41953" target="_blank" rel="noopener"><em>Sea Ice Nomenclature</em></a> (SIN) from the World Meteorological Organization (WMO) and the <a title="SeaVoX Polygons" href="https://doi.org/10.14284/590" target="_blank" rel="noopener"><em>Polygon data set of water body extent from the SeaVoX Salt and Fresh Water Body Gazetteer</em></a> by the British Oceanographic Data Centre (BODC). We grouped the type of sources into primary sources (DS), secondary sources (A, ER), and tertiary sources (IM, CT). RESICE enhances reusability of the included sea ice core data through enrichment. RESICE provides a comprehensive resource for sea ice modeling applications that aim at using information from compiled sea ice core data. Some examples are calibration and validation of physics-based process models addressing the generation and evolution of sea ice or the training of data-driven models that rely on harmonized training data. As data and metadata are combined from many sources, each entry in the data set needs to be traceable to the original source and its corresponding DOI or URL. Where appropriate, we refer to the original excerpt, figure or table of the original source or comment on inconsistencies or required changes to transparently communicate the entries origin. This is the general information on the database. Please find <a title="RESICE Part A" href="https://www.doi.org/10.5281/zenodo.14745035" target="_blank" rel="noopener">Part A</a> of the database that can be reused under license CC-BY, and <a href="https://www.doi.org/10.5281/zenodo.14744942">Part B</a> of the database that can be reused under license CC-BY-SA. RESICE can be interactively viewed, analyzed and plotted in the <a title="MOSAiC webODV" href="https://mvre.webodv.cloud.awi.de/DataExploration/id/DVevtE7c">MOSAiC webODV</a> instance. RESICE can be reproduced and extended with the <a title="pyresice Python package" href="https://doi.org/10.5281/zenodo.11198658" target="_blank" rel="noopener">pyresice</a> Python package available on <a title="pyresice gitLab" href="https://git.rwth-aachen.de/mbd/pyresice/" target="_blank" rel="noopener">gitLab</a>.</div> </div>
Data from: Bratzel et al. (2022) Target-enrichment sequencing reveals for the first time a well-resolved phylogeny of the core Bromelioideae (Bromeliaceae). Taxon
<p>DNA sequence alignments used for phylogenetic analyses in Bratzel et al. (2022) Target-enrichment sequencing reveals for the first time a well-resolved phylogeny of the core Bromelioideae (Bromeliaceae). Taxon.</p>
Linked collectors and determiners for: Combining target enrichment and Sanger sequencing data to clarify the systematics of the diverse Neotropical butterfly subtribe Euptychiina (Nymphalidae, Satyrinae).
Natural history specimen data linked to collectors and determiners held within, "Combining target enrichment and Sanger sequencing data to clarify the systematics of the diverse Neotropical butterfly subtribe Euptychiina (Nymphalidae, Satyrinae)". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/bfb878f3-8a74-46d3-a104-36485c32aaba">https://bionomia.net/dataset/bfb878f3-8a74-46d3-a104-36485c32aaba</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/bfb878f3-8a74-46d3-a104-36485c32aaba">https://gbif.org/dataset/bfb878f3-8a74-46d3-a104-36485c32aaba</a>. Formatted as a Frictionless Data package.
Fig. 4 in Phylogenomics of the tropical plant family Ochnaceae using targeted enrichment of nuclear genes and 250+ taxa
Fig. 4. Distribution of specimen ages and the number of loci recovered in the phylogenomic study of Ochnaceae. A, Histogram of the collection years of all Ochnaceae specimens; B & C, Relationship between the year of collection of the specimens and the number of loci recovered for tissue obtained from herbarium material (excluding specimens with silica-dried leaf material), analysed for Ochneae and all the remaining Ochnaceae separately, either using the consensus-alignment (B) or the sample-specific (C) reference-based assembly approach. Pearson correlation coefficients and confidence intervals are given for each group.
Fig. 2 in Phylogenomics of the tropical plant family Ochnaceae using targeted enrichment of nuclear genes and 250+ taxa
Fig. 2. RAxML trees based on the concatenated nuclear loci of Ochnaceae. A, Early-diverging branches of Ochnaceae and relationships within Quiinoideae based on the FAM dataset; B, Phylogenetic relationships of Sauvagesieae, Luxemburgieae and Testuleeae based on the SLT dataset. Numbers on the branches are bootstrap values>50%. Numbers in parentheses after species names correspond to the specimen IDs (only for species with multiple accessions). The indicated classification of subfamilies and tribes follows Schneider & al. (2014).
Fig. 1 in Phylogenomics of the tropical plant family Ochnaceae using targeted enrichment of nuclear genes and 250+ taxa
Fig. 1. Overview of the phylogenetic relationships of the major clades of Ochnaceae based on the FAM dataset together with images of representative species. The classification follows Schneider & al. (2014). Ochninae is by far the most species-rich clade comprising about two-thirds of the family's species and six genera (Brack. = Brackenridgea; Cmp. = Campylospermum, clades A and B; I. = Idertia; Ochna; Ouratea; Rh. = Rhabdophyllum). Letters around the tree refer to the photos (mostly flowers except where indicated) and the relative position of the displayed taxa on the tree. A, Medusagyne oppositifolia (Medusagynoideae); B, Froesia venezuelensis (Quiinoideae); C, Luxemburgia schwackeana (Luxemburgieae); D, Rhytidanthera sulcata; E, Cespedesia spathulata; F, Poecilandra retusa; G, Godoya antioquiensis; H, Wallacea insignis; I, Sauvagesia semicylindrifolia; J, Sauvagesia erecta (Sauvagesieae); K, Infructescence of Lophira lanceolata with accrescent sepals (Lophirinae); L, Flower of Elvasia kollmannii (Elvasiinae); M, Perissocarpa umbellifera; N, Fruiting Rhabdophyllum arnoldianum; O, Brackenridgea nitida; P, Campylospermum glaberrimum; Q, Ochna serrulata; R, Fruit of Ochna integerrima with drupelets sitting on enlarged receptacle; S, Fruit of Ouratea sp. with enlarged red receptable; T, Ouratea sp. — Photos: A, K & N from www.africanplants.senckenberg.de (Dressler & al., 2014–); B by Julio Schneider; C by William Milliken/ Royal Botanic Gardens, Kew; D by Sandra Reinales; E by Reinaldo Aguilar; F, H & M by Francisco Farroñay; G by John Clark; I, J, S & T by Domingos Cardoso; L by Claudio Nicoletti de Fraga; O by John Elliott; P by Warran McCleland; Q by Marja Broersma; R by Pierre Grard.
Pan-cancer Proteomics Analysis to Identify Tumor-Enriched and Highly Expressed Cell Surface Antigens as Potential Targets for Cancer Therapeutics
<p>CPTAC PAN-cancer Data Repository</p> <p>Welcome to the CPTAC PAN-cancer Data Repository! This repository serves as a data repository for the CPTAC PAN-cancer effort, which focuses on cancer target discovery. It contains various data sets related to protein abundance estimation, derived TMT-TPA, iBAQ, iBAQ-derived copy number, and differential protein expression for CPTAC ten indications.</p> <p>## Contents</p> <p>The repository includes the following data:</p> <p>- FragPipe Output: Protein abundance estimation data generated using the FragPipe software.<br> - Derived TMT-TPA: Data derived from Tandem Mass Tag (TMT) based Total Protein Approach (TPA).<br> - iBAQ: Data representing intensity-based absolute quantification (iBAQ) of proteins.<br> - iBAQ-derived Copy Number: Data derived from iBAQ analysis for copy number estimation.<br> - Differential Protein Expression: Data indicating differential expression of proteins between tumor and NAT.</p> <p>## Data Organization</p> <p>The data in this repository is organized in a structured manner to facilitate easy access and navigation. The repository structure is as follows:</p> <p>FragPipe/<br> [fragpipe_data_files]<br> Derived_TMT_TPA/<br> [derived_tmt_tpa_data_files]<br> iBAQ/<br> [ibaq_data_files]<br> iBAQ-derived_copy_number/<br> [ibaq_copy_number_data_files]<br> Differential_protein_expression/<br> [differential_expression_data_files]</p>
Data from: Targeted enrichment of large gene families for phylogenetic inference: phylogeny and molecular evolution of photosynthesis genes in the Portullugo clade (Caryophyllales)
Hybrid enrichment is an increasingly popular approach for obtaining hundreds of loci for phylogenetic analysis across many taxa quickly and cheaply. The genes targeted for sequencing are typically single-copy loci, which facilitate a more straightforward sequence assembly and homology assignment process. However, this approach limits the inclusion of most genes of functional interest, which often belong to multi-gene families. Here we demonstrate the feasibility of including large gene families in hybrid enrichment protocols for phylogeny reconstruction and subsequent analyses of molecular evolution, using a new set of bait sequences designed for the "portullugo" (Caryophyllales), a moderately sized lineage of flowering plants (∼2200 species) that includes the cacti and harbors many evolutionary transitions to C4 and CAM photosynthesis. Including multi-gene families allowed us to simultaneously infer a robust phylogeny and construct a dense sampling of sequences for a major enzyme of C4 and CAM photosynthesis, which revealed the accumulation of adaptive amino acid substitutions associated with C4 and CAM origins in particular paralogs. Our final set of matrices for phylogenetic analyses included 75–218 loci across 74 taxa, with ∼50% matrix completeness across datasets. Phylogenetic resolution was greatly improved across the tree, at both shallow and deep levels. Concatenation and coalescent-based approaches both resolve the sister lineage of the cacti with strong support: Anacampserotaceae + Portulacaceae, two lineages of mostly diminutive succulent herbs of warm, arid regions. In spite of this congruence, BUCKy concordance analyses demonstrated strong and conflicting signals across gene trees. Our results add to the growing number of examples illustrating the complexity of phylogenetic signals in genomic-scale data.
A target enrichment probe set for resolving phylogenetic relationships in the coffee family, Rubiaceae
<p><em>Rubiaceae </em>is among the most species-rich, morphologically and geographically diverse plant families. Phylogenies have been inferred for many different groups across the family, however these have mostly relied on few genomic and plastid loci, as opposed to large-scale genomic data. Target enrichment provides the ability to generate sequence data for hundreds to thousands of phylogenetically informative, single-copy loci, which often leads to improved phylogenetic resolution at both shallow and deep taxonomic scales; however, a publicly accessible <em>Rubiaceae</em>-specific probe set that allows for comparable phylogenetic inference across clades is lacking. Here, we use publicly accessible genomic resources to identify putatively single copy nuclear loci for target enrichment in two <em>Rubiaceae </em>tribes: Hillieae (<em>Cinchonoideae</em>) and Palicoureeae+Psychotrieae (<em>Rubioideae</em>). We sequenced 2270 exons corresponding to 1059 supercontigs in our target clades, and generated in silico target enrichment sequences for other <em>Rubiaceae </em>taxa using our designed probe set. Our probe set, which we call <em>Rubiaceae </em>2270, was effective for targeting loci in species across and even outside of <em>Rubiaceae</em>. This probe set will facilitate phylogenomic studies in <em>Rubiaceae </em>and advance systematics and macroevolutionary studies in the family.</p>
RESICE - Reusability-targeted Enriched Sea Ice Core Database - Interactive data2source Traceability
<div> <div> <p>The <em>interactive_data2source_traceability.svg</em> file illustrates data and metadata availability as well as their relation to their original sources for 287 sea ice cores that are part of the <a title="RESICE" href="https://doi.org/10.5281/zenodo.10866346" target="_blank" rel="noopener">Reusability-targeted Enriched Sea Ice Core Database (RESICE)</a> as it is published in Zenodo.</p> <p>(a) shows the availability of selected data and metadata features (x-axis) for each sea ice core as indexed on the y-axis. The color map indicates the type of availability. Primary availability refers to data or metadata extracted directly from the data set. Secondary availability refers to data or metadata are extracted from articles or reports that provide supplementary information about the data sets. Tertiary availability means that additional resources unrelated to the data sets, such as instrument manuals, were used to obtain the metadata of interest.</p> <p>(b) allows to trace the data and metadata collected per RESICE sea ice core back to their original data sources. Gray dots indicate the sea ice core - again as indexed on the y-axis. Green, orange and pink dots represent primary sources (data sets), secondary sources (articles and expedition reports) and tertiary sources (instrument manuals). <br>Each of the dots is interactive, so you can click on it, and it will link you to the original doi or url of the source or to the YAML file in the RESICE living database provided on gitLab.</p> <p>The code used to generate this figure is available in the <a title="pyresice " href="https://doi.org/10.5281/zenodo.11198658" target="_blank" rel="noopener">pyresice</a> Python package available on <a title="https://git.rwth-aachen.de/mbd/pyresice" href="https://git.rwth-aachen.de/mbd/pyresice">GitLab</a>.</p> </div> </div>
RESICE - Reusability-targeted Enriched Sea Ice Core Database - Part B
<div> <div> <div> </div> RESICE is described in detail in the article <em>Reusability-targeted enrichment of sea ice core data</em> published on 2025-03-20 in Scientific Data (DOI: <a href="https://doi.org/10.1038/s41597-025-04665-x" target="_blank" rel="noopener">10.1038/s41597-025-04665-x</a>). This is Part B of RESICE. <em>RESICE_PartB.csv</em> contains all data including profile data (several rows per core), RESICE_PartB_cores.csv provides all data excuding profile data (one row per core) and <em>sources_PartB.csv</em> provides a list of all sources. The database including its <a title="RE-SICE Part A" href="https://www.doi.org/10.5281/zenodo.14745035" target="_blank" rel="noopener">Part A</a> is described in the <a title="RE-SICE General Information" href="https://www.doi.org/10.5281/zenodo.14744912" target="_blank" rel="noopener">general information.</a> Part A and Part B had to be separated due to different licenses of the orginal data sources.</div> </div>
Data for a preliminary molecular phylogeny of the family Hydroptilidae (Trichoptera): exploring the combination of targeted enrichment data and legacy Sanger sequence data
<p><span>The purpose of this study is to provide a proof-of-concept that the use of molecular data, particularly targeted enrichment data, and statistically supported methods of analysis can result in the construction of a stable phylogenetic framework for the microcaddisflies (Trichoptera: Hydroptilidae). Here, we use a combination of targeted enrichment data for ca. 300 nuclear protein-coding genes and legacy (Sanger-based) sequence data for the mitochondrial COI gene and partial sequence from the 28S rRNA gene.</span></p>
Pathways enriched in downstream target genes of miRNAs associated with NT-proBNP and OPN
<p>The full set of 30 Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways that were significantly enriched (FDR-adjusted p<0.01) in the predicted mRNA targets of either the OPN- or NT-proBNP-associated miRNAs; 21 of these pathways were significantly enriched in both sets of targets.</p> <p>KEGG: Kyoto Encyclopedia of Genes and Genomes; miRNAs: microRNAs; NT-proBNP: N-terminal pro <a href="http://www.mayomedicallaboratories.com/test-catalog/Clinical+and+Interpretive/83873">B-type natriuretic peptide</a>; OPN: osteopontin</p>
Fig. 3 in Phylogenomics of the tropical plant family Ochnaceae using targeted enrichment of nuclear genes and 250+ taxa
Fig. 3. Continues. For caption, see next part.
A new pipeline for removing paralogs in target enrichment data
<p><span><span>Target enrichment (such as Hyb-Seq) is a well-established high throughput sequencing method that has been increasingly used for phylogenomic studies. Unfortunately, current widely used pipelines for analysis of target enrichment data do not have a vigorous procedure to remove paralogs in target enrichment data. In this study, we develop a pipeline we call Putative Paralogs Detection (PPD) to better address putative paralogs from enrichment data. The new pipeline is an add-on to the existing HybPiper pipeline, and the entire pipeline applies criteria in both sequence similarity and heterozygous sites at each locus in the identification of paralogs. Users may adjust the thresholds of sequence identity and heterozygous sites to identify and remove paralogs according to the level of phylogenetic divergence of their group of interest. The new pipeline also removes highly polymorphic sites attributed to errors in sequence assembly and gappy regions in the alignment. We demonstrated the value of the new pipeline using empirical data generated from Hyb-Seq and the Angiosperm 353 kit for two woody genera <i>Castanea</i> (Fagaceae, Fagales) and <i>Hamamelis</i> (Hamamelidaceae, Saxifragales). Comparisons of datasets showed that the PPD identified many more putative paralogs than the popular method HybPiper. Comparisons of tree topologies and divergence times showed evident differences between data from HybPiper and data from our new PPD pipeline. We further evaluated the accuracy and error rates of PPD by BLAST mapping of putative paralogous and orthologous sequences to a reference genome sequence of<i> Castanea mollissima</i>. Compared to HybPiper alone, PPD identified substantially more paralogous gene sequences that mapped to multiple regions of the reference genome (31 genes for PPD compared with 4 genes for HybPiper alone). In conjunction with HybPiper, paralogous genes identified by both pipelines can be removed resulting in the construction of more robust orthologous gene datasets for phylogenomic and divergence time analyses. Our study demonstrates the value of Hyb-Seq with data derived from the Angiosperm 353 probe set for elucidating species relationships within a genus, and argues for the importance of additional steps to filter paralogous genes and poorly aligned regions (e.g., as occur through assembly errors), such as our new PPD pipeline described in this study.</span></span></p>
Target enrichment of long open reading frames and ultraconserved elements to link microevolution and macroevolution in non-model organisms
<p>Despite the increasing accessibility of high-throughput sequencing, obtaining high-quality genomic data on non-model organisms without proximate well-assembled and annotated genomes remains challenging. Here we describe a workflow that takes advantage of distant genomic resources and ingroup transcriptomes to select and jointly enrich long open reading frames (ORFs) and ultraconserved elements (UCEs) from genomic samples for integrative studies of microevolutionary and macroevolutionary dynamics. This workflow is applied to samples of the African unionid bivalve tribe Coelaturini (Parreysiinae) at basin and continent-wide scales. Our results indicate that ORFs are efficiently captured without prior identification of intron-exon boundaries. The enrichment of UCEs was less successful but nevertheless produced substantial datasets. Exploratory continent-wide phylogenetic analyses with ORF supercontigs (> 515,000 parsimony informative sites) resulted in a fully resolved phylogeny, the backbone of which was also retrieved with UCEs (> 11,000 informative sites). Variant calling on ORFs and UCEs of Coelaturini from the Malawi Basin produced ~2,000 SNPs per population pair. Estimates of nucleotide diversity and population differentiation were similar for ORFs and UCEs. They were low compared to previous estimates in mollusks, but comparable to those in recently diversifying Malawi cichlids and other taxa at an early stage of speciation. Skimming off-target sequence data from the same enriched libraries of Coelaturini from the Malawi Basin, we reconstructed the maternally-inherited mitogenome, which displays the gene order inferred for the most recent common ancestor of Unionidae. Overall, our workflow and results provide exciting perspectives for integrative genomic studies of microevolutionary and macroevolutionary dynamics in non-model organisms.</p>
Raw target enrichment of conserved element sequence data for 24 black coral species
<p><span>Deep-sea lineages are generally thought to arise from shallow-water ancestors, but this hypothesis is based on a relatively small number of taxonomic groups. Anthozoans, which include corals and sea anemones, are significant contributors to the faunal diversity of the deep sea, but the timing and mechanisms of their invasion into this biome remain elusive. Here, we reconstruct a fully resolved, time-calibrated phylogeny of 83 species in the order Antipatharia (black coral) to investigate their bathymetric evolutionary history. Our reconstruction indicates that extant black coral lineages first diversified in continental slope depths (~250–3,000 m) during the early Silurian (~437 Ma) and subsequently radiated into, and diversified within, both continental shelf (<250 m) and abyssal (>3,000 m) habitats. Ancestral state reconstruction analysis suggests that the appearance of morphological features that enhanced the ability of black corals to acquire nutrients coincided with their invasion of novel depths. Our findings have important conservation implications for anthozoan lineages, as the loss of "source" slope lineages could threaten millions of years of evolutionary history and confound future invasion events, thereby warranting protection. </span></p>
Data for a preliminary molecular phylogeny of the family Hydroptilidae (Trichoptera): exploring the combination of targeted enrichment data and legacy Sanger sequence data
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.