Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

660

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

660 results for “genome assembly”

Learn how ShareScore rates datasets ↗
dryad36/100

Data from: high repeat content in the genomes of sparrows: the importance of genome assembly completeness for transposable element discovery

<p>Transposable elements (TE) play critical roles in shaping genome evolution. However, the highly repetitive sequence content of TEs is a major source of assembly gaps. This makes it difficult to decipher the impact of these elements on the dynamics of genome evolution. The increased capacity of long-read sequencing technologies to span highly repetitive regions of the genome should provide novel insights into patterns of TE diversity. Here we report the generation of highly contiguous reference genomes using PacBio long read and Omni-C technologies for three species of sparrows in the family Passerellidae. To assess the influence of sequencing technology on TE annotation, we compared these assemblies to three chromosome-level sparrow assemblies recently generated by the Vertebrate Genomes Project and nine other sparrow species generated using a variety of short- and long-read technologies. All long-read based assemblies were longer in length (range: 1.12-1.41 Gb) than short-read assemblies (0.91-1.08 Gb). Assembly length was strongly correlated with the amount of repeat content, with longer genomes showing much higher levels of repeat content than typically reported for the avian order Passeriformes. Repeat content for the Bell's sparrow (31.2% of genome) was the highest level reported to date for a songbird genome assembly and was more in line with woodpecker (order Piciformes) genomes. CR1 LINE elements retained from an expansion that occurred 25-30 million years ago were the most abundant TEs in the song sparrow genome. Although the other five sparrow species also exhibit evidence for a spike in CR1 LINE activity at 25-30 million years ago, LTR elements stemming from more recent expansions were the most abundant elements in these species. LTRs were uniquely abundant in the Bell's sparrow genome deriving from two recent peaks of activity. Higher levels of repeat content (79.2-93.7%) were found on the W chromosome relative to the Z (20.7-26.5) or autosomes (16.1-30.9%). These patterns support a dynamic model of transposable element expansion and contraction underpinning the seemingly constrained and small sized genomes of birds. Our work highlights how the resolution of difficult-to-assemble regions of the genome with new sequencing technologies promises to transform our understanding of avian genome evolution.</p>

opencc-zeroDec 2023View details →
zenodo36/100

Non-redundant metagenome-assembled genomes of activated sludge reactors at different disturbances and scales

<p>Metagenome-assembled genomes (MAGs) are microbial genomes reconstructed from metagenomic data and can be assigned to known taxa or lead to uncovering novel ones. MAGs can provide insights into how microbes interact with the environment. Here, we performed genome-resolved metagenomics on sequencing data from four studies using sequencing batch reactors at microcosm (~25 mL) and mesocosm (~4 L) scales inoculated with sludge from full-scale wastewater treatment plants. These studies investigated how microbial communities in such plants respond to two environmental disturbances: the presence of toxic 3-chloroaniline and changes in organic loading rate. We report 839 non-redundant MAGs with at least 50% completeness and 10% contamination (MIMAG medium-quality criteria). From these, 399 are of putative high-quality, while sixty-seven meet the MIMAG high-quality criteria. MAGs in this catalogue represent the microbial communities in sixty-eight laboratory-scale reactors used for the disturbance experiments, and in the full-scale wastewater treatment plant which provided the source sludge. This dataset can aid meta-studies aimed at understanding the responses of microbial communities to disturbances, particularly as ecosystems confront rapid environmental changes.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

The HIFI reads and assembled contigs used to construct the mitochondrial genome assembly of Glyphodes pyloalis

<p><span>Glyphodes pyloalis (Lepidoptera; Crambidae; Spilomelinae), known as the mulberry pyralid, is a notorious insect pest belonging to the Lepidopteran that threatens mulberry cultivation across China. <span>The mitochondrial genome is an invaluable genetic resource owing to its maternal inheritance, rapid evolution, and lack of recombination. A previous study documented the G. pyloalis mitochondrial genome (mtDNA) as NC_025933 using short-read next-generation sequences that, while pioneering, leaves room for improvement with modern long-read sequencing. PacBio&rsquo;s high-fidelity circular consensus sequencing (HIFI CCS) generates extra-long sequences sufficient to produce complete mtDNA assemblies with base pair accuracy.</span></span></p>

opencc-by-4.0Dec 2023View details →
dryad36/100

A pan-cetacean MHC amplicon sequencing panel developed and evaluated in combination with genome assemblies

<p>The major histocompatibility complex (MHC) is a highly polymorphic gene family that is crucial in immunity, and its diversity can be effectively used as a fitness marker for populations. Despite this, MHC remains poorly characterised in non-model species (e.g., cetaceans: whales, dolphins and porpoises) as high gene copy number variation, especially in the fast-evolving class I region, makes analyses of genomic sequences difficult. To date, only small sections of class I and IIa genes have been used to assess functional diversity in cetacean populations. Here, we undertook a systematic characterisation of the MHC class I and IIa regions in available cetacean genomes. We extracted full-length gene sequences to design pan-cetacean primers that amplified the complete exon2 from MHC class I and IIa genes in one combined sequencing panel. We validated this panel in 19 cetacean species and described 354 alleles for both classes.  Furthermore, we identified likely assembly artefacts for many MHC class I assemblies based on the presence of class I genes in the amplicon data compared to missing genes from genomes. Finally, we investigated MHC diversity using the panel in 25 humpback and 30 southern right whales, including four paternity trios for humpback whales. This revealed copy-number variable class I haplotypes in humpback whales, which is likely a common phenomenon across cetaceans. These MHC alleles will form the basis for a cetacean branch of the Immuno-Polymorphism Database (IPD-MHC), a curated resource intended to aid in the systematic compilation of MHC alleles across several species, to support conservation initiatives.</p>

opencc-zeroJan 2024View details →
zenodo36/100

De novo genome assembly of rice varieties using Nanopore long reads

<p>Genome sequences for Sugimura et al. (2024) of the rice (O. sativa) varieties 'Hitomebore' and 'Arroz da Terra.'</p> <p>Yusaku Sugimura, Kaori Oikawa, Yu Sugihara, Hiroe Utsushi, Eiko Kanzaki, Kazue Ito, Yumiko Ogasawara, Tomoaki Fujioka, Hiroki Takagi, Motoki Shimizu, Hiroyuki Shimono, Ryohei Terauchi, Akira Abe. Impact of rice GENERAL REGULATORY FACTOR14h (GF14h) on low-temperature seed germination and its application to breeding. PLoS Genet 20(8): e1011369. https://doi.org/10.1371/journal.pgen.1011369</p> <p>bioRxiv doi: https://doi.org/10.1101/2024.02.16.580620</p>

opencc-by-4.0Jan 2024View details →
dryad36/100

Data from: Chromosome-scale genome assembly of bread wheat's wild relative Triticum timopheevii

<p>Wheat (<em>Triticum aestivum</em>) is one of the most important food crops with an urgent need for increase in its production to feed the growing world. Wheat's wild relative species provide a hugely untapped reservoir of genetic diversity for wheat improvement. <em>Triticum timopheevii</em> (2n = 4x = 28) is a tetraploid wheat wild relative species containing the A<sup>t</sup> and G genomes that has been exploited in many wheat pre-breeding programmes over the last few decades. In this study, we report the generation of a chromosome-scale reference genome assembly of <em>T. timopheevii</em> accession PI 94760 based on PacBio HiFi reads and chromosome conformation capture (Hi-C). The total assembly size was 9.35 Gb with a contig N50 of 42.4 Mb. In total, 166,325 gene models were predicted. Comparative genome analysis confirmed previously known chromosomal translocations and indicated new chromosome rearrangements. Analysis of the genomic distribution of DNA methylation showed that the G genome had on average more methylated bases than the A<sup>t</sup> genome. The G genome was also more closely related to the<em> </em>S genome of <em>Aegilops speltoides</em> than to the B genome of hexaploid or tetraploid wheat. In summary, the <em>T. timopheevii</em> genome assembly provides a valuable resource for genome-informed discovery and cloning of agronomically important genes for future food security.</p>

opencc-zeroJan 2024View details →
dryad36/100

Ion Torrent data for the genome assembly and phylogenomic placement of mitochondrial genomes with a focus on houndsharks (Chondrichthyes: Triakidae)

<p>Here, we present the Ion Torrent® next-generation sequencing (NGS) data for five houndsharks (Chondrichthyes: Triakidae), which include <em>Galeorhinus galeus</em> (17,487 bp; GenBank accession number ON652874), <em>Mustelus asterias</em> (16,708; ON652873), <em>Mustelus mosis</em> (16,755; ON075077), <em>Mustelus palumbes</em> (16,708; ON075076), and <em>Triakis megalopterus</em> (16,746 bp; ON075075). All assembled mitogenomes encode 13 protein-coding genes (PCGs), two ribosomal (r)RNA genes, and 22 transfer (t)RNA genes (<em>tRNA<sup>Leu</sup></em><sup> </sup>and <em>tRNA<sup>Ser</sup> </em>are duplicated), except for <em>G</em>. <em>galeus</em> which contains 23 tRNA genes where <em>tRNA<sup>Thr</sup> </em>is duplicated. The data presented in this paper can assist other researchers in further elucidating the diversification of triakid species and the phylogenetic relationships within Carcharhiniformes (groundsharks) as mitogenomes accumulate in public repositories.</p>

opencc-zeroJan 2024View details →
dryad36/100

Annotated genome assemblies for Geoscapheus dilatatus, Panesthia cribrata and Neogeoscapheus hanni

<p>Genetic changes that enabled the evolution of eusociality have long captivated biologists. More recently, attention has focussed on the consequences of eusociality on genome evolution. Studies have reported higher molecular evolutionary rates in eusocial hymenopteran insects compared with their solitary relatives. To investigate the genomic consequences of eusociality in termites, we sequenced genomes from three of their non-eusocial cockroach relatives. Using a phylogenomic approach, we found that termite genomes experienced lower rates of synonymous mutations than those of cockroaches, possibly as a result of longer generation times. We identified higher rates of nonsynonymous mutations in termite genomes than in cockroach genomes, and identified pervasive relaxed selection in the former (24–31% of the genes analysed) compared with the latter (2–4%). We infer that this is due to a reduction in effective population size, rather than gene-specific effects (e.g., indirect selection of caste-biased genes). We found no obvious signature of increased genetic load in termites, and postulate efficient purging of deleterious alleles at the colony level. Additionally, we identified genomic adaptations that may underpin caste formation, such as genes involved in post-translational modifications. Our results provide insights into the evolution of termites and the genomic consequences of eusociality more broadly.</p>

opencc-zeroFeb 2024View details →
dryad36/100

Long read genome assembly of Automeris io (Lepidoptera: Saturniidae) an emerging model for the evolution of deimatic displays

<p>Automeris moths are a morphologically diverse group with 145 described species that have a geographic range that spans from the New World temperate zone to the Neotropics. Many Automeris have hindwing eyespots that are thought to deter or disrupt the attack of potential predators, allowing the moth time to escape. Some species in the genus have vestigial eyespots or lack them completely, suggesting that this trait may provide a selective benefit. The Io moth (Automeris io), known for its striking eyespots, is the most widely studied species within the genus and is an emerging model system to study the evolution of deimatism, a predatory defense that combines visual stimuli and movement. Here we present a high-quality, PacBio HiFi genome assembly for Io moth to aid existing research on the molecular development of eyespots. Genomic research is needed to address questions involving antipredatory defenses and eyespot pattern development. BUSCO analysis for this genome shows a completeness of 98.4%, and N50 of 15.</p>

opencc-zeroFeb 2024View details →
dryad36/100

A nuclear genome assembly of an extinct flightless bird, the little bush moa

<p>We present a draft genome of the little bush moa (<em>Anomalopteryx didiformis</em>) - one of approximately nine species of extinct flightless birds from Aotearoa, New Zealand - using ancient DNA recovered from a fossil bone from the South Island. We recover a complete mitochondrial genome at 249.9X depth of coverage and almost 900 Mb of a male moa nuclear genome at ~4-5X coverage, with sequence contiguity sufficient to identify more than 85% of avian universal single-copy orthologs. We describe a diverse landscape of transposable elements and satellite repeats; estimate a long-term effective population size of ~240,000; identify a diverse suite of olfactory receptor genes and an opsin repertoire with sensitivity in the UV range; show that the wingless moa phenotype is likely not attributable to gene loss or pseudogenization; and identify potential function-altering coding sequence variants in moa that could be synthesized for future functional assays. This genomic resource should support further studies of avian evolution and morphological divergence.</p>

opencc-zeroMar 2024View details →
zenodo36/100

Genome assembly and annotation for temperate coral Astrangia poculata

<p>Initial release of the genome assembly, gene prediction, and functional annotation for the temperate coral <em>Astrangia poculata.&nbsp;</em></p> <p>The genomic resources made available here are described in <a href="https://www.biorxiv.org/content/10.1101/2023.09.22.558704v1.full">Stankiewicz et al., 2023</a>. The files are as follows:</p> <p>&nbsp;</p> <p>Genome assembly:</p> <ul> <li>apoculata.genome.fasta.gz --unmasked version of the genome assembly</li> <li>apoculata.genome.masked.fasta.gz --masked version of the genome assembly</li> <li> <p><span>apoculata_circularized_mitogenome.fasta --circularized mitochondrial genome assembly</span></p> </li> </ul> <p>Annotation:</p> <ul> <li>apoculata_GeneAnnotation_combined.txt.gz: function annotation</li> <li>apoculata.gff3.gz: gene predictions in GFF3&nbsp;&nbsp;</li> <li>apoculata.gtf.gz: gene predictions in GTF&nbsp;&nbsp;</li> <li>apoculata_cds.fasta.gz: coding sequences&nbsp;</li> <li>apoculata_longest_trans_cds.fasta.gz: coding sequences filtered to just the longest translatable per gene&nbsp;&nbsp;</li> <li>apoculata_mrna.fasta.gz: mRNA sequences&nbsp;&nbsp;</li> <li>apoculata_proteins.fasta.gz: protein sequences</li> </ul>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Chromosome-scale genome assembly and de novo annotation of Alopecurus aequalis.

<p><em>Alopecurus aequalis</em> is a winter annual or short-lived perennial bunchgrass which has in recent years emerged as the dominant agricultural weed of barley and wheat in certain regions of China and Japan, causing significant yield losses. Its robust tillering capacity and high fecundity, combined with the development of both target and non-target-site resistance to herbicides means it is a formidable challenge to food security. Here we report on a chromosome-scale assembly of <em>A. aequalis</em> with a genome size of 2.83 Gb. The genome contained 33,758 high-confidence protein-coding genes with functional annotation. Comparative genomics revealed that the genome structure of <em>A. aequalis</em> is more similar to <em>Hordeum vulgare </em>rather than the more closely related <em>Alopecurus myosuroides</em>. The datasets provided here are the assembly FASTA file (lpAloAequ1.1.prim.cur.20230912.fasta.gz), the high-confidence protein-coding genes (Alaeq_EIv0.2.release_HC_genes.gff3.gz) and the full annotation which includes both low and high confidence features of all biotypes (Alaeq_EIv0.2.release.gff3.gz)&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Metagenome-Assembled Genomes of 2_2_Ac_Mat

<p>The dataset is featured in the data report titled "MAGnificent Microbes: Metagenome-Assembled Genomes of Marine Microorganisms in Mats from a Submarine Groundwater Discharge Site in Mabini, Batangas, Philippines." The study utilized shotgun metagenomics to examine the diversity and functional profiles of marine microorganisms in microbial mats from an SGD-influenced site in Mabini. The dataset includes extracted metagenome-assembled genomes (MAGs) along with their annotations using RAST.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Supplemental Material for Astrangia poculata genome assembly

<p><span>Supplementary Table S<strong>5</strong>: GO annotations for the 143 terms enriched in the gene families unique to <em>A. poculata</em> relative to all other cnidarians included in this analysis (see Supplementary Table S2).&nbsp;</span><span>&nbsp;</span></p> <p><span>Supplementary Table S<strong>6</strong>: The gene annotations for all genes included in the gene families that were significantly different in size between <em>A. poculata</em> and <em>A. millepora</em>.</span></p> <p><span>Supplememtary Table S<strong>7</strong>: GO annotations for significantly enriched terms in the 73 gene families that were significantly larger in <em>A. millepora </em>relative to <em>A. poculata</em>.</span></p> <p><span>Supplememtary Table S<strong>8</strong>: GO annotations for significantly enriched terms in the 97 gene families that were significantly larger in <em>A. poculata </em>relative to <em>A. millepora</em>.</span></p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Zostera marina leaf associated bacterial metagenome assembled genomes

<p>Metagenome assembled genomes (MAGs) associated with:</p> <p>A genomic resource for exploring bacterial-viral dynamics in seagrass ecosystems</p> <p>Analysis, code, intermediate and supporting files are archived here: <a href="https://doi.org/10.5281/zenodo.14226514">10.5281/zenodo.14226514</a></p> <p>Viral sequences from this work are archived here: <a href="https://doi.org/10.5281/zenodo.14226038">10.5281/zenodo.14226038</a><br><br>This archive contains:<br>(i) Fifty-six fasta files representing the MAGs described in the above titled work with &gt; 80% completion and &lt; 10% contamination based on CheckM2 metrics<br>(ii) Metadata file describing the MAGs (i.e., subset of Table S3 from the above work)</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Draft genomes of 972 carbapenem-resistant Pseudomonas aeruginosa isolates (shovill assemblies of Reyes et al 2023 dataset raw reads)

<p>This dataset contains shovill assemblies of raw reads released under&nbsp;NCBI BioProject PRJNA824880 generated in Reyes J et al (The Lancet Microbe. Volume 4 Issue 3 Pages e159-e170 (March 2023); DOI: 10.1016/S2666-5247(22)00329-9)</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Data From: Oatk - a de novo assembly tool for complex plant organelle genomes

<p>This reposity hosts the data for 195 plant organelle genome assemblies generated in the manuscript "Oatk: a de novo assembly tool for complex plant organelle genomes". The sequence data were produced by the Tree of Life programme at the Sanger Institute, mostly from the Darwin Tree of Life (DToL) project, including 24 monocots, 154 eudicots, 16 mosses and one liverwort. See SAMPLE_LIST file for descriptions of these species.</p> <p>In each species subfolder, below files are included.</p> <ol> <li><code>PLTD.fasta</code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Plastome assembly file in FASTA format</li> <li><code>PLTD.annot.bed</code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Plastome assembly annotation file in BED format</li> <li><code>MITO.fasta</code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Mitogenome assembly file in FASTA format</li> <li><code>MITO.annot.bed</code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Mitogenome assembly annotation file in BED format</li> <li><code>MBG.gfa</code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Genome assembly file in GFA format generated with MBG</li> <li><code>PMAT.gfa</code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Genome assembly file in GFA format generated with OATK</li> <li><code>OATK.gfa</code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Genome assembly file in GFA format generated with PMAT (may not exist)</li> </ol> <p>&nbsp;</p> <p>Updates in the New Version:</p> <p>In the previous version, our raw PacBio HiFi read pre-processing pipeline had screened out some reads that it erroneously thought contained HiFi adapter sequence, which led to the gaps in the Hibiscus plastomes. We now fixed this and have rerun all the assemblies that led to any linear organelle components (37 species). All plastomes remain unchanged except for the three Hibiscuses, which are now also circular. Thirteen mitogenomes changed, with six of them now becoming circular.</p>

opencc-by-4.0Oct 2024View details →
dryad36/100

A roadmap to durable BCTV resistance using long-read genome assembly of genetic stock KDH13

<p>PacBio Sequence data associated with genetic stock KDH13.Datasets include genome resources and annotated files associated with the manuscript "Long-read genome assembly of Double Haploid Sugar Beet KDH13 provides roadmap for durable genetic resistance to Beet Curly Top Virus". This includes genome assembly, ordered genome assembly, protein predictions, variant call format files for an F1 hybrid (KDH13xKDH19-17).</p>

opencc-zeroNov 2021View details →
zenodo36/100

MACIE scores for human genome assembly GRCh37 Part 1 (Chr1 - Chr3)

<p>MACIE (Multi-dimensional Annotation Class Integrative Estimation) is an unsupervised multivariate mixed model framework to assess multi-dimensional functional impacts for both coding and non-coding variants in the human genome. MACIE integrates a variety of functional annotations, including protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics, and estimates the joint posterior probabilities of each genetic variant being functional.</p> <p>For each non-synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of &ldquo;not damaging protein functional and evolutionarily conserved&rdquo; (MACIE01); &ldquo;damaging protein functional and not evolutionarily conserved&rdquo; (MACIE10); &ldquo;not damaging protein functional and not evolutionarily conserved&rdquo; (MACIE00); &ldquo;both damaging protein functional and evolutionarily conserved&rdquo; (MACIE11). MACIE_protein is the estimated posterior probability of &ldquo;damaging protein functional&rdquo;, which is the sum of MACIE10 and MACIE11; MACIE_conserved is the estimated posterior probability of &ldquo;evolutionarily conserved&rdquo;, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of &ldquo;damaging protein functional&rdquo; or &ldquo;evolutionarily conserved&rdquo;, which is the sum of MACIE01, MACIE10, and MACIE11.</p> <p>For each non-coding and synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of &ldquo;not evolutionarily conserved and regulatory functional&rdquo; (MACIE01); &ldquo;evolutionarily conserved and not regulatory functional&rdquo; (MACIE10); &ldquo;not evolutionarily conserved and not regulatory functional&rdquo; (MACIE00); &ldquo;both evolutionarily conserved and regulatory functional (MACIE11). MACIE_conserved is the estimated posterior probability of &ldquo;evolutionarily conserved&rdquo;, which is the sum of MACIE10 and MACIE11; MACIE_regulatory is the estimated posterior probability of &ldquo;regulatory functional&rdquo;, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of &ldquo;evolutionarily conserved&rdquo; or &ldquo;regulatory functional&rdquo;, which is the sum of MACIE01, MACIE10, and MACIE11.</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

MACIE scores for human genome assembly GRCh37 Part 4 (Chr14 - Chr22)

<p>MACIE (Multi-dimensional Annotation Class Integrative Estimation) is an unsupervised multivariate mixed model framework to assess multi-dimensional functional impacts for both coding and non-coding variants in the human genome. MACIE integrates a variety of functional annotations, including protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics, and estimates the joint posterior probabilities of each genetic variant being functional.</p> <p>For each non-coding and synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of &ldquo;not evolutionarily conserved and regulatory functional&rdquo; (MACIE01); &ldquo;evolutionarily conserved and not regulatory functional&rdquo; (MACIE10); &ldquo;not evolutionarily conserved and not regulatory functional&rdquo; (MACIE00); &ldquo;both evolutionarily conserved and regulatory functional (MACIE11). MACIE_conserved is the estimated posterior probability of &ldquo;evolutionarily conserved&rdquo;, which is the sum of MACIE10 and MACIE11; MACIE_regulatory is the estimated posterior probability of &ldquo;regulatory functional&rdquo;, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of &ldquo;evolutionarily conserved&rdquo; or &ldquo;regulatory functional&rdquo;, which is the sum of MACIE01, MACIE10, and MACIE11.</p>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record