Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
168
datasets available to search
ShareScore release 0.7.1
Dataset results
168 results for “chromosomal assembly”
A chromosome-scale high-contiguity genome assembly of the threatened cheetah (Acinonyx jubatus)
<p><span>The cheetah (<em>Acinonyx</em> <em>jubatus</em>, SCHREBER 1775) is a large felid and is considered the fastest land animal. Historically, it inhabited open grassland across Africa, the Arabian Peninsula, and southwestern Asia; however, only small and fragmented populations remain today. Here, we present a de novo genome assembly of the cheetah based on PacBio continuous long reads and Hi-C proximity ligation data. The final assembly (VMU_Ajub_asm_v1.0) has a total length of 2.38 Gb, of which 99.7% are anchored into the expected 19 chromosome-scale scaffolds. The contig and scaffold N50 values of 96.8 Mb and 144.4 Mb, respectively, a BUSCO completeness of 95.4% and a k-mer completeness of 98.4%, emphasize the high quality of the assembly. Furthermore, annotation of the assembly identified 23,622 genes and a repeat content of 40.4%. This new highly contiguous and chromosome-scale assembly will greatly benefit conservation and evolutionary genomic analyses and will be a valuable resource, e.g., to gain a detailed understanding of the function and diversity of immune response genes in felids.</span></p>
Chromosome-scale genome assembly of the African spiny mouse (Acomys cahirinus)
<p>Genomic DNA was extracted from blood from a single male A. cahirinus animal using a Monarch HMW DNA Extraction Kit for Cells & Blood (T3050, New England Biolabs, Ipswich MA) following the manufacturer’s recommended protocol. DNA was quantified prior to library construction using the Qubit DNA HS Assay (ThermoFischer, Waltham MA) and DNA fragment lengths were assessed using the Agilent Femto Pulse System (Santa Clara, CA). Libraries were prepared for sequencing using the Oxford Nanopore ligation kit (SQK-LSK110) following the manufacturers’ instructions, except that DNA repair and A-tailing was performed for 30 min and the ligation was allowed to continue for 1 hr. Prepared libraries were quantified using a Qubit fluorometer and 30 fmol of the library was loaded onto a Nanopore version R.9.4.1 flow cell and loaded on a PromethION running MinKNOW version (21.05.20). To increase output, the flow cell was washed after approximately 24 hr of sequencing then an additional 12 fmol of library was added to the flow cell and run for an additional 48 hr. Basecalling was performed using Guppy 5.0.12 (Oxford Nanopore) using the superior model (dna_r9.4.1_450bps_sup_prom.cfg). FASTQ files for assembly were extracted from unaligned bam files using samtools (Li et al. 2009) then Flye version 2.9 for assembly using the --nano-hq flag (Kolmogorov et al. 2019). Haplotigs and overlaps in the assembly were purged using purge_dups (https://github.com/dfguan/purge_dups). The assembly was then polished using Medaka version 1.4.2 (https://github.com/nanoporetech/medaka) followed by a second polishing step with pilon version 1.24 (Walker et al. 2014). Assembly statistics at each step were generated using Quast (Gurevich et al. 2013) and BUSCO (Simão et al. 2015) (Table S2). The primary contigs assembled from the Nanopore data were anchored to chromosomes using 505,210,505 read pairs of a Hi-C library isolated from another A. cahirinus individual of unknown sex downloaded from the NCBI Short Read Archive (SRX13258644) (Wang et al. 2022). After aligning the Hi-C reads with the ArimaHi-C Mapping Pipeline (https://github.com/ArimaGenomics/mapping_pipeline), YaHS v1.0 (Zhou et al. 2023) was used with default error correction for scaffolding, and Juicebox v1.11.08 (Dudchenko et al. 2018) was used to generate a Hi-C contact map. Progressive Cactus was used (Armstrong et al. 2020) to perform a whole-genome alignment of the A. cahirinus draft assembly to the Mus musculus GRCm39 reference genome (RefSeq GCF_000001635.27_GRCm39). Comparative annotation of the draft genomes was then performed using the Comparative Annotation Toolkit (CAT) (Fiddes et al. 2018). Briefly, the M. musculus RefSeq annotation GFF was parsed and validated with the “parse_ncbi_gff3” and “validate_gff3” programs (respectively) from CAT. The M. musculus reference transcript cDNA sequences were downloaded and mapped to the M. musculus draft genome with minimap2 (Li 2018) and provided to CAT as long-read RNA-seq reads in the “[ISO_SEQ_BAM]” field of the configuration file. For A. cahirinus, bulk RNA-seq data obtained from multiple pooled organs were downloaded from NCBI SRA BioProject PRJNA342864 (Bellofiore et al. 2017) and mapped to the draft assembly with STAR (Dobin et al. 2013) then provided to CAT in the “[BAMS]” field. CpG islands were identified using the cpg_lh utility from the UCSC suite of tools (Kent et al. 2002).</p>
Chromosome-level assemblies of the Pieris mannii butterfly genome suggest Z-origin and rapid evolution of the W chromosome
<p><span>The insect order Lepidoptera (butterflies and moths) represents the largest group of organisms with ZW/ZZ sex determination. While the origin of the Z chromosome predates the evolution of the Lepidoptera, the W chromosomes are considered younger, but their origin is debated. To shed light on the origin of the lepidopteran W, we here produce chromosome-level genome assemblies for the butterfly <em>Pieris</em> <em>mannii</em>, and compare the sex chromosomes within and between <em>P. mannii </em>and its sister species <em>P. rapae</em>. Our analyses clearly indicate a common origin of the W chromosomes of the two <em>Pieris</em> species, and reveal similarity between the Z and W in chromosome sequence and structure. This supports the view that the W in these species originates from Z-autosome fusion rather than from a redundant B chromosome. We further demonstrate the extremely rapid evolution of the W relative to the other chromosomes and argue that this may preclude reliable conclusions about the origins of W chromosomes based on comparisons among distantly related Lepidoptera. Finally, we find that sequence similarity between the Z and W chromosomes is greatest toward the chromosome ends, perhaps reflecting selection for the maintenance of recognition sites essential to chromosome segregation. Our study highlights the utility of long-read sequencing technology for illuminating chromosome evolution.</span></p>
Chromosome-scale assembly of the wild wheat relative Aegilops umbellulata
<p><span>Wild wheat relatives have been explored in plant breeding to increase the genetic diversity of bread wheat, one of the most important food crops. <em>Aegilops umbellulata</em> is a diploid U genome-containing grass species that serves as a genetic reservoir for wheat improvement. In this study, we report the construction of a chromosome-scale reference assembly of <em>Ae. umbellulata</em> accession TA1851 based on corrected PacBio HiFi reads and chromosome conformation capture. The total assembly size was 4.25 Gb with a contig N50 of 17.7 Mb. In total, 36,268 gene models were predicted. We benchmarked the performance of hifiasm and LJA, two of the most widely used assemblers using standard and corrected HiFi reads, revealing a positive effect of corrected input reads. Comparative genome analysis confirmed substantial chromosome rearrangements in <em>Ae. umbellulata</em> compared to bread wheat. In summary, the <em>Ae. umbellulata</em> assembly provides a resource for comparative genomics in Triticeae and for the discovery of agriculturally important genes.</span></p>
Chromosome-scale assembly of the wild wheat relative Aegilops umbellulata
Open the record for dataset details and reuse information.
Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C
Open the record for dataset details and reuse information.
Chromosome-level assemblies of the Pieris mannii butterfly genome suggest Z-origin and rapid evolution of the W chromosome
Open the record for dataset details and reuse information.
A chromosome-scale genome assembly of the okapi (Okapia johnstoni)
Open the record for dataset details and reuse information.
Data from: A prelude to conservation genomics: First chromosome-level genome assembly of a flying squirrel (Pteromyini: Pteromys volans)
Open the record for dataset details and reuse information.
A chromosome-level genome assembly of the snow leopard, Panthera uncia
Open the record for dataset details and reuse information.
Data from: A chromosomal-scale genome assembly of Tectona grandis reveals the importance of tandem gene duplication and enables discovery of genes in natural product biosynthetic pathways
Open the record for dataset details and reuse information.
Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis
Open the record for dataset details and reuse information.
Chromosome-level genome assembly of Poropuntius huangchuchieni
Open the record for dataset details and reuse information.
Chromosome-level genome assembly of the bay scallop Argopecten irradians
Open the record for dataset details and reuse information.
Chromosome-level genome assembly of Dynastes reidi reveals structural evolution of autosomes and the sex chromosomes in Hercules Beetles
Open the record for dataset details and reuse information.
A chromosome-scale assembly of the quinoa genome provides insights into the structure and dynamics of its subgenomes
Open the record for dataset details and reuse information.
Chromosome-level assembly and annotation of Lates japonicus
Open the record for dataset details and reuse information.
Data from: Chromosome-scale genome assembly of bread wheat’s wild relative Triticum timopheevii
Open the record for dataset details and reuse information.
A chromosome-level genome assembly of the highly heterozygous sea urchin Echinometra sp. EZ reveals adaptation in the regulatory regions of stress response genes
Open the record for dataset details and reuse information.
Chromosome-level assembly of two pearl millet (Cenchrus americanus) genomes, functional annotation and transcriptomes
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.