Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
168
datasets available to search
ShareScore release 0.7.1
Dataset results
168 results for “chromosomal assembly”
Chromosome-Level genome assembly and transcriptome analysis of the ural owl, Strix uralensis Pallas, 1771
Open the record for dataset details and reuse information.
A chromosome-scale high-contiguity genome assembly of the threatened cheetah (Acinonyx jubatus)
Open the record for dataset details and reuse information.
Data from: Chromosome-level genome assembly of a cyprinid fish Onychostoma macrolepis by integration of Nanopore Sequencing, Bionano and Hi-C technology
<p><i>Onychostoma macrolepis</i> is an emerging commercial cyprinid fish species. It is a model system for studies of sexual dimorphism and genome evolution. Here, we report the chromosome-level assembly of the<i> O.macrolepis</i> genome obtained from the integration of Nanopore long-read sequencing with physical maps produced using Bionano and Hi-C technology. A total of 87.9 Gb of Nanopore sequence provided approximately 100-fold coverage of the genome. The preliminary genome assembly was 883.2 Mb in size with a contig N50 size of 11.2 Mb. The 969 corrected contigs obtained from Bionano optical mapping were assembled into 853 scaffolds and produced an assembly of 886.5 Mb with a scaffold N50 of 16.5 Mb. Finally, using the Hi-C data, 881.3 Mb (99.4% of genome) in 526 scaffolds were anchored and oriented in 25 chromosomes ranging in size from 25.27 to 56.49 Mb. In total, 24,770 protein-coding genes were predicted in the genome, and ~96.85% of the genes were functionally annotated. The annotated assembly contains 93.3% complete genes from the BUSCO reference set. In addition, we identified 409 Mb (46.23% of the genome) of repetitive sequence, and 11,213 non-coding RNAs, in the genome. Evolutionary analysis revealed that <i>O.macrolepis</i> diverged from common carp approximately 24.25 million years ago. The chromosomes of <i>O.macrolepis</i> showed an unambiguous correspondence to the chromosomes of zebrafish. The high-quality genome assembled in this work provides a valuable genomic resource for further biological and evolutionary studies of <i>O. macrolepis</i>.</p>
Generation of a chromosome-scale genome assembly of the insect-repellant terpenoid-producing Lamiaceae species, Callicarpa americana
<p>Background: Plants exhibit wide chemical diversity due to production of specialized metabolites which function as pollinator attractants, defensive compounds, and signaling molecules. Lamiaceae (mints) are known for their chemodiversity and have been cultivated for use as culinary herbs and as sources of insect repellents, health-promoting compounds, and fragrance. Findings: We report the chromosome-scale genome assembly of <em>Callicarpa americana</em> L. (American beautyberry), a species within the early diverging Callicarpoideae clade of the Lamiaceae, known for its metallic purple fruits and use as an insect repellent due to its production of terpenoids. Using long reads and Hi-C scaffolding, we generated a 506.1 Mb assembly spanning 17 pseudomolecules with an N50 contig and N50 scaffold size of 7.5 Mb and 29.0 Mb, respectively. A total of 32,164 genes was annotated including 53 candidate terpene synthases and 47 putative clusters of specialized metabolite biosynthetic pathways. Whole genome duplication analyses revealed three putative events, which together with local tandem duplication events, contributed to gene family expa, American beautyberransion of terpene synthases. Kolavenyl diphosphate is a gateway to many of <em>C. americana</em>'s bioactive terpenoids; experimental validation confirmed that CamTPS2 encodes kolavenyl diphosphate synthase. Syntenic analyses with <em>Tectona grandis</em> L. f. (teak), a member of the Tectonoideae clade of Lamiaceae known for exceptionally strong wood resistant to insects, revealed 963 collinear blocks and 21,297 <em>C. americana</em> syntelogs. Conclusions: Access to the <em>C. americana</em> genome provides a roadmap for rapid discovery of genes encoding plant-derived agrichemicals and a key resource to understand the evolution of chemical diversity in Lamiaceae. </p> <p> </p>
Construction of a chromosome-scale long-read reference genome assembly for potato
<p><span><strong>Background:</strong> Worldwide, the cultivated potato, <i>Solanum tuberosum </i>L<i>.</i>, is the number one vegetable crop and a critical food security crop. The genome sequence of DM1-3 516 R44, a doubled monoploid clone of S. <i>tuberosum </i>Group Phureja, was published in 2011 using a whole-genome shotgun sequencing approach with short read sequence data. Current advanced sequencing technologies now permit generation of near-complete, high-quality chromosome-scale genome assemblies at a minimal cost. </span></p> <p><span><strong>Findings: </strong>Here, we present an updated version of the DM1-3 516 R44 genome sequence (v6.1) using Oxford Nanopore Technologies long reads coupled with proximity-by-ligation scaffolding (Hi-C) yielding a chromosome-scale assembly. The new (v6.1) assembly represents 741.6 Mb of sequence (87.8 %) of the estimated 844 Mb genome, of which, 741.5 Mb is non-gapped with 731.2 Mb anchored to the 12 chromosomes. Use of Oxford Nanopore Technologies full-length cDNA sequencing enabled annotation of 32,917 high-confidence protein-coding genes encoding 44,851 gene models that had a significantly improved representation of conserved orthologs compared to the previous annotation. The new assembly has improved contiguity with a 595-fold increase in N50 contig size, 99% reduction in the numbersof contigs, a 44-fold increase in N50 scaffold size, and an LTR Assembly Index score of 13.56, placing it in the category of reference genome quality. The improved assembly also permitted annotation of the centromeres via alignment to sequencing reads derived from CENH3 nucleosomes. </span></p> <p><span><strong>Conclusions: </strong>Access to advanced sequencing technologies and improved software permitted generation of a high-quality, long-read, chromosome-scale assembly and improved annotation dataset for the reference genotype of potato that will facilitate research aimed at improving agronomic traits and understanding genome evolution.</span></p>
Data from: Chromosome-level assembly of Southern catfish (Silurus meridionalis) provides insights into visual adaptation to the nocturnal and benthic lifestyles
Southern catfish (<i><span>Silurus meridionalis</span></i>) is a nocturnal and benthic freshwater fish endemic to the Yangtze River and its tributaries with an important economic value but a drastically declining wild population. In this study, we constructed a chromosome-level draft genome of <i><span>S. meridionalis</span></i> using 69.7 Gb Nanopore long reads and 49.5 Gb Illumina short reads. The genome assembly was 741.2 Mb in size with a contig N50 of 13.19 Mb. An additional 116.4 Gb of Bionano and 77.4 Gb of Hi-C data were applied to assemble contigs into scaffolds and further into 29 chromosomes, resulting in a 738.9 Mb genome with a scaffold N50 of 28.04 Mb. A total of 22,965 protein-coding genes were predicted from the genome with 22,519 (98.06%) genes functionally annotated. Comparative genomic and transcriptomic analyses revealed that catfish possess a rod-dominated visual system which is responsible for scotopic vision. The absence of cone opsins SWS1 and SWS2 resulted in the lack of UV and violet sensitivity. Mutations at key amino acid sites of RH1.1, RH1.2 and RH2 resulted in spectral tuning good for dim light vision and narrow color vision. A higher expression level of rod phototransduction genes than that of cone genes and higher rod-to-cone ratio leaded to higher optical sensitivity under dim light conditions. In addition, analysis of the genes involved in eye morphogenesis and development revealed the loss of conserved noncoding elements (CNEs), which might be associated with the small eyes in catfish. Taken together, our study provided important clues for the adaptation of the catfish visual system to a benthic lifestyle. The draft genome of <i><span>S. meridionalis</span></i> represents a valuable resource for elucidation of molecular mechanism of ecological adaptation, as well as genetic breeding in aquaculture.
Data from: Chromosome-level genome assembly of Paralithodes platypus provides insights into evolution and adaptation of king crabs
<p>The blue king crab, <i>Paralithodes platypus</i>, which belongs to the Lithodidae family, is a commercially and ecologically important species. However, a high-quality reference genome for the king crab has not yet been reported. Here, we assembled the first chromosome-level blue king crab genome, which contains 104 chromosomes and an N50 length of 51.15 Mb. Furthermore, we determined that the large genome size can be attributed to the insertion of long interspersed nuclear elements and long tandem repeats. Genome assembly assessment showed that 96.54% of the assembled transcripts could be aligned to the assembled genome. Phylogenetic analysis showed the blue king crab to have a close relationship with the Eubrachyura crabs, from which it diverged 272.5 million years ago. Population history analyses indicated that the effective population of the blue king crab declined sharply and then gradually increased from the Cretaceous and Neogene periods, respectively. Furthermore, gene families related to developmental pathways, steroid and thyroid hormone synthesis, and inflammatory regulation, were expanded in the genome, suggesting that these genes contributed substantially to the environmental adaptation and unique body plan evolution of the blue king crab. The high-quality reference genome reported here provides a solid molecular basis for further study of the blue king crab's development and environmental adaptation.</p>
Chromosome‐level genome assembly of Lethenteron reissneri provides insights into lamprey evolution
<p>The reissner lamprey<i> Lethenteron reissneri</i> belonging to Cyclostomata, serves as a bridge between invertebrates and jawed vertebrates, and is considered the most direct ancestor of vertebrates. However, the genetic mechanisms underlying the adaptive evolution of lampreys remain unclear. Here, we supplied the genome data and annotation data of <em>Lethenteron reissneri</em>. Total 5 files were uploaded, including the assembled genome of <i>Lethenteron reissneri</i>, the gene annotation file in gff format, the gene function annotation file, the LIP gene sequences of 50 species used in the article, and the readme file. This study not only provides the first chromosome-level reference genome in Cyclostomata, but also indicates the unique biology and adaptive evolution of lampreys.</p>
Data from: Heterochromatin-enriched assemblies reveal the sequence and organization of the Drosophila melanogaster Y chromosome
Heterochromatic regions of the genome are repeat-rich and poor in protein coding genes, and are therefore underrepresented in even the best genome assemblies. One of the most difficult regions of the genome to assemble are sex-limited chromosomes. The Drosophila melanogaster Y chromosome is entirely heterochromatic, yet has wide-ranging effects on male fertility, fitness, and genome-wide gene expression. The genetic basis of this phenotypic variation is difficult to study, in part because we do not know the detailed organization of the Y chromosome. To study Y chromosome organization in D. melanogaster, we develop an assembly strategy involving the in silico enrichment of heterochromatic long single-molecule reads and use these reads to create targeted de novo assemblies of heterochromatic sequences. We assigned contigs to the Y chromosome using Illumina reads to identify male-specific sequences. Our pipeline extends the D. melanogaster reference genome by 11.9 Mb, closes 43.8% of the gaps, and improves overall contiguity. The addition of 10.6 MB of Y-linked sequence permitted us to study the organization of repeats and genes along the Y chromosome. We detected a high rate of duplication to the pericentric regions of the Y chromosome from other regions in the genome. Most of these duplicated genes exist in multiple copies. We detail the evolutionary history of one sex-linked gene family—crystal-Stellate. While the Y chromosome does not undergo crossing over, we observed high gene conversion rates within and between members of the crystal-Stellate gene family, Su(Ste), and PCKR, compared to genome-wide estimates. Our results suggest that gene conversion and gene duplication play an important role in the evolution of Y-linked genes.
Data from: Chromosome-level reference genome assembly and gene editing of the dead-leaf butterfly Kallima inachus
<p class="p">The leaf resemblance of <i><span>Kallima</span></i> (Nymphalidae) butterflies<i> </i>is an important ecological adaptive mechanism that increases survival. However, the genetic mechanism underlying ecological adaptation remains unclear owing to a dearth of genomic information. Herein, we revealed the karyotype (n = 31) of the dead-leaf butterfly <i><span>Kallima inach</span></i><i><span>us</span></i>, assembled its high-quality chromosome-level reference genome (568.92 Mb; contig N50: 19.20 Mb), and identified its Z and candidate W chromosomes. To our knowledge, this is the first study to report on these aspects of this species. In the assembled genome, 15,309 protein-coding genes and 49.86% repeat elements were annotated. Phylogenetic analysis showed that <i><span>K. inachus</span></i> diverged from <i><span>Melitaea cinxia </span></i>(no leaf resemblance), both of which are in Nymphalinae, around 40 million years ago. Demographic analysis indicated that the effective population size of <i><span>K. inach</span></i><i><span>us</span></i> decreased during the last interglacial period in the Pleistocene. The wings of adults with the pigmentary gene <i><span>ebony</span></i> knocked out using CRISPR/Cas9 showed phenotypes in which the orange dorsal region and entire ventral surface darkened, suggesting its vital role in the ecological adaption of dead-leaf butterflies. Our results provide important genome resources for investigating the genetic mechanism underlying protective resemblance in dead-leaf butterflies and insights into the molecular basis of protective coloration.</p>
Data from: Major improvements to the Heliconius melpomene genome assembly used to confirm 10 chromosome fusion events in 6 million years of butterfly evolution
The Heliconius butterflies are a widely studied adaptive radiation of 46 species spread across Central and South America, several of which are known to hybridize in the wild. Here, we present a substantially improved assembly of the Heliconius melpomene genome, developed using novel methods that should be applicable to improving other genome assemblies produced using short read sequencing. First, we whole-genome-sequenced a pedigree to produce a linkage map incorporating 99% of the genome. Second, we incorporated haplotype scaffolds extensively to produce a more complete haploid version of the draft genome. Third, we incorporated ∼20x coverage of Pacific Biosciences sequencing, and scaffolded the haploid genome using an assembly of this long-read sequence. These improvements result in a genome of 795 scaffolds, 275 Mb in length, with an N50 length of 2.1 Mb, an N50 number of 34, and with 99% of the genome placed, and 84% anchored on chromosomes. We use the new genome assembly to confirm that the Heliconius genome underwent 10 chromosome fusions since the split with its sister genus Eueides, over a period of about 6 million yr.
Chromosome-level genome assembly of Triticum turgidum var 'Kronos'
<p> </p> <h2><strong>This data is made available under the Toronto Agreement. </strong></h2> <p><strong>All of the data listed here is available under the prepublication data sharing principle of the </strong><a href="https://www.nature.com/articles/461168a"><strong>Toronto agreement</strong></a><strong> (1). By using this data, you agree to:</strong></p> <ul> <li><strong>respect the rights of the data producers and contributors to analyze and publish the first global analyses and certain other reserved analyses of this data set in a peer-reviewed publication.</strong></li> <li><strong>not redistribute, release, or otherwise provide access to the data to anyone outside of the group, until the data has been published & submitted to the public data repositories.</strong></li> <li><strong>contact the authors to discuss any plans to publish data or analyses that utilize this data to avoid the overlap of any planned analyses.</strong></li> <li><strong>fully cite the prepublication data along with any applicable versioning details.</strong></li> <li><strong>understand that this data as accessed is precompetitive and is not patentable in its present state.</strong></li> </ul> <p><strong>This agreement does not expire by time but only upon publication of the first global analysis by the data producers and contributors.</strong><br><strong>(1) Toronto International Data Release Workshop Authors. Prepublication data sharing. </strong><em><strong>Nature</strong></em><strong> 461, 168–170 (2009). </strong><a href="https://doi.org/10.1038/461168a"><strong>https://doi.org/10.1038/461168a</strong></a></p> <p> </p> <ul> <li><em>If you have questions about <strong>the use</strong> of this dataset, please contact Ksenia Krasileva: kseniak [at] berkeley.edu</em></li> </ul> <p> </p> <p><strong>Updates in Zenodo v7</strong></p> <p>This update includes annotations of non-coding RNAs. Please refer to our <a href="https://github.com/s-kyungyong/Kronos">github</a> to understand how this datasets were produced. Please check additional datasets here: <a href="https://zenodo.org/records/15801566">Chromosome-level genome assembly of Triticum turgidum var 'Kronos' additional datasets.</a></p> <p> </p> <p><strong>Acknowledgement</strong></p> <p>This work has been funded by the United States Department of Agriculture - National Institute for Food and Agriculture Award (2021-67013-35726). <br><br><br></p>
The de novo assembly of a European wild boar genome revealed unique patterns of chromosomal structural variations and segmental duplications
<div> <div> <p><a href="https://onlinelibrary.wiley.com/doi/10.1111/age.13181">https://onlinelibrary.wiley.com/doi/10.1111/age.13181</a></p> <h1>The de novo assembly of a European wild boar genome revealed unique patterns of chromosomal structural variations and segmental duplications</h1> <div> </div> <div> <div> <div> <div><a href="https://onlinelibrary.wiley.com/authored-by/Chen/Jianhai">Jianhai Chen</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Zhong/Jie">Jie Zhong</a>, <a href="https://onlinelibrary.wiley.com/authored-by/He/Xuefei">Xuefei He</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Li/Xiaoyu">Xiaoyu Li</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Ni/Pan">Pan Ni</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Safner/Toni">Toni Safner</a>, <a href="https://onlinelibrary.wiley.com/authored-by/%C5%A0prem/Nikica">Nikica Šprem</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Han/Jianlin">Jianlin Han</a></div> </div> </div> </div> <p>The rapid progress of sequencing technology has greatly facilitated the de novo genome assembly of pig breeds. However, the assembly of the wild boar genome is still lacking, hampering our understanding of chromosomal and genomic evolution during domestication from wild boars into domestic pigs. Here, we sequenced and de novo assembled a European wild boar genome (ASM2165605v1) using the long-range information provided by 10× Linked-Reads sequencing. We achieved a high-quality assembly with contig N50 of 26.09 Mb. Additionally, 1.64% of the contigs (222) with lengths from 107.65 kb to 75.36 Mb covered 90.3% of the total genome size of ASM2165605v1 (~2.5 Gb). Mapping analysis revealed that the contigs can fill 24.73% (93/376) of the gaps present in the orthologous regions of the updated pig reference genome (Sscrofa11.1). We further improved the contigs into chromosome level with a reference-assistant scaffolding method. Using the ‘assembly-to-assembly’ approach, we identified intra-chromosomal large structural variations (SVs, length >1 kb) between ASM2165605v1 and Sscrofa11.1 assemblies. Interestingly, we found that the number of SV events on the X chromosome deviated significantly from the linear models fitting autosomes (<em>R</em><sup>2</sup> > 0.64, <em>p</em> < 0.001). Specifically, deletions and insertions were deficient on the X chromosome by 66.14 and 58.41% respectively, whereas duplications and inversions were excessive on the X chromosome by 71.96 and 107.61% respectively. We further used the large segmental duplications (SDs, >1 kb) events as a proxy to understand the large-scale inter-chromosomal evolution, by resolving parental-derived relationships for SD pairs. We revealed a significant excess of SD movements from the X chromosome to autosomes (<em>p</em> < 0.001), consistent with the expectation of meiotic sex chromosome inactivation. Enrichment analyses indicated that the genes within derived SD copies on autosomes were significantly related to biological processes involving nervous system, lipid biosynthesis and sperm motility (<em>p</em> < 0.01). Together, our analyses of the de novo assembly of ASM2165605v1 provides insight into the SVs between European wild boar and domestic pig, in addition to the ongoing process of meiotic sex chromosome inactivation in driving inter-chromosomal interaction between the sex chromosome and autosomes.</p> </div> </div> <div>The work has been pulished here: https://onlinelibrary.wiley.com/doi/full/10.1111/age.13181</div> <div> </div> <div>The current dataset include the genome annotation files.</div> <div> </div> <div>For the whole-genomic assembly, please check NCBI: </div> <div>https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_021656055.1/</div> <div> <table> <tbody> <tr> <th> </th> <th>GenBank</th> </tr> </tbody> <tbody> <tr> <td>Genome size</td> <td>2.5 Gb</td> </tr> <tr> <td>Total ungapped length</td> <td>2.4 Gb</td> </tr> <tr> <td>Number of scaffolds</td> <td>12,642</td> </tr> <tr> <td>Scaffold N50</td> <td>28.3 Mb</td> </tr> <tr> <td>Scaffold L50</td> <td>25</td> </tr> <tr> <td>Number of contigs</td> <td>41,323</td> </tr> <tr> <td>Contig N50</td> <td>157.9 kb</td> </tr> <tr> <td>Contig L50</td> <td>4,562</td> </tr> <tr> <td>GC percent</td> <td>42</td> </tr> <tr> <td>Genome coverage</td> <td>56.0x</td> </tr> <tr> <td>Assembly level</td> <td>Scaffold</td> </tr> </tbody> </table> <p> </p> <h2>Assembly methods</h2> <div>Sequencing technology 10xgenomics Assembly method Supernova v. 2.1.1 <p> </p> <p>part_** are genome fasta for the GCA_021656055.1</p> <p>You could use the following to combine and uncompress.</p> </div> </div> <div> <div> <div><code><span>cat</span> part_* > archive_combined.zip </code></div> </div> <div> <div> </div> <div><code>unzip archive_combined.zip</code></div> </div> </div> <div> </div> <div> </div>
GALA: a computational framework for de novo chromosome-by-chromosome assembly with long reads
<p>C.elegans, O.sativa and Human assemblies produced by GALA software</p>
The assembly of caprine Y chromosome sequence reveals a unique paternal phylogenetic pattern and improves our understanding the origin of domestic goat
<p>The mammalian Y chromosome offers a unique perspective on the male reproduction and paternal evolutionary histories. However, further understanding of the Y chromosome biology for most mammals is hindered by the lack of a Y chromosome assembly. This study presents an integrated <i>in silico</i> strategy for identifying and assembling the goat Y-linked scaffolds <span>using existing data</span>. A total of 11.5 Mb Y-linked sequences were clustered into 33 scaffolds, and 187 protein-coding genes were annotated. We also identified high abundance of repetitive elements. A 5.84 Mb subset was further ordered into an assembly with the evidence from the goat Radiation-Hybrid map (RH map). The existing whole-genome re-sequencing data of 96 goats (worldwide distribution) were utilized to exploit the paternal relationships among bezoars and domestic goats. Goat paternal lineages were clearly divided into two clades (Y1 and Y2), predating the goat domestication. Demographic history analyses indicated that maternal lineages experienced a bottleneck effect around 2,000 YBP (years before present), after which goats belonging to the A haplogroup spread worldwide from the Near East. As opposed to this, paternal lineages experienced a population decline around the 10,000 YBP. The evidence from the Y chromosome suggests that male goats were not affected by the A haplogroup worldwide transmission, which implies sexually unbalanced contribution to the goat trade and population expansion in post-Neolithic period.</p>
Data from: Chromosome scale genome assemblies and annotations for Poales species Carex cristatella, Carex scoparia, Juncus effusus and Juncus inflexus
<p>The majority of sequenced genomes in the Monocots are from species belonging to the Poaceae, which includes many commercially important crops. Here, we expand the number of sequenced genomes from the Monocots to include the genomes of four related Cyperids: <em>Carex cristatella</em> and <em>Carex scoparia</em> from Cyperaceae and <em>Juncus effusus</em> and <em>Juncus inflexus</em> from Juncaceae. The high-quality, chromosome-scale genome sequences from these four Cyperids were assembled by combining whole-genome shotgun sequencing of Nanopore long reads, Illumina short reads, and Hi-C sequencing data. Some members of the Cyperaceae and Juncaceae are known to possess holocentric chromosomes. We examined the repeat landscapes in our sequenced genomes to search for potential repeats associated with centromeres. Several large satellite repeat families, comprising 3.2% to 9.5% of our sequenced genomes, showed dispersed distribution of large repeat clusters across all <em>Carex</em> chromosomes, with few instances of these repeats clustering in the same chromosomal regions. In contrast, most large <em>Juncus</em> satellite repeats were clustered in a single location on each chromosome, with sporadic instances of large satellite repeats throughout the Juncus genomes. Recognizable transposable elements account for about 20% of the assemblies, with the <em>Carex</em> genomes containing more DNA transposons than retrotransposons while the converse is true for the <em>Juncus</em> genomes. These genome sequences and annotations will facilitate better comparative analysis within monocots.</p>
A chromosome-level genome assembly of the orange wheat blossom midge, Sitodiplosis mosellana Géhin (Diptera: Cecidomyiidae)
<p><span>The Orange wheat blossom midge <i>Sitodiplosis mosellana </i>Géhin (Diptera: Cecidomyiidae), an important insect pest, has caused serious yield losses in most wheat-growing areas worldwide in the past half-century. In this study, we assembled the first chromosomal level genome for <i>S. mosellana</i> using PacBio long-read, Illumina short-read sequences and high-throughput chromatin conformation capture (Hi-C) genome scaffolding techniques. The final genome assembly was 180.69 Mb, with contig and scaffold N50 sizes of 998.71 kb and 44.56 Mb, respectively. Hi-C scaffolding reliably anchored four pseudochromosomes, accounting for 99.67% of the assembled genome. The assembly showed high integrity and quality, with 91.7% of short reads mapped to the genome and a coverage rate of 99.8%. The assembly quality was evaluated using Core Eukaryotic Genes Mapping Approach and Benchmarking Universal Single-Copy Orthologs. In total, 12,269 protein-coding genes were predicted, of which 91% were functionally annotated. Phylogenetic analysis indicated that <i>S. mosellana</i> and its close relative the swede midge <i>Contarinia nasturtii</i> diverged about 32.7 million years ago. <i>S. mosellana</i> genome showed high chromosomal synteny with the genome of <i>Drosophila melanogaster</i> and <i>Anopheles gambiae</i>. The key gene families involved in chemosensation and detoxification of plant secondary chemistry were analysed<i>.</i> The high-quality <i>S. mosellana</i> genome data will provide an invaluable resource for research in a broad range of areas, including the biology, ecology, genetics, and evolution of midges as well as insect-plant interactions and co-evolution, and their relatives more generally.</span></p>
A chromosome-level genome assembly of Gekko japonicus
<p>We assembled and annotated a chromosome-level genome of <em>Gekko japonicus.</em></p>
A chromosome-scale reference genome assembly of the great sand eel, Hyperoplus lanceolatus
<p><span>Despite increasing sequencing efforts, numerous fish families still lack a reference genome, which complicates genetic research. One such understudied family is the sand lances (Ammodytidae, literally: 'sand burrower'), a globally distributed clade of over 30 fish species that tend to avoid tidal currents by burrowing into the sand. Here, we present the first annotated chromosome-level genome assembly of the great sand eel (<em>Hyperoplus</em> <em>lanceolatus</em>). The genome assembly was generated using Oxford Nanopore Technologies long sequencing reads and Illumina short reads for polishing. The final assembly has a total length of 808.5 Mbp, of which 97.1% were anchored into 24 chromosome-scale scaffolds using proximity-ligation scaffolding. The assembly is highly contiguous with a scaffold and contig N50 of 33.7 Mbp and 31.3 Mbp, respectively, and has a BUSCO completeness score of 96.9%. The presented genome assembly is a valuable resource for future studies of sand lances, as they are of great ecological and commercial importance and may also contribute to studies aiming to resolve the suprafamiliar taxonomy of bony fishes.</span></p>
Genome annotation associated with the publication "Chromosome-level genome assembly of the Cape cliff lizard (Hemicordylus capensis)"
<p>Genome annotation associated with the publication "Chromosome-level genome assembly of the Cape cliff lizard (<em>Hemicordylus capensis</em>)"</p> <p>rHemCap1.1.gff3 - Genome annotation in GFF3 format<br> rHemCap1.1.proteins.fa - Multi-fasta file of protein coding genes<br> rHemCap1.1.cds-transcripts.fa - Multi-fasta file of transcripts (CDS)</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.