Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
68
datasets available to search
ShareScore release 0.9.0
Dataset results
68 results for “De novo genome assembly”
Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C
<p><i>Telopea speciosissima, </i>the New South Wales waratah, is an Australian endemic woody shrub in the family Proteaceae. Waratahs have great potential as a model clade to better understand processes of speciation, introgression and adaptation, and are significant from a horticultural perspective. Here, we report the first chromosome-level genome for <i>T. speciosissima</i>. Combining Oxford Nanopore long-reads, 10x Genomics Chromium linked-reads and Hi-C data, the assembly spans 823 Mb (scaffold N50 of 69.0 Mb) with 97.8 % of Embryophyta BUSCOs 'Complete'. We present a new method in Diploidocus (<a href="https://github.com/slimsuite/diploidocus">https://github.com/slimsuite/diploidocus</a>) for classifying, curating and QC-filtering scaffolds, which combines read depths, <i>k</i>-mer frequencies and BUSCO predictions. We also present a new tool, DepthSizer (<a href="https://github.com/slimsuite/depthsizer">https://github.com/slimsuite/depthsizer</a>), for genome size estimation from the read depth of single-copy orthologues and estimate the genome size to be approximately 900 Mb. The largest 11 scaffolds contained 94.1 % of the assembly, conforming to the expected number of chromosomes (2<i>n</i> = 22). Genome annotation predicted 40,158<code> </code>protein-coding genes, 351 rRNAs and 728 tRNAs. We investigated <i>CYCLOIDEA </i>(<i>CYC</i>)<i> </i>genes, which have a role in determination of floral symmetry, and confirm the presence of two copies in the genome. Read depth analysis of 180 'Duplicated' BUSCO genes using a new tool, DepthKopy (<a href="https://github.com/slimsuite/depthkopy">https://github.com/slimsuite/depthkopy</a>), suggests almost all are real duplications, increasing confidence in the annotation and highlighting a possible need to revise the BUSCO set for this lineage. The chromosome-level <i>T. speciosissima</i> reference genome (Tspe_v1) provides an important new genomic resource of Proteaceae to support the conservation of flora in Australia and further afield.</p>
De novo assembly of 20 chicken genomes reveals the undetectable phenomenon for thousands of core genes on micro-chromosomes and sub-telomeric regions
<p>The gene numbers and evolutionary rates of birds were assumed to be much lower than those of mammals, which is in sharp contrast to the huge species number and morphological diversity of birds. It is therefore necessary to construct a complete avian genome and analyze its evolution. We constructed a chicken pan-genome from 20 <em>de novo</em> assembled genomes with high sequencing depth, and identified 1,335 protein-coding genes and 3,011 long noncoding RNAs not found in GRCg6a. The majority of these novel genes were detected across most individuals of the examined transcriptomes but were seldomly measured in each of the DNA sequencing data regardless of Illumina or PacBio technology. Furthermore, different from previous pan-genome models, most of these novel genes were overrepresented on chromosomal sub-telomeric regions and micro-chromosomes, surrounded by extremely high proportions of tandem repeats, which strongly blocks DNA sequencing. These hidden genes were proved to be shared by all chicken genomes, included many housekeeping genes, and enriched in immune pathways. Comparative genomics revealed the novel genes had three-fold elevated substitution rates than known ones, updating the knowledge about evolutionary rates in birds. Our study provides a framework for constructing a better chicken genome, which will contribute towards the understanding of avian evolution and improvement of poultry breeding.</p>
De novo genome assembly of Kallima inachus
<p><span>Oakleaf butterflies in the genus <em>Kallima</em> have a polymorphic wing phenotype, enabling these insects to masquerade as dead leaves. By studying mechanisms that shape the genetic and species diversity of these butterflies, a new perspective can be provided to understand the evolutionary innovation driven by geographic changes and natural selection.</span></p> <p><span>We found that leaf wing polymorphism in <em>Kallima</em> butterflies is controlled by the wing patterning gene cortex. We hypothesized that multiple mechanisms may independently lead to the reduction or suppression of recombination among different cortex haplotypes. To test this hypothesis, w</span>e performed Nanopore re-sequencing and <em>de novo</em> genome assembly for 4 <em>Kallima inachus</em> individuals and obtained 4 individual genomes. We identified two chromosomal inversions spanning these haplotypes.</p>
Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C
Open the record for dataset details and reuse information.
De novo genome assembly of Kallima inachus
Open the record for dataset details and reuse information.
De novo genome assembly of human cell line CHM13 nanopore ultra-long reads using Shasta
Open the record for dataset details and reuse information.
De novo genome assembly of Tectona grandis (Teak) with 2993 scaffolds
<p>Teak (<em>Tectona grandis</em> L. f.) is one of the precious bench mark tropical hardwood having qualities of durability, strength and visual pleasantries. Natural teak populations harbour a variety of characteristics that determine their economic, ecological and environmental importance. Sequencing of whole nuclear genome of teak provides a platform for functional analyses and development of genomic tools in applied tree improvement. A draft genome of 317 Mb was assembled at 151× coverage and annotated 36, 172 protein-coding genes. Approximately about 11.18% of the genome was repetitive. Microsatellites or simple sequence repeats (SSRs) are undoubtedly the most informative markers in genotyping, genetics and applied breeding applications. We generated 182,712 SSRs at the whole genome level, of which, 170,574 perfect SSRs were found; 16,252 perfect SSRs showed <em>in silico</em> polymorphisms across six genotypes suggesting their promising use in genetic conservation and tree improvement programmes. Genomic SSR markers developed in this study have high potential in advancing conservation and management of teak genetic resources. Phylogenetic studies confirmed the taxonomic position of the genus <em>Tectona</em> within the family Lamiaceae. Interestingly, estimation of divergence time inferred that the Miocene origin of the <em>Tectona</em> genus to be around 21.4508 million years ago.</p>
Data from: "De novo transcriptome assembly of the mountain fly Drosophila nigrosparsa using short RNA-seq reads" in Genomic Resources Notes Accepted 1 August 2014-30 September 2014
Drosophila (Drosophila) nigrosparsa is a habitat specialist restricted to the European montane/alpine zone (Bächli 2008). Mountain biodiversity is considered highly vulnerable to ongoing climate warming (IPCC 2013), and organisms at high altitudes have only limited possibility to shift to cooler habitats at elevations above (Pertoldi & Bach 2007). For such species, rapid evolution may offer a solution for long-term survival. We are establishing D. nigrosparsa as a model system to test the extent and tempo of adaptive evolution under thermal stress in the laboratory. In this study, we used Illumina high-throughput sequencing to assemble the species' transcriptome using the pooled mRNA from 22 developmental and physiological stages.
Data from: "De novo assembly transcriptome for the rostrum dace (Leuciscus burdigalensis, Cyprinidae: fish) naturally infected by a copepod ectoparasite" in Genomic Resources Notes accepted 1 December 2014 to 31 January 2015
The emergence of pathogens represents substantial threats to public health, livestock, domesticated animals, and biodiversity. How wild populations respond to emerging pathogens has generated a lot of interest in the last two decades. With the recent advent of high-throughput sequencing technologies it is now possible to develop large transcriptomic resources for non-model organisms, hence allowing new research avenues on the immune responses of hosts from a large taxonomic spectra. We here focused on a wild population of the rostrum dace (Leuciscus burgiladensis) that is infected by Tracheliastes polycolpus, an emerging freshwater ectoparasite copepod. We used next generation Illumina sequencing technology to sequence the transcriptome of eight L. burdigalensis adult individuals collected in natura from the same sampling site. Four individuals were non-infected and four individuals were infected by T. polycolpus. We specifically focused on the spleen, the head kidney and epithelial cells and mucus from the fins, three tissues known to be involved in the immune response of fish. We used the Trinity methodology to reconstruct a de novo full-length transcriptome for L. burdigalensis. The resulting transcriptome will serve as an important broad-scale genomic resource for further studying the response of local population of L. burdigalensis to T. polycolpus pressures.
Data from: "De novo assembled transcriptome of organs involved in reproduction in an endangered endemic Iberian cyprinid fish (Squalius pyrenaicus)" in Genomic Resources Notes Accepted 1 June 2015 to 31 July 2015
Sex determination systems are diverse, especially among fish, and include genetic and/or environmental components. Unexpectedly for such a basic aspect of development, sex determination systems change rapidly during evolution and gonadal fate is not ultimate, being actively maintained lifelong. Here, sequences of expressed genes involved in maintenance of gonad identity and reproduction processes were obtained through transcriptome assembly of the brain-gonadal axis tissues of a freshwater fish inhabiting highly variable environments, the gonochoristic Iberian fish Squalius pyrenaicus. Through Illumina total RNA-sequencing, male and female transcriptomes of brain and gonad tissues were assembled with Trans-ABySS software and merged to produce a more comprehensive S. pyrenaicus transcriptome. Coding sequences (CDS) predicted by TransDecoder were annotated using blastx. By means of read mapping against the reference transcriptome and CDS datasets, using Bowtie2, the accuracy of read mapping was assessed. This first endemic Iberian cyprinid transcriptome of organs involved in reproduction processes may serve as a valuable genomic resource for studying sexual mechanisms and other aspects of evolution, such as speciation and responses to environmental changes, and may be a useful tool for conservation studies since S. pyrenaicus is an endangered species.
Data from: Likelihood-based inference of population history from low coverage de novo genome assemblies
Short-read sequencing technologies have in principle made it feasible to draw detailed inferences about the recent history of any organism. In practice, however, this remains challenging due to the difficulty of genome assembly in most organisms and the lack of statistical methods powerful enough to discriminate among recent, non-equilibrium histories. We address both the assembly and inference challenges. We develop a bioinformatic pipeline for generating outgroup-rooted alignments of orthologous sequence blocks from de novo low-coverage short-read data for a small number of genomes, and show how such sequence blocks can be used to fit explicit models of population divergence and admixture in a likelihood framework. To illustrate our approach, we reconstruct the Pleistocene history of an oak-feeding insect (the oak gallwasp Biorhiza pallida) which, in common with many other taxa, was restricted during Pleistocene ice ages to a longitudinal series of southern refugia spanning theWestern Palaearctic. Our analysis of sequence blocks sampled from a single genome from each of three major glacial refugia reveals support for an unexpected history dominated by recent admixture. Despite the fact that 80% of the genome is affected by admixture during the last glacial cycle, we are able to infer the deeper divergence history of these populations. These inferences are robust to variation in block length, mutation model, and the sampling location of individual genomes within refugia. This combination of de novo assembly and numerical likelihood calculation provides a powerful framework for estimating recent population history that can be applied to any organism without the need for prior genetic resources.
De novo genome assembly of Meloidogyne chitwoodi
<p>A whole genome of <em>M. chitwoodi</em> was <em>de novo</em> assembled by empirically optimizing k-mer sizes </p>
The de novo assembly of a European wild boar genome revealed unique patterns of chromosomal structural variations and segmental duplications
<div> <div> <p><a href="https://onlinelibrary.wiley.com/doi/10.1111/age.13181">https://onlinelibrary.wiley.com/doi/10.1111/age.13181</a></p> <h1>The de novo assembly of a European wild boar genome revealed unique patterns of chromosomal structural variations and segmental duplications</h1> <div> </div> <div> <div> <div> <div><a href="https://onlinelibrary.wiley.com/authored-by/Chen/Jianhai">Jianhai Chen</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Zhong/Jie">Jie Zhong</a>, <a href="https://onlinelibrary.wiley.com/authored-by/He/Xuefei">Xuefei He</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Li/Xiaoyu">Xiaoyu Li</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Ni/Pan">Pan Ni</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Safner/Toni">Toni Safner</a>, <a href="https://onlinelibrary.wiley.com/authored-by/%C5%A0prem/Nikica">Nikica Šprem</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Han/Jianlin">Jianlin Han</a></div> </div> </div> </div> <p>The rapid progress of sequencing technology has greatly facilitated the de novo genome assembly of pig breeds. However, the assembly of the wild boar genome is still lacking, hampering our understanding of chromosomal and genomic evolution during domestication from wild boars into domestic pigs. Here, we sequenced and de novo assembled a European wild boar genome (ASM2165605v1) using the long-range information provided by 10× Linked-Reads sequencing. We achieved a high-quality assembly with contig N50 of 26.09 Mb. Additionally, 1.64% of the contigs (222) with lengths from 107.65 kb to 75.36 Mb covered 90.3% of the total genome size of ASM2165605v1 (~2.5 Gb). Mapping analysis revealed that the contigs can fill 24.73% (93/376) of the gaps present in the orthologous regions of the updated pig reference genome (Sscrofa11.1). We further improved the contigs into chromosome level with a reference-assistant scaffolding method. Using the ‘assembly-to-assembly’ approach, we identified intra-chromosomal large structural variations (SVs, length >1 kb) between ASM2165605v1 and Sscrofa11.1 assemblies. Interestingly, we found that the number of SV events on the X chromosome deviated significantly from the linear models fitting autosomes (<em>R</em><sup>2</sup> > 0.64, <em>p</em> < 0.001). Specifically, deletions and insertions were deficient on the X chromosome by 66.14 and 58.41% respectively, whereas duplications and inversions were excessive on the X chromosome by 71.96 and 107.61% respectively. We further used the large segmental duplications (SDs, >1 kb) events as a proxy to understand the large-scale inter-chromosomal evolution, by resolving parental-derived relationships for SD pairs. We revealed a significant excess of SD movements from the X chromosome to autosomes (<em>p</em> < 0.001), consistent with the expectation of meiotic sex chromosome inactivation. Enrichment analyses indicated that the genes within derived SD copies on autosomes were significantly related to biological processes involving nervous system, lipid biosynthesis and sperm motility (<em>p</em> < 0.01). Together, our analyses of the de novo assembly of ASM2165605v1 provides insight into the SVs between European wild boar and domestic pig, in addition to the ongoing process of meiotic sex chromosome inactivation in driving inter-chromosomal interaction between the sex chromosome and autosomes.</p> </div> </div> <div>The work has been pulished here: https://onlinelibrary.wiley.com/doi/full/10.1111/age.13181</div> <div> </div> <div>The current dataset include the genome annotation files.</div> <div> </div> <div>For the whole-genomic assembly, please check NCBI: </div> <div>https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_021656055.1/</div> <div> <table> <tbody> <tr> <th> </th> <th>GenBank</th> </tr> </tbody> <tbody> <tr> <td>Genome size</td> <td>2.5 Gb</td> </tr> <tr> <td>Total ungapped length</td> <td>2.4 Gb</td> </tr> <tr> <td>Number of scaffolds</td> <td>12,642</td> </tr> <tr> <td>Scaffold N50</td> <td>28.3 Mb</td> </tr> <tr> <td>Scaffold L50</td> <td>25</td> </tr> <tr> <td>Number of contigs</td> <td>41,323</td> </tr> <tr> <td>Contig N50</td> <td>157.9 kb</td> </tr> <tr> <td>Contig L50</td> <td>4,562</td> </tr> <tr> <td>GC percent</td> <td>42</td> </tr> <tr> <td>Genome coverage</td> <td>56.0x</td> </tr> <tr> <td>Assembly level</td> <td>Scaffold</td> </tr> </tbody> </table> <p> </p> <h2>Assembly methods</h2> <div>Sequencing technology 10xgenomics Assembly method Supernova v. 2.1.1 <p> </p> <p>part_** are genome fasta for the GCA_021656055.1</p> <p>You could use the following to combine and uncompress.</p> </div> </div> <div> <div> <div><code><span>cat</span> part_* > archive_combined.zip </code></div> </div> <div> <div> </div> <div><code>unzip archive_combined.zip</code></div> </div> </div> <div> </div> <div> </div>
Data from: De Novo Genome assembly of the Caucasian dwarf goby Knipowitschia cf. caucasica, a new alien Gobiidae invading the River Rhine
<p><strong>Background (Abstract from Paper)</strong></p> <p>The Caucasian dwarf goby <em>Knipowitschia</em> cf. <em>caucasica</em> is a new invasive alien Gobiidae spreading in the<br>Lower Rhine since 2019. Little is known about the invasion biology of the species and further investiga-<br>tions to reconstruct the invasion history are lacking genomic resources. We assembled a high-quality<br>chromosome-scale reference genome of <em>Knipowitschia</em> cf. <em>caucasica</em> by combining PacBio, Omni-C and<br>Illumina technologies. The size of the assembled genome is 956.58 Mb with a N50 scaffold length of 43 Mb,<br>which includes 92.3 % complete vertebrate/Actinopterygii Benchmarking Universal Single-Copy Orthologs.<br>98.96 % of the assembly sequence was assigned to 23 chromosome-level scaffolds, with a GC-content of<br>42.83 %. Repetitive elements account for 53.08 % of the genome. The chromosome-level genome contained<br>49,622 transcripts with 42,926 multi-exons, of which 45,512 genes were functionally annotated. In summary,<br>the high-quality genome assembly provides a fundamental basis to understand the adaptive advantage of<br>the species.<br><br>The file provided here is the <strong>non-redundant repeat library</strong> (2,812 consensus sequences of repeat families). </p> <p><strong>Method:</strong> Repetitive elements were identified de novo with RepeatModeler version 2.0.1. Repetitive DNA and soft-masking was performed with RepeatMasker version 4.1.1 (Smit et al., 2013) using the repeat library previously identified via RepeatModeler and skipping the bacterial insertion element check (-no_is) and run with rmblastn version 2.10.0+ (Flynn et al., 2020). </p>
De novo assembly and annotation of parasitic trematode genomes
<p>Contained in this release are 19 genome assemblies and annotations of parasitic trematodes, encompassing 13 species. This included representatives of the <em>Schistosoma</em> (<em>n</em> = 13 assemblies), <em>Trichobilharzia</em> (<em>n</em> = 2 assemblies), <em>Heterobilharzia americana</em> (<em>n</em> = 2 assemblies) and <em>Dicrocoelium dendriticum </em>(<em>n </em>= 1 assembly). The <em>Schistosoma curassoni</em> assembly has been released previously (10.5281/zenodo.6594833) but a new annotation is included with the original assembly here. </p> <p>These genomes were assembled from a variety of sources including stored parasites from museum collections, established laboratory strains and wild-caught isolates sampled from natural hosts in endemic regions. Using a combination of DNA sequencing approaches, all genomes were assembled into chromosomal-scale scaffolds. This was followed by genome annotation based on short-read RNA sequencing (RNA-seq) and long-read isoform sequencing (Iso-seq) transcriptomic data.</p> <p>Included here are the primary assemblies (representing a non-redundant haploid genome) for each species (*.primary.fa), alternate loci (alternate representations of loci found in a largely haploid assembly; *.haplotypes.fa) and annotations (*.gff3). Metadata for each assembly can be found in the included spreadsheets (metadata.xlsx). </p> <p>This data is part of a pre-publication release. For information on the proper use of pre-publication data shared by the Wellcome Trust Sanger Institute (including details of any publication moratoria), please see https://www.sanger.ac.uk/about/research-policies/open-access-science/.</p> <p>This repository will be updated with a complete list of collaborators/authors prior to publication. Please contact Duncan Berger (db22@sanger.ac.uk) with questions regarding pre-publication use of this dataset. </p>
De novo genome assembly for Eulemur rufifrons
<p>As one of the most threatened mammalian taxa, lemurs of Madagascar are facing unprecedented anthropogenic pressures. To address conservation imperatives such as this, researchers have increasingly relied on conservation genomics to identify populations of particular concern. However, many of these genomic approaches necessitate high-quality genomes. While the advent of next generation sequencing technologies and the resulting reduction of associated costs have led to the proliferation of genomic data and high-quality reference genomes, global discrepancies in genomic sequencing capabilities often result in biological samples from biodiverse host countries being exported to facilities in the Global North, creating inequalities in access and training within genomic research. Here, we present the first reference genome for the endangered red-fronted brown lemur (Eulemur rufifrons) from sequencing efforts conducted entirely within the host country using portable Oxford Nanopore sequencing. Using an archived E. rufifrons specimen, we conducted long-read, nanopore sequencing at the Centre ValBio Research Station near Ranomafana National Park, in rural Madagascar, generating over 750 Gb of sequencing data from 10 MinION flow cells. Exclusively using this long-read data, we assembled 2.215 gigabase, 20,330-contig assembly with an N50 of 98.9 Mb and a 17,108 bp mitogenome. The nuclear assembly had 31x average coverage and was comparable in completeness to other primate reference genomes, with a 95.51% BUSCO completeness score for primate-specific genes. As the first reference genome for E. rufifrons and the only annotated genome available for the speciose Eulemur genus, this resource will prove vital for conservation genomic studies while our efforts exhibit the potential of this protocol to address research inequalities and build genomic capacity. </p>
Assemblies generated in the manuscript "Geometric deep learning framework for de novo genome assembly"
<p>Assemblies evaluated in the manuscript "Geometric deep learning framework for de novo genome assembly". All the assemblies were generated by us, except CHM13.ONT.Flye-2.9.fa.gz which was generated by <a href="https://www.nature.com/articles/s41587-019-0072-8">Kolmogorov et al. (2019)</a>.</p>
Quast outputs for "When do longer reads matter? A benchmark of long read de novo assembly tools for eukaryotic genomes"
<p>Quast outputs for "When do longer reads matter? A benchmark of long read de novo assembly tools for eukaryotic genomes"</p>
Data from: "De novo transcriptome assembly of the mountain fly Drosophila nigrosparsa using short RNA-seq reads" in Genomic Resources Notes Accepted 1 August 2014-30 September 2014
Open the record for dataset details and reuse information.
Data from: "De novo transcriptome assembly and polymorphism detection in ecological important widely distributed Neotropical toads from the Rhinella marina species complex (Anura: Bufonidade)" in Genomic Resources Notes Accepted 1 August 2014-30 September 2014
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.