Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
181
datasets available to search
ShareScore release 0.9.0
Dataset results
181 results for “de novo assembly”
De novo assembly of a long-read Amblyomma americanum genome (NCBI/Genbank deposited genome)
<p>Genome assembly of Amblyomma americanum generated from PacBio HiFi sequencing of 50 individual female ticks. This repository contains the phased pseudo-haploid tick genome generated after assembly using Flye, phasing using Purge_Dups, and clean-up using custom python scripts generated in-house. </p> <p>NCBI Bioproject: PRJNA932813</p>
Assemblies for "Linear time complexity de novo long read genome assembly with GoldRush"
<p>GoldRush is a <em>de novo</em> genome assembly algorithm with linear time complexity in the number of input long sequencing reads. We tested GoldRush on Oxford Nanopore Technologies datasets with different base error profiles describing the genomes of three human cell lines (NA24385, HG01243 and HG02055), Oryza sativa (rice), and Solanum lycopersicum (tomato). Here, we provide the assemblies for the GoldRush, Flye, Redbean and Shasta assemblies of these long read datasets.</p>
De novo transcriptome assembly and discovery of drought-responsive genes in eastern white spruce (Picea glauca)
Open the record for dataset details and reuse information.
Draft de novo genome assembly of the elusive jaguarundi, Puma yagouaroundi
Open the record for dataset details and reuse information.
Data from: A de novo chromosome-level genome assembly of Coregonus sp. “Balchen”: one representative of the Swiss Alpine whitefish radiation
Open the record for dataset details and reuse information.
De novo genome assembly of the Tobacco Hornworm moth (Manduca sexta)
<p><strong>We present the new reference genome for M sexta, JHU_Msex_v1.0, applying a combination of modern technologies in a de novo assembly to increase continuity, accuracy, and completeness. The assembly is 470 Mb and is ~25x more continuous than the original assembly, with scaffold N50 >14 Mb. We annotated the assembly by lifting over existing annotations and supplementing with additional supporting RNA-based data for a total of 25,256 genes. The new reference assembly is accessible in annotated form for public use.</strong></p>
de novo genome assembly of the LNCaP human prostate cancer cell line
<p>Whole-genome sequencing reads from the LNCaP human prostate cancer cell line were used to generate a <em>de novo </em>assembly with SGA v0.10.15. Please see https://github.com/sciseim/PCaWGS for associated scripts. Library preparation was performed using a TruSeq Nano DNA kit (Illumina) with a target insert size of 350bp. Paired-end libraries (150bp) were sequenced using a HiSeqX sequencer (Illumina).</p>
de novo genome assembly of the PC3 human prostate cancer cell line
<p>Whole-genome sequencing reads from the PC3 human prostate cancer cell line were used to generate a <em>de novo </em>assembly with SGA v0.10.15. Please see https://github.com/sciseim/PCaWGS for associated scripts. Library preparation was performed using a TruSeq Nano DNA kit (Illumina) with a target insert size of 350bp. Paired-end libraries (150bp) were sequenced using a HiSeqX sequencer (Illumina).</p> <p> </p> <p> </p> <p> </p>
De novo genome assembly of rice varieties using Nanopore long reads
<p>Genome sequences for Sugimura et al. (2024) of the rice (O. sativa) varieties 'Hitomebore' and 'Arroz da Terra.'</p> <p>Yusaku Sugimura, Kaori Oikawa, Yu Sugihara, Hiroe Utsushi, Eiko Kanzaki, Kazue Ito, Yumiko Ogasawara, Tomoaki Fujioka, Hiroki Takagi, Motoki Shimizu, Hiroyuki Shimono, Ryohei Terauchi, Akira Abe. Impact of rice GENERAL REGULATORY FACTOR14h (GF14h) on low-temperature seed germination and its application to breeding. PLoS Genet 20(8): e1011369. https://doi.org/10.1371/journal.pgen.1011369</p> <p>bioRxiv doi: https://doi.org/10.1101/2024.02.16.580620</p>
Chromosome-scale genome assembly and de novo annotation of Alopecurus aequalis.
<p><em>Alopecurus aequalis</em> is a winter annual or short-lived perennial bunchgrass which has in recent years emerged as the dominant agricultural weed of barley and wheat in certain regions of China and Japan, causing significant yield losses. Its robust tillering capacity and high fecundity, combined with the development of both target and non-target-site resistance to herbicides means it is a formidable challenge to food security. Here we report on a chromosome-scale assembly of <em>A. aequalis</em> with a genome size of 2.83 Gb. The genome contained 33,758 high-confidence protein-coding genes with functional annotation. Comparative genomics revealed that the genome structure of <em>A. aequalis</em> is more similar to <em>Hordeum vulgare </em>rather than the more closely related <em>Alopecurus myosuroides</em>. The datasets provided here are the assembly FASTA file (lpAloAequ1.1.prim.cur.20230912.fasta.gz), the high-confidence protein-coding genes (Alaeq_EIv0.2.release_HC_genes.gff3.gz) and the full annotation which includes both low and high confidence features of all biotypes (Alaeq_EIv0.2.release.gff3.gz) </p>
De novo transcriptome assembly Aegilops cylindrica
<p>De novo transcriptome assembly of Aegilops cylindrica was created using trinity v. 2.15.1. The assembly contains 285000 transcripts encoding for 174040 genes.</p>
De novo transcriptome assembly of the rockrose Helianthemum marifolium
<p>Illumina paired-end RNA sequences from leaves of 16 individuals of Helianthemum marifolium (four individuals from each of the four recognised taxonomic subspecies) were cleaned and assembled de novo using Trinity and Oases. EvidentialGene provides the curated assembled transcriptome containing 122002 transcripts, which are contained in the file Helianthemum_marifolium_transcriptome.fa.</p> <p>The predicted peptide sequences obtained with TransDecoder are in the file Helianthemum_marifolium_fa_transdecoder.pep.</p> <p>Gene ontology annotations from Trinotate are in the file Helianthemum_marifolium_trinotate_annotation.xls.</p>
Data From: Oatk - a de novo assembly tool for complex plant organelle genomes
<p>This reposity hosts the data for 195 plant organelle genome assemblies generated in the manuscript "Oatk: a de novo assembly tool for complex plant organelle genomes". The sequence data were produced by the Tree of Life programme at the Sanger Institute, mostly from the Darwin Tree of Life (DToL) project, including 24 monocots, 154 eudicots, 16 mosses and one liverwort. See SAMPLE_LIST file for descriptions of these species.</p> <p>In each species subfolder, below files are included.</p> <ol> <li><code>PLTD.fasta</code> Plastome assembly file in FASTA format</li> <li><code>PLTD.annot.bed</code> Plastome assembly annotation file in BED format</li> <li><code>MITO.fasta</code> Mitogenome assembly file in FASTA format</li> <li><code>MITO.annot.bed</code> Mitogenome assembly annotation file in BED format</li> <li><code>MBG.gfa</code> Genome assembly file in GFA format generated with MBG</li> <li><code>PMAT.gfa</code> Genome assembly file in GFA format generated with OATK</li> <li><code>OATK.gfa</code> Genome assembly file in GFA format generated with PMAT (may not exist)</li> </ol> <p> </p> <p>Updates in the New Version:</p> <p>In the previous version, our raw PacBio HiFi read pre-processing pipeline had screened out some reads that it erroneously thought contained HiFi adapter sequence, which led to the gaps in the Hibiscus plastomes. We now fixed this and have rerun all the assemblies that led to any linear organelle components (37 species). All plastomes remain unchanged except for the three Hibiscuses, which are now also circular. Thirteen mitogenomes changed, with six of them now becoming circular.</p>
Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C
<p><i>Telopea speciosissima, </i>the New South Wales waratah, is an Australian endemic woody shrub in the family Proteaceae. Waratahs have great potential as a model clade to better understand processes of speciation, introgression and adaptation, and are significant from a horticultural perspective. Here, we report the first chromosome-level genome for <i>T. speciosissima</i>. Combining Oxford Nanopore long-reads, 10x Genomics Chromium linked-reads and Hi-C data, the assembly spans 823 Mb (scaffold N50 of 69.0 Mb) with 97.8 % of Embryophyta BUSCOs 'Complete'. We present a new method in Diploidocus (<a href="https://github.com/slimsuite/diploidocus">https://github.com/slimsuite/diploidocus</a>) for classifying, curating and QC-filtering scaffolds, which combines read depths, <i>k</i>-mer frequencies and BUSCO predictions. We also present a new tool, DepthSizer (<a href="https://github.com/slimsuite/depthsizer">https://github.com/slimsuite/depthsizer</a>), for genome size estimation from the read depth of single-copy orthologues and estimate the genome size to be approximately 900 Mb. The largest 11 scaffolds contained 94.1 % of the assembly, conforming to the expected number of chromosomes (2<i>n</i> = 22). Genome annotation predicted 40,158<code> </code>protein-coding genes, 351 rRNAs and 728 tRNAs. We investigated <i>CYCLOIDEA </i>(<i>CYC</i>)<i> </i>genes, which have a role in determination of floral symmetry, and confirm the presence of two copies in the genome. Read depth analysis of 180 'Duplicated' BUSCO genes using a new tool, DepthKopy (<a href="https://github.com/slimsuite/depthkopy">https://github.com/slimsuite/depthkopy</a>), suggests almost all are real duplications, increasing confidence in the annotation and highlighting a possible need to revise the BUSCO set for this lineage. The chromosome-level <i>T. speciosissima</i> reference genome (Tspe_v1) provides an important new genomic resource of Proteaceae to support the conservation of flora in Australia and further afield.</p>
De novo assembly of 20 chicken genomes reveals the undetectable phenomenon for thousands of core genes on micro-chromosomes and sub-telomeric regions
<p>The gene numbers and evolutionary rates of birds were assumed to be much lower than those of mammals, which is in sharp contrast to the huge species number and morphological diversity of birds. It is therefore necessary to construct a complete avian genome and analyze its evolution. We constructed a chicken pan-genome from 20 <em>de novo</em> assembled genomes with high sequencing depth, and identified 1,335 protein-coding genes and 3,011 long noncoding RNAs not found in GRCg6a. The majority of these novel genes were detected across most individuals of the examined transcriptomes but were seldomly measured in each of the DNA sequencing data regardless of Illumina or PacBio technology. Furthermore, different from previous pan-genome models, most of these novel genes were overrepresented on chromosomal sub-telomeric regions and micro-chromosomes, surrounded by extremely high proportions of tandem repeats, which strongly blocks DNA sequencing. These hidden genes were proved to be shared by all chicken genomes, included many housekeeping genes, and enriched in immune pathways. Comparative genomics revealed the novel genes had three-fold elevated substitution rates than known ones, updating the knowledge about evolutionary rates in birds. Our study provides a framework for constructing a better chicken genome, which will contribute towards the understanding of avian evolution and improvement of poultry breeding.</p>
De novo genome assembly of Kallima inachus
<p><span>Oakleaf butterflies in the genus <em>Kallima</em> have a polymorphic wing phenotype, enabling these insects to masquerade as dead leaves. By studying mechanisms that shape the genetic and species diversity of these butterflies, a new perspective can be provided to understand the evolutionary innovation driven by geographic changes and natural selection.</span></p> <p><span>We found that leaf wing polymorphism in <em>Kallima</em> butterflies is controlled by the wing patterning gene cortex. We hypothesized that multiple mechanisms may independently lead to the reduction or suppression of recombination among different cortex haplotypes. To test this hypothesis, w</span>e performed Nanopore re-sequencing and <em>de novo</em> genome assembly for 4 <em>Kallima inachus</em> individuals and obtained 4 individual genomes. We identified two chromosomal inversions spanning these haplotypes.</p>
De novo assembly of chimpanzee ONT reads
<p><em>De novo</em> assembly of a western chimpanzee genome (Pan troglodytes verus) from LCL cell line (AG18359), generated with Oxford Nanopore Technologies (ONT) reads basecalled with Guppy v5.0.11 and assembled with Shasta software v0.7.0 for Linux. Nanopore raw reads were originally published in PMID: 32143403.</p>
De novo assembly of SNPs in VCF format for 112 individualss of Campylorhynchus in western Ecuador
<p>Climate variability has a significant impact on the evolution of biodiversity, which results in genetic and phenotypic diversity within species. A balance between gene flow and selection maintains changes in the frequency of genetic and phenotypic variants that occur along an environmental gradient. Here, we investigate a hybrid zone in western Ecuador involving C. zonatus, C. fasciatus, and admixed populations. We hypothesized that different ecological preferences and geographical distances result in limited dispersal between populations along the precipitation gradient in western Ecuador.</p> <p>In the context of testing IBE and IBD shaping distributions of C. zonatus, C. fasciatus, and potential hybrids, we asked (1) Is there evidence of genetic admixture and introgression between these taxa in Western Ecuador? And (2) What is the relative contribution of IBE and IBD on patterns of genetic differentiation and admixture patterns? We analyzed 4409 SNPs from the blood of 112 individuals sequenced using ddRadSeq. The most likely clusters ranged from K=2-4, corresponding to categories defined by geographic origins, known phylogenetics, and physical or ecological constraints. Evidence for IBE was weak but stronger for IBD. We observed gradual changes in genetic admixture between C. f. pallescens and C. zonatus along the environmental gradient. Genetic differentiation of the two populations of C. f. pallescens could be driven by a previously undescribed potential physical barrier near the center of western Ecuador. Lowland habitats in this region may be limited due to the proximity of the Andes to the coastline, limiting dispersal and gene flow, particularly among dry-habitat specialists.</p>
Data from: RAD sequencing, genotyping error estimation and de novo assembly optimization for population genetic inference
Restriction site-associated DNA sequencing (RADseq) provides researchers with the ability to record genetic polymorphism across thousands of loci for non-model organisms, potentially revolutionising the field of molecular ecology. However, as with other genotyping methods, RADseq is prone to a number of sources of error that may have consequential effects for population genetic inferences, and these have received only limited attention in terms of the estimation and reporting of genotyping error rates. Here we use individual sample replicates, under the expectation of identical genotypes, to quantify genotyping error in the absence of a reference genome. We then use sample replicates to (1) optimize de novo assembly parameters within the program Stacks, by minimizing error and maximizing the retrieval of informative loci, and; (2) quantify error rates for loci, alleles and SNPs. As an empirical example we use a double digest RAD dataset of a non-model plant species, Berberis alpina, collected from high altitude mountains in Mexico.
Error, noise and bias in de novo transcriptome assemblies
<p><i>De novo</i> transcriptome assembly is a powerful tool, widely used over the last decade for making evolutionary inferences. However, it relies on two implicit assumptions: that the assembled transcriptome is an unbiased representation of the underlying expressed transcriptome, and that expression estimates from the assembly are good, if noisy approximations of the relative abundance of expressed transcripts. Using publicly available data for model organisms, we demonstrate that, across assembly algorithms and data sets, these assumptions are consistently violated. Bias exists at the nucleotide level, with genotyping error rates ranging from 30-83%. As a result, diversity is underestimated in transcriptome assemblies, with consistent under-estimation of heterozygosity in all but the most inbred samples. Even at the gene level, expression estimates show wide deviations from map-to-reference estimates, and positive bias at lower expression levels. Standard filtering of transcriptome assemblies improves the robustness of gene expression estimates but leads to the loss of a meaningful number of protein-coding genes, including many that are highly expressed. We demonstrate a computational method, length-rescaled CPM, to partly alleviate noise and bias in expression estimates. Researchers should consider ways to minimize the impact of bias in transcriptome assemblies.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.