Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
181
datasets available to search
ShareScore release 0.9.0
Dataset results
181 results for “de novo assembly”
Data from: Two low coverage bird genomes and a comparison of reference-guided versus de novo genome assemblies
Open the record for dataset details and reuse information.
Data from: Likelihood-based inference of population history from low coverage de novo genome assemblies
Open the record for dataset details and reuse information.
Data from: De novo assembly of a tadpole shrimp (Triops newberryi) transcriptome and preliminary differential gene expression analysis
Open the record for dataset details and reuse information.
Data from: The de novo genome assembly and annotation of a female domestic dromedary of North African origin
Open the record for dataset details and reuse information.
Data from: Performance of gene expression analyses using de novo assembled transcripts in polyploid species
Motivation: Quality of gene expression analyses using de novo assembled transcripts in species that experienced recent polyploidization remains unexplored. Results: Differential gene expression (DGE) analyses using putative genes inferred by Trinity, Corset and Grouper performed slightly differently across five plant species that experienced various poly-ploidy histories. In species that lack recent polyploidy events that occurred in the past several millions of years, DGE analyses using de novo assembled transcriptomes identified 54–82% of the differen-tially expressed genes recovered by mapping reads to the reference genes. However, in species that experienced more recent polyploidy events, the percentage decreased to 21–65%. Gene co-expression network analyses using de novo assemblies vs. mapping to the reference genes recov-ered the same module that significantly correlated with treatment in one species that lacks recent polyploidization.
Data from: Evaluation of the impact of RNA preservation methods of spiders for de novo transcriptome assembly
With advances in high-throughput sequencing technologies, de novo transcriptome sequencing and assembly has become a cost-effective method to obtain comprehensive genetic information of a species of interest, especially in non-model species with large genomes such as spiders. However, high-quality RNA is essential for successful sequencing and sample preservation conditions require careful consideration for the effective storage of field-collected samples. To this end, we report a streamlined feasibility study of various storage conditions and their effects on de novo transcriptome assembly results. The storage parameters considered include temperatures ranging from room temperature to -80°C; preservatives, including ethanol, RNAlater, TRIzol, and RNAlater-ICE; and sample submersion states. As a result, intact RNA was extracted and assembly was successful when samples were preserved at low temperatures regardless of the type of preservative used. The assemblies as well as the gene expression profiles were shown to be robust to RNA degradation, when 30 million 150 bp paired-end reads are obtained. The parameters for sample storage, RNA extraction, library preparation, sequencing, and in silico assembly considered in this work provide a guideline for the study of field-collected samples of spiders.
Data from: De novo assembly and comparative analysis of the Ceratodon purpureus transcriptome
The bryophytes are a morphologically and ecologically diverse group of plants that have recently emerged as major model systems for a variety of biological processes. In particular, the genome sequence of the moss, Physcomitrella patens, has significantly enhanced our understanding of the evolution of developmental processes in land plants. However, to fully explore the diversity within bryophytes, we need additional genomic resources. Here we describe analyses of the transcriptomes of a male and a female isolate of the moss, C. purpureus, generated using the 454 FLX technology. Comparative analyses between C. purpureus and P. patens indicated that this strategy generated nearly complete coverage of the protonemal transcriptome. An analysis of the overlap in gene sets between C. purpureus and P. patens provides new insights into the evolution of gene family composition across the land plants. In spite of the overall transcriptomic similarity between the two species, Ka/Ks analysis of P. patens and C. purpureus suggest considerable physiological and developmental divergence. Additionally, while the codon usage was very similar between these two mosses, C. purpureus genes showed a slightly greater codon usage bias than P. patens genes potentially because of the contrasting mating system of the two species. Finally, we found evidence of a genome doubling ~65-76 MYA that likely coincided with the contemporaneous polyploidy event inferred for P. patens but postdates the divergence of P. patens and C. purpureus. The powerful laboratory tools now available for C. purpureus will enable the research community to fully exploit these genomic resources.
Data from: De novo transcriptome assemblies of four accessions of the metal hyperaccumulator plant Noccaea caerulescens
Noccaea caerulescens of the Brassicaceae family has become the key model plant among the metal hyperaccumulator plants. Populations/accessions of N. caerulescens from geographic locations with different soil metal concentrations differ in their ability to hyperaccumulate and hypertolerate metals. Comparison of transcriptomes in several accessions provides candidates for detailed exploration of the mechanisms of metal accumulation and tolerance and local adaptation. This can have implications in the development of plants for phytoremediation and improved mineral nutrition. Transcriptomes from root and shoot tissues of four N. caerulescens accessions with contrasting Zn, Cd and Ni hyperaccumulation and tolerance traits were sequenced with Illumina Hiseq2000. Transcriptomes were assembled using the Trinity de novo assembler and were annotated and the protein sequences predicted. The comparison against the BUSCO plant early release dataset indicated high-quality assemblies.The predicted protein sequences have been clustered into ortholog groups with closely related species. The data serve as important reference sequences in whole transcriptome studies, in analyses of genetic differences between the accessions and other species, and for primer design.
Data from: De novo genome assembly and annotation of rice sheath rot fungus Sarocladium oryzae reveals genes involved in Helvolic acid and Cerulenin biosynthesis pathways
Background: Sheath rot disease caused by Sarocladium oryzae is an emerging threat for rice cultivation at global level. However, limited information with respect to genomic resources and pathogenesis is a major setback to develop disease management strategies. Considering this fact, we sequenced the whole genome of highly virulent Sarocladium oryzae field isolate, Saro-13 with 82x sequence depth. Results: The genome size of S. oryzae was 32.78 Mb with contig N50 18.07 Kb and 10526 protein coding genes. The functional annotation of protein coding genes revealed that S. oryzae genome has evolved with many expanded gene families of major super family, proteinases, zinc finger proteins, sugar transporters, dehydrogenases/reductases, cytochrome P450, WD domain G-beta repeat and FAD-binding proteins. Gene orthology analysis showed that around 79.80 % of S. oryzae genes were orthologous to other Ascomycetes fungi. The polyketide synthase dehydratase, ATP-binding cassette (ABC) transporters, amine oxidases, and aldehyde dehydrogenase family proteins were duplicated in larger proportion specifying the adaptive gene duplications to varying environmental conditions. Thirty-nine secondary metabolite gene clusters encoded for polyketide synthases, nonribosomal peptide synthase, and terpene cyclases. Protein homology based analysis indicated that nine putative candidate genes were found to be involved in helvolic acid biosynthesis pathway. The genes were arranged in cluster and structural organization of gene cluster was similar to helvolic acid biosynthesis cluster in Metarhizium anisophilae. Around 9.37 % of S. oryzae genes were identified as pathogenicity genes, which are experimentally proven in other phytopathogenic fungi and enlisted in pathogen-host interaction database. In addition, we also report 13212 simple sequences repeats (SSRs) which can be deployed in pathogen identification and population dynamic studies in near future. Conclusions: Large set of pathogenicity determinants and putative genes involved in helvolic acid and cerulenin biosynthesis will have broader implications with respect to Sarocladium disease biology. This is the first genome sequencing report globally and the genomic resources developed from this study will have wider impact worldwide to understand Rice-Sarocladium interaction.
Data from: De novo genome assembly of Geosmithia morbida, the causal agent of thousand cankers disease
Geosmithia morbida is a filamentous ascomycete that causes thousand cankers disease in the eastern black walnut tree. This pathogen is commonly found in the western U.S.; however, recently the disease was also detected in several eastern states where the black walnut lumber industry is concentrated. G. morbida is one of two known phytopathogens within the genus Geosmithia, and it is vectored into the host tree via the walnut twig beetle. We present the first de novo draft genome of G. morbida. It is 26.5 Mbp in length and contains less than 1% repetitive elements. The genome possesses an estimated 6,273 genes, 277 of which are predicted to encode proteins with unknown functions. Approximately 31.5% of the proteins in G. morbida are homologous to proteins involved in pathogenicity, and 5.6% of the proteins contain signal peptides that indicate these proteins are secreted. Several studies have investigated the evolution of pathogenicity in pathogens of agricultural crops; forest fungal pathogens are often neglected because research efforts are focused on food crops. G. morbida is one of the few tree phytopathogens to be sequenced, assembled and annotated. The first draft genome of G. morbida serves as a valuable tool for comprehending the underlying molecular and evolutionary mechanisms behind pathogenesis within the Geosmithia genus.
Data from: De novo assembly and characterization of the Hucho taimen transcriptome
Taimen (Hucho taimen) is an important ecological and economic species that is classified as vulnerable by the IUCN Red List of Threatened Species; however, limited genomic information is available on this species. RNA-Seq is a useful tool for obtaining genetic information and developing genetic markers for non-model species in addition to its application in gene expression profiling. In this study, we performed a comprehensive RNA-Seq analysis of taimen. We obtained 157 M clean reads (14.7 Gb) and used them to de novo assemble a high-quality transcriptome with a N50 size of 1060 bp. In the assembly, 82% of the transcripts were annotated using several databases, and 14,666 of the transcripts contained a full open reading frame. The assembly covered 75% of the transcripts of Atlantic salmon and 57.3% of the protein-coding genes of rainbow trout. To learn about the genome evolution, we performed a systematic comparative analysis across 11 teleosts including 8 salmonids, and found 313 unique gene families in taimen. Using Atlantic salmon and rainbow trout transcriptomes as the background, we identified 250 positive selection transcripts. The pathway enrichment analysis revealed a unique characteristic of taimen: it possesses more immune-related genes than Atlantic salmon and rainbow trout; moreover, some genes have undergone strong positive selection. We also developed a pipeline for identifying microsatellite marker genotypes in samples, and successfully identified 24 polymorphic microsatellite markers for taimen. These data and tools are useful for studying conservation genetics, phylogenetics, evolution among salmonids and selective breeding for threatened taimen.
Data from: De novo assembly of a chromosome-level reference genome of red spotted grouper (Epinephelus akaara) using nanopore sequencing and Hi-C
The red spotted grouper Epinephelus akaara (E. akaara) is one of the most economically important marine fish in China, Japan and Southeast Asia, and is a threatened species. The species is also considered a good model for studies of sex-inversion, development, genetic diversity and immunity. Despite its importance, molecular resources for E. akaara remain limited and no reference genome has been published to date. In this study, we constructed a chromosome-level reference genome of E. akaara by taking advantage of long-read single molecule sequencing and de novo assembly by Oxford Nanopore Technologies (ONT) and Hi-C. A red-spotted grouper genome of 1.135 Gb was assembled from a total of 106.29 Gb polished Nanopore sequence (GridION, ONT), equivalent to 96-fold genome coverage. The assembled genome represents 96.8% completeness (BUSCO) with a contig N50 length of 5.25 Mb and a longest contig of 25.75 Mb. The contigs were clustered and ordered onto 24 pseudo-chromosomes covering approximately 95.55% of the genome assembly with Hi-C data, with a scaffold N50 length of 46.03 Mb. The genome contained 43.02% repeat sequences and 5,480 non-coding RNAs. Furthermore, after mining several RNA-seq datasets, 23,809 (99.5%) genes were functionally annotated from a total of 23,924 predicted protein-coding sequences. The high-quality chromosome-level reference genome of E. akaara was assembled for the first time and will be a valuable resource for molecular breeding and functional genomics studies of red-spotted grouper in the future.
Data from: De novo transcriptome assembly databases in the butterfly orchid Phalaenopsis equestris
Orchids are renowned for their spectacular flowers and ecological adaptations. After the sequencing of the genome of the tropical epiphytic orchid Phalaenopsis equestris, we combined Illumina HiSeq2000 for RNA-Seq and Trinity for de novo assembly to characterize the transcriptomes for 11 diverse P. equestris tissues representing the root, stem, leaf, flower buds, column, lip, petal, sepal and three developmental stages of seeds. Our aims were to contribute to a better understanding of the molecular mechanisms driving the analysed tissue characteristics and to enrich the available data for P. equestris. Here, we present three databases. The first dataset is the RNA-Seq raw reads, which can be used to execute new experiments with different analysis approaches. The other two datasets allow different types of searches for candidate homologues. The second dataset includes the sets of assembled unigenes and predicted coding sequences and proteins, enabling a sequence-based search. The third dataset consists of the annotation results of the aligned unigenes versus the Nonredundant (Nr) protein database, Kyoto Encyclopaedia of Genes and Genomes (KEGG) and Clusters of Orthologous Groups (COG) databases with low e-values, enabling a name-based search.
Exploring the venom gland transcriptome of Bothrops asper and Bothrops jararaca: de novo assembly and analysis of novel toxic proteins
Open the record for dataset details and reuse information.
Assemblies generated in the paper "phasebook: haplotype-aware de novo assembly of diploid genomes from long reads"
<p>Assemblies generated in the paper "phasebook: haplotype-aware de novo assembly of diploid genomes from long reads"</p>
Yak k-mer dumps for partition human chrX/Y in de novo assemblies
<p>See <a href="https://github.com/lh3/yak">yak</a> for details.</p>
A chromosome-scale de novo genome assembly of the dwarf tomato variety Micro-Tom
<p>The cultivated tomato (<em>Solanum lycopersicum</em>) is an important crop and model species for genetics and plant molecular biology research. The dwarf tomato variety Micro-Tom is used extensively in research because it is rapid flowering, easy to grow in high volumes in minimal space, and is amenable to genetic transformation. Here we provide a de novo chromosome-scale genome assembly of Micro-Tom that was generated using PacBio HiFi reads and scaffolded using chromosome confirmation capture data. The HiFi data was assembled using the Hifiasm assembler and OmniC data was used for scaffolding using Salsa and several rounds of manual curation and validation.</p>
Data from: De novo sequencing, assembly, and annotation of four threespine stickleback genomes based on microfluidic partitioned DNA libraries
Open the record for dataset details and reuse information.
Data from: De novo transcriptome assemblies of four accessions of the metal hyperaccumulator plant Noccaea caerulescens
Open the record for dataset details and reuse information.
Data from: De novo assembly and characterization of four anthozoan (phylum Cnidaria) transcriptomes
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.