Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
181
datasets available to search
ShareScore release 0.9.0
Dataset results
181 results for “de novo assembly”
De novo genome assembly for Eulemur rufifrons
<p>As one of the most threatened mammalian taxa, lemurs of Madagascar are facing unprecedented anthropogenic pressures. To address conservation imperatives such as this, researchers have increasingly relied on conservation genomics to identify populations of particular concern. However, many of these genomic approaches necessitate high-quality genomes. While the advent of next generation sequencing technologies and the resulting reduction of associated costs have led to the proliferation of genomic data and high-quality reference genomes, global discrepancies in genomic sequencing capabilities often result in biological samples from biodiverse host countries being exported to facilities in the Global North, creating inequalities in access and training within genomic research. Here, we present the first reference genome for the endangered red-fronted brown lemur (Eulemur rufifrons) from sequencing efforts conducted entirely within the host country using portable Oxford Nanopore sequencing. Using an archived E. rufifrons specimen, we conducted long-read, nanopore sequencing at the Centre ValBio Research Station near Ranomafana National Park, in rural Madagascar, generating over 750 Gb of sequencing data from 10 MinION flow cells. Exclusively using this long-read data, we assembled 2.215 gigabase, 20,330-contig assembly with an N50 of 98.9 Mb and a 17,108 bp mitogenome. The nuclear assembly had 31x average coverage and was comparable in completeness to other primate reference genomes, with a 95.51% BUSCO completeness score for primate-specific genes. As the first reference genome for E. rufifrons and the only annotated genome available for the speciose Eulemur genus, this resource will prove vital for conservation genomic studies while our efforts exhibit the potential of this protocol to address research inequalities and build genomic capacity. </p>
De novo assemblies for the manuscrip "Candida albicans isolates contain frequent heterozygous structural variants and transposable elements within genes and centromeres"
Open the record for dataset details and reuse information.
Assemblies generated in the manuscript "Geometric deep learning framework for de novo genome assembly"
<p>Assemblies evaluated in the manuscript "Geometric deep learning framework for de novo genome assembly". All the assemblies were generated by us, except CHM13.ONT.Flye-2.9.fa.gz which was generated by <a href="https://www.nature.com/articles/s41587-019-0072-8">Kolmogorov et al. (2019)</a>.</p>
Data from: De novo assembly of a tadpole shrimp (Triops newberryi) transcriptome and preliminary differential gene expression analysis
Next-generation sequencing techniques, such as RNA sequencing, have provided a wealth of genomic information for nonmodel species. Transcriptomic information can be used to quantify the patterns of gene expression, which can identify how environmental differences invoke organismal stress responses and provide a gauge in predicting species adaptability. In our study, we used RNA sequencing to characterize the first transcriptome from a naupliar tadpole shrimp (Triops newberryi) to identify the genes expressed during the early life history stages and which could be important for future genomic studies. RNA was extracted from naupliar T. newberryi that were reared in a laboratory-controlled setting and in two different water types, a native and a non-native condition. A total of six replicates, three per condition, were sequenced with the Illumina Hi-Seq 2000 achieving 365 M 50-nt reads. High-quality reads were produced and de novo assembly was used to construct a T. newberryi transcriptome that was approximately 24.8 M base pairs. More than 10 000 peptides were predicted from the assembly, and genes were sorted into gene ontology categories. The use of different water conditions allowed for a preliminary differential gene expression analysis in order to compare the changes in gene expression between conditions. There were 299 differentially expressed genes between water conditions that might serve as a focal point for future genomic studies of Triops acclimation to different environments. The Triops transcriptome could serve as vital genomic information for additional studies on Branchiopod crustaceans.
Data from: De novo sequencing and assembly of Azadirachta indica fruit transcriptome
Azadirachta indica (neem) is a unique, versatile and important tree species. Many parts of the plant are traditionally used as pesticide, insecticide, fungicide and for other medicinal purposes. Azadirachta fruits and seeds, a good source of oil, are widely used for agriculturally important pest management. Neem oil and its derivatives also support multiple cottage industries in India. Past efforts have been mostly concentrated towards identifying, characterizing and synthesizing one of its principal components, i.e. azadirachtin from seed kernels. Despite diverse use of the neem plant, a modern drug-development programme which systematically exploits the therapeutic ability of Azadirachta fruits remains to be fully established. Next generation sequencing technology that helps decode genomes and transcriptomes has transformational impact on medicine, agriculture, bio-fuel and biodiversity studies. Here, we report sequencing, assembly and analysis of Azadirachta fruit transcriptome using next-generation sequencing technology. We believe that our study shall offer valuable insights towards realizing the larger vision of understanding the key medicinally active compounds and their pathways.
Data associated with the publication "Assembling membraneless organelles from de novo designed proteins"
<p>Raw data used in the manuscript "Assembling membraneless organelles from de novo designed proteins"</p>
Quast outputs for "When do longer reads matter? A benchmark of long read de novo assembly tools for eukaryotic genomes"
<p>Quast outputs for "When do longer reads matter? A benchmark of long read de novo assembly tools for eukaryotic genomes"</p>
Data from: De novo and reference transcriptome assembly of transcripts expressed during flowering provide insight into seed setting in tetraploid red clover
Open the record for dataset details and reuse information.
Data from: "De novo transcriptome assembly of the mountain fly Drosophila nigrosparsa using short RNA-seq reads" in Genomic Resources Notes Accepted 1 August 2014-30 September 2014
Open the record for dataset details and reuse information.
Data from: "De novo transcriptome assembly and polymorphism detection in ecological important widely distributed Neotropical toads from the Rhinella marina species complex (Anura: Bufonidade)" in Genomic Resources Notes Accepted 1 August 2014-30 September 2014
Open the record for dataset details and reuse information.
Data from: De novo sequencing and assembly of Azadirachta indica fruit transcriptome
Open the record for dataset details and reuse information.
Data from: "De novo assembly transcriptome for the rostrum dace (Leuciscus burdigalensis, Cyprinidae: fish) naturally infected by a copepod ectoparasite" in Genomic Resources Notes accepted 1 December 2014 to 31 January 2015
Open the record for dataset details and reuse information.
Data from: De novo assembly and characterization of leaf and floral transcriptomes of the hybridizing bromeliad species (Pitcairnia spp.) adapted to Neotropical Inselbergs
Open the record for dataset details and reuse information.
Data from: "De novo assembled transcriptome of organs involved in reproduction in an endangered endemic Iberian cyprinid fish (Squalius pyrenaicus)" in Genomic Resources Notes Accepted 1 June 2015 to 31 July 2015
Open the record for dataset details and reuse information.
Data from: De novo assembly of the transcriptome of an invasive snail and its multiple ecological applications
Open the record for dataset details and reuse information.
Data from: De novo transcriptome assembly for the lobster Homarus americanus and characterization of differential gene expression across nervous system tissues
Open the record for dataset details and reuse information.
De novo genome assembly of Leptodactylus fuscus
Open the record for dataset details and reuse information.
De novo genome assembly for Eulemur rufifrons
Open the record for dataset details and reuse information.
De novo genome assembly of Tectona grandis (Teak) with 2993 scaffolds
Open the record for dataset details and reuse information.
Data from: De novo assembly and characterization of the skeletal muscle and electric organ transcriptomes of the African weakly-electric fish Campylomormyrus compressirostris (Mormyridae, Teleostei)
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.