Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
76
datasets available to search
ShareScore release 0.9.0
Dataset results
76 results for “de novo sequencing”
Data from: De novo sequencing and assembly of Azadirachta indica fruit transcriptome
Open the record for dataset details and reuse information.
Data from: Restriction site associated DNA (RAD) for de novo sequencing and marker discovery in sugarcane borer, Diatraea saccharalis Fab. (Lepidoptera: Crambidae)
Open the record for dataset details and reuse information.
Data from: Dealing with the adaptive immune system during de novo evolution of genes from intergenic sequences
Open the record for dataset details and reuse information.
Data from: De novo sequencing and variant calling with nanopores using PoreSeq
The accuracy of sequencing single DNA molecules with nanopores is continually improving, but de novo genome sequencing and assembly using only nanopore data remain challenging. Here we describe PoreSeq, an algorithm that identifies and corrects errors in nanopore sequencing data and improves the accuracy of de novo genome assembly with increasing coverage depth. The approach relies on modeling the possible sources of uncertainty that occur as DNA transits through the nanopore and finds the sequence that best explains multiple reads of the same region. PoreSeq increases nanopore sequencing read accuracy of M13 bacteriophage DNA from 85% to 99% at 100× coverage. We also use the algorithm to assemble Escherichia coli with 30× coverage and the λ genome at a range of coverages from 3× to 50×. Additionally, we classify sequence variants at an order of magnitude lower coverage than is possible with existing methods.
Data from: De novo transcriptome characterization and development of genomic tools for Scabiosa columbaria L. using next-generation sequencing techniques.
Next-generation sequencing (NGS) technologies are increasingly applied in many organisms, including non-model organisms that are important for ecological and conservation purposes. Illumina and 454 sequencing are among the most used NGS technologies and have been shown to produce optimal results at reasonable costs when used together. Here, we describe the combined application of these two NGS technologies to characterize the transcriptome of a plant species of ecological and conservation relevance for which no genomic resource is available, Scabiosa columbaria. We obtained 528,557 reads from a 454 GS-FLX run and a total of 28,993,627 reads from two lanes of an Illumina GAII single run. After reads trimming, the de novo assembly of both types of reads produced 109,630 contigs. Both the contigs and the >75 bp remaining singletons were blasted against Uniprot/Swissprot database, resulting in 29,676 and 10,515 significant hits, respectively. Based on sequence similarity with known gene products, these sequences represent at least 12,516 unique genes, most of which are well covered by contig sequences. In addition, we identified 4,320 microsatellite loci, of which 856 had flanking sequences suitable for PCR primer design. We also identified 75,054 putative SNPs. This annotated sequence collection and the relative molecular markers represent a main genomic resource for S. columbaria which should contribute to future research in conservation and population biology studies. Our results demonstrate the utility of NGS technologies as starting point for the development of genomic tools in nonmodel but ecologically important species.
Data from: Using Illumina Next Generation Sequencing Technologies to sequence multigene families in de novo species
The advent of Next Generation Sequencing Technology (NGST) has revolutionized molecular biology research, allowing for rapid gene/genome sequencing from a multitude of diverse species. As high throughput sequencing becomes more accessible, more efficient workflows must be developed to deal with the amounts of data produced and better assemble the genomes of de novo lineages. We combine traditional laboratory methods with Illumina NGST to amplify and sequence the largest mammalian multigene family, the Olfactory Receptor gene family, for species with and without a reference genome. We develop novel assembly methods to annotate and filter these data, which can be utilized for any gene family or any species. We find no significant difference between the ratio of genes within their respective gene families of our data compared with available genomic data. Using simulated data we explore the limitations of short-read sequence data and our assembly in recovering this gene family. We highlight the benefits and shortcomings of these methods. Compared with data generated from traditional polymerase chain reaction, cloning and Sanger sequencing methodologies, sequence data generated using our pipeline increases yield and sequencing efficiency without reducing the number of unique genes amplified. A cloning step is not required, therefore shortening data generation time. The novel downstream methodologies and workflows described provide a tool to be utilized by many fields of biology, to access and analyze the vast quantities of data generated. By combining laboratory and in silico methods, we provide a means of extracting genomic information for multigene families without complete genome sequencing.
Data from: De novo assembly of a chromosome-level reference genome of red spotted grouper (Epinephelus akaara) using nanopore sequencing and Hi-C
The red spotted grouper Epinephelus akaara (E. akaara) is one of the most economically important marine fish in China, Japan and Southeast Asia, and is a threatened species. The species is also considered a good model for studies of sex-inversion, development, genetic diversity and immunity. Despite its importance, molecular resources for E. akaara remain limited and no reference genome has been published to date. In this study, we constructed a chromosome-level reference genome of E. akaara by taking advantage of long-read single molecule sequencing and de novo assembly by Oxford Nanopore Technologies (ONT) and Hi-C. A red-spotted grouper genome of 1.135 Gb was assembled from a total of 106.29 Gb polished Nanopore sequence (GridION, ONT), equivalent to 96-fold genome coverage. The assembled genome represents 96.8% completeness (BUSCO) with a contig N50 length of 5.25 Mb and a longest contig of 25.75 Mb. The contigs were clustered and ordered onto 24 pseudo-chromosomes covering approximately 95.55% of the genome assembly with Hi-C data, with a scaffold N50 length of 46.03 Mb. The genome contained 43.02% repeat sequences and 5,480 non-coding RNAs. Furthermore, after mining several RNA-seq datasets, 23,809 (99.5%) genes were functionally annotated from a total of 23,924 predicted protein-coding sequences. The high-quality chromosome-level reference genome of E. akaara was assembled for the first time and will be a valuable resource for molecular breeding and functional genomics studies of red-spotted grouper in the future.
Squeegee: de novo identification of reagent and laboratory induced microbial contaminants in low biomass microbiomes, simulation dataset 0.5% spike-in contaminant sequences
<p>Computational analysis of host-associated microbiomes has opened the door to numerous discoveries relevant to human health and disease. However, contaminant sequences in metagenomic samples can potentially impact the interpretation of findings reported in microbiome studies, especially in low biomass environments. Our hypothesis is that contamination from DNA extraction kits or sampling lab environments will leave taxonomic "bread crumbs” across multiple distinct sample types, allowing for the detection of microbial contaminants when negative controls are unavailable. To test this hypothesis we implemented Squeegee, a de novo contamination detection tool. We tested Squeegee on simulated and real low biomass metagenomic datasets. On the low biomass samples, we compared Squeegee predictions to experimental negative control data and show that Squeegee accurately recovers known contaminants. We also analyzed 749 metagenomic datasets from the Human Microbiome Project and identified likely previously unreported kit contamination. Collectively, our results highlight that Squeegee can identify microbial contaminants with high precision.</p> <p> </p> <p>Simulation Dataset 0.5% contaminant spike-in. </p>
Traning set for "Accurate De Novo Peptide Sequencing Using Fully Convolutional Neural Networks"
Open the record for dataset details and reuse information.
Data from: De novo sequencing, assembly, and annotation of four threespine stickleback genomes based on microfluidic partitioned DNA libraries
Open the record for dataset details and reuse information.
Data from: Using Illumina Next Generation Sequencing Technologies to sequence multigene families in de novo species
Open the record for dataset details and reuse information.
Data from: De novo assembly of a chromosome-level reference genome of red spotted grouper (Epinephelus akaara) using nanopore sequencing and Hi-C
Open the record for dataset details and reuse information.
Data from: De novo transcriptome characterization and development of genomic tools for Scabiosa columbaria L. using next-generation sequencing techniques.
Open the record for dataset details and reuse information.
Data from: De novo sequencing and variant calling with nanopores using PoreSeq
Open the record for dataset details and reuse information.
Data from: De novo assembly of genomes from long sequence reads reveals uncharted territories of Propionibacterium freudenreichii
Open the record for dataset details and reuse information.
DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers [Human oligo UMI-STARR-seq]
GEO Series GSE183938. Homo sapiens; synthetic construct. 4 samples. Type: Other.
RNA Sequencing with de novo assembly of earthworm transcriptome by the treatment of lindane with different exposure and aging duration.
GEO Series GSE118678. Eisenia fetida. 15 samples. Type: Expression profiling by high throughput sequencing.
De novo RNA-sequencing and transcriptome analysis of Aconitum carmichaelii to analyze key genes involved in the biosynthesis of diterpene alkaloids
GEO Series GSE106247. Aconitum carmichaelii. 4 samples. Type: Expression profiling by high throughput sequencing.
DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers [Drosophila genome-wide UMI-STARR-seq]
GEO Series GSE183936. Drosophila melanogaster; synthetic construct. 6 samples. Type: Other.
Deep sequencing and de novo assembly of the mouse oocyte transcriptome define the contribution of transcription to the DNA methylation landscape.
GEO Series GSE70116. Mus musculus. 4 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.