Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
76
datasets available to search
ShareScore release 0.9.0
Dataset results
76 results for “de novo sequencing”
De novo sequencing of phages T4 and T7
<p>Raw data and assemblies of the <em>de novo </em>sequencing of the "model" phages T4 and T7, performed with Illumina NextSeq 2x150 and Oxford Nanopore, generated for the "Phage Annotation Workshop" held online on November 2021:</p> <p>https://github.com/quadram-institute-bioscience/phage-annotation-workshop/</p>
Squeegee: de novo identification of reagent and laboratory induced microbial contaminants in low biomass microbiomes, simulation dataset 0.25% spike-in contaminant sequences
<p>Computational analysis of host-associated microbiomes has opened the door to numerous discoveries relevant to human health and disease. However, contaminant sequences in metagenomic samples can potentially impact the interpretation of findings reported in microbiome studies, especially in low biomass environments. Our hypothesis is that contamination from DNA extraction kits or sampling lab environments will leave taxonomic "bread crumbs” across multiple distinct sample types, allowing for the detection of microbial contaminants when negative controls are unavailable. To test this hypothesis we implemented Squeegee, a de novo contamination detection tool. We tested Squeegee on simulated and real low biomass metagenomic datasets. On the low biomass samples, we compared Squeegee predictions to experimental negative control data and show that Squeegee accurately recovers known contaminants. We also analyzed 749 metagenomic datasets from the Human Microbiome Project and identified likely previously unreported kit contamination. Collectively, our results highlight that Squeegee can identify microbial contaminants with high precision.</p> <p> </p> <p>Simulation Dataset 0.25% contaminant spike-in.</p>
Squeegee: de novo identification of reagent and laboratory induced microbial contaminants in low biomass microbiomes, simulation dataset 1% spike-in contaminant sequences
<p>Computational analysis of host-associated microbiomes has opened the door to numerous discoveries relevant to human health and disease. However, contaminant sequences in metagenomic samples can potentially impact the interpretation of findings reported in microbiome studies, especially in low biomass environments. Our hypothesis is that contamination from DNA extraction kits or sampling lab environments will leave taxonomic "bread crumbs” across multiple distinct sample types, allowing for the detection of microbial contaminants when negative controls are unavailable. To test this hypothesis we implemented Squeegee, a de novo contamination detection tool. We tested Squeegee on simulated and real low biomass metagenomic datasets. On the low biomass samples, we compared Squeegee predictions to experimental negative control data and show that Squeegee accurately recovers known contaminants. We also analyzed 749 metagenomic datasets from the Human Microbiome Project and identified likely previously unreported kit contamination. Collectively, our results highlight that Squeegee can identify microbial contaminants with high precision.</p> <p> </p> <p>Simulation Dataset 1% contaminant spike-in.</p>
De novo nanopore sequencing overrepresents RNA modification landscape
<p>RNA modifications are critical to the functional diversity and regulatory complexity of the transcriptome. With increasing frequency, direct nanopore RNA sequencing is applied to identify RNA modifications de novo. Here, we directly compare the MS2 phage genome RNA modification profiles determined using nanopore to orthogonal LC-MS/MS assays. The results reveal very different views of the modification landscape, suggesting caution when calling new RNA modifications using nanopore alone.</p>
Data from: Affordable de novo generation of fish mitogenomes using amplification-free enrichment of mitochondrial DNA and deep sequencing of long fragments
<p>Biomonitoring surveys from environmental DNA make use of metabarcoding tools to describe the community composition. These studies match their sequencing results against public genomic databases to identify the species. However, mitochondrial genomic reference data are yet incomplete, only a few genes may be available, or the suitability of existing sequence data is suboptimal for species-level resolution. Here we present a dedicated and cost-effective workflow with no DNA amplification for generating complete fish mitogenomes for the purpose of strengthening fish mitochondrial databases. Two different long-fragment sequencing approaches using Oxford Nanopore sequencing coupled with mitochondrial DNA enrichment were used. One where the enrichment is achieved by preferential isolation of mitochondria followed by DNA extraction and nuclear DNA depletion ('mitoenrichment'). A second enrichment approach takes advantage of the CRISPR-Cas9 targeted scission on previously dephosphorylated DNA ('targeted mitosequencing'). The sequencing results varied between tissue, species, and integrity of the DNA. The mitoenrichment method yielded 0.17-12.33 % of sequences on target and a mean coverage ranging from 74.9 to 805-fold. The targeted mitosequencing experiment from native genomic DNA yielded 1.83-55 % of sequences on target and a 38 to 2123-fold mean coverage. This produced complete the mitogenome of species with homopolymeric regions, tandem repeats, and gene rearrangements. We demonstrate that deep sequencing of long fragments of native fish DNA is possible and can be achieved with low computational resources in a cost-effective manner, opening the discovery of mitogenomes of non-model or understudied fish taxa to a broad range of laboratories worldwide.</p>
Deep learning-driven fragment ion series classification enables highly precise and sensitive de novo peptide sequencing
<p>This Zenodo record contains the dataset and model weights for "Deep learning-driven fragment ion series classification enables highly precise and sensitive de novo peptide sequencing".</p> <p> </p> <p>This repository contains the following files:</p> <ul> <li> <p>For the human dataset by Wang et al.:</p> <ul> <li> <p>train_val_test_split.csv containing the mapping of the correct peptide by MaxQuant to either train, validation or test set</p> </li> <li> <p>psms_train_val_test.csv containing the mapping of correct PSMs (scan number, raw file and correct peptide by MaxQuant) to either train, validation or test set</p> </li> <li> <p>updated_spectralis_test_out.csv as before containing Spectralis-EA predictions and scores on test set, as well as initial peptides and scores by Casanovo and Novor and now containing also correct peptides by MaxQuant and Spectralis-scores on the combination of Casanovo and Novor sequences (column named spectralis_score_onlyRescoring)</p> </li> <li> <p>spectralis_test_out_heart_analysis.csv subset of 20220822_spectralis_test_out.csv containing only PSMs for the tissue heart with the computation of precision and recall values</p> </li> <li> <p>spectralis_test_out_pointnovo_deepnovo.csv containing predictions by DeepNovo and PointNovo with original scores and Spectralis-score, as well as correct peptides by MaxQuant</p> </li> </ul> </li> </ul> <p> </p> <ul> <li> <p>For the nine-species dataset by Tran et al.:</p> <ul> <li> <p>spectralis_ninespecies_out.csv containing spectrum identifiers, correct peptides by PEAKSDB, predicted peptides by the different de novo sequencing tools as well as original scores and Spectralis-scores for the different PSMs.</p> </li> </ul> </li> </ul>
Data for 'NovoRank: Refinement for De Novo Peptide Sequencing Based on Spectral Clustering and Deep Learning'
<p>The uploaded files include two archives for <a href="https://pubs.acs.org/doi/10.1021/acs.jproteome.4c00300" target="_blank" rel="noopener">NovoRank: Refinement for De Novo Peptide Sequencing Based on Spectral Clustering and Deep Learning</a>. The '<em>mgf_data</em>' archive contains all MGF files used in the study, while the '<em>sample_data</em>' archive includes sequencing data, clustering data generated using <code>MSCluster</code>, and XCorr calculation data computed with <code>CometX</code>, all of which were used in the research.</p>
Data from: RAD sequencing, genotyping error estimation and de novo assembly optimization for population genetic inference
Restriction site-associated DNA sequencing (RADseq) provides researchers with the ability to record genetic polymorphism across thousands of loci for non-model organisms, potentially revolutionising the field of molecular ecology. However, as with other genotyping methods, RADseq is prone to a number of sources of error that may have consequential effects for population genetic inferences, and these have received only limited attention in terms of the estimation and reporting of genotyping error rates. Here we use individual sample replicates, under the expectation of identical genotypes, to quantify genotyping error in the absence of a reference genome. We then use sample replicates to (1) optimize de novo assembly parameters within the program Stacks, by minimizing error and maximizing the retrieval of informative loci, and; (2) quantify error rates for loci, alleles and SNPs. As an empirical example we use a double digest RAD dataset of a non-model plant species, Berberis alpina, collected from high altitude mountains in Mexico.
Accuracy of de novo assembly of DNA sequences from double‐digest libraries varies substantially among software
Open the record for dataset details and reuse information.
Data from: RAD sequencing, genotyping error estimation and de novo assembly optimization for population genetic inference
Open the record for dataset details and reuse information.
Data from: Affordable de novo generation of fish mitogenomes using amplification-free enrichment of mitochondrial DNA and deep sequencing of long fragments
Open the record for dataset details and reuse information.
CRAG: De novo characterization of cell-free DNA fragmentation hotspots in plasma whole-genome sequencing
<p><strong>Supplementary dataset for the manuscript:</strong></p> <p><strong>CRAG: De novo characterization of cell-free DNA fragmentation hotspots in plasma whole-genome sequencing</strong></p> <p> Xionghui Zhou1,*, Haizi Zheng1,*, Hailu Fu1,*, Kelsey L. Dillehay McKillip2-3, Susan M. Pinney2,4, Yaping Liu1-2,5-7 #</p> <p>Affiliations:</p> <p>1 Division of Human Genetics, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH 45229</p> <p>2 University of Cincinnati Cancer Center, Cincinnati, OH 45229</p> <p>3 Department of Pathology & Laboratory Medicine, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>4 Department of Environmental and Public Health Sciences, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>5 Division of Biomedical Informatics, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH 45229</p> <p>6 Department of Pediatrics, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>7 Department of Electrical Engineering and Computing Sciences, University of Cincinnati College of Engineering and Applied Science, Cincinnati, OH 45229</p> <p>* These authors contributed equally</p> <p># Email: lyping1986@gmail.com</p>
additional file 1 and 2 for De novo transcriptome sequencing of Serangium japonicum (Coleoptera: Coccinellidae)
<p>file1:some comman statastic results of transcriptome sequences of 6 S.japonicum samples. file2:GO classed different expression genes.</p>
additional file 3 for De novo transcriptome sequencing of Serangium japonicum (Coleoptera: Coccinellidae)
<p>The different expression unigenes between in summer and in winter. TPM: transcripts per million reads. pValue: statistic test p value. qValue: corrected p value after multiple tests. </p>
Data from: De novo sequencing and assembly of Azadirachta indica fruit transcriptome
Azadirachta indica (neem) is a unique, versatile and important tree species. Many parts of the plant are traditionally used as pesticide, insecticide, fungicide and for other medicinal purposes. Azadirachta fruits and seeds, a good source of oil, are widely used for agriculturally important pest management. Neem oil and its derivatives also support multiple cottage industries in India. Past efforts have been mostly concentrated towards identifying, characterizing and synthesizing one of its principal components, i.e. azadirachtin from seed kernels. Despite diverse use of the neem plant, a modern drug-development programme which systematically exploits the therapeutic ability of Azadirachta fruits remains to be fully established. Next generation sequencing technology that helps decode genomes and transcriptomes has transformational impact on medicine, agriculture, bio-fuel and biodiversity studies. Here, we report sequencing, assembly and analysis of Azadirachta fruit transcriptome using next-generation sequencing technology. We believe that our study shall offer valuable insights towards realizing the larger vision of understanding the key medicinally active compounds and their pathways.
De novo nanopore sequencing overrepresents RNA modification landscape, part 2
<p>RNA modifications are critical to the functional diversity and regulatory complexity of the transcriptome. With increasing frequency, direct nanopore RNA sequencing is applied to identify RNA modifications de novo. Here, we directly compare the MS2 phage genome RNA modification profiles determined using nanopore to orthogonal LC-MS/MS assays. The results reveal very different views of the modification landscape, suggesting caution when calling new RNA modifications using nanopore alone.</p>
De novo nanopore sequencing overrepresents RNA modification landscape, part 3
<p>RNA modifications are critical to the functional diversity and regulatory complexity of the transcriptome. With increasing frequency, direct nanopore RNA sequencing is applied to identify RNA modifications de novo. Here, we directly compare the MS2 phage genome RNA modification profiles determined using nanopore to orthogonal LC-MS/MS assays. The results reveal very different views of the modification landscape, suggesting caution when calling new RNA modifications using nanopore alone.</p>
De novo reconstruction of satellite repeat units from sequence data
<p>Results generated for preprint 'De novo reconstruction of satellite repeat units from sequence data".</p>
Data from: Rapid microsatellite isolation from a butterfly by de novo transcriptome sequencing: performance and a comparison with AFLP-derived distances
Open the record for dataset details and reuse information.
Data from: Microsatellites for the marsh Fritillary butterfly: de novo transcriptome sequencing, and a comparison with amplified fragment length polymorphism (AFLP) markers
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.