Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
46
datasets available to search
ShareScore release 0.9.0
Dataset results
46 results for “sequencing quality”
MACREL software benchmark data set: Simulated metagenomes with sequencing quality, errors profile and abundance distributions derived from real samples
<p>These metagenomes were used in the benchmarking of FACS pipeline, and were designed after NGLess benchmark dataset (doi.org/10.5281/zenodo.2560288). Metagenomes were simulated with <a href="https://www.niehs.nih.gov/research/resources/software/biostatistics/art/index.cfm">ART-bin-MountRainier-2016.06.05</a> using real abundance profiles (.abund files) available <a href="https://doi.org/10.5281/zenodo.2560288">elsewhere</a>, and <a href="http://progenomes1.embl.de/data/repGenomes/representatives.contigs.fasta.gz">proGenomes' representative contigs</a> as reference genomes. There are available metagenomes with 40, 60 and 80 M (million of reads) based in the reference genomes and abundances of the following samples:</p> <pre><code>SAMEA2466916 SAMEA2466953 SAMEA2466965 SAMEA2621107 SAMEA2621229 SAMEA2621247</code></pre> <p>To convert them from the CRAM format back to fastq files:</p> <pre><code> ## 1. converting from cram to bam format: samtools view -b -T refgenome.fa -o file.bam file.cram ## 2. sorting the bam file: samtools sort -n file.bam -o input_sorted.bam # sort reads by identifier-name (-n) ## 3. converting from bam to fastq format: bedtools bamtofastq -i input_sorted.bam -fq output_r1.fastq -fq2 output_r2.fastq </code></pre> <p> </p>
L'Aquila 2009 seismic sequence: integrated dataset of automatic first motion polarities focal mechanisms and RMT with HypoDD high quality relative earthquake locations
<p>This dataset is related to the L'Aquila 2009 seismic sequence that happened in Central Apennines (Italy).</p> <p>It contains:</p> <ul> <li>2782 quality selected focal mechanisms produced with the standard software FPFIT based on automatically determined first motion polarities of automatically detected and analyzed foreshocks and aftershocks recorded from January 2009 to December 2009 (flag <strong>fty</strong> in the header is MP)</li> <li>475 (out of 627) quality selected focal mechanisms produced with the standard software FPFIT also based on automatically determined first motion polarities but for only 3204 M<sub>L</sub> >= 1.9 earthquakes and by using take-off angles calculated within a local 3d tomographic velocity model (<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2011GL047365">Di Stefano et al., 2011</a>) , published and released in <a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2011JB008352">Chiaraluce et al., 2011</a> (flag <strong>fty</strong> in the header is JG)</li> <li>165 (out of 181) Regional Moment Tensors determined for earthquakes M<sub>L</sub> >= 3.0 based on broadband waveform inversion of ground velocities and published by <a href="https://pubs.geoscienceworld.org/ssa/bssa/article-abstract/101/3/975/349796/Regional-Moment-Tensors-of-the-2009-L-Aquila">Hermann et al., 2011</a> (flag <strong>fty</strong> in the header is HM)</li> <li>The hypocenters of the total 3422 earthquakes reported in the present focal solutions dataset have been taken from the very high quality double difference locations of the about 64000 aftershocks reported in <a href="https://agupubs.onlinelibrary.wiley.com/doi/pdf/10.1002/jgrb.50130">Valoroso et al., 2013</a> and published, as part of the full dataset, <a href="https://doi.org/10.5281/zenodo.4036248">on Zenodo</a>. </li> </ul> <p>The association between the focal solutions and the HypoDD hypocenters has been performed through the direct use of the HypoDD event identifier where possible (the whole MP dataset) and through spatial and temporal earthquakes coordinates matching in all the other case by using the capability of a MySQL database. </p> <p>Two files are uploaded, one in plain text with blank separator, the second in plain text with ";" separator and .csv extension.</p> <p>Here below the header is explained.</p> <p><strong>OT_Date:</strong> date of the origin time in the format YYYY-MM-DD</p> <p><strong>OT_Time:</strong> time of the origin time in the format HH:mm:ss.dcm</p> <p><strong>lat:</strong> hypocenter latitude expressed in degrees </p> <p><strong>lon:</strong> hypocenter longitude east of Greenwich, expressed in degrees</p> <p><strong>dep:</strong> hypocenter depth expressed in km </p> <p><strong>ML:</strong> local magnitude (pure number) from <a href="https://agupubs.onlinelibrary.wiley.com/doi/pdf/10.1002/jgrb.50130">Valoroso et al., 2013</a> (see last column notes also)</p> <p> </p> <p><strong>id_dd:</strong> the hypoDD event identifier, allowing to directly connect to the <a href="https://agupubs.onlinelibrary.wiley.com/doi/pdf/10.1002/jgrb.50130">Valoroso et al., 2013</a> full dataset</p> <p><strong>IMPORTANT NOTE about st1 and st2 (below): </strong>the focal solutions are presented here based on the convention they where produced or published, so there are two different (but compatible) conventions for the fault plains orientation in the 3d space</p> <p><strong>st1:</strong></p> <ul> <li><strong>for fty=</strong>HM or JG this is the strike of plane 1 (CMT convention)</li> <li><strong>for fty=</strong>MP this is the <strong>strike of the dip direction </strong>of plain 1 (FPFIT convention)</li> </ul> <p><strong>dip1: </strong>dip of plane 1</p> <p><strong>rk1: </strong>rake of plane 1</p> <p><strong>st2:</strong></p> <ul> <li><strong>for fty=</strong>HM or JG this is the strike of plane 2 (CMT convention)</li> <li><strong>for fty=</strong>MP this is the <strong>strike of the dip direction </strong>of plain 2 (FPFIT convention)</li> </ul> <p><strong>dip2: </strong>dip of plane 2</p> <p><strong>rk2: </strong>rake of plane 2</p> <p><strong>fty:</strong> flag to distinguish the type of solution, CMT=HM or JG, FPFIT=MP</p> <p><strong>MW:</strong> only for HM, this columns reports also MW from <a href="https://pubs.geoscienceworld.org/ssa/bssa/article-abstract/101/3/975/349796/Regional-Moment-Tensors-of-the-2009-L-Aquila">Hermann et al., 2011</a></p>
High-quality large curated dataset of protein sequences (1.83 million) and their corresponding Position Specific Scoring Matrices
<p>As part of his master thesis at the Rostlab, which is located at the Technical University of Munich (TUM), Mr. Issar Arab developed the first language model that encodes evolutionary information of proteins explicitly. The pre-training involved the creation of a novel high-quality dataset of protein sequences (around 1.83 million proteins, or ~0.8 Billion amino acids) with their corresponding Position Specific Scoring Matrices (PSSMs). Those matrices reflect the relative frequency of each amino acid at each position in a protein and is derived from evolutionarily related proteins.</p> <p>Mr. Arab makes this work publicly available to help other researchers speed up their work to leverage AI to learn the representation of protein evolutionary information more explicitly. The set of sequences was derived by extracting all PSSMs from the <a href="https://predictprotein.org/">PredictProtein</a> (PP) cache, which were also part o the UniProt Reference Cluster with 50% sequence identity (uniref50 2019_12). The overlap between PP and uniref50 was further filtered to only include high-quality samples, e.g. only multiple sequence alignments with a certain number of aligned sequences were considered. The processing led to a training set of 1.83 Million sequences, a validation set of 879 instances, and a test set of 879 entries. The training data of proteins is reduced to 40% sequence identity, with respect to the validation/test sets, and contains sequences ranging between 18 and 9858 residues in length.</p> <p>Refer to the Jupyter notebook for a detailed description of the files' structure and a Python code snippet to correctly manipulate this data.</p> <p>To access the full original work, please visit the following link: <a href="https://mediatum.ub.tum.de/node?id=1579236">Manuscript</a> <br><br><strong>Note:</strong> The dataset was recently used to fine tune a protein sequence language model (<a href="https://github.com/issararab/PEvoLM">PEvoLM</a>). The work was presented at the CIBCB'23 conference. If you use PEvoLM or this dataset in your work, please cite the following publication:</p> <p>- Issar Arab, <strong>PEvoLM: Protein Sequence Evolutionary Information Language Model</strong>, <em>IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), Eindhoven, Netherlands</em>, (2023), pp. 1-8, doi:<a href="https://ieeexplore.ieee.org/document/10264890">10.1109/CIBCB56990.2023.10264890</a></p>
Literature consistency of bioinformatics sequence databases is effective for assessing record quality
<p>Bioinformatics sequence databases such as Genbank or UniProt contain hundreds of millions of records of genomic data. These records are derived from direct submissions from individual laboratories, as well as from bulk submissions from large-scale sequencing centres; their diversity and scale means that they suffer from a range of data quality issues including errors, discrepancies, redundancies, ambiguities, incompleteness and inconsistencies with the published literature. In this work, we seek to investigate and analyze the data quality of sequence databases from the perspective of a curator, who must detect anomalous and suspicious records. Specifically, we emphasize the detection of inconsistent records with respect to the literature. Focusing on GenBank, we propose a set of 24 quality indicators, which are based on treating a record as a query into the published literature, and then use query quality predictors. We then carry out an analysis that shows that the proposed quality indicators and the quality of the records have a mutual relationship, in which one depends on the other. We propose to represent record literature consistency as a vector of these quality indicators. By reducing the dimensionality of this representation for visualization purposes using principal component analysis, we show that records which have been reported as inconsistent with the literature fall roughly in the same area, and therefore share similar characteristics. By manually analyzing records not previously known to be erroneous that fall in the same area than records know to be inconsistent, we show that one record out of four is inconsistent with respect to the literature. This high density of inconsistent record opens the way towards the development of automatic methods for the detection of faulty records. We conclude that literature inconsistency is a meaningful strategy for identifying suspicious records.</p>
Data for: Beyond signal quality: The value of unmaintained pH, dissolved oxygen, and oxidation-reduction potential sensors for remote performance monitoring of on-site sequencing batch reactors
<p>Sensor maintenance is time-consuming and is a bottleneck for monitoring on-site wastewater treatment systems. Hence, we compare maintained and unmaintained sensors to monitor the biological performance of a small-scale sequencing batch reactor (SBR). The sensor types are ion-selective pH, optical dissolved oxygen (DO), and oxidation-reduction potential (ORP) with platinum electrode. We created soft sensors using engineered features: ammonium valley for pH, oxidation ramp for DO, and nitrite ramp for the ORP. Four soft sensors based on unmaintained pH sensors correctly identified the completion of the ammonium oxidation (89 to 91 out of 107 cycles), about as many times as soft sensors based on a maintained pH sensor (91 out of 107 cycles). In contrast, the DO soft sensor using data from a maintained sensor showed slightly better (89 out of 96 cycles) detection performance than that using data from two unmaintained sensors (77, respectively 82 out of 96 correct). Furthermore, the DO soft sensor using maintained data is much less sensitive to the optimisation of cut-off frequency and slope tolerance than the soft sensor using unmaintained data. The nitrite ramp provided no useful information on the state of nitrite oxidation, so no comparison of maintained and unmaintained ORP sensors was possible in this case. We identified two hurdles when designing soft sensors for unmaintained sensors: i) Sensors' type- and design-specific deterioration affects performance. ii) Feature engineering for soft sensors is sensor type specific, and the outcome is strongly influenced by operational parameters such as the aeration rate. In summary, the results with the provided soft sensors show that frequent sensor maintenance is not necessarily needed to monitor the performance of SBRs. Without sensor maintenance monitoring smalls-scale SBRs becomes practicable, which could improve the reliability of unstaffed on-site treatment systems substantially.</p>
List of known SNP positions (based on SNP chip data) for base quality score recalibration of alignments for whole-genome resequencing and whole-genome bisulfite sequencing data from great tits (Parus major)
<p>The profiling of epigenetic marks like DNA methylation has become a central aspect of studies in evolution and ecology. Bisulfite sequencing is commonly used for assessing genome-wide DNA methylation at single nucleotide resolution but these data can also provide information on genetic variants like single nucleotide polymorphisms (SNPs). However, bisulfite conversion causes unmethylated cytosines to appear as thymines, complicating the alignment and subsequent SNP calling. Several tools have been developed to overcome this challenge, but there is no independent evaluation of such tools for non-model species, which often lack genomic references. Here, we used whole-genome bisulfite sequencing (WGBS) data from four female great tits (<i>Parus major</i>) to evaluate the performance of seven tools for SNP calling from bisulfite sequencing data. We used SNPs from whole-genome resequencing data of the same samples as baseline SNPs to assess common performance metrics like sensitivity, precision, and the number of true positive, false positive, and false negative SNPs for the full range of variant and genotype quality values. We found clear differences between the tools in either optimizing precision (Bis-SNP), sensitivity (biscuit), or a compromise between both (all other tools). Overall, the choice of SNP caller strongly depends on which performance parameter should be maximized and whether ascertainment bias should be minimized to optimize downstream analysis, highlighting the need for studies that assess such differences.</p>
179 high quality metagenome-assembled genomes sequences and annotations
<p>We analyzed seven sediment samples collected adjacent to ferromanganese nodules from the Clarion–Clipperton Fracture Zone (CCFZ) in the eastern Pacific Ocean. Through deep metagenomic sequencing, assembly, and binning, we reconstructed 179 high quality metagenome-assembled genomes (MAGs). This archive contains these genomes sequences and annotations. </p>
Isolation, biochemical characterization, and genome sequencing of two high-quality genomes of a novel chitinolytic Jeongeupia species
<p>Raw data sets for the microbiologyOpen research article "Isolation, biochemical characterization, and genome sequencing of two high-quality genomes of a novel chitinolytic Jeongeupia species" including ClustalW Trees and the respective input file; canB 3.0 results; TYGS phylogenetic tree results; PGAP genome annotation files and Canu 2.0 assembly reports of the two genomes.</p>
Dataset for: mRNA vaccine quality analysis using RNA sequencing
<p>The success of mRNA vaccines has been realised, in part, by advances in manufacturing that enabled billions of doses to be produced at sufficient quality and safety. However, mRNA vaccines must be rigorously analysed to measure their integrity and detect contaminants that reduce their effectiveness and induce side-effects. Currently, mRNA vaccines and therapies are analysed using a range of time-consuming and costly methods. Here we describe a streamlined method to analyse mRNA vaccines and therapies using long-read nanopore sequencing. Compared to other industry-standard techniques, VAX-seq can comprehensively measure key mRNA vaccine quality attributes, including sequence, length, integrity, and purity. We also show how direct RNA sequencing can analyse mRNA chemistry, including the detection of nucleoside modifications. To support this approach, we provide supporting software to automatically report on mRNA and plasmid template quality and integrity. Given these advantages, we anticipate that RNA sequencing methods, such as VAX-seq, will become central to the development and manufacture of mRNA drugs.</p>
16S rRNA sequencing data: Altered microbiota, impaired quality of life, malabsorption, infection, and inflammation in CVID patients with diarrhoea
Open the record for dataset details and reuse information.
Dataset for: mRNA vaccine quality analysis using RNA sequencing
Open the record for dataset details and reuse information.
List of known SNP positions (based on SNP chip data) for base quality score recalibration of alignments for whole-genome resequencing and whole-genome bisulfite sequencing data from great tits (Parus major)
Open the record for dataset details and reuse information.
Data from: Finding the right coverage: The impact of coverage and sequence quality on SNP genotyping error rates
Restriction-enzyme-based sequencing methods enable the genotyping of thousands of single nucleotide polymorphism (SNP) loci in non-model organisms. However, in contrast to traditional genetic markers, genotyping error rates in SNPs derived from restriction-enzyme-based methods remain largely unknown. Here, we estimated genotyping error rates in SNPs genotyped with double digest RAD sequencing from Mendelian incompatibilities in known mother-offspring dyads of Hoffman's two-toed sloth (Choloepus hoffmanni) across a range of coverage and sequence quality criteria, for both reference-aligned and de novo-assembled datasets. Genotyping error rates were more sensitive to coverage than sequence quality and low coverage yielded high error rates, particularly in de novo-assembled datasets. For example, coverage ≥5 yielded median genotyping error rates of ≥0.03 and ≥0.11 in reference-aligned- and de novo-assembled datasets, respectively. Genotyping error rates declined to ≤0.01 in reference-aligned datasets with a coverage >30, but remained >0.04 in the de novo-assembled datasets. We observed approximately 10- and 13-fold declines in the number of loci sampled in the reference-aligned and de novo-assembled datasets when coverage was increased from >5 to >30 at quality score ≥30, respectively. Finally, we assessed the effects of genotyping coverage on a common population genetic application, parentage assignments, and showed that the proportion of incorrectly assigned maternities was relatively high at low coverage. Overall, our results suggest that the tradeoff between sample size and genotyping error rates be considered prior to building sequencing libraries, reporting genotyping error rates become standard practice, and that effects of genotyping errors on inference be evaluated in restriction-enzyme-based SNP studies.
Community Medical Center, Continuous Quality Improvement Project, Rapid Sequence Intubation
ClinicalTrials.gov study NCT05505799. IPD Sharing: NO. Countries: 1. Publications: 8.
Data from: Restriction site-associated DNA sequencing generates high-quality single nucleotide polymorphisms for assessing hybridization between bighead and silver carp in the United States and China
Open the record for dataset details and reuse information.
Data from: Finding the right coverage: The impact of coverage and sequence quality on SNP genotyping error rates
Open the record for dataset details and reuse information.
Data from: High quality whole genome sequence of an abundant Holarctic odontocete, the harbour porpoise (Phocoena phocoena)
Open the record for dataset details and reuse information.
Data from: PCR-Free enrichment of mitochondrial DNA from human blood and cell lines for high quality next-generation DNA sequencing
Recent advances in sequencing technology allow for accurate detection of mitochondrial sequence variants, even those in low abundance at heteroplasmic sites. Considerable sequencing cost savings can be achieved by enriching samples for mitochondrial (relative to nuclear) DNA. Reduction in nuclear DNA (nDNA) content can also help to avoid false positive variants resulting from nuclear mitochondrial sequences (numts). We isolate intact mitochondrial organelles from both human cell lines and blood components using two separate methods: a magnetic bead binding protocol and differential centrifugation. DNA is extracted and further enriched for mitochondrial DNA (mtDNA) by an enzyme digest. Only 1 ng of the purified DNA is necessary for library preparation and next generation sequence (NGS) analysis. Enrichment methods are assessed and compared using mtDNA (versus nDNA) content as a metric, measured by using real-time quantitative PCR and NGS read analysis. Among the various strategies examined, the optimal is differential centrifugation isolation followed by exonuclease digest. This strategy yields >35% mtDNA reads in blood and cell lines, which corresponds to hundreds-fold enrichment over baseline. The strategy also avoids false variant calls that, as we show, can be induced by the long-range PCR approaches that are the current standard in enrichment procedures. This optimization procedure allows mtDNA enrichment for efficient and accurate massively parallel sequencing, enabling NGS from samples with small amounts of starting material. This will decrease costs by increasing the number of samples that may be multiplexed, ultimately facilitating efforts to better understand mitochondria-related diseases.
Data from: Diversity measures in environmental sequences are highly dependent on alignment quality—data from ITS and new LSU primers targeting basidiomycetes
The ribosomal DNA comprised of the ITS1-5.8S-ITS2 regions is widely used as a fungal marker in molecular ecology and systematics but cannot be aligned with confidence across genetically distant taxa. In order to study the diversity of Agaricomycotina in forest soils, we designed primers targeting the more alignable 28S (LSU) gene, which should be more useful for phylogenetic analyses of the detected taxa. This paper compares the performance of the established ITS1F/4B primer pair, which targets basidiomycetes, to that of two new pairs. Key factors in the comparison were the diversity covered, off-target amplification, rarefaction at different Operational Taxonomic Unit (OTU) cutoff levels, sensitivity of the method used to process the alignment to missing data and insecure positional homology, and the congruence of monophyletic clades with OTU assignments and BLAST-derived OTU names. The ITS primer pair yielded no off-target amplification but also exhibited the least fidelity to the expected phylogenetic groups. The LSU primers give complementary pictures of diversity, but were more sensitive to modifications of the alignment such as the removal of difficult-to align stretches. The LSU primers also yielded greater numbers of singletons but also had a greater tendency to produce OTUs containing sequences from a wider variety of species as judged by BLAST similarity. We introduced some new parameters to describe alignment heterogeneity based on Shannon entropy and the extent and contents of the OTUs in a phylogenetic tree space. Our results suggest that ITS should not be used when calculating phylogenetic trees from genetically distant sequences obtained from environmental DNA extractions and that it is inadvisable to define OTUs on the basis of very heterogeneous alignments.
Supplementary material 1 from: Nilsson RH, Sánchez-García M, Ryberg M, Abarenkov K, Wurzbacher C, Kristiansson E (2017) Read quality-based trimming of the distal ends of public fungal DNA sequences is nowhere near satisfactory. MycoKeys 26: 13-24. https://doi.org/10.3897/mycokeys.26.14591
Details on the fungal genomes/contigs targeted : Data type: Excel spreadsheet
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.