Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
253
datasets available to search
ShareScore release 0.9.0
Dataset results
253 results for “amplicons”
Rapid and real-time identification of fungi up to the species level with long amplicon Nanopore sequencing from clinical samples
<p>Samples collected from fungal cultures, skin of dogs and ZymoBIOMICS<sup>TM </sup>mock community (which includes <em>Saccharomyces cerevisiae</em> and <em>Cryptococcus neoformans</em>). The amplicons length of the fungal cultures and ZymoBIOMICS<sup>TM </sup>mock community is 3,5 Kb and 6 Kb, while the <em>Malassezia spp</em> samples used as control is 3,5 Kb. The amplicons length of the four samples from the skin is 3,5 Kb.</p>
Amplicon sequence variants by sample table from Antarctic methane seeps
<p>Antarctica is estimated to contain as much as a quarter of earth's marine methane, however we have not discovered an active Antarctic methane seep limiting our understanding of the methane cycle. In 2011, an expansive (70m x 1m) microbial mat formed at 10m water depth in the Ross Sea, Antarctica and we carried out 16S rRNA gene analysis on samples collected one year and five years after the methane seep formed. The data set attached is the resulting Amplicon Sequence Variant table by sample that we used to track the community composition change during this time and in comparison to other sampling in the McMurdo sound region. </p>
Consensus calling of MinION amplicon reads improves metabarcoding results
<p><strong>Background</strong>: Metabarcoding environmental DNA with high-throughput sequencing is a state-of-the-art method to assess biodiversity and to uncover dark taxa. MinION is the first handheld sequencer that can be taken into the field for on-site metabarcoding. This research aims to answer if bioinformatics can solve the issues that arise because of the higher error rate of MinION data.</p> <p><strong>Results</strong>: Biodiverse samples with a presumed large portion of dark taxa were selected from the Dutch Caribbean. The cytochrome oxidase 1 gene (CO1) is used as a barcode for identification at the species level or higher levels. Generating a consensus sequence from closely related sequences resulted in minimized random errors and increased species identification of 175% compared to unclustered MinION data. Additional to the formulation of the workflow, an ecological analysis was conducted that revealed co-occurrence of species in similar habitats, and that the proportion of dark taxa in the sampled region is 81.87%.</p> <p><strong>Conclusion</strong>: Although the workflow did not attain the results that the existing Illumina workflows do, the potential is evident. The high proportion of dark taxa in the sampled region of Statia and the Saba Bank indicates the need for continued barcoding of species in the Dutch Caribbean to resolve database limitations.</p>
Data from: Algorithm for post-clustering curation of DNA amplicon data yields reliable biodiversity estimates
DNA metabarcoding is promising for cost-effective biodiversity monitoring, but reliable diversity estimates are difficult to achieve and validate. Here we present and validate a method, called LULU, for removing erroneous molecular operational taxonomic units (OTUs) from community data derived by high-throughput sequencing of amplified marker genes. LULU identifies errors by combining sequence similarity and co-occurrence patterns. To validate the LULU method, we use a unique data set of high quality survey data of vascular plants paired with plant ITS2 metabarcoding data of DNA extracted from soil from 130 sites in Denmark spanning major environmental gradients. OTU tables are produced with several different OTU definition algorithms and subsequently curated with LULU, and validated against field survey data. LULU curation consistently improves α-diversity estimates and other biodiversity metrics, and does not require a sequence reference database; thus, it represents a promising method for reliable biodiversity estimation.
Non-Perennial rivers and streams under hydrological stress – comparing the effectiveness of amplicon sequencing and digital microscopy for diatom biodiversity appraisal
<p>ABC is project leader, data collector, contact person, etc.</p>
A pan-cetacean MHC amplicon sequencing panel developed and evaluated in combination with genome assemblies
<p>The major histocompatibility complex (MHC) is a highly polymorphic gene family that is crucial in immunity, and its diversity can be effectively used as a fitness marker for populations. Despite this, MHC remains poorly characterised in non-model species (e.g., cetaceans: whales, dolphins and porpoises) as high gene copy number variation, especially in the fast-evolving class I region, makes analyses of genomic sequences difficult. To date, only small sections of class I and IIa genes have been used to assess functional diversity in cetacean populations. Here, we undertook a systematic characterisation of the MHC class I and IIa regions in available cetacean genomes. We extracted full-length gene sequences to design pan-cetacean primers that amplified the complete exon2 from MHC class I and IIa genes in one combined sequencing panel. We validated this panel in 19 cetacean species and described 354 alleles for both classes. Furthermore, we identified likely assembly artefacts for many MHC class I assemblies based on the presence of class I genes in the amplicon data compared to missing genes from genomes. Finally, we investigated MHC diversity using the panel in 25 humpback and 30 southern right whales, including four paternity trios for humpback whales. This revealed copy-number variable class I haplotypes in humpback whales, which is likely a common phenomenon across cetaceans. These MHC alleles will form the basis for a cetacean branch of the Immuno-Polymorphism Database (IPD-MHC), a curated resource intended to aid in the systematic compilation of MHC alleles across several species, to support conservation initiatives.</p>
(Extended Data) Amplicon deep sequencing of ama1 and mdr1 to track within-host P. falciparum diversity throughout treatment in a clinical drug trial
<p>These extended data accompany the manuscript: Targeted Amplicon deep sequencing of ama1 and mdr1 to track within-host <em>P. falciparum</em> diversity throughout treatment in a clinical drug trial</p> <p><strong>Table S1: Concentration ratios and resulting parasitemia in artificial dna mixtures of P. falciparum Lab Isolates 3D7 and Dd2.</strong> This table presents the parasitemia for the artificial mixtures of P. falciparum lab isolates 3D7 and Dd2. Each mixture was prepared at varying ratios of 3D7 to Dd2, starting from equal proportions to a complete presence of only 3D7. The original concentration of each isolate was approximately 50,000 parasites per microliter (pf/μl), and the table displays the proportion of each strain in the mixture and the resulting total parasitemia concentration.</p> <p><strong>Table S2. List of PCR and deep sequencing primers.</strong> This table shows the list of forward and reverse primers used for deep sequencing. In boldface are the MID tags, while in the regular face are the forward primers</p> <p><strong>Table S3. The relative frequencies of each ama1 variant and the number of samples with each variant.</strong> The relative frequencies (%) of the 33 AMA1 variants in pre-and post-treatment samples (n = 330) are shown as a 33 amino acid sequence. The frequencies were calculated by dividing the number of reads of each microhaplotype by the total number of reads obtained per sample (116,187,131).</p> <p><strong>Table S4. Distribution of microhaplotypes among samples.</strong> This table shows the occurrence of microhaplotypes across all participants, both with monoclonal and multiclonal ama1 infections. It presents the ama1 clonality – monoclonal or multiclonal (column 1) - participant IDs (column 2), microhaplotype IDs (column 3), and the relative frequencies of these microhaplotypes across timepoints from 0 to 1008 hours (day 42) (column 3). Dashes represent time points where microhaplotypes were missing or were not detected.</p> <p><strong>Table S5. Distribution of rare microhaplotypes among samples.</strong> This table shows the occurrence of rare microhaplotypes in various samples. It presents participant IDs (column 1), microhaplotype IDs (column 2), and the relative frequencies of these microhaplotypes across time points from 0 to 1008 hours (day 42) (column 3). Samples containing rare microhaplotypes - specifically from PID10, PID32, PID38, PID40, PID49, PID60, PID63, and PID65 - are shown in orange, along with the corresponding rare microhaplotypes and their time points of occurrence. Furthermore, participants are categorised by shared microhaplotypes to indicate instances of rarity and commonality. Except for one microhaplotype unique to PID30, rare microhaplotypes were detected in several samples, frequently exceeding a 5% relative frequency. Dashes represent time points where microhaplotypes were missing or were not detected.</p> <p><strong>Table S6. The parasitemia levels associated with each ama1 microhaplotype per timepoint.</strong> This table shows the parasitemia for each ama1 microhaplotype per timepoint and each participant. “Patient ID” represents the patient ID, “AMA1 COI at 0h” represents the complexity of infection (COI) for each participant at baseline, based on ama1 while subsequent columns represent the parasitemia for each ama1 microhaplotype from timepoint 0h to 1008h. Parasitemia was back-calculated using the COI and total parasitemia for each time point. For time points with a COI > 1, parasitemia for the respective ama1 microhaplotypes are separated by commas, cells in red indicate timepoints without sequencing data (ND = not determined). In contrast, cells in grey indicate time points where microhaplotypes were detected below 10 parasites/μl, hence at risk of falling below the sampling limit.</p> <p><strong>Figure S1. Performance of AmpSeq in the sequencing controls.</strong> Six aliquots were prepared for each control set to ensure sufficient control data in case of PCR or sequencing failure. The median read depth in the lab controls was 5,658 (range 4,310 – 12,603) and 704 (291 – 1,676). The x-axis represents the aliquot identifier across the five mixtures, starting from 1 to 6, while the y-axis represents the proportions of each variant across all aliquots. For ama1 (A), two variants (3D7 and Dd2) were detected, whereas in mdr1 (B), two variants were detected YY, FY and NY following amplification of Dd2 Copy I, Dd2 Copy II and 3D7, respectively. For ama1, sequencing failed for aliquot 6 of control set 1, while for mdr1, sequencing failed for aliquot 2 and 6 of control set 3, aliquots 1 and 6 of control set 4 and aliquots 1 and 5 of control set 5. Under the mdr1 control set 4, the Dd2 copy II (86F, 184Y) was not identified, possibly due to having very low concentrations that were not picked up in this aliquot. Based on our control mixtures, the minimum variant frequency we could detect was 5%.</p> <p><strong>Figure S2. Heatmaps of the successfully PCR amplified and sequenced samples for ama1 (A) and mdr1 (B).</strong> The rows represent the study participants, while the columns represent time in hours. Successfully sequenced samples are shown in blue, those that failed PCR are shown in red and those that failed sequencing are in black. The timepoint “ Rec” represents unscheduled visits where a recurrent sample was collected. The unshaded areas with "-" are time points where samples were not collected. For each time point, the number of samples successfully sequenced (n Successful) is indicated in the last row of each panel. The table in panel C shows the groupings of samples based on parasitemia, high (> 5,000), moderate (100-5,000) and low (< 100 parasites per microlitre). Many samples collected between 0h-12h had high parasitemia, samples collected between 18h–30h had moderate parasitemia, while samples collected after 30h were primarily of low parasitemia.</p> <p><strong>Figure S3. The mean complexity of infection (COI) by AMA1 throughout treatment.</strong> The mean COI (red diamonds) appeared to be stable (between 1.5 - 2) from baseline (0h) up to 72h and thereafter fluctuated due to the small sample sizes (<5) in the post-treatment samples. The black dots represent the COI per sample.</p> <p> </p>
eDNA replicates, polymerase and amplicon size impact inference of richness across habitat
<p>Environmental DNA-based monitoring has been increasingly used in the last decade to monitor biodiversity in aquatic and terrestrial systems. Molecular-based surveys now allow quick and reliable production of baseline knowledge of species community composition on a large scale, allowing better understanding of ecosystem function and mitigation of stressors linked to anthropogenic activities. Despite this, technical hurdles often remain, and the impact of replicates, PCR polymerases and amplicon size on the recovered species richness is still poorly understood. Here, we conducted a large controlled experiment, with bulk samples collected from terrestrial, marine and freshwater environments to assess the impact of natural and technical replicates, PCR polymerases with different degrees of fidelity or proofreading activity, as well as amplicon size on species richness recovery across habitats. In this study, we consistently found variations in sample species richness depending on PCR polymerase choice. We further demonstrate the dissimilar impacts between natural and technical replicates on species richness recovery, and the necessity of increasing natural replications in eDNA based surveys. We highlight the benefits and limitations of replication strategies, polymerase choice and amplicon size across terrestrial, marine and freshwater habitats, and provide recommendations to increase the reliability of future eDNA-based metabarcoding studies.</p> <p> </p>
Raw data: multispecies amplicon sequencing (Loera, Studer, and Kölliker, 2021, Molecular Ecology Resources)
<p>Grasslands cover close to two fifths of Earth's land. They provide many ecosystem services related to the maintenance of soil integrity, and the regulation of water, carbon and nitrogen flows. Grasslands constitute the basis for sustainable roughage production for ruminant feeding. In Switzerland, grasslands cover more than 70% of the total agricultural land, which highlights their importance in the domestic food production chains.</p> <p>Plant genetic diversity (PGD), a component of biodiversity, influences ecosystem functioning in grasslands. High levels of grassland PGD are related to resistance against invasive plants and yield stabilization during environmental stress (e.g., drought or frost). The PGD of grasses and legumes —the two most economically relevant plant families found in grasslands, which naturally grow in a wide climate spectrum— harbors valuable genetic resources for forage breeding. Nevertheless, most PGD studies of natural or semi-natural grasslands (i.e., grasslands that are not sown) focus on a single or a few related species. Traditional PGD monitoring methods (e.g., simple sequence repeats, or SSRs) are ill-suited for large-scale, multispecies assessments. This limits our ability to study the ecological effects of grassland PGD, its spatiotemporal patterns, and its significance for grassland management.</p> <p>Looking to provide cost-effective tools for multispecies PGD monitoring in grasslands, we performed a sequence capture assay targeting 611 single-copy nuclear loci, followed by multispecies amplicon sequencing (i.e., amplicon sequencing using primer pairs that can be used in multiple species) on eleven selected loci.</p> <p>Our results indicate that multispecies amplicon sequencing is a cost-effective tool for genetic diversity assessment in grassland plant species. Furthermore, the sequence capture data provides the means to extend the number of multispecies amplicons for further research.</p>
Amplicon structure creates collateral therapeutic vulnerability in cancer
<p>Raw RNA sequencing data expressed in KELLY cells with vs. without ectopic DDX1 expression</p> <p>AH_YB_021 DDX1, AH_YB_023 DDX1, AH_YB_025 DDX1, : DDX1 -</p> <p>AH_YB_022 DDX1, AH_YB_024 DDX1, AH_YB_026 DDX1, : DDX1 +</p>
Shallow shotgun sequencing of the microbiome recapitulates 16S amplicon results and provides functional insights
<p>Prevailing 16S rRNA gene-amplicon methods for characterizing the bacterial microbiome of wildlife are economical, but result in coarse taxonomic classifications, are subject to primer and 16S copy number biases, and do not allow for direct estimation of microbiome functional potential. While deep shotgun metagenomic sequencing can overcome many of these limitations, it is prohibitively expensive for large sample sets. We evaluated the ability of shallow shotgun metagenomic sequencing to characterize taxonomic and functional patterns in the fecal microbiome of a model population of feral horses (Sable Island, Canada). Since 2007, this unmanaged population has been the subject of an individual-based, long-term ecological study. Using deep shotgun metagenomic sequencing, we determined the sequencing depth required to accurately characterize the horse microbiome. In comparing conventional versus high-throughput shotgun metagenomic library preparation techniques, we validate the use of more cost-effective lab methods. Finally, we characterize similarities between 16S amplicon and shallow shotgun characterization of the microbiome and demonstrate that the latter recapitulates biological patterns first described in a published amplicon dataset. Unlike amplicon data, we further demonstrate how shallow shotgun metagenomic data provide useful insights about microbiome functional potential which support previously hypothesized diet effects in this study system.</p>
SNP amplicons results of 50 hybrid Chinook-Coho salmon
<p>These SNP panel results confirm the hybrid origin in 50 Chinook-Coho Salmon individuals. The SNP panel is composed of two amplicons and five diagnostic SNPs. DNA amplicons OkiOts_120255 and Oki_RAD41030 have SNP sites fixed for alternate base pairs in Chinook and Coho salmon (Beacham and Wallace 2019). The panel examined genotypes at one diagnostic position in OkiOts_120255 SNP: 113 (Reference=A, Variant=C) and four positions in Oki_RAD41030: 45 (TC), 51(CG), 195 (GA), and 198 (TG) called via Proton software Variant Caller®. The hybrid salmon were heterozygous for a Chinook and Coho haplotype at both SNP loci, confirming these as the parental species involved in the hybridization and consistent with all being F1 or higher order (F2 or back-cross) hybrid individuals. </p> <p>Beacham, T. D. & Wallace, C. G. (2019). Salmon species identification via direct DNA sequencing of single amplicons. <i>Conservation Genetics Resources,</i> 1-7<i>. </i><a href="https://doi.org/10.1007/s12686-o19-01102-1">https://doi.org/10.1007/s12686-o19-01102-1</a>.</p>
Tutorial output for Tourmaline amplicon sequence processing workflow
<p>Tutorial output for the <a href="https://github.com/aomlomics/tourmaline">Tourmaline</a> amplicon sequence processing workflow.</p> <p>Tourmaline was run on the test data provided in the directory <a href="https://github.com/aomlomics/tourmaline/tree/master/00-data">00-data</a>, which were downloaded along with the rest of the repository using this command:</p> <pre><code>git clone https://github.com/aomlomics/tourmaline</code></pre> <p>Reference data were downloaded and symlinked using these commands:</p> <pre><code>cd tourmaline/01-imported wget https://data.qiime2.org/2021.2/common/silva-138-99-seqs-515-806.qza wget https://data.qiime2.org/2021.2/common/silva-138-99-tax-515-806.qza ln -s silva-138-99-seqs-515-806.qza refseqs.qza ln -s silva-138-99-tax-515-806.qza reftax.qza</code></pre> <p>Paths in 00-data/manifest_pe.csv and 00-data/manifest_se.csv were edited to match the local paths.</p> <p>Output for all modes of the workflow were then generated in series:</p> <pre><code>conda activate qiime2-2021.2 snakemake dada2_pe_report_unfiltered snakemake dada2_pe_report_filtered snakemake dada2_se_report_unfiltered snakemake dada2_se_report_filtered snakemake deblur_se_report_unfiltered snakemake deblur_se_report_filtered</code></pre> <p> </p>
Improved library preparation protocols for amplicon sequencing-based noninvasive fetal genotyping for RHD-positive D antigen-negative alleles
<p>We aimed to simplify our fetal <i>RHD</i> genotyping protocol by changing the method to attach Illumina's sequencing adaptors to PCR products from the ligation-based method to a PCR-based method, and to improve its quantitative accuracy by introducing unique molecular indexes, which allow us to count the numbers of DNA fragments used as PCR templates and to minimize the effects of PCR and sequencing errors. Both of the newly established protocols reduced time and cost compared with our conventional protocol. Removal of PCR duplicates using UMIs reduced the frequencies of erroneously mapped sequences reads likely generated by PCR and sequencing errors. The modified protocols will help us facilitate implementing fetal <i>RHD</i> genotyping for East Asian populations into clinical practice.</p>
Detection of SARS-CoV-2 variants by genomic analysis of wastewater ampliconic samples (Galaxy Training Material)
<p>The tutorial aims to train how to run workflows to analyze lineages abundances in SAR-CoV-2 wastewater ampliconic samples. (https://training.galaxyproject.org/training-material/)</p>
Population admixtures in medaka inferred by multiple arbitrary amplicon sequencing
<p>Cost-effective genotyping can be achieved by sequencing PCR amplicons. Short 3–10 base primers can arbitrarily amplify thousands of loci using only a few primers. To improve the sequencing efficiency of the multiple arbitrary amplicon sequencing (MAAS) approach, we designed new primers and examined their efficiency in sequencing and genotyping. To demonstrate the effectiveness of our method, we applied it to examine the population structure of the small freshwater fish, medaka (<em>Oryzias</em> <em>latipes</em>). We obtained 2,987 informative SNVs with no missing genotype calls for 67 individuals from 15 wild populations and three artificial strains. The estimated phylogenic and population genetic structures of the wild populations were consistent with previous studies, corroborating the accuracy of our genotyping method. We also attempted to reconstruct the genetic backgrounds of a commercial orange mutant strain, Himedaka, which has caused a genetic disturbance in wild populations. Our admixture analysis focusing on Himedaka showed that at least two wild populations had genetically contributed to the nuclear genome of this mutant strain. Our genotyping methods and results will be useful in quantitative assessments of genetic disturbance by this commercially available strain.</p>
[DATA] Amplicon-based gene expression profiling of blood from mice stimulated with R848 and treated with TLR7/8 inhibitor MHV370
<p>129/Sv mice (5 per group) were treated with daily R848 i.p. (2.5 mg/kg) for 14 days. From day 7 onward, one group received oral doses of 15 mg/kg MHV370 b.i.d., one group vehicle. At day 14, both groups were compared to naive 129/Sv mice. Gene expression profiling of blood was performed via amplicon sequencing (Ion AmpliSeq Transcriptome Mouse Gene Expression Kit, Thermo Fisher Scientific).</p> <p>Raw amplicon counts, sample metadata and gene information are provided in tab-separated value (TSV) files counts.tsv, samples.tsv and genes.tsv, respectively.</p> <p>The entire dataset is also provided as a DGEList object saved as an RDS file (dge_unfiltered.rds) that can be explored using R (package edgeR).</p> <p>The respective code for analysis can be found under DOI: 10.5281/zenodo.7575672</p>
Amplicon sequence variants (ASV) of gut pathogens in hooded cranes and domestic geese
<p>Driven by habitat loss from anthropogenic activities, wintering migratory birds forage together with poultry in paddy fields, and thus impose risks of cross transmitting pathogens. To date, there is little evidence for such risks of pathogen transmission between wild birds and poultry. Using the high-throughput sequencing, we report on detected potential pathogens of both wild hooded cranes <em>Grus monacha</em> and sympatric domestic geese <em>Anser</em> <em>anser</em> <em>domesticus</em> during the wintering period and infer the possibility of cross-species pathogen transmission. The results revealed that the number of shared amplicon sequence variants (ASVs) of potential pathogens between the gut microbiota of the two species was low during the early wintering stage (17.2%; 5 ASVs shared) but increased to 56.3% (18 ASVs shared) during the late wintering stage. That is, potential pathogens in the gut microbial communities of the two species became more similar through co-foraging in paddy fields, supporting cross-transmission of pathogens between hooded cranes and domestic geese during the wintering period. Importantly, transmission appeared to be largely from wild hooded cranes to domestic geese, although some potential pathogens may have become specialized to the domestic goose in late wintering stage. Humans are also facing the risks of contracting these potential pathogens from migratory birds through their frequent contact with domestic poultry. It is, therefore, necessary to closely monitor this pathway of pathogen transmission from wild birds to domestic animals and even to humans.</p>
Community assembly amplicon sequences, with pipeline to get asv table for "Spatial structure drives compositional convergence between nutrient environments in experimental microbial communities"
<p>Community assembly amplicon sequences, with pipeline to get asv table for "Spatial structure drives compositional convergence between nutrient environments in experimental microbial communities"</p> <p> </p> <p>compressed FASTA files for 16s amplicon sequences relating to two separate projects, "Spatial structure drives compositional convergence between nutrient environments in experimental microbial communities" and "Habitat filtering leads to phylogenetic clustering in synthetic microbial communities". DADA22 pipeline is included, which pools all samples for better accuracy. A Julia script bioinfo.jl is then used to select only the samples relevant to spatial structure project.</p> <p> </p> <p>All csv filenames are appended with "_q" indicating an increase in the stringency of quality filtering parameters (also increasing minimum hamming distance used in DADA2 algorithm to 5) to produce a taxa table with a sensible number of ASVs (given a known number of input strains) with each ASV uniquely aligning to an individual sequence from colony PCR of said input strains.</p> <p> </p>
Inferring microbial co-occurrence network from amplicon data: a systematic evaluation
<p>Supporting data for the manuscript "<em>Inferring microbial co-occurrence network from amplicon data: a systematic evaluation</em>".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.