Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

253

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

253 results for “amplicons”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Adapterama II: universal amplicon sequencing on Illumina platforms (TaggiMatrix)

Next-generation sequencing (NGS) of amplicons is used in a wide variety of contexts. In many cases, NGS amplicon sequencing remains overly expensive and inflexible, with library preparation strategies relying upon the fusion of locus-specific primers to full-length adapter sequences with a single identifying sequence or ligating adapters onto PCR products. In Adapterama I, we presented universal stubs and primers to produce thousands of unique index combinations and a modifiable system for incorporating them into Illumina libraries. Here, we describe multiple ways to use the Adapterama system and other approaches for amplicon sequencing on Illumina instruments. In the variant we use most frequently for large-scale projects, we fuse partial adapter sequences (TruSeq or Nextera) onto the 5' end of locus-specific PCR primers with variable-length tag sequences between the adapter and locus-specific sequences. These fusion primers can be used combinatorially to amplify samples within a 96-well plate (eight forward primers + 12 reverse primers yield 8 x 12 = 96 combinations), and the resulting amplicons can be pooled. The initial PCR products then serve as template for a second round of PCR with dual-indexed iTru or iNext primers (also used combinatorially) to make full-length libraries. The resulting quadruple-indexed amplicons have diversity at most base positions and can be pooled with any standard Illumina library for sequencing. The number of sequencing reads from the amplicon pools can be adjusted, facilitating deep sequencing when required or reducing sequencing costs per sample to an economically trivial amount when deep coverage is not needed. We demonstrate the utility and versatility of our approaches with results from six projects using different implementations of our protocols. Thus, we show that these methods facilitate amplicon library construction for Illumina instruments at reduced cost with increased flexibility. A simple web page to design fusion primers compatible with iTru primers is available at: http://baddna.uga.edu/tools-taggi.html. A fast and easy to use program to demultiplex amplicon pools with internal indexes is available at: https://github.com/lefeverde/Mr_Demuxy.

opencc-zeroSep 2020View details →
dryad28/100

Data from: Proportion methylation at a set of CpGs from 4 amplicons

<p><span>The age structure of populations, or the ageing rate of individuals, impacts aspects of animal ecology, epidemiology and conservation. Yet for many wild organisms, age is an inaccessible trait. In many cases measuring age or ageing rates in the wild requires molecular biomarkers of age. Epigenetic clocks based on DNA methylation have been shown to accurately estimate the age of humans and laboratory mice, but they also show variable ticking rates that are associated with mortality risk above and beyond that predicted by chronological age. Thus, epigenetic clocks are proving to be useful markers of both chronological and biological age, and they are beginning to be applied to wild mammals and birds. We have acquired strong evidence that an accurate clock will be possible for the wood mouse <i>Apodemus sylvaticus </i>by adapting epigenetic information from the laboratory mouse. <i>Apodemus sylvaticus is</i> a well-studied field system that is amenable to experimental perturbations and longitudinal sampling of individuals across their lives, and these features of the wood mouse offer opportunities to disentangle causal relationships between ageing rates and environmental stress. Our wood mouse epigenetic clock is PCR-based, and so requires tiny amounts of tissue and non-destructive sampling. We quantified methylation using Oxford Nanopore sequencing technology and present a new bioinformatics pipeline for data analysis. We thus describe a new and generalizable system that should enable ecologists and other field biologists to go from tiny tissue samples to an epigenetic clock for their study animal.   </span></p>

opencc-zeroOct 2020View details →
dryad28/100

Data from: Quantifying sequence proportions in a DNA-based diet study using Ion Torrent amplicon sequencing: which counts count?

A goal of many environmental DNA barcoding studies is to infer quantitative information about relative abundances of different taxa based on sequence read proportions generated by high-throughput sequencing. However, potential biases associated with this approach are only beginning to be examined. We sequenced DNA amplified from faeces (scats) of captive harbour seals (Phoca vitulina) to investigate whether sequence counts could be used to quantify the seals' diet. Seals were fed fish in fixed proportions, a chordate-specific mitochondrial 16S marker was amplified from scat DNA and amplicons sequenced using an Ion Torrent PGM™. For a given set of bioinformatic parameters, there was generally low variability between scat samples in proportions of prey species sequences recovered. However, proportions varied substantially depending on sequencing direction, level of quality filtering (due to differences in sequence quality between species) and minimum read length considered. Short primer tags used to identify individual samples also influenced species proportions. In addition, there were complex interactions between factors; for example, the effect of quality filtering was influenced by the primer tag and sequencing direction. Resequencing of a subset of samples revealed some, but not all, biases were consistent between runs. Less stringent data filtering (based on quality scores or read length) generally produced more consistent proportional data, but overall proportions of sequences were very different than dietary mass proportions, indicating additional technical or biological biases are present. Our findings highlight that quantitative interpretations of sequence proportions generated via high-throughput sequencing will require careful experimental design and thoughtful data analysis.

opencc-zeroDec 2012View details →
dryad28/100

Data from: From benchtop to desktop: important considerations when designing amplicon sequencing workflows

Amplicon sequencing has been the method of choice in many high-throughput DNA sequencing (HTS) applications. To date there has been a heavy focus on the means by which to analyse the burgeoning amount of data afforded by HTS. In contrast, there has been a distinct lack of attention paid to considerations surrounding the importance of sample preparation and the fidelity of library generation. No amount of high-end bioinformatics can compensate for poorly prepared samples and it is therefore imperative that careful attention is given to sample preparation and library generation within workflows, especially those involving multiple PCR steps. This paper redresses this imbalance by focusing on aspects pertaining to the benchtop within typical amplicon workflows: sample screening, the target region, and library generation. Empirical data is provided to illustrate the scope of the problem. Lastly, the impact of various data analysis parameters is also investigated in the context of how the data was initially generated. It is hoped this paper may serve to highlight the importance of pre-analysis workflows in achieving meaningful, future-proof data that can be analysed appropriately. As amplicon sequencing gains traction in a variety of diagnostic applications from forensics to environmental DNA (eDNA) it is paramount workflows and analytics are both fit for purpose.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Amplicon pyrosequencing late Pleistocene permafrost: the removal of putative contaminant sequences and small-scale reproducibility

DNA sequencing of ancient permafrost samples can be used to reconstruct past plant, animal and bacterial communities. In this study, we assess the small-scale reproducibility of taxonomic composition obtained from sequencing four molecular markers (mitochondrial 12S ribosomal DNA (rDNA), prokaryote 16S rDNA, mitochondrial cox1 and chloroplast trnL intron) from two soil cores sampled 10 cm apart. In addition, sequenced control reactions were used to produce a contaminant library that was used to filter similar sequences from sample libraries. Contaminant filtering resulted in the removal of 1% of reads or 0.3% of operational taxonomic units. We found similar richness, overlap, abundance and taxonomic diversity from the 12S, 16S and trnL markers from each soil core. Jaccard dissimilarity across the two soil cores was highest for metazoan taxa detected by the 12S and cox1 markers. Taxonomic community distances were similar for each marker across the two soil cores when the chi-squared metric was used; however, the 12S and cox1 markers did not cluster well when the Goodall similarity metric was used. A comparison of plant macrofossil vs. read abundance corroborates previous work that suggests eastern Beringia was dominated by grasses and forbs during cold stages of the Pleistocene, a habitat that is restricted to isolated sites in the present-day Yukon.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Slippage of degenerate primers can cause variation in amplicon length

It is well understood that homopolymer regions should be avoided for primer binding to prevent off-target amplification. However, in metabarcoding, it is often difficult to avoid primer degeneracy in order to maximize taxa detection. We here investigate primer binding specificity using different primer sets from several invertebrate metabarcoding studies. Our results indicate that primers frequently bound 1-2 bp upstream in taxa where a homopolymer region was present in the amplification direction. Primer binding 1 bp downstream was observed less frequently. This primer slippage leads to taxon-specific length variation in amplicons and subsequent length variation in recovered sequences. Some widely used primer sets were severely affected by this bias, while others were not. While this variation will only have small impacts on the designation of Operational Taxonomic Units (OTUs) by clustering algorithms that ignore terminal gaps, primer sets employed in metabarcoding projects should be evaluated for their sensitivity to slippage. Moreover, steps should be taken to reduce slippage by improving protocols for primer design. For example, the flanking region adjacent to the 3′ end of the primer is not considered by current primer development software although GC clamps in this position could mitigate slippage.

opencc-zeroDec 2017View details →
zenodo28/100

Pre-processed B-cell receptor amplicon sequencing data from SRR1842411

<p>An example dataset containing B-cell receptor (BCR) gene sequences. This dataset is intended to be used for testing software tools developed to annotate (i.e. map Variable, Diversity and Joining segments) and perform clonal analysis of BCR sequencing data.</p> <p><strong>Sequencing:</strong></p> <p>Libraries prepared using 5'RACE from PBMCs of a healthy donor. Input molecules were tagged with unique molecular identifiers (UMIs). Sequencing was ran on MiSeq , 300+300bp reads.</p> <p><strong>Contents:</strong></p> <p>The dataset contains both raw sequencing reads and high-quality consensus sequences assembled using unique molecular tagging (UMI) approach. Consensus assembly corrects for sequencing errors and eliminates sequencing artifacts.</p> <ul> <li>age_ig_s7_R1.fastq.gz and age_ig_s7_R2.fastq.gz contain raw reads</li> <li>age_ig_s7_R1.t10.cf.fastq.gz and age_ig_s7_R2.t10.cf.fastq.gz contain consensus sequences</li> </ul> <p>All files contain an UMI tag sequence in their header, in form UMI:NNNN:QQQQ where N is the base character and Q is the quality character (for assembled consensuses the total number of reads is given instead of Q string).</p> <p>Note that consensus sequences were assembled using only raw sequences that correspond to UMI tags supported by at least 10 sequencing reads. That means that consensus sequence files contain a subset of all UMI tags found in raw sequences. Thus, if one wants to assess software performance on raw sequencing reads using assembled consensus sequences as a high-quality data standard, raw sequencing reads should be filtered to contain only those UMI tags that are present in consensus sequence file.</p> <p><strong>Citations:</strong></p> <p>The whole dataset was used to benchmark MiXCR software and was originally referenced in Bolotin DA, et al. MiXCR: software for comprehensive adaptive immunity profiling Nature methods 12(5):380-381, 2015.</p> <p>Data pre-processing was carried out using MIGEC software, Shugay M et al. Towards error-free profiling of immune repertoires. Nature Methods 11(6):653-655, 2014.</p> <p><strong>Contributors:</strong></p> <p>The dataset was generated in Prof. Chudakov lab (Adaptive Immunity Group in Masaryk University, Brno and Genomics of Adaptive Immunity Lab in Institute of Bioorganic Chemistry, Moscow). Sample preparation and sequencing was performed by Dr. Olga Britanova and Dr. Maria Turchaninova. Raw sequencing reads were pre-processed and uploaded by Dr. Mikhail Shugay.</p>

opencc-by-4.0Jun 2017View details →
zenodo28/100

Supplementary material 2 from: Reid BN, Servis JA, Timmers M, Rohwer F, Naro-Maciel E (2022) 18S rDNA amplicon sequence data (V1–V3) of the Palmyra Atoll National Wildlife Refuge, Central Pacific. Metabarcoding and Metagenomics 6: e78762. https://doi.org/10.3897/mbmg.6.78762

Figure S2

opencc-zeroApr 2022View details →
zenodo28/100

Supplementary material 5 from: Reid BN, Servis JA, Timmers M, Rohwer F, Naro-Maciel E (2022) 18S rDNA amplicon sequence data (V1–V3) of the Palmyra Atoll National Wildlife Refuge, Central Pacific. Metabarcoding and Metagenomics 6: e78762. https://doi.org/10.3897/mbmg.6.78762

Table S3

opencc-zeroApr 2022View details →
zenodo28/100

Supplementary material 6 from: Reid BN, Servis JA, Timmers M, Rohwer F, Naro-Maciel E (2022) 18S rDNA amplicon sequence data (V1–V3) of the Palmyra Atoll National Wildlife Refuge, Central Pacific. Metabarcoding and Metagenomics 6: e78762. https://doi.org/10.3897/mbmg.6.78762

Table S4

opencc-zeroApr 2022View details →
zenodo28/100

Supplementary material 1 from: Reid BN, Servis JA, Timmers M, Rohwer F, Naro-Maciel E (2022) 18S rDNA amplicon sequence data (V1–V3) of the Palmyra Atoll National Wildlife Refuge, Central Pacific. Metabarcoding and Metagenomics 6: e78762. https://doi.org/10.3897/mbmg.6.78762

Figure S1

opencc-zeroApr 2022View details →
zenodo28/100

Supplementary material 4 from: Reid BN, Servis JA, Timmers M, Rohwer F, Naro-Maciel E (2022) 18S rDNA amplicon sequence data (V1–V3) of the Palmyra Atoll National Wildlife Refuge, Central Pacific. Metabarcoding and Metagenomics 6: e78762. https://doi.org/10.3897/mbmg.6.78762

Table S2

opencc-zeroApr 2022View details →
zenodo28/100

Supplementary material 3 from: Reid BN, Servis JA, Timmers M, Rohwer F, Naro-Maciel E (2022) 18S rDNA amplicon sequence data (V1–V3) of the Palmyra Atoll National Wildlife Refuge, Central Pacific. Metabarcoding and Metagenomics 6: e78762. https://doi.org/10.3897/mbmg.6.78762

Table S1

opencc-zeroApr 2022View details →
zenodo28/100

Rapid multiplexed nanopore amplicon sequencing to distinguish Plasmodium falciparum recrudescence from new infection in antimalarial drug trials

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo28/100

Supplementary material 4 from: Zafeiropoulos H, Gargan L, Hintikka S, Pavloudi C, Carlsson J (2021) The Dark mAtteR iNvestigator (DARN) tool: getting to know the known unknowns in COI amplicon data. Metabarcoding and Metagenomics 5: e69657. https://doi.org/10.3897/mbmg.5.69657

Figure S2

opencc-zeroNov 2021View details →
zenodo28/100

Supplementary material 3 from: Zafeiropoulos H, Gargan L, Hintikka S, Pavloudi C, Carlsson J (2021) The Dark mAtteR iNvestigator (DARN) tool: getting to know the known unknowns in COI amplicon data. Metabarcoding and Metagenomics 5: e69657. https://doi.org/10.3897/mbmg.5.69657

Figure S1

opencc-zeroNov 2021View details →
zenodo28/100

Supplementary material 2 from: Zafeiropoulos H, Gargan L, Hintikka S, Pavloudi C, Carlsson J (2021) The Dark mAtteR iNvestigator (DARN) tool: getting to know the known unknowns in COI amplicon data. Metabarcoding and Metagenomics 5: e69657. https://doi.org/10.3897/mbmg.5.69657

Table S2

opencc-zeroNov 2021View details →
zenodo28/100

Supplementary material 1 from: Zafeiropoulos H, Gargan L, Hintikka S, Pavloudi C, Carlsson J (2021) The Dark mAtteR iNvestigator (DARN) tool: getting to know the known unknowns in COI amplicon data. Metabarcoding and Metagenomics 5: e69657. https://doi.org/10.3897/mbmg.5.69657

Table S1

opencc-zeroNov 2021View details →
zenodo28/100

Figure 1 from: Gueidan C, Li L (2022) A long-read amplicon approach to scaling up the metabarcoding of lichen herbarium specimens. MycoKeys 86: 195-212. https://doi.org/10.3897/mycokeys.86.77431

Figure 1 Examples of lichen herbarium specimens used for this study AParmotrema perlatum, specimen J.A. Elix 43686 (CANB790817) BEndocarpon pusillum, specimen H. Streiman 45100 (CBG9011273) CBuellia albula, specimen J.A. Elix 45138 (CANB810791) DCatillaria sp., specimen J.A Elix 37142 (CANB872684). Scale bar: 1 cm. Photos C. Gueidan.

opencc-by-4.0Feb 2022View details →
zenodo28/100

Figure 2 from: Gueidan C, Li L (2022) A long-read amplicon approach to scaling up the metabarcoding of lichen herbarium specimens. MycoKeys 86: 195-212. https://doi.org/10.3897/mycokeys.86.77431

Figure 2 Sequencing success for different morphological groups of taxa included in this study. Specimens were grouped into three main morphological categories: 1Buellia, Catillaria and other crustose saxicolous taxa 2Endocarpon and other squamulose terricolous taxa 3 the foliose corticolous genus Parmotrema. In the graph, stalked columns show successful samples (sequence generated for the target species) in dark grey and unsuccessful samples (no sequence generated or generated sequences not from the target species) in light grey. The total number of samples (N) is indicated below each corresponding column.

opencc-by-4.0Feb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record