Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

12

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

12 results for “Primer Selection”

Learn how ShareScore rates datasets ↗
dryad36/100

A fast machine-learning-guided primer design pipeline for selective whole genome amplification

<p>Addressing many of the major outstanding questions in the fields of microbial evolution and pathogenesis will require analyses of populations of microbial genomes. Although population genomic studies provide the analytical resolution to investigate evolutionary and mechanistic processes at fine spatial and temporal scales – precisely the scales at which these processes occur – microbial population genomic research is currently hindered by the practicalities of obtaining sufficient quantities of the relatively pure microbial genomic DNA necessary for next-generation sequencing. Here we present swga2.0, an optimized and parallelized pipeline to design selective whole genome amplification (SWGA) primer sets. Unlike previous methods, swga2.0 incorporates active and machine learning methods to evaluate the amplification efficacy of individual primers and primer sets. Additionally, swga2.0 optimizes primer set search and evaluates strategies, including parallelization at each stage of the pipeline, to dramatically decrease program runtime from weeks to minutes. Here we describe the swga2.0 pipeline, including the empirical data used to identify primer and primer set characteristics, that improve amplification performance. Additionally, we evaluated the novel swga2.0 pipeline by designing primers sets that successfully amplify <em>Prevotella melaninogenica</em>, an important component of the lung microbiome in cystic fibrosis patients, from samples dominated by human DNA.</p>

opencc-zeroAug 2022View details →
dryad36/100

Dataset for: How eDNA data filtration, sequence coverage, and primer selection influence assessment of fish communities in northern temperate lakes

<p><span>For nearly 15 years now, environmental DNA</span><span> has demonstrated</span><span> its effectiveness in monitoring biodiversity. Methodological and technical improvements have significantly enhanced the field. However, the effect of factors such as sequence coverage, bioinformatic filtration and primer choice have been less explored or need to be optimized according </span><span>to </span><span>specific survey objectives and </span><span>study </span><span>site characteristics. We evaluated these factors </span><span>to </span><span>help optimize monitoring fish biodiversity in North American temperate lakes. We sampled water for fish community eDNA analysis in 12 lakes from southwestern Québec, Canada. The lakes were selected to encompass a wide range of surface areas and species richness. We sampled water from a total of </span><span>520</span><span> sites (25 to 50 per lake) and analyzed three mitochondrial DNA regions (12S rRNA; 16S rRNA; and cytb) using NovaSeq</span><span> sequencing. Our results, based on rarefied count matrices (from a sequencing depth of 100,000 to a minimum </span><span>depth </span><span>of 1,000 reads per sample), </span><span>showed</span><span> that </span><span>keeping only</span><span> species </span><span>in each sample if they</span><span> represented </span><span>at least one thousandth (species </span><span>minimum </span><span>read proportion threshold =</span><span> 0.001</span><span>)</span><span> of the </span><span>sample's</span><span> reads was adequate to remove false positives </span><span>and had a limited negative</span><span> impact on true positives</span><span> with low read counts. The</span><span> sequencing depth </span><span>was found to have</span><span> a negligible impact </span><span>on the accuracy</span><span> of fish </span><span>community assessment in a given lake. With the same sequencing depth and a complete local reference database for each primer set, </span><span>a single primer set </span><span>produced</span><span> similar species richness medians than the combination of two or three primer sets. Overall, 12S and 16S detected more species and provided more consistent community profiles than cytb. </span><span>Based on our observations, we suggest using the 12S MiFish-U primer set and applying a minimum proportion of 0.001 reads per species and site to monitor north-temperate lentic freshwater fish communities.</span></p>

opencc-zeroJun 2023View details →
dryad36/100

A fast machine-learning-guided primer design pipeline for selective whole genome amplification

Open the record for dataset details and reuse information.

publicAug 2022View details →
dryad36/100

Dataset for: How eDNA data filtration, sequence coverage, and primer selection influence assessment of fish communities in northern temperate lakes

Open the record for dataset details and reuse information.

publicJun 2023View details →
dryad32/100

Data from: Determining diet from faeces: selection of metabarcoding primers for the insectivore Pyrenean desman (Galemys pyrenaicus)

Molecular techniques allow non-invasive dietary studies from faeces, providing an invaluable tool to unveil ecological requirements of endangered or elusive species. They contribute to progress on important issues such as genomics, population genetics, dietary studies or reproductive analyses, essential knowledge for conservation biology. Nevertheless, these techniques require general methods to be tailored to the specific research objectives, as well as to substrate- and species-specific constraints. In this pilot study we test a range of available primers to optimise diet analysis from metabarcoding of faeces of a generalist aquatic insectivore, the endangered Pyrenean desman (Galemys pyrenaicus, É. Geoffroy Saint-Hilaire, 1811, Talpidae), as a step to improve the knowledge of the conservation biology of this species. Twenty-four faeces were collected in the field, DNA was extracted from them, and fragments of the standard barcode region (COI) were PCR amplified by using five primer sets (Brandon-Mong, Gillet, Leray, Meusnier and Zeale). PCR outputs were sequenced on the Illumina MiSeq platform, sequences were processed, clustered into OTUs (Operational Taxonomic Units) using UPARSE algorithm and BLASTed against the NCBI database. Although all primer sets successfully amplified their target fragments, they differed considerably in the amounts of sequence reads, rough OTUs, and taxonomically assigned OTUs. Primer sets consistently identified a few abundant prey taxa, probably representing the staple food of the Pyrenean desman. However, they differed in the less common prey groups. Overall, the combination of Gillet and Zeale primer sets were most cost-effective to identify the widest taxonomic range of prey as well as the desman itself, which could be further improved stepwise by adding sequentially the outputs of Leray, Brandon-Mong and Meusnier primers. These results are relevant for the conservation biology of this endangered species as they allow a better characterization of its food and habitat requirements.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Improving PCR detection of prey in molecular diet studies: importance of group-specific primer set selection and extraction protocol performances

While morphological identification of prey remains in feces of predators is the method most commonly used to study trophic interactions, many studies indicate that this method does not detect all consumed prey. Polymerase Chain Reaction based methods are increasingly used to detect prey DNA in the predator food bolus and have proved themselves efficient, with high accuracy. When studying complex diet samples, the extraction of total DNA is a critical step, as PCR inhibitors may be co-extracted. Another critical step consist in carefully select suitable group-specific primer sets that should only amplify prey DNA from the targeted taxon. In this study, the food boluses of five Rattus rattus and seven Rattus exulans were analyzed using both morphological and molecular methods. We tested a panel of 30 PCR specific primer sets targeting Bird, Invertebrate and Plant sequences and four were finally selected to be use as group-specific primer pairs in PCR protocols. The performances of four DNA extraction protocols (QIAamp DNA stool mini kit, DNeasy mericon food kit and two CTAB-based methods) were compared using four variables: DNA concentration, A260/A280 absorbance ratio, food compartment analyzed (stomach or fecal contents), total number of prey specific PCR amplification per sample. Our results clearly indicate that the A260/A280 absorbance ratio, which varies between extraction protocols, is positively correlated to the number of PCR amplifications of each prey taxon. We recommend using the DNeasy mericon food kit (Qiagen), which yielded results very similar to those achieved with the morphological approach.

opencc-zeroDec 2011View details →
dryad32/100

Data from: Tissue storage and primer selection influence pyrosequencing-based inferences of diversity and community composition of endolichenic and endophytic fungi

Next-generation sequencing technologies have provided unprecedented insights into fungal diversity and ecology. However, intrinsic biases and insufficient quality control in next-generation methods can lead to difficult-to-detect errors in estimating fungal community richness, distributions, and composition. The aim of this study was to examine how tissue storage prior to DNA extraction, primer design, and various quality-control approaches commonly used in 454 amplicon pyrosequencing might influence ecological inferences in studies of endophytic and endolichenic fungi. We first contrast 454 data sets generated contemporaneously from subsets of the same plant and lichen tissues that were stored in CTAB buffer, dried in silica gel, or freshly frozen prior to DNA extraction. We show that storage in silica gel markedly limits the recovery of sequence data and yields a small fraction of the diversity observed by the other two methods. Using lichen mycobiont sequences as internal positive controls, we next show that despite careful filtering of raw reads and utilization of current best-practice OTU clustering methods, homopolymer errors in sequences representing rare taxa artificially increased estimates of richness ca. 15-fold in a model data set. Third, we show that inferences regarding endolichenic diversity can be improved by using a novel primer that reduces amplification of the mycobiont. Together, our results provide a rationale for selecting tissue treatment regimes prior to DNA extraction, demonstrate the efficacy of reducing mycobiont amplification in studies of the fungal microbiomes of lichen thalli, and highlight the difficulties in differentiating true information about fungal biodiversity from methodological artifacts.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Tissue storage and primer selection influence pyrosequencing-based inferences of diversity and community composition of endolichenic and endophytic fungi

Open the record for dataset details and reuse information.

publicMar 2014View details →
dryad32/100

Data from: Improving PCR detection of prey in molecular diet studies: importance of group-specific primer set selection and extraction protocol performances

Open the record for dataset details and reuse information.

publicOct 2012View details →
dryad32/100

Data from: Determining diet from faeces: selection of metabarcoding primers for the insectivore Pyrenean desman (Galemys pyrenaicus)

Open the record for dataset details and reuse information.

publicDec 2018View details →
geo24/100

Subfamily-Selective PCR primers for the Human LINE1 L1PA Lineage

GEO Series GSE301759. Homo sapiens. 6 samples. Type: Other.

openGEO-OpenJul 2025View details →
geo20/100

Multiplexed RNA structure characterization with selective 2'-hydroxyl acylation analyzed by primer extension sequencing (SHAPE-Seq)

GEO Series GSE31573. Bacillus subtilis. 2 samples. Type: Other.

openGEO-OpenAug 2011View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record