Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
253
datasets available to search
ShareScore release 0.9.0
Dataset results
253 results for “amplicons”
(Fastq Files) Amplicon sequencing of ama1 and mdr1 to track within-host P. falciparum diversity in Kilifi, KENYA
<p>These data were generated from amplicon sequencing of <em>Plasmodium falciparum</em> <em>ama1 </em>and<em> </em><em>mdr1</em> genes in samples collected from Kilifi, at the coast of Kenya.</p> <p>The two papers that reference these data will soon be included here:</p> <ol> <li> The Journal of Infectious Diseases - https://doi.org/10.1093/infdis/jiac144</li> <li>Wellcome Open Research - https://wellcomeopenresearch.org/articles/7-95</li> </ol> <p>Two objectives were explored:</p> <ol> <li>To determine temporal changes in the genetic diversity of malaria parasites in asymptomatic and febrile infections.</li> <li>To track within-host parasite diversity, throughout treatment in a clinical drug trial.</li> </ol>
Benchmarking bioinformatic tools for amplicon-based sequencing of norovirus
<p>This repository contains associated datasets and accession numbers for a study entitled '<strong>Benchmarking bioinformatic tools for amplicon-based sequencing of norovirus'</strong>. The scripts for this project can be found on the GitHub project<a href="https://github.com/ahfitzpa/Benchmarking-bioinformatics-norovirus-amplicons"> page</a>. </p> <p>Expected composition tsv files are the OTU tables for each simulation performed (001-010). OTU IDs in this case are the expected taxonomy with the associated accession numbers. Samples are numbered 1-40, including the simulation number. Expected sequences fasta files contain the sequences used as input for each simulation, without primers or Illumina adapter sequences.</p> <p>Amplicons were generated using the following primers:</p> <p><strong>GI Primers </strong><br> GISKF: CTG CCC GAA TTY GTA AAT GA 4<br> GISKR: CCA ACC CAR CCA TTR TAC A 5<br> <br> <strong>GII Primers </strong><br> G2SKF: CNT GGG AGG GCG ATC GCAA 8<br> G2SKR: CCR CCN GCA TRH CCR TTR TAC AT</p> <p>In this study, three databases and multiple classifiers were compared. Here we include the taxonomy and fasta files for each database; noronet =NoroNet RIVM, calicinet= HuCat CDC and custom, randomly generated database. Fasta files for the classifiers include the GI/GII primers listed above in a 5-3 orientation. </p> <p>The tags.txt file contains the Illumina adapters used for the simulation component of the study.</p>
Inventory of High-resolution phylogenetic profiles of the planktonic microbial communities (via 16S and 18S rRNA gene amplicons) from Shark River Slough and Taylor Slough, Everglades National Park (FCE LTER), Florida, USA, 2017 - ongoing
Planktonic microbial communities mediate many vital biogeochemical processes in wetland ecosystems, yet compared to other aquatic ecosystems, like oceans, lakes, rivers, or estuaries, they remain relatively underexplored. Our study site, the Florida Everglades (USA)—a vast iconic wetland consisting of a slow-moving system of shallow rivers connecting freshwater marshes with coastal mangrove forests and seagrass meadows—is a highly threatened model ecosystem for studying salinity and nutrient gradients, as well as the effects of sea level rise and saltwater intrusion. This dataset provides the first high-resolution phylogenetic profiles of planktonic bacterial and eukaryotic microbial communities (using 16S and 18S rRNA gene amplicons) from these environments. The dataset contains 16S and 18S rRNA data from 2017, and contains 16S rRNA data for monthly (2019) and quarterly water samples (2020-ongoing). The 2017 data are published in Laas et al. 2022. A detailed list of sequence data and their accession numbers in GenBank is provided and will be updated as more data are published. This data package is an inventory of sequence read archive (SRA) entries available through GenBank BioProject PRJNA525456 (at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA525456) and BioProject PRJNA1018945 (at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1018945). This data package is associated with the following publication: Laas, P., Ugarelli, K., Travieso, R., Stumpf, S., Gaiser, E. E., Kominoski, J. S., & Stingl, U. (2022). Water column microbial communities vary along salinity gradients in the Florida Coastal Everglades wetlands. Microorganisms, 10(2), 215. https://doi.org/10.3390/microorganisms10020215 Instead of citing this package, which is an inventory, please cite the original GenBank data or journal article, as appropriate. Citation guidance for the journal article is available on the respective publisher's website.
BAMBI ITS - Analysis of the fungal component (via ITS amplicon sequencing) of stool samples from preterm babies
<p>Amplicon analysis of ITS amplicons from preterm babies.</p> <p>Associated GitHub repository: <a href="https://github.com/quadram-institute-bioscience/bambi-its">https://github.com/quadram-institute-bioscience/bambi-its</a></p>
EMA-amplicon-based taxonomic characterisation of the viable bacterial community present in untreated and SODIS treated roof-harvested rainwater
<p>Dataset for publication: EMA-amplicon-based taxonomic characterisation of the viable bacterial community present in untreated and SODIS treated roof-harvested rainwater, Strauss et al. (2018). DOI: 10.1039/c8ew00613j.</p>
Identification of grapevine clones via high-throughput amplicon sequencing: a proof-of-concept study VCF files
<p>VCF files used and cited in the article: Identification of grapevine clones via high-throughput amplicon sequencing: a proof-of-concept study</p>
Inventory of soil prokaryotic microbiome (via 16S based on rRNA gene amplicons) in freshwater and brackish water marshes following saltwater intrusion along Shark River Slough boundary, Everglades National Park (FCE LTER), Florida, USA, September 2018
Global sea-level rise is transforming coastal ecosystems, especially freshwater wetlands, in part due to increased episodic or chronic saltwater exposure, leading to shifts in microbial communities and related ecological services. Soil prokaryotes play a fundamental role in regulating important biogeochemical processes in coastal wetland ecosystem. Yet, it is still difficult to predict how soil prokaryotic communities respond to the saltwater exposure because of poorly understood prokaryotic sensitivity within complex wetland soil microbial communities, as well as the high heterogeneity of wetland soils and saltwater exposure. To address this, a four-year experimental simulation of saltwater intrusion in a pristine freshwater site and a previously saltwater-impacted site was conducted. The saltwater addition started in October 2014 on a monthly basis and continued through October 2018. The dataset contains amplicon sequencing date of 16S rRNA gene obtained from saltwater-exposed soils and unmanipulated native soils in both sites (collected in September 2018). The 2018 data are published in Zhao et al. 2023. A detailed list of sequence data and their accession numbers in GenBank is provided, and data collection is complete. This data package is an inventory of sequence read archive (SRA) entries available through GenBank BioProject PRJNA804545 (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA804545). This data package is associated with the following publication: Zhao, J., Chakrabarti, S., Chambers, R., Weisenhorn, P., Travieso, R., Stumpf, S., Standen, E., Briceno, H., Troxler, T., Gaiser, E., Kominoski, J., Dhillon, B., & Martens-Habbena, W. (2023). Year-around survey and manipulation experiments reveal differential sensitivities of soil prokaryotic and fungal communities to saltwater intrusion in Florida Everglades wetlands. Science of The Total Environment, 858, 159865. https://doi.org/10.1016/j.scitotenv.2022.159865 Instead of citing this package, which is an
Inventory of soil prokaryotic and fungal microbiome (via 16S rRNA gene amplicons and ITS sequencing) from Shark River Slough and Taylor Slough, Everglades National Park (FCE LTER), Florida, USA, February 2019 - October 2020
Global sea-level rise is transforming coastal ecosystems, especially freshwater wetlands, in part due to increased saltwater exposure, leading to change in soil microbial communities and many important biogeochemical processes. Given the high spatial and temporal heterogeneity in coastal wetlands, especially in tropical or subtropical climates characterized by seasonal temperature, precipitation, and tidal fluctuations, it remains unclear which environmental factors influence the compositions of soil microbial communities in wetlands affected by varying degrees of sea-water intrusion. To address this, a two-year survey was conducted on microbial community structure in submerged surface soils from 14 wetland sites across the Florida Everglades, representing three major ecosystem types, i.e. freshwater marshes, mangrove forests, and seagrass meadows. Bulk surface soil samples of each site were collected from February 2019 to October 2020 to cover dry and wet seasons. In addition to bulk soil samples, soil cores were collected from each site in August 2020 to assess vertical gradients of microbial communities. The dataset contains amplicon sequencing data of 16S rRNA gene (both bulk soil and soil cores) and ITS gene (only the bulk soil). The 2019 to 2020 data are published in Zhao et al. 2023. A detailed list of sequence data and their accession numbers in GenBank is provided, and data collection is complete. This data package is an inventory of sequence read archive (SRA) entries available through GenBank BioProject PRJNA804243 (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA804243), PRJNA804246 (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA804246), and PRJNA804228 (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA804228). This data package is associated with the following publication: Zhao, J., Chakrabarti, S., Chambers, R., Weisenhorn, P., Travieso, R., Stumpf, S., Standen, E., Briceno, H., Troxler, T., Gaiser, E., Kominoski, J., Dhillon, B., & Martens-H
Code for generating figures and analyzing amplicon sequencing of human mRNA and reporter mRNA targeted with type III-A CRISPR complex from Streptococcus thermophiles
<p>This dataset contains code for analyzing amplicon sequencing data and generating figures in the manuscript by Anna Nemudraia, Artem Nemudryi, and Blake Wiedenheft (2024), "Repair of CRISPR-guided RNA breaks enables site-specific RNA excision in human cells." </p> <p>Amplicon sequencing data has been deposited to NCBI Sequence Read Archive (SRA) under BioProject PRJNA1099688. The description of read files deposited to SRA can be found in the spreadsheet ./code_for_sequencing_data_analysis/SRA_read_files_description.xlsx</p> <p>The code for analyzing amplicon sequencing data can be found in the archive "code_for_sequencing_data_analysis.tar.gz." Output files from this analysis were used to generate figures. Figures were generated using the ggplot2 package in R and finalized in CorelDRAW.</p> <p>Code for generating figures can be found in the archive "code_for_generating_figures.tar.gz". </p> <p>Any questions or requests regarding the data or the code should be addressed to Dr. Artem Nemudryi at artem.nemudryi@gmail.com.</p>
The processed clean data of 16S rRNA V4 amplicon sequecnces for the six stage of phenolic microbiome domestication
Open the record for dataset details and reuse information.
Data from: Pitfalls and pointers: an accessible guide to marker gene amplicon sequencing in ecological applications
<p>Next Generation Sequencing (NGS) is a powerful tool that has been rapidly adopted by many ecologists studying microbial communities. Despite the exciting demonstration of NGS technology as a tool for ecological research, cryptic pitfalls inherent to its use can obscure correct interpretation of NGS data. Here, we provide an accessible overview of a NGS process that uses marker gene amplicon sequences (MGAS) that will allow scientists, particularly community ecologists, to make appropriate methodological choices and understand limits on inference about community composition and diversity that can be drawn from MGAS data.</p> <p>We describe the MGAS pipeline, focusing specifically on cryptic sources of variation that have received less emphasis in the ecological literature, but which may substantially impact inference about microbial community diversity and composition. By simulating communities from published microbiome data, we demonstrate how these sources of variation can generate inaccurate or misleading patterns.</p> <p>We specifically highlight sample dilution without researcher awareness and lane-to-lane variability, two cryptic sources of variation arising during the MGAS pipeline. These sources of variation affect estimates of species presence and relative abundance, particularly for species with moderate to low abundances. Each of these sources of bias can lead to errors in the estimation of both absolute and relative abundance within, and turnover among, microbial communities.</p> <p>Awareness and understanding of what happens and, specifically, why it happens during MGAS generation is key to generating a strong data set and building a robust community matrix. Requesting sample dilution information from the sequencing center, including technical replicates across sequencing lanes, and understanding how sampling intensity and community taxa distribution patterns shape the measurement of community richness, evenness, and diversity are critical for drawing correct ecological inferences using MGAS data.</p>
Bacteria and archaea of the Columbia and Willamette Rivers, 16S rRNA gene amplicon library metadata
<p>Bacterial and archaeal communities in the Columbia and Willamette Rivers in the Portland, OR, USA, region were characterized by 16S rRNA gene amplicon sequencing as part of the Lewis & Clark College spring 2022 Microbial Ecology course. Whole-water (>0.2 µm) samples were collected from: the Willamette River; the Columbia River above the confluence with the Willamette; and the Columbia River just downstream of the confluence with the Willamette.</p> <p>This dataset provides additional metadata to supplement the DNA sequences archived with the NCBI SRA at <a href="https://www.ncbi.nlm.nih.gov/sra/PRJNA865380">https://www.ncbi.nlm.nih.gov/sra/PRJNA865380</a></p>
Assessment of microphytobenthos communities in the Kinzig catchment using photosynthesis-related traits, digital light microscopy and 18S-V9 amplicon sequencing
<p>This folder contains the datasets used in the article submitted to Frontiers in Ecology and Evolution in which we investigated the functional and compositional responses of microphytobenthos communities to surrounding land uses in the Kinzig River catchment, central Germany. We measured photosynthetic biomass using a Benthotorch, and analysed the diatom community using a newly developed digital light microscopy approach and 18S-V9 amplicon sequencing to characterise the whole protistan assemblages at sampling sites located in rural vs. urban areas.</p> <p>The folder contain the following datasets:</p> <p>kinzig2021_18SV9_filtered.csv # microphytobenthos 18SV9 amplicon sequencing data</p> <p>kinzig_paper.csv # OMNIDIA output of diatom data from microscopy and diatom subset from 18SV9 amplicon sequencing (including 4 letters OMINIDIA taxa codes for diatoms)</p> <p>kinzig2021_benthotorch.csv # Photosynthetic biomass (BenthoTorch data)</p> <p>Kinzig2021_fieldData.csv # environmental data</p> <p>env_dataKinz2021.csv # environment dataset</p> <p>Kinzig2021_diatDM&18S.R # Rscript used for analysis</p>
MinION sequence data: MinION sequencing of colorectal cancer tumor microbiomes – a comparison with amplicon-based and RNA-Sequencing
<p>MinION sequencing data that was unmapped by minimap2 for the 11 samples using in the "MinION sequencing of colorectal cancer tumor microbiomes – a comparison with amplicon-based and RNA-Sequencing" paper.</p>
Simultaneous genotyping of snails and infecting trematode parasites using high-throughput amplicon sequencing.
<p>Several methodological issues currently hamper the study of entire trematode communities within populations of their intermediate snail hosts. Here we develop a new workflow using high-throughput amplicon sequencing to simultaneously genotype snail hosts and their infecting trematode parasites. We designed primers to amplify 4 snail and 5 trematode markers in a single multiplex PCR. While also applicable to other genera, we focused on medically and economically important snail genera within the Superorder Hygrophila and targeted a broad taxonomic range of parasites within the Class Trematoda. We tested the workflow using 417 <i>Biomphalaria glabrata </i>specimens experimentally infected with <i>Schistosoma rodhaini</i>, two strains of<i> Schistosoma mansoni</i>,<i> </i>and combinations thereof. We evaluated the reliability of infection diagnostics, the robustness of the workflow, its specificity related to host and parasite identification, and the sensitivity to detect co-infections, immature infections, and changes of parasite biomass during the infection process. Finally, we investigated its applicability in wild-caught snails of other genera naturally infected with diverse trematode assemblages. After stringent quality control the workflow allows the identification of snails to species level, and of trematodes to taxonomic levels ranging from family to strain. It is sensitive to detect immature infections and changes in parasite biomass described in previous experimental studies. Co-infections were successfully identified, opening the possibility to examine parasite-parasite interactions such as interspecific competition. Altogether, these results demonstrate that our workflow provides a powerful tool to analyze the processes shaping trematode communities within natural snail populations.</p>
UMI SSU rRNA amplicon datasets
<p>Analysis of 721 SSU rRNA amplicon data from 58 stations in the Pacific Ocean. This item contains following files.</p> <p>abundance.csv</p> <p>- Read count of 155906 OTUs in UMI dataset</p> <p>centroid_seqs.fa</p> <p>- Centroid sequence of each OTU</p> <p>OTU_module_taxon.csv</p> <p>- Results of clustering by WGCNA analysis and assigned taxonomy using the SILVA database</p> <p>Thaumarchaeota.fasta</p> <p>- Sequenced used for phylogenetic analysis of Thaumarchaeota OTUs</p>
FASTA consensus sequences obtained using amplicon-based genome sequencing of SARS-CoV-2
<p>Set of 22 FASTA consensus sequences that were produced during routine SARS-CoV-2 sequencing obtained using amplicon-based sequencing (ARTIC protocol). Those sequences were compared to those generated in NASCarD applications.</p>
Data from: Pitfalls and pointers: an accessible guide to marker gene amplicon sequencing in ecological applications
Open the record for dataset details and reuse information.
Simultaneous genotyping of snails and infecting trematode parasites using high-throughput amplicon sequencing.
Open the record for dataset details and reuse information.
Raw Fast5 data for "Microbiota profiling with long amplicons using Nanopore sequencing: full-length 16S rRNA gene and the 16S-ITS-23S of the rrn operon" - PART I
<p>Raw Fast5 data for "Microbiota profiling with long amplicons using Nanopore sequencing: full-length 16S rRNA gene and the 16S-ITS-23S of the rrn operon". See Supplementary Table 2 for associating each sample to its barcode.</p> <p>- FC1_1 includes data for the HM mock community from BEI resources and skin microbiome of the chin in dogs.</p> <p>- FC1_2 includes data for the dorsal skin samples</p> <p>- FC2 includes data for the Zymobiomics mock community and Staphylococcus pseudintermedius isolate</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.