Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

33

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

33 results for “OTU”

Learn how ShareScore rates datasets ↗
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Taxonomic Profiling, OTU Tables, and Visualisation

<p>Data and conda software environment file for the chapter &#39;Taxonomic Profiling, OTU Tables, and Visualisation&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

18S V4 rDNA sequences organized at the OTU level for the SOMLIT-Astan time-series (2009-2016)

<p>The present file includes metadata for each 18S V4<strong> rDNA OTU</strong> from the SOMLIT-Astan time series (2009-2016) including the following fields: <strong>amplicon</strong> = identifier of the representative (most abundant) sequence; <strong>total</strong> = total number of reads; <strong>spread </strong>= number of samples in which the OTU has been found; <strong>cloud </strong>= number of unique sequences constituting the OTU;&nbsp; <strong>sequence</strong> =&nbsp; nucleic acid sequence of the representative sequence; <strong>length</strong> = length of the representative sequence; <strong>quality </strong>= minimum expected error observed for the representative sequence, divided by sequence length;&nbsp;<strong> taxonomy</strong> = taxonomic path assigned to the representative sequence; <strong>identity</strong> = percentage of identity of the representative sequence to the closest reference sequence from PR2; <strong>references</strong> = best hit reference sequence(s) ;&nbsp; <strong>RA090107_02:RA161222_3 </strong>= 375 samples from January 2009 to December 2016, the first two number are the year followed by the month and the day (sampling twice a month during 8 years). Values after &ldquo;_&rdquo; indicate the size of the filter used for the filtration: 02 for 0.2 &micro;m and 3 for 3 &micro;m.</p> <p>Generation of 18S V4 rDNA Operational Taxonomic Units (OTUs) from the raw sequencing reads and their assembly into a OTUtable was obtained according to the following pipeline (https://doi.org/10.5281/zenodo.5791089). The V4 region was extracted from the 18S rDNA reference sequences from PR2 v4.12 (Guillou et al., 2013) with Cutadapt. The representative sequences of each OTU were compared to these V4 reference sequences by pairwise global alignment (usearch_global VSEARCH&rsquo;s command). Each OTU inherits the taxonomy of the best hit or the last common ancestor in case of ties. OTUs with a score below 80% similarity were considered as unassigned (Mah&eacute; et al., 2017; Stoeck et al., 2010).</p> <p>The final dataset (filtered OTU table) contains 375 samples (sampled twice per month from 2009 to 2016) with a total of ~30 million sequence reads and 21,418 OTUs.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

HEE Burn-Control Site OTU FASTA

<p>Paper: Diversity and composition of fungal soil communities across prescribed burn areas in temperate hardwood forests</p> <p>Authors: S.D. Russell &amp; M.C. Aime</p> <p>FASTA file containing the sequences of each OTU that was recovered from the burn and control survey areas using the methodology described in the paper above.</p>

opencc-by-4.0Feb 2020View details →
zenodo40/100

OTU level 16s sequence data for "Algae drive convergent bacterial community assembly when nutrients are scarce"

<p>16s sequence data at the OTU level for the experiments conducted in&nbsp;&quot;Algae drive convergent bacterial community assembly when nutrients are scarce&quot;</p> <p>The file is in fasta format, which can be read by many software packages including biopython, R, and SILVA&#39;s alignment, classification and tree service.</p> <p>The sequence ids can be used to find the phylogeny and OTU ID of the sequences from Supplementary dataset 4.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

OTU table - Joli et al. Scientific Reports - Janus Gateway

<p>This texte file represent the OTU table from the paper Joli et al. published in Scientific Reports as &quot;Need for focus on microbial species following ice melt and changing freshwater regimes in a Janus Arctic Gateway&quot;. The first column contains the OTU number and each following columns contain the number of reads for each sample. The last column gives the taxonomy after confronting dataset to a eukaryotic 18S database.</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

Supplement data for the article "The impact of OTU sequence similarity threshold on diatom-based bioassessment: A case study of the rivers of Mayotte (France, Indian Ocean)", in preparation

<p>These are supplement data for the article &quot;The impact of OTU sequence similarity threshold on diatom-based bioassessment: A case study of the rivers of Mayotte (France, Indian Ocean)&quot;, in preparation</p> <p>The folowing files are available:</p> <ul> <li>Supplement 1. Map of Mayotte with the sampling sites and the rivers.</li> <li>Supplement 2. <em>rbcL</em> primers, reaction mixture, and conditions used for the PCR of the 312-bp <em>rbcL</em> fragment. The information provided is for a single reaction with a final volume of 25&micro;L.</li> <li>Supplement 3. The 20 fastq files containing the demultiplexed DNA reads.</li> <li>Supplement 4. Number of sequence reads for each sample before and after the trimming procedure.</li> <li>Supplement 5. The 20 OTU lists, corresponding to the 20 SSTs, including the number of DNA reads within the 90 samples and their assigned taxonomy.</li> <li>Supplement 6. Sampling site description with sample codes, names of rivers, year, number of raw DNA reads and GPS coordinates.</li> <li>Supplement 7. Values and summary statistics for the environmental variables.</li> <li>Supplement 8. The script used in Mothur for the bioinformatic analysis from trimming to the used OTU lists.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Sep 2018View details →
zenodo40/100

OTU table of genus from Tibet WWTPs

<p>Activated sludge was collected from two municipal wastewater treatment plants (WWTPs) located in a high altitude Plateau in Tibet, China (~3650 m above the sea level). T1 is applied for cyclic activated sludge system (CASS), while T2 is operated at anaerobic-anoxic-aerobic (A2O) process. DNA samples were taken from bioreactor of T1, aerobic and anaerobic bioreactors of T2 (marked as T2AE and T2AN). The 16S rRNA gene amplicons of V4-V5 regions that were sequenced on an Illumina HiSeq PE250 platform, and then annotated using the 16S-Silva database. Data present here are normalized OTU relative number of each sample are presented at the taxonomic levels of genus.</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

Tara Prokaryote Annotated 16s OTU Counts

<p>&quot;Tara Oceans systematically collected ~35,000 samples for morphological, genetic, and environmental analyses using standardized protocols across multiple depths at global scale, aiming to facilitate a holistic study on how environmental factors and biogeochemical cycles affect oceanic life.&quot; (1) &quot;For each prokaryote-enriched sample (N=139), we extracted metagenomic merged Illumina reads (miTAGs) that contained signatures of the 16S rRNA gene (Logares et al. 2013). These fragments were mapped to a set of 16S reference sequences that were downloaded from the SILVA database (Release 115: Quast et al. 2013) and clustered into 97% operational taxonomic units. The OTU count table was summarized at different taxonomic levels.&quot; (2)<br> (1): https://www-science-org.offcampus.lib.washington.edu/doi/full/10.1126/science.1261359<br> (2): http://ocean-microbiome.embl.de/companion.html</p>

opencc-by-4.0Jul 2023View details →
edi40/100

Operational taxonomic unit (OTU) table characterizing water track and adjacent soil microbial communities in Taylor Valley, Antarctica during the 2012-13 austral summer

This data package includes the abundance of microbial operational taxonomic units (OTUs) for samples collected during the austral summer of 2012-2013 in the Lake Hoare and Goldman Glacier Basins of Taylor Valley, Antarctica. A total of twenty samples from on- and off-water track soils were collected and analyzed. Samples were collected from the Lake Hoare Basin on 27 December 2012 and from the Goldman Glacier Basin on 4 January 2013. The aim of the study was to identify how variation in the measured physical and chemical environment of water tracks within the two water track systems influenced soil microbial community structure and diversity. Soil bacterial biodiversity was assessed using cultivation independent 16S rRNA gene sequencing.

openOpenDec 2020View details →
zenodo36/100

OTU-Taxid Mapping File for gg_13_5_99_otus tree

<p>File containing mapping information of 99_otus tree from gg_13_5. First column: OTU. Second column: Taxid. Third Column: accession.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

OTUs with valid matched taxid on gg_13_15 99_otu tree

<p>A mapping file between OTUs and Taxids on the 99_otus tree of gg_13_5 data package, for reproducibility of the&nbsp;WGSUniFrac project.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

Plankton Planet Pilot Project rDNA 18S V9 OTU tables

<p>This repository contains the rDNA 18S V9 OTU table, its rarefied version, and the related contextual data from Plankton Planet Pilot Project.</p> <p>In the file <strong>P2_TO_18SV9_otu_table.tsv.gz</strong>, each OTU, one per row, is described by the following fields: <strong>amplicon</strong> = identifier of the representative (most abundant) sequence of the swarm; <strong>total</strong> = total number of reads in the entire dataset; <strong>cloud</strong> =&nbsp; number of unique sequences constituting the OTU; <strong>length</strong> = length of the representative sequence; <strong>spread</strong> = number of samples in which the OTU has been found; <strong>quality</strong> = minimum expected error observed for the representative sequence, divided by sequence length; <strong>sequence</strong> = nucleic acid sequence of the representative sequence; <strong>identity</strong> = percentage of identity of the representative sequence to the closest reference sequence from PR2_V9 (https://doi.org/10.5281/zenodo.3768951); <strong>references</strong> = best hit reference sequence(s); <strong>taxonomy</strong> = taxonomic path assigned to the representative sequence; <strong>taxogroup</strong> = high-taxonomic level assignation of the representative barcode; <strong>chloroplast</strong> = <em>yes</em>: presence of permanent chloroplast / <em>no</em>: absence of permanent chloroplast / <em>NA</em>: undetermined; <strong>symb_small</strong> = <em>parasite</em>: the species is a parasite / <em>commensal</em>: the species is a commensal / <em>mutualist</em>: the species is a mutualist symbiont, most often a microalgal taxa involved in photosymbiosis / <em>no</em>: the species is not involved in a symbiosis as small partner / <em>NA</em>: undetermined; <strong>symbiont_host</strong> = <em>photo</em>: the host species relies on a mutualistic microalgal photosymbiont to survive (obligatory photosymbiosis) / <em>photo_falc</em>: same as photo, but facultative relationship / <em>photo_klep</em>: the host species maintains chloroplasts from microalgal prey(s) to survive / <em>photo_klep_falc</em>: same as <em>photo_klep</em>, but facultative / <em>Nfix</em> = the host species must interact with a mutualistic symbiont providing N2 fixation to survive / <em>Nfix_falc</em> = same as Nfix, but facultative / <em>no</em>: the species is not involved in any mutualistic symbioses; <em>NA</em>: undetermined; <strong>silicification</strong> = <em>yes</em>: the species has a silicified skeleton / <em>no</em>: it does not / <em>NA</em>: undetermined; <strong>calcification</strong> = <em>yes</em>: the species has a calcified skeleton / <em>no</em>: it does not / <em>NA</em>: undetermined; <strong>strontification</strong> = <em>yes</em>: the species has a skeleton made of strontium / <em>no</em>: it does not / <em>NA</em>: undetermined; <strong>PPXXX</strong> = number of reads in each of the 214 Plankton Planet samples; <strong>TARA_XXXXXXXXXX</strong> = number of reads in each of the 386 <em>Tara</em> Oceans samples.</p> <p>The file <strong>P2_TO_18SV9_otu_table_raref_313539.tsv.gz</strong> contains the same fields but with number of reads (total and per sample) obtained after random subsampling (313,539 reads per sample).</p> <p>In the file <strong>P2_TO_18SV9_context.tsv.gz</strong>, each sample is described by the following fields: <strong>sample</strong> = identifier of the sample; <strong>lower_size_fraction</strong> = lower limit of the size fraction in &micro;m; <strong>upper_size_fraction</strong> = lower limit of the size fraction in &micro;m; <strong>event_date</strong> = date (year-month-day); <strong>event_latitude</strong> = geographic position (latitude in DD); <strong>event_longitude</strong> = geographic position (longitude in DD); <strong>depth</strong> = depth in meters; <strong>temperature</strong> = sea water temperature in &deg;C</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

OTU table, lefse result and ARGs abundance data.

<ul> <li> <p><strong>Table_otu.raw.txt.xls</strong>: OTU table processed by Uparse(97% similarity) on the&nbsp;<a href="http://cloud.magigene.com/" rel="nofollow">Magigene Cloud Platform</a>. Chloroplasts, mitochondria, Unclassified, archaea otus were removed.</p> </li> <li> <p><strong>rpkm.type.txt</strong>: Abundance of Antibiotic-resistant genes (ARGs) grouped by types processed by&nbsp;<a href="https://github.com/xinehc/args_oap">ARGs-OAP v3.0</a></p> </li> <li> <p><strong>rpkm.subtype.txt</strong>: ARGs grouped by subtypes processed by&nbsp;<a href="https://github.com/xinehc/args_oap">ARGs-OAP v3.0</a></p> </li> <li> <p><strong>rpkm.genes.txt</strong>: ARGs grouped by genes processed by&nbsp;<a href="https://github.com/xinehc/args_oap">ARGs-OAP v3.0</a></p> </li> <li> <p><strong>lefse_results.txt.xls</strong>: lefse analysis result produced by&nbsp;<a href="http://cloud.magigene.com/" rel="nofollow">Magigene Cloud Platform</a>.</p> </li> <li> <p><strong>otus.fa</strong>: Representative otu sequences processed by Uparse(97% similarity) on the&nbsp;<a href="http://cloud.magigene.com/" rel="nofollow">Magigene Cloud Platform</a>. Chloroplasts, mitochondria, Unclassified, archaea otus were removed.</p> </li> <li> <p><strong>otus_aligned.fasta</strong>: Aligned otu sequences by mafft.</p> <p><code>Bash: mafft otus.fa &gt; otus_aligned.fa </code></p> </li> <li> <p><strong>otus_fasttree.tre</strong>: Construct phylogenic tree from&nbsp;<code>otus_aligned.fa</code> by fasttree.</p> </li> </ul>

opencc-by-4.0Jun 2024View details →
dryad36/100

Achieving bio-protection in New Zealand ecosystems mesocosm fungal pathogen OTU table

<p>We established 80 experimental ecosystems (mesocosms), manipulated interactions between plants and soil biota in a fully factorial design. Each mesocosm was grown in a 125 L pot (575 mm diameter), and comprised one of 20 unique, eight-species plant communities varying orthogonally in the proportion of exotic and woody shrub/tree species (0-100% and 0-63%, respectively). These plants were taken from a pool of 20 exotic and 19 native/endemic New Zealand plant species. Soil biota were manipulated using a modified plant-soil feedback approach, where each plant species was grown in monoculture in 10 L pots containing field-collected soil for 9-10 months, allowing the conditioning of typical associated soil biota for each of the plant species. We created 'home' soils by taking the conditioned soil from each of the eight representative species in a mesocosm and mixing it together to create a single inoculum. Each 'home' soil mixture was also used as an 'away soil' in a different mesocosm that did not contain any of the representative plants in that inoculum. These soils were intended to increase the relative biomass in inocula of specialized and preferred interaction partners of the resident (or non-resident) plant species. After approximately one year of growth, we harvested all plants from each mesocosm, took root samples from each individual plant (n=491), extracted DNA and sequenced the fungi in the roots. Fungal sequences were paired and clustered into operational taxonomic units (OTUs) at 97% similarity. We assigned functional attributes to fungal OTUs using the FUNGUILD database and retained only the taxa assigned as "probable" or "highly probable" plant pathogens.</p>

opencc-zeroJul 2024View details →
zenodo36/100

OTU sequences - Joli et al. Scientific Reports - Janus Gateway

<p>Fasta file representing the sequences of the representative OTUs from the paper Joli et al. published Scientific Reports as&nbsp;<strong>Need for focus on microbial species following ice melt and changing freshwater regimes in a Janus Arctic Gateway.&nbsp;</strong></p> <p>Those sequences have been selected based on 99% identity between reads.</p>

opencc-by-4.0Mar 2018View details →
zenodo36/100

OTU table: Ecosystems and Networks Integrated with Genes and Molecular Assemblies (ENIGMA)

<p>This dataset contains the OTU table and associated metadata from the&nbsp;ENIGMA study which was used in &quot;A Practical Guide to Methods Controlling False Discoveries in Computational Biology&quot; (Korthauer, K. and Kimes, P., et al.&nbsp;2018; associated github: https://github.com/pkimes/benchmark-fdr/).</p> <p>The original raw data is available on MG-RAST (project mgp8190):&nbsp;https://www.mg-rast.org/mgmain.html?mgpage=project&amp;project=mgp8190</p> <p>These data were processed as described in Korthauer &amp; Kimes et al, 2018. The OTU table was provided by Renmao Tian and generated&nbsp;with the following pipeline:&nbsp;<a href="http://zhoulab5.rccc.ou.edu/pipelines/ASAP_web/pipeline_asap.php">http://zhoulab5.rccc.ou.edu/pipelines/ASAP_web/pipeline_asap.php</a>.&nbsp;</p> <p>More information about the ENIGMA project can be found at&nbsp;http://enigma.lbl.gov/</p>

opencc-by-4.0Oct 2018View details →
zenodo36/100

Lake Hazen watershed OTU table

<p>This is a processed OTU table used for all analyses featured in a manuscript whose publication is currently pending at the FEMS Journal. This table includes OTU data for three main watershed compartments found within the Lake Hazen watershed (Northern Ellesmere Island, NU, Canada), comprising of: glacial rivers, Lake Hazen itself,&nbsp;an active layer thaw-fed continuum, as well as snow and some snowmelt samples obtained from the same watershed.</p>

opencc-by-4.0Aug 2019View details →
dryad36/100

Raw sequence data and OTU tables of soil microorganisms obtained across a summit in the Lesotho highlands

<p>Mountain regions represent unique environments characterized by strong topographical diversity which drive climatic and environmental variability within these environments. These regions thus provide an opportunity to explore the relationships between various environmental factors and soil microorganisms. In this study, we investigated the impact of micro-topographical (i.e., north/south-facing slope aspects and flat plateau between them) variations on microbial diversity and community structures across a Lesotho mountain summit.</p> <p>Raw sequenced data were generated using the Illumina MiSeq platform on DNA extracted from soil samples collected across the plateau, north- and south-facing slopes. This data was then used for taxonomic classification of the bacterial and fungal OTUs for the determination of the alpha- and beta-diversity across the slopes. These analyses revealed that a relatively greater bacterial and fungal diversity could be observed for the north-facing slope compared to the south-facing slope and plateau. While there was no difference in group variance of bacterial and fungal community structures across the plateau, north- and south-facing slopes.</p> <p>Multiple comparison analyses were conducted to determine the impact of various abiotic and geographical factors on bacterial and fungal diversity and community structures. These analyses indicated that the slope aspect significantly affects bacterial and fungal community structures at this location. These results provide an original insight into soil microbial diversity in the Lesotho highlands and offer an opportunity to investigate the response of soil microorganisms to changes in environmental and climatic factors in highly variable mountain environments such as the Lesotho highlands.</p>

opencc-zeroFeb 2023View details →
zenodo36/100

18S rDNA OTU table of fungal community in a tropical forest

<p>This study aims to elucidate how fungal community responses to N deposition</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Tara Eukaryote Annotated 18s OTU Counts

<p>&quot;Marine plankton support global biological and geochemical processes. Surveys of their biodiversity have hitherto been geographically restricted and have not accounted for the full range of plankton size. We assessed eukaryotic diversity from 334 size-fractionated photic-zone plankton communities collected across tropical and temperate oceans during the circumglobal Tara Oceans expedition. We analyzed 18S ribosomal DNA sequences across the intermediate plankton-size spectrum from the smallest unicellular eukaryotes (protists, &gt;0.8 micrometers) to small animals of a few millimeters. Eukaryotic ribosomal diversity saturated at ~150,000 operational taxonomic units, about one-third of which could not be assigned to known eukaryotic groups... Most eukaryotic plankton biodiversity belonged to heterotrophic protistan groups.&quot;</p> <p>https://www.science.org/doi/10.1126/science.1261605</p>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record