Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
390
datasets available to search
ShareScore release 0.9.0
Dataset results
390 results for “rDNA”
18S V4 rDNA sequences organized at the OTU level for the SOMLIT-Astan time-series (2009-2016)
<p>The present file includes metadata for each 18S V4<strong> rDNA OTU</strong> from the SOMLIT-Astan time series (2009-2016) including the following fields: <strong>amplicon</strong> = identifier of the representative (most abundant) sequence; <strong>total</strong> = total number of reads; <strong>spread </strong>= number of samples in which the OTU has been found; <strong>cloud </strong>= number of unique sequences constituting the OTU; <strong>sequence</strong> = nucleic acid sequence of the representative sequence; <strong>length</strong> = length of the representative sequence; <strong>quality </strong>= minimum expected error observed for the representative sequence, divided by sequence length; <strong> taxonomy</strong> = taxonomic path assigned to the representative sequence; <strong>identity</strong> = percentage of identity of the representative sequence to the closest reference sequence from PR2; <strong>references</strong> = best hit reference sequence(s) ; <strong>RA090107_02:RA161222_3 </strong>= 375 samples from January 2009 to December 2016, the first two number are the year followed by the month and the day (sampling twice a month during 8 years). Values after “_” indicate the size of the filter used for the filtration: 02 for 0.2 µm and 3 for 3 µm.</p> <p>Generation of 18S V4 rDNA Operational Taxonomic Units (OTUs) from the raw sequencing reads and their assembly into a OTUtable was obtained according to the following pipeline (https://doi.org/10.5281/zenodo.5791089). The V4 region was extracted from the 18S rDNA reference sequences from PR2 v4.12 (Guillou et al., 2013) with Cutadapt. The representative sequences of each OTU were compared to these V4 reference sequences by pairwise global alignment (usearch_global VSEARCH’s command). Each OTU inherits the taxonomy of the best hit or the last common ancestor in case of ties. OTUs with a score below 80% similarity were considered as unassigned (Mahé et al., 2017; Stoeck et al., 2010).</p> <p>The final dataset (filtered OTU table) contains 375 samples (sampled twice per month from 2009 to 2016) with a total of ~30 million sequence reads and 21,418 OTUs.</p>
EukRibo: a manually curated eukaryotic 18S rDNA reference database
<p>EukRibo is a manually curated database of reference small-subunit ribosomal RNA gene (18S rDNA) sequences of eukaryotes, specifically aimed at taxonomic annotation of high-throughput metabarcoding datasets. Unlike other reference databases of ribosomal genes, it is not meant to exhaustively capture all publicly available 18S rDNA sequences from the INSDC repositories, but to represent a subset of highly trustable sequences covering the whole known diversity of eukaryotes, with a focus on protists, manually verified taxonomic identifications, and relatively low genetic redundancy.</p> <p>EukRibo is part of a suite of public resources generated by the UniEuk project (www.unieuk.org), which are all designed to follow a common taxonomic framework for maximal interoperability. The high level of taxonomic accuracy of EukRibo, together with a newly designed, phylogenetically-informed annotation approach, allow high confidence in the taxonomic annotation of environmental metabarcodes, as well as identification of new eukaryotic diversity at various taxonomic levels using a connected components approach.</p> <p>* * *</p> <p>Accompanying preprint available at <a href="https://doi.org/10.1101/2022.11.03.515105">https://doi.org/10.1101/2022.11.03.515105</a>.</p> <p>* * *</p> <p><strong>EukRibo ReadMe file, versions 1 and 2</strong></p> <p>Each EukRibo release consists of <strong>4 files</strong>:<br> - a <strong>tsv table </strong>containing the taxonomic and other information about the 18S rDNA sequences included in the release<br> - a <strong>fasta file </strong>containing the <strong>full sequences </strong>as retrieved from the INSDC repositories (NCBI, EMBL-EBI/ENA, DDBJ)<br> - a <strong>fasta file </strong>containing the <strong>variable region V4 </strong>extracted from all these sequences (based on the fragment amplified with the Tara-Oceans V4 primers)<br> - a <strong>fasta file </strong>containing the <strong>variable region V9 </strong>extracted from the subset of sequences where it is present (based on the fragment amplified with the Tara-Oceans V9 primers)</p> <p>The primary goal of EukRibo was to be used to annotate the EukBank meta-dataset of available V4 metabarcoding datasets, and therefore all sequences included in EukRibo contain the variable region V4.<br> Only a subset of these sequences (about 75%) also contain the variable region V9; this is because many 18S rDNA sequences in the INSDC repositories stop before the V9 fragment.</p> <p>Sequences with slightly incomplete V4 or V9 fragments were kept if phylogenetically useful - i.e. if they are the only available representatives of a certain taxonomic lineage.<br> <strong>V4 </strong>We allowed up to 50 missing positions in the relatively conserved area at the 5' end of the V4 fragment (for an average fragment length of about 380 bp); no sequence incomplete at the 3' end of the V4 fragment is included.<br> <strong>V9 </strong>We allowed up to 30 missing positions in the relatively conserved area at the 3' end of the V9 fragment (for an average length of about 135 bp); no sequence incomplete at the 5' end of the V9 fragment is included.<br> We allowed a higher proportion of missing positions for the V9 region because being more conservative would imply losing too many sequences, including entire taxonomic lineages.</p> <p><strong>Version 1 of EukRibo</strong><br> This is the starting version of EukRibo that was used for the taxonomic annotation of the EukBank dataset, with taxonomy strings that were fixed as of October 2020.<br> - Contains 46,345 sequences with a sufficiently complete V4 region; 46,299 with the actual complete V4 region and 46 (about 0.1%) with missing positions at the 5' end.<br> - Of these, 34,438 also include a sufficiently complete V9 region; 23,226 with the actual complete V9 region and 11,206 (about 33%) with missing positions at the 3' end.</p> <p><strong>Version 2 of EukRibo</strong><br> This is a version of EukRibo that was made taxonomically compatible with version 3 of the EukProt database (<a href="https://doi.org/10.1101/2020.06.30.180687">https://doi.org/10.1101/2020.06.30.180687</a>), with taxonomic revisions as of July 2022 as well as additional information on the included selection of sequences that was not provided in the tsv file of version 1.<br> - Contains the exact same selection of sequences as in version 1, with the addition of genus <em>Meteora</em>, the last remaining known supergroup-level eukaryotic lineage for which an 18S rDNA was not previously available. (The <em>Meteora </em>sequence contains the full V4 fragment but does not include a sufficiently complete V9 fragment.)<br> - Only 34,432 sequences with a sufficiently complete V9 region are now retained because of 6 previously unrecognised chimeric sequences where the V9 fragment does not originate from the same organism as the V4 fragment.</p> <p><strong>Files in EukRibo version 1</strong>:<br> 46345_EukRibo.tsv.gz<br> 46345_EukRibo_full_seqs.fas.gz<br> 46345_EukRibo_V4.fas.gz<br> 34438_EukRibo_V9.fas.gz</p> <p>The tsv file contains 6 columns:<br> <strong>gb_accession </strong>- INSDC accession number of the sequence<br> <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2 </strong>- binning of the taxa into strictly monophyletic clades of evolutionary and/or ecological significance<br> <strong>UniEuk_taxonomy_string </strong>- full UniEuk-compatible taxonomic annotation of the sequence<br> - an unlimited number of levels is allowed (going down to strain for isolated organisms or to clone for environmental sequences)<br> - informal names are used for phylogenetically supported clades without formal name<br> <strong>V9 </strong>- presence ('Y') or absence ('N') of a sufficiently complete V9 fragment in the sequence</p> <p><strong>Files in EukRibo version 2</strong>:<br> 46346_EukRibo-02.tsv.gz<br> 46346_EukRibo-02_full_seqs.fas.gz<br> 46346_EukRibo-02_V4.fas.gz<br> 34432_EukRibo-02_V9.fas.gz</p> <p>The tsv file now contains 12 columns:<br> <strong>gb_accession</strong>, <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2</strong>, <strong>UniEuk_taxonomy_string</strong><br> - same columns as in version 1<br> <strong>alternative_strain_names </strong>(new) - provides alternative strain/isolate names when known to help cross-linking genetic data coming from the same organism<br> <strong>V4 </strong>(new) - indicates whether the V4 fragment is complete ('yes - complete') or missing positions at the 5' end ('yes - partial')<br> <strong>V9 </strong>(emended content) - now contains more precise information than in version 1 about whether it is complete ('yes - complete'), missing positions at the 3' end ('yes - partial'), or was excluded, and the 6 possible reasons why ('no - missing', 'no - too incomplete', 'no - chimera', 'no - bad quality', 'no - deletion in V9', 'no - Ns in V9')<br> <strong>EukProt_ID_same_strain </strong>(new) - accession of EukProt datasets from the same isolate<br> <strong>EukProt_ID_different_strain </strong>(new) - accession of EukProt datasets from a different isolate of the same species<br> <strong>columns_modified_since_previous_version </strong>(new) - lists all of the 6 pre-existing columns that have a modified content compared to version 1<br> <strong>remarks </strong>(new) - additional information such as presence of an intron in the V9 fragment, taxonomic identity of the two parts of chimeric sequences, or the presence of Ns or a deletion in the V4 or the V9 fragment (but insufficient to warrant exclusion)</p>
PR2_V9, a SSU V9 rDNA reference database with functional annotations.
<p>The present data set provides a tab separated text file compressed in a gzip archive. The file includes 63,401 18S V9 rDNA reference sequences for 44,084 unique eukaryotic taxa and 9,759 16S V9 rDNA references sequences for 9,661 unique prokaryotic taxa. It includes the following fields : sequence = nucleic acid sequence of reference; lineage = taxonomic path of the reference sequence; refs = original accession numbers corresponding to the reference sequence; name = reference sequence identifier; taxogroup = high-taxonomic level assignation of the reference sequence. The file also includes six categories of functional annotations: (1) chloroplast: yes, presence of permanent chloroplast; no, absence of permanent chloroplast ; NA, undetermined. (2) symbiont (small partner): parasite, the species is a parasite; commensal, the species is a commensal; mutualist, the species is a mutualist symbiont, most often a microalgal taxon involved in photosymbiosis; no the species is not involved in a symbiosis as small partner; NA, undetermined. (3) symbiont (host): photo, the host species relies on a mutualistic microalgal photosymbiont to survive (obligatory photosymbiosis); photo_falc, same as photo, but facultative relationship; photo_klep, the host species maintains chloroplasts from microalgal prey(s) to survive; photo_klep_falc, same as photo_klep, but facultative; Nfix, the host species must interact with a mutualistic symbiont providing N2 fixation to survive; Nfix_falc, same as Nfix, but facultative; no, the species is not involved in any mutualistic symbioses; NA, undetermined. (4) silicification; yes, the species has a silicified skeleton; no, it does not; NA, undetermined. (5) calcification; yes, the species has a calcified skeleton; no, it does not; NA, undetermined. (6) strontification; yes, the species has a skeleton made of strontium; no, it does not; NA, undetermined.</p>
FIGURE 2 in A new species of Auriculostoma (Trematoda: Allocreadiidae) from the intestine of Brycon guatemalensis (Characiformes: Bryconidae) from the Usumacinta River Basin, Mexico, based on morphology and 28 S rDNA sequences, with a key to species of the genus
FIGURE 2. Scanning electron micrographs of Auriculostoma lobata n. sp. (A) Anterior end with oral sucker bearing muscular lobe on either side, genital atrium and ventral sucker. (B) Pair of muscular lobes with a short posterior ‘‘ free’ ’ end, and 4 apical dome-like papillae (white arrows). (C) Lateral view of muscular lobe. (D) Distribution of dome-like papillae over oral sucker, 6 anterior (white arrows), 4 on the inner surface (white triangle), and 5 on the outer surface (white arrows).
rDNA 18S V9 metabarcoding tables (Swarm) for Tara Oceans Expedition (2009-2013), including Tara Polar Circle Expedition (2013)
<p>Reads were grouped into OTUs using the following swarm-based pipeline: paired-end reads were merged with vsearch’s --fastq_mergepairs command (version 2.15.1, allowing for staggered reads; Rognes et al., 2016), and trimmed with cutadapt (version 3.0; Martin, 2011), keeping only reads containing both forward and reverse primers. After trimming, the expected error per read was estimated with vsearch’s command --fastq_filter and the option --eeout. Each sample was then de-replicated, i.e. strictly identical reads were merged, using vsearch’s command --derep_fulllength, and converted into fasta format. Clustering was performed at the sample level with swarm 3.0 using default parameters (Mahé et al., 2015). Prior to global clustering, individual fasta files (one per sample) were pooled and further dereplicated with vsearch. Files containing per-read expected error values were also dereplicated to retain only the lowest expected error for each unique sequence. Global clustering was performed with swarm (using the fastidious option). Cluster representative sequences were then searched for chimeras with vsearch’s command --uchime_denovo using default parameters (Edgar et al., 2011).</p> <p>Clustering results, expected error values, taxonomic assignments, and chimera detection results were used to build a “raw” occurrence table. Reads without primers, reads shorter than 32 nucleotides and reads with uncalled bases (“N”) were discarded. For a “filtered” occurrence table, non-chimeric sequences, sequences with an expected error per nucleotide below 0.0002, and clusters containing at least 2 reads were retained. Since primer trimming is not perfect, some sequences can still contain primer fragments or be excessively trimmed. These sub- or super-sequences were identified using vsearch and merged with their closest, most abundant perfectly trimmed sequence. Finally, occurrence patterns throughout our sample collection were used to further refine the occurrence table. Clusters that contain sub-clusters with only a single-nucleotide difference but with different ecological patterns (defined here as uncorrelated abundance values in at least 5% of the samples) were turned into distinct clusters (https://github.com/frederic-mahe/fred-metabarcoding-pipeline). On the other hand, clusters with similar sequences that had correlated abundance values in at least 95% of the samples, were merged using a re-implementation of lulu's method (Frøslev et al. 2017; https://github.com/frederic-mahe/mumu).</p> <p> </p>
Fig. 7 in Ultrastructure and 28S rDNA Phylogeny of Two Gregarines: Cephaloidophora cf. communis and Heliospora cf. longissima with Remarks on Gregarine Morphology and Phylogenetic Analysis
Fig. 7. Relative rates of molecular evolution in long-branch apicomplexans: SSU rDNA (white columns) and LSU rDNA (black columns), calculated as ratio of the length of the current branch to average branch length of the non-long-branch apicomplexans (see the text for more explanations). Relative rates of LSU rDNA evolution are lower than those of SSU rDNA, especially in gregarines.
Fig. 2 in Ultrastructure and 28S rDNA Phylogeny of Two Gregarines: Cephaloidophora cf. communis and Heliospora cf. longissima with Remarks on Gregarine Morphology and Phylogenetic Analysis
Fig. 2. Light microscopy of the gregarine studied: free individuals (gamonts) of Cephaloidophora cf. communis (A, common light microsopy; B, DIC microscopy); a free gamont (C) and a syzygy (D) of Heliospora cf. longissima. Epimerite (ep), promerite (pr), deutomerite (de), septum between poto- and deutomerite (s1), and septum between proto- and epimerite (s2) are visible.
Fig. 1 in Ultrastructure and 28S rDNA Phylogeny of Two Gregarines: Cephaloidophora cf. communis and Heliospora cf. longissima with Remarks on Gregarine Morphology and Phylogenetic Analysis
Fig. 1. Layout of ribosomal operon fragment amplifications. Up- per part, schematic ribosomal operon with approximate positions of the direct and reverse primers used. Lower part, the amplified fragments of ribosomal DNA aligned with the ribosomal operon (above). Numbers indicate the length of the overlapping regions. Roman numerals denote the fragments discussed in this paper. SSU rDNA fragments analyzed previously by Rueckert et al. (2011b) have no numerical designations.
Fig. 7 in Comprehensive Analysis of the Jellyfish (Goette, 1886) (Semaeostomeae: Pelagiidae) with Description of the Complete rDNA Sequence.
Fig. 7. Nucleotide divergences of the cnidarians 18S and 28S rDNAs (datasets used in Table 1) based on corrected p-distances. Genetic distances between each paired sequence were calculated by the Kimura 2-parameter model, where a total of 16 cnidarian species were compared. Statistical analysis showed that the 18S rDNA divergences were significantly different from those of 28S rDNA (Student t-test, P <0.05, N = 66).
Fig. 6 in Comprehensive Analysis of the Jellyfish (Goette, 1886) (Semaeostomeae: Pelagiidae) with Description of the Complete rDNA Sequence.
Fig. 6. Phylogenetic relationships of the family Pelagiidae, including the genera Chrysaora, Pelagia and Sanderia, inferred from 18S rDNA (A), 28S rDNA (B) and morphological characters (C), which were redrawn from Fig. 95 in Morandini and Marques (2010). Phylogenetic trees of the rDNAs were constructed using the maximum-likelihood (ML) algorithms with the GTR+G model. A jellyfish Cyanea capillata (the family Cyaneidae) was used as the outgroup. Additional Bayesian trees generated similar branch patterns. The first and second numbers at the nodes display bootstrap proportions (BP) and posterior probabilities (PP) obtained in the ML and Bayesian analyses, respectively. Branch lengths are proportional to the scale given. Thick lines represent congruent branches between 18S and 28S, and morphological systematics.
Fig. 4. A in Comprehensive Analysis of the Jellyfish (Goette, 1886) (Semaeostomeae: Pelagiidae) with Description of the Complete rDNA Sequence.
Fig. 4. A dot matrix comparison of rDNA sequences between Chrysaora pacifica (KY 212123) and Aurelia coerulea (EU276014). Color scale bars represent consecutive sequence length of some regions detected similarly between the two sequence pairs. The open boxes in matrices indicate rDNA coding regions such as 18S, 5.8S, and 28S.
Fig. 2. A in Comprehensive Analysis of the Jellyfish (Goette, 1886) (Semaeostomeae: Pelagiidae) with Description of the Complete rDNA Sequence.
Fig. 2. A schematic representation of the single unit of rDNA (A), and GC content (%), nucleic acid distribution (% thymine), sequence complexity, and entropy (dS) in 100-bp windows across the entire rDNA nucleotides of Chrysaora pacifica (B). In the full rDNA (A), solid boxes indicate the ribosomal RNA genes and thin lines represent ITS or IGS. Nucleotide sequences in length and GC composition of each locus are represented near a line by calculation from a single unit of rDNA. The putative transcription start site is represented by an arrow; solid inverted-triangles represent sub-repeats in IGS.
Fig. 5. Phylogenetic relationships between jellyfishes within the order Semaeostomeae inferred from nearly complete 18S in Comprehensive Analysis of the Jellyfish (Goette, 1886) (Semaeostomeae: Pelagiidae) with Description of the Complete rDNA Sequence.
Fig. 5. Phylogenetic relationships between jellyfishes within the order Semaeostomeae inferred from nearly complete 18S rDNA (A) and partial 28S rDNA sequences (B) with maximum-likelihood (ML) algorithms. ML analyses of 18S and 28S were used as the nucleotide substitution model of GTR+G. Two hydrozoans (Hydractinia echinata and Podocoryne carnea for 18S rDNA; Astrohydra japonica and Melicertissa sp. for 28S) were included as the outgroups. Additional Bayesian analysis generated similar topology of the tree compared with the ML tree. Posterior probabilities (PP) from the analyses were incorporated into the ML tree to support the strength of each branch. The first and second numbers at the nodes display bootstrap proportions (BP) (> 50%) in ML and PP (> 0.50) in Bayesian, respectively. Branch lengths are proportional to the scale given. *Represents controversial species names, because they were suspected as different species by Bayha et al. (2017).
Fig. 1 in Comprehensive Analysis of the Jellyfish (Goette, 1886) (Semaeostomeae: Pelagiidae) with Description of the Complete rDNA Sequence.
Fig. 1. Live Chrysaora pacifica in natural habitat: basolateral (A and B), lateral (C) and apical view (D).
Fig. 3. Maximum likelihood tree constructed from 38 nuclear rDNA ITS1 and ITS2 sequences from Apiaceae genus Daucus and relatives using a in Molecular phylogeny of Daucus (Apiaceae): Evidence from nuclear ribosomal DNA ITS sequences
Fig. 3. Maximum likelihood tree constructed from 38 nuclear rDNA ITS1 and ITS2 sequences from Apiaceae genus Daucus and relatives using a transition/transversion rate ratio of 1.6. Branch lengths are proportional to the number of expected nucleotide substitutions per site.
Fig. 4 in The distribution and host-association of a haemoparasite of damselfishes (Pomacentridae) from the eastern Caribbean based on a combination of morphology and 18S rDNA sequences
Fig. 4. Phylogenetic analysis of the Haemohormidium-like parasite based on 18S rDNA sequences. Bayesian inference (BI) analysis showing the phylogenetic relationships for 8 Haemohormidium-like parasite isolates, 6 from the present study (GenBank: MH401637-42) (in bold) and 2 from Renoux et al. (2017), isolated from three species of Stegastes including Stegastes adustus, Stegastes diencaeus and Stegastes planifrons, from 5 sites in the eastern Caribbean. Comparative sequences representing known coccidia, with Adelina dimidiata (DQ096835) as outgroup, were downloaded from the GenBank database. Nodal support values> 50% are represented on the tree.
Fig. 2 in The distribution and host-association of a haemoparasite of damselfishes (Pomacentridae) from the eastern Caribbean based on a combination of morphology and 18S rDNA sequences
Fig. 2. Peripheral blood stages of the Haemohormidium-like parasite infecting species of Stegastes. Giemsa stained light micrographs of the Haemohormidium-like parasite as observed in the peripheral blood of Stegastes diencaeus from St Thomas, eastern Caribbean (Genbank accession number MH401641). A. rare possible trophozoite stage. B. possible meront stages undergoing transverse binary fission. C. possible meront stages undergoing longitudinal binary fission. Scale bar = 10 μm.
Fig. 3 in The distribution and host-association of a haemoparasite of damselfishes (Pomacentridae) from the eastern Caribbean based on a combination of morphology and 18S rDNA sequences
Fig. 3. Prevalence of infection differences among six Stegastes spp., averaged across six study sites. 95% confidence intervals calculated using the Wilson procedure with continuity corrections. Different lower-case letters above each bar indicates a significant (p ≤ 0.05) difference between species, as indicated by a binomial logistic regression (GLMM results shown in Table 1).
Fig. 1 in The distribution and host-association of a haemoparasite of damselfishes (Pomacentridae) from the eastern Caribbean based on a combination of morphology and 18S rDNA sequences
Fig. 1. Map of the Eastern Caribbean region showing collection sites for the current study and Cook et al., 2015.
Lecanora markjohnstonii rDNA tree
<p>This is a phylogenetic tree comprising 75 rDNA contigs of lichen mycobionts. The tree was used to infer placement of the new species Lecanora markjohnstonii from a broad sampling of lichenized fungi.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.