Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
FIGURE 2. Adelphydraena amazonica n in Three new species and DNA sequence data of the rare South American water beetle genus Adelphydraena Perkins, 1989 (Coleoptera: Hydraenidae)
FIGURE 2. Adelphydraena amazonica n. sp., habitus of holotype.
Data from: Deciduous dentition and dental eruption sequence in Interatheriinae (Notoungulata, Interatheriidae): implications in the systematics of the group
Studies focused on deciduous dentition, ontogenetic series, and tooth eruption and replacement patterns in fossil mammals have lately increased due to the recognized taxonomic and phylogenetic weight of these aspects. A study of the deciduous and permanent dentition of Interatherium and Protypotherium (Interatheriinae) is presented, mainly based on unpublished materials. Deciduous cheek teeth are brachydont and placed covering the apex of the respective permanent tooth; in addition, some morphological and metrical differences are observed along the crown height. Five dental ontogenetic stages are distinguished among the juvenile specimens, based on the degree of wear, the replacement of the deciduous premolars, and the eruption of the molars. The crown height and the wear degree of different Interatheriinae taxa show: (1) eruption pattern of molars in an anterior-posterior direction (M/m1 to M/m3); (2) pattern of replacement of deciduous premolars and eruption of permanent premolars in a posterior-anterior direction (dP/dp4 to dP/dp2 and P/p4 to P/p2); and (3) eruption of M/m3 before the replacement of dP/dp4. Results allow evaluating the diagnostic dental characteristics used to describe some interatheriines, as well as reinterpreting some taxonomic assumptions: the holotype of Protypotherium diversidens is recognized as a juvenile of another species of the genus and the species is not validated, considering it as Protypotherium sp.; the holotype of Eudiastatus lingulatus falls in the variability of Protypotherium, becoming P. lingulatus, tentatively maintaining the species, and implying the synonymy between Eudiastatus and Protypotherium; and the holotype of Eopachyrucos ranchoverdensis is reinterpreted as bearing deciduous premolars.
Differential analysis of binarized single-cell RNA sequencing data captures biological variation
<p>Processed datasets used for binary differential analysis experiments.</p>
Data from: Robust DNA isolation and high-throughput sequencing library construction for herbarium specimens
Herbaria are an invaluable source of plant material that can be used in a variety of biological studies. The use of herbarium specimens is associated with a number of challenges including sample preservation quality, degraded DNA, and destructive sampling of rare specimens. In order to more effectively use herbarium material in large sequencing projects, a dependable and scalable method of DNA isolation and library preparation is needed. This paper demonstrates a robust, beginning-to-end protocol for DNA isolation and high-throughput library construction from herbarium specimens that does not require modification for individual samples. This protocol is tailored for low quality dried plant material and takes advantage of existing methods by optimizing tissue grinding, modifying library size selection, and introducing an optional reamplification step for low yield libraries. Reamplification of low yield DNA libraries can rescue samples derived from irreplaceable and potentially valuable herbarium specimens, negating the need for additional destructive sampling and without introducing discernible sequencing bias for common phylogenetic applications. The protocol has been tested on hundreds of grass species, but is expected to be adaptable for use in other plant lineages after verification. This protocol can be limited by extremely degraded DNA, where fragments do not exist in the desired size range, and by secondary metabolites present in some plant material that inhibit clean DNA isolation. Overall, this protocol introduces a fast and comprehensive method that allows for DNA isolation and library preparation of 24 samples in less than 13 hours, with only 8 hours of active hands-on time with minimal modifications.
Data from: High throughput sequencing reveals alterations in the recombination signatures with diminishing spo11 activity
Spo11 is the topoisomerase-like enzyme responsible for the induction of the meiosis-specific double strand breaks (DSBs), which initiate the recombination events responsible for proper chromosome segregation. Nineteen PCR-induced alleles of SPO11 were identified and characterized genetically and cytologically. Recombination, spore viability and synaptonemal complex (SC) formation were decreased to varying extents in these mutants. Arrest by ndt80 restored these events in two severe hypomorphic mutants, suggesting that ndt80-arrested nuclei are capable of extended DSB activity. While crossing-over, spore viability and synaptonemal complex (SC) formation defects correlated, the extent of such defects was not predictive of the level of heteroallelic gene conversions (prototrophs) exhibited by each mutant. High throughput sequencing of tetrads from spo11 hypomorphs revealed that gene conversion tracts associated with COs are significantly longer and gene conversion tracts unassociated with COs are significantly shorter than in wild type. By modeling the extent of these tract changes, we could account for the discrepancy in genetic measurements of prototrophy and crossover association. These findings provide an explanation for the unexpectedly low prototroph levels exhibited by spo11 hypomorphs and have important implications for genetic studies that assume an unbiased recovery of prototrophs, such as measurements of CO homeostasis. Our genetic and physical data support previous observations of DSB-limited meioses, in which COs are disproportionally maintained over NCOs (CO homeostasis).
Data from: Prevention, diagnosis, and treatment of high-throughput sequencing data pathologies
High Throughput Sequencing (HTS) technologies generate millions of sequence reads from DNA/RNA molecules rapidly and cost-effectively, enabling single investigator laboratories to address a variety of "omics" questions in non-model organisms, fundamentally changing the way genomic approaches are used to advance biological research. One major challenge posed by HTS is the complexity and difficulty of data quality control (QC). While QC issues associated with sample isolation, library preparation, and sequencing are well known and protocols for their handling are widely available, the QC of the actual sequence reads generated by HTS is often overlooked. HTS-generated sequence reads can contain various errors, biases, and artefacts whose identification and amelioration can greatly impact subsequent data analysis. However, a systematic survey on QC procedures for HTS data is still lacking. In this review, we begin by presenting standard "health check-up" QC procedures recommended for HTS datasets and establishing what "healthy" HTS data look like. We next proceed by classifying errors, biases and artifacts present in HTS data into three major types of "pathologies", discussing their causes and symptoms, and illustrating with examples their diagnosis and impact on downstream analyses. We conclude this review by offering examples of successful "treatment" protocols and recommendations on standard practices and treatment options. Notwithstanding the speed with which HTS technologies–and consequently their pathologies–change, we argue that careful QC of HTS data is an important–yet often neglected–aspect of their application in molecular ecology, and lay the groundwork for developing a HTS data QC "best practices" guide.
Data from: Low-coverage, whole-genome sequencing of Artocarpus camansi (Moraceae) for phylogenetic marker development and gene discovery
Premise of the study: We used moderately low-coverage (17×) whole-genome sequencing of Artocarpus camansi (Moraceae) to develop genomic resources for Artocarpus and Moraceae. Methods and Results: A de novo assembly of Illumina short reads (251,378,536 pairs, 2 × 100 bp) accounted for 93% of the predicted genome size. Predicted coding regions were used in a three-way orthology search with published genomes of Morus notabilis and Cannabis sativa. Phylogenetic markers for Moraceae were developed from 333 inferred single-copy exons. Ninety-eight putative MADS-box genes were identified. Analysis of all predicted coding regions resulted in preliminary annotation of 49,089 genes. An analysis of synonymous substitutions for pairs of orthologs (Ks analysis) in M. notabilis and A. camansi strongly suggested a lineage-specific whole-genome duplication in Artocarpus. Conclusions: This study substantially increases the genomic resources available for Artocarpus and Moraceae and demonstrates the value of low-coverage de novo assemblies for nonmodel organisms with moderately large genomes.
Data from: Plastid genome sequences of legumes reveal parallel inversions and multiple losses of rps16 in papilionoids
To date, publicly available plastid genomes of legumes have for the most part been limited to the subfamily Papilionoideae. Here we report 13 new plastid genomes of legumes spanning all three subfamilies. The genomes representing Caesalpinioideae and Mimosoideae are highly conserved in gene content and gene order, similar to the ancestral angiosperm genome organization. Genomes within the Papilionoideae, however, have reduced sizes due to deletions in nine intergenic spacers primarily in the large single copy region. Our study also indicates that rps16 has been independently lost at least five times in legumes, with additional gene and intron losses scattered among the papilionoids. Additionally, genera from two distinct lineages within the papilionoids, Lupinus and Robinia, have a parallel inversion of 36 kb and 39 kb, respectively. This parallel inversion is novel as it appears to be caused by a 29 bp repeat within two trnS genes. This repeat is present in all available legume plastid genomes indicating that there is the potential for this inversion to be present in more species. This case of a homoplasious inversion is also evidence that some inversion events may not be reliable phylogenetic markers.
Data from: Lineage-specific sequence evolution and exon edge conservation partially explain the relationship of evolutionary rate and expression level in A. thaliana
Rapidly evolving proteins can aid the identification of genes underlying phenotypic adaptation across taxa, but functional and structural elements of genes can also affect evolutionary rates. In plants, the 'edges' of exons, flanking intron junctions, are known to contain splice enhancers and to have a higher degree of conservation compared to the remainder of the coding region. However, the extent to which these regions may be masking indicators of positive selection or account for the relationship between dN/dS and other genomic parameters is unclear. We investigate the effects of exon edge conservation on the relationship of dN/dS to various sequence characteristics and gene expression parameters in the model plant Arabidopsis thaliana. We also obtain lineage-specific dN/dS estimates, making use of the recently sequenced genome of Thellungiella parvula, the second closest sequenced relative after the sister species Arabidopsis lyrata. Overall, we find that the effect of exon edge conservation, as well as the use of lineage-specific substitution estimates, upon dN/dS ratios partly explains the relationship between the rates of protein evolution and expression level. Furthermore, the removal of exon edges shifts dN/dS estimates upwards, increasing the proportion of genes potentially under adaptive selection. We conclude that lineage-specific substitutions and exon edge conservation have an important effect on dN/dS ratios and should be considered when assessing their relationship with other genomic parameters..
Sequencing data from: High levels of primary biogenic organic aerosols are driven by only a few plant-associated microbial taxa
<p>Primary biogenic organic aerosols (PBOA) represent a major fraction of coarse organic matter (OM) in air. Despite their implication in many atmospheric processes and human health problems, we surprisingly know little about PBOA characteristics (i.e., composition, dominant sources, and contribution to airborne-particles). In addition, specific primary sugar compounds (SCs) are generally used as markers of PBOA associated with bacteria and fungi but our knowledge of microbial communities associated with atmospheric particulate matter (PM) remains incomplete. This work aimed at providing a comprehensive understanding of the microbial fingerprints associated with SCs in PM<sub>10</sub> (particles smaller than 10µm) and their main sources in the surrounding environment (soils and vegetation). An intensive study was conducted on PM<sub>10</sub> collected at rural background site located in an agricultural area in France. We combined high-throughput sequencing of bacteria and fungi with detailed physicochemical characterization of PM<sub>10</sub>, soils and plant samples, and monitored meteorology and agricultural activities throughout the sampling period. Results shows that in summer SCs in PM<sub>10</sub> are a major contributor of OM in air, representing 0.8 to 13.5% of OM mass. SCs concentrations are clearly determined by the abundance of only a few specific airborne fungi and bacteria taxa. The temporal fluctuations in the abundance of only 4 predominant fungal genera, namely <i>Cladosporium</i>, <i>Alternaria</i>, <i>Sporobolomyces</i> and <i>Dioszegia</i> reflect the temporal dynamics in SC concentrations. Among bacteria taxa, the abundance of only <i>Massilia</i>, <i>Pseudomonas</i>, <i>Frigoribacterium</i> and <i>Sphingomonas</i> are positively correlated with SC species. These microbial are significantly enhanced in leaf over soil samples. Interestingly, the overall community structure of bacteria and fungi are similar within PM<sub>10</sub> and leaf samples and significantly distinct between PM<sub>10</sub> and soil samples, indicating that surrounding vegetation are the major source of SC-associated microbial taxa in PM<sub>10</sub> on rural area of France.</p>
Data from: Reconstructing phylogeny from reduced-representation genome sequencing data without assembly or alignment
Reduced-representation genome sequencing such as RADseq aids the analysis of genomes by reducing the quantity of data, thereby lowering both sequencing costs and computational burdens. RADseq was initially designed for studying genetic variation across genomes at the population level, but has also proved to be suitable for interspecific phylogeny reconstruction. RADseq data pose challenges for standard phylogenomic methods, however, due to incomplete coverage of the genome and large amounts of missing data. Alignment-free methods are both efficient and accurate for phylogenetic reconstructions with whole genomes and are especially practical for non-model organisms; nonetheless, alignment-free methods have not been applied with reduced genome sequencing data. Here, we test a full-genome assembly and alignment-free method, AAF, in application to RADseq data and propose two procedures for reads selection to remove reads from restriction sites that were not found in taxa being compared. We validate these methods using both simulations and real datasets. Reads selection improved the accuracy of phylogenetic construction in every simulated scenario and the two real datasets, making AAF as good or better than a comparable alignment-based method, even though AAF had much lower computational burdens. We also investigated the sources of missing data in RADseq and their effects on phylogeny reconstruction using AAF. The AAF pipeline modified for RADseq or other reduced-representation sequencing data, phyloRAD, is available on github (https://github.com/fanhuan/phyloRAD).
Data from: Quantifying sequence proportions in a DNA-based diet study using Ion Torrent amplicon sequencing: which counts count?
A goal of many environmental DNA barcoding studies is to infer quantitative information about relative abundances of different taxa based on sequence read proportions generated by high-throughput sequencing. However, potential biases associated with this approach are only beginning to be examined. We sequenced DNA amplified from faeces (scats) of captive harbour seals (Phoca vitulina) to investigate whether sequence counts could be used to quantify the seals' diet. Seals were fed fish in fixed proportions, a chordate-specific mitochondrial 16S marker was amplified from scat DNA and amplicons sequenced using an Ion Torrent PGM™. For a given set of bioinformatic parameters, there was generally low variability between scat samples in proportions of prey species sequences recovered. However, proportions varied substantially depending on sequencing direction, level of quality filtering (due to differences in sequence quality between species) and minimum read length considered. Short primer tags used to identify individual samples also influenced species proportions. In addition, there were complex interactions between factors; for example, the effect of quality filtering was influenced by the primer tag and sequencing direction. Resequencing of a subset of samples revealed some, but not all, biases were consistent between runs. Less stringent data filtering (based on quality scores or read length) generally produced more consistent proportional data, but overall proportions of sequences were very different than dietary mass proportions, indicating additional technical or biological biases are present. Our findings highlight that quantitative interpretations of sequence proportions generated via high-throughput sequencing will require careful experimental design and thoughtful data analysis.
Data from: A phylogenomic approach based on PCR target enrichment and high throughput sequencing: resolving the diversity within the South American species of Bartsia l. (Orobanchaceae)
Advances in high-throughput sequencing (HTS) have allowed researchers to obtain large amounts of biological sequence information at speeds and costs unimaginable only a decade ago. Phylogenetics, and the study of evolution in general, is quickly migrating towards using HTS to generate larger and more complex molecular datasets. In this paper, we present a method that utilizes microfluidic PCR and HTS to generate large amounts of sequence data suitable for phylogenetic analyses. The approach uses the Fluidigm Access Array System (Fluidigm, San Francisco, CA, USA) and two sets of PCR primers to simultaneously amplify 48 target regions across 48 samples, incorporating sample-specific barcodes and HTS adapters (2,304 unique amplicons per Access Array). The final product is a pooled set of amplicons ready to be sequenced, and thus, there is no need to construct separate, costly genomic libraries for each sample. Further, we present a bioinformatics pipeline to process the raw HTS reads to either generate consensus sequences (with or without ambiguities) for every locus in every sample or—more importantly—recover the separate alleles from heterozygous target regions in each sample. This is important because it adds allelic information that is well suited for coalescent-based phylogenetic analyses that are becoming very common in conservation and evolutionary biology. To test our approach and bioinformatics pipeline, we sequenced 576 samples across 96 target regions belonging to the South American clade of the genus Bartsia L. in the plant family Orobanchaceae. After sequencing cleanup and alignment, the experiment resulted in ~25,300bp across 486 samples for a set of 48 primer pairs targeting the plastome, and ~13,500bp for 363 samples for a set of primers targeting regions in the nuclear genome. Finally, we constructed a combined concatenated matrix from all 96 primer combinations, resulting in a combined aligned length of ~40,500bp for 349 samples.
Data from: RAD-sequencing reveals within-generation polygenic selection in response to anthropogenic organic and metal contamination in North Atlantic Eels
Measuring the effects of selection on the genome imposed by human-altered environment is currently a major goal in ecological genomics. Given the polygenic basis of most phenotypic traits, quantitative genetic theory predicts that selection is expected to cause subtle allelic changes among covarying loci rather than pronounced changes at few loci of large effects. The goal of this study was to test for the occurrence of polygenic selection in both North Atlantic eels (European Eel, Anguilla anguilla and American Eel, A. rostrata), using a method that searches for covariation among loci that would discriminate eels from "control" vs. "polluted" environments and be associated with specific contaminants acting as putative selective agents. RAD-seq libraries resulted in 23,659 and 14,755 filtered loci for the European and American Eels respectively. A total of 142 and 141 covarying markers discriminating European and American Eels from "control" vs. "polluted" sampling localities were obtained using the Random Forest algorithm. Distance-based redundancy analyses (db-RDAs) were used to assess the relationships between these covarying markers and concentration of 34 contaminants measured for each individual eel. PCB153, 4'4'DDE and selenium were associated with covarying markers for both species, thus pointing to these contaminants as major selective agents in contaminated sites . Gene enrichment analyses suggested that sterol regulation plays an important role in the differential survival of eels in "polluted" environment. This study illustrates the power of combining methods for detecting signals of polygenic selection and for associating variation of markers with putative selective agents in studies aiming at documenting the dynamics of selection at the genomic level, and particularly so in human altered environments.
Data from: Exploring the potential of small RNA subunit and ITS sequences for resolving phylogenetic relationships within the phylum Ctenophora
Ctenophores are a phylum of non-bilaterian marine (mostly planktonic) animals, characterised by several unique synapomorphies (e.g. comb rows, apical organ). Relationships between and within the nine recognised ctenophore orders are far from understood, notably due to a paucity of phylogenetically-informative anatomical characters. Previous attempts to address ctenophore phylogeny using molecular data (18S rRNA) led to poorly resolved trees but demonstrated the paraphyly of the order Cydippida. Here we compiled an updated 18S rRNA data set, notably including a few newly-sequenced species representing previously unsampled families (Lampeidae, Euryhamphaeidae), and we built up an additional more rapidly-evolving ITS1+5.8SrRNA+ITS2 alignment. These data sets have been analysed separately and in combination under a probabilistic framework, using different methods (Maximum Likelihood, Bayesian inference) and models (e.g. doublet model to accommodate secondary structure; data partitioning). An important lesson from our exploration of these datasets is that the fast-evolving ITS regions are useful markers for reconstructing high-level relationships within ctenophores. Our results confirm the paraphyly of the order Cydippida (and thus a "cyddipid-like" ctenophore common ancestor) and suggest that the family Mertensiidae could be the sister-group of all other ctenophores. The family Lampeidae (also part of the former "Cydippida") is probably the sister-group of the order Platyctenida (benthic ctenophores). The order Beroida might not be monophyletic, due to the position of Beroe abyssicola outside of a clade grouping the other Beroe species and members of the "Cydippida" family Haeckeliidae. Many relationships (i.e. between Pleurobrachiidae, Beroida, Cestida, Lobata, Thalassocalycida) remain unresolved. Future progress in understanding ctenophore phylogeny will come from the use of additional rapidly-evolving markers and improvement of taxonomic sampling.
Data from: Universal and blocking primer mismatches limit the use of high-throughput DNA sequencing for the quantitative metabarcoding of arthropods
The quantification of the biological diversity in environmental samples using high-throughput DNA sequencing is hindered by the PCR bias caused by variable primer–template mismatches of the individual species. In some dietary studies, there is the added problem that samples are enriched with predator DNA, so often a predator-specific blocking oligonucleotide is used to alleviate the problem. However, specific blocking oligonucleotides could coblock nontarget species to some degree. Here, we accurately estimate the extent of the PCR biases induced by universal and blocking primers on a mock community prepared with DNA of twelve species of terrestrial arthropods. We also compare universal and blocking primer biases with those induced by variable annealing temperature and number of PCR cycles. The results show that reads of all species were recovered after PCR enrichment at our control conditions (no blocking oligonucleotide, 45 °C annealing temperature and 40 cycles) and high-throughput sequencing. They also show that the four factors considered biased the final proportions of the species to some degree. Among these factors, the number of primer–template mismatches of each species had a disproportionate effect (up to five orders of magnitude) on the amplification efficiency. In particular, the number of primer–template mismatches explained most of the variation (~3/4) in the amplification efficiency of the species. The effect of blocking oligonucleotide concentration on nontarget species relative abundance was also significant, but less important (below one order of magnitude). Considering the results reported here, the quantitative potential of the technique is limited, and only qualitative results (the species list) are reliable, at least when targeting the barcoding COI region.
Data from: Raw whole Drosophila genome sequence traces have contaminant sequences from bacterial symbionts
Many Drosophila genomes have been sequenced and assembled recently, and many more genome sequencing projects are in progress. However, Drosophila have bacterial, fungal, and protozoan symbionts, and the DNA of these symbionts may be isolated in the process of sequencing Drosophila genomes. Here, we assess how much sequence is isolated from these symbionts and if the sequence contamination affected how these Drosophila genomes were assembled. We do find raw sequence from bacterial symbionts and humans in Drosophila genome sequence traces analyzed. Surprisingly, the four most-common contaminant species were shared among the Drosophila genomes. However, we do not find evidence of bacterial sequences in two published Drosophila genome assemblies.
Data from: A comprehensive analysis of teleost MHC class I sequences
Background: MHC class I (MHCI) molecules are the key presenters of peptides generated through the intracellular pathway to CD8-positive T-cells. In fish, MHCI genes were first identified in the early 1990′s, but we still know little about their functional relevance. The expansion and presumed sub-functionalization of cod MHCI and access to many published fish genome sequences provide us with the incentive to undertake a comprehensive study of deduced teleost fish MHCI molecules. Results: We expand the known MHCI lineages in teleosts to five with identification of a new lineage defined as P. The two lineages U and Z, which both include presumed peptide binding classical/typical molecules besides more derived molecules, are present in all teleosts analyzed. The U lineage displays two modes of evolution, most pronouncedly observed in classical-type alpha 1 domains; cod and stickleback have expanded on one of at least eight ancient alpha 1 domain lineages as opposed to many other teleosts that preserved a number of these ancient lineages. The Z lineage comes in a typical format present in all analyzed ray-finned fish species as well as lungfish. The typical Z format displays an unprecedented conservation of almost all 37 residues predicted to make up the peptide binding groove. However, also co-existing atypical Z sub-lineage molecules, which lost the presumed peptide binding motif, are found in some fish like carps and cavefish. The remaining three lineages, L, S and P, are not predicted to bind peptides and are lost in some species. Conclusions: Much like tetrapods, teleosts have polymorphic classical peptide binding MHCI molecules, a number of classical-similar non-classical MHCI molecules, and some members of more diverged MHCI lineages. Different from tetrapods, however, is that in some teleosts the classical MHCI polymorphism incorporates multiple ancient MHCI domain lineages. Also different from tetrapods is that teleosts have typical Z molecules, in which the residues that presumably form the peptide binding groove have been almost completely conserved for over 400 million years. The reasons for the uniquely teleost evolution modes of peptide binding MHCI molecules remain an enigma.
Data from: Genome sequence of Striga asiatica provides insight into the evolution of plant parasitism
Parasitic plants in the genus Striga, commonly known as witchweeds, cause major crop losses in sub-Saharan Africa and pose a threat to agriculture worldwide. An understanding of Striga parasite biology, which could lead to agricultural solutions, has been hampered by the lack of genome information. Here we report the draft genome sequence of Striga asiatica with 34,577 predicted protein-coding genes, which reflects gene family contractions and expansions that are consistent with a three-phase model of parasitic plant genome evolution. Striga seeds germinate in response to host-derived strigolactones (SLs) and then develop a specialised penetration structure, the haustorium, to invade the host root. A family of SL receptors has undergone a striking expansion, suggesting a molecular basis for the evolution of broad host range among Striga spp. We found that genes involved in lateral root development in non-parasitic model species are coordinately induced during haustorium development in Striga, suggesting a pathway that was partly co-opted during the evolution of the haustorium. In addition, we found evidence for horizontal transfer of host genes as well as retrotransposons, indicating gene flow to S. asiatica from hosts. Our results provide valuable insights into the evolution of parasitism and a key resource for the future development of Striga control strategies.
Data from: Diet assessment of the Atlantic Sea Nettle Chrysaora quinquecirrha in Barnegat Bay, New Jersey, using next-generation sequencing
Next generation sequencing (NGS) methodologies have proven useful in deciphering the food items of generalist predators, but have yet to be applied to gelatinous animal gut and tentacle content. NGS can potentially supplement traditional methods of visual identification. Chrysaora quinquecirrha (Atlantic sea nettle) has progressively become more abundant in Mid-Atlantic United States' estuaries including Barnegat Bay (New Jersey), potentially having detrimental effects on both marine organisms and human enterprises. Full characterization of this predator's diet is essential for a comprehensive understanding of its impact on the food web and its management. Here we tested the efficacy of NGS for prey item determination in the Atlantic sea nettle. We implemented a NGS "shotgun" approach to randomly sequence DNA fragments isolated from gut lavages and gastric pouch/tentacle picks of 8 and 84 sea nettles, respectively. These results were verified by visual identification and co-occurring plankton tows. Over 550,000 contigs were assembled from ~110 million paired-end reads. Of these, 100 contigs were confidently assigned to 23 different taxa, including soft bodied organisms previously undocumented as prey species, including copepods, fish, ctenophores, anemones, amphipods, barnacles, shrimp, polychaete worms, flukes, flatworms, echinoderms, gastropods, bivalves, and hemichordates. Our results not only indicate that a "shotgun" NGS approach can supplement visual identification methods, but targeted enrichment of a specific amplicon/gene is not a prerequisite for identifying Atlantic sea nettle prey items.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.