Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,076
datasets available to search
ShareScore release 0.7.1
Dataset results
1,076 results for “Metabarcoding”
Niche overlap between two large herbivores across landscape variability using dietary eDNA metabarcoding: Raw sequences of European bison and red deer – Bialowieza
<p><span>Understanding the trophic ecology of herbivore species is key to assess their environmental requirements and to improve management policies, but measuring their trophic interactions remains challenging. Among the methods available, quantifying the plant composition of a species' diet provides a detailed picture of how species exploit the resources in their environment and their associated niche overlap. Yet, most studies focusing on herbivore trophic ecology ignore the influence that landscape variability may have. Here, we studied how landscape variability influences trophic interactions through niche partitioning. We used eDNA metabarcoding to quantify the diet composition of two large herbivores of the Bialowieza Forest, red deer (<em>Cervus elaphus</em>) and European bison (<em>Bison bonasus</em>, hereafter referred to as bison) to investigate how increasing habitat quality and predation risk in their environment influence their diet composition and niche partitioning. We found red deer to have an overall greater diet variability and lower niche overlap within species compared to bison. Moreover, our findings indicate herbivore interactions are non-homogeneous across the landscape. Higher habitat quality was associated with higher niche overlap only within bison. We also detected an increase in niche overlap with increasing predation risk within red deer, indicating they modify their diet choice as a reaction to wolf predation risk. This study provides evidence of eDNA dietary metabarcoding as a useful tool for wildlife management to assess the status of species in an ecosystem based on known environmental factors. We suggest future studies to integrate the environments' variability when studying trophic ecology of herbivores to capture the whole extent of species interaction in order to improve conservation and management guidelines. </span></p>
Row fastq files generated by 16S rRNA sequencing for the metabarcoding analysis of Hidalgo-Villeda et al. paper
<p>Fastq files used for the microbiome profiling of the terminal ileal and caecum content.</p> <p>Sequencing of the terminal ileal content was performed at the Institute Hospitalo-Universitaire Méditerannée Infection, Marseille, France.</p> <p>Sequencing of the caecum content was performed at Genoscreen, Lille, France</p> <p>The sequencing methodology is described on the methods part of the paper "A physiological mouse model of severe acute malnutrition reveals prolonged dysbiosis and altered immunity under nutritional intervention, Hidalgo-Villeda et al.".</p> <p>The metadata.xlsx table present the description and correspondance of each fastq files.</p>
Sample data for tephritid metabarcoding workshop
<p>Raw data files and illumina sample sheets for tephritid metabarcoding workshop hosted at Agribio in September 2022</p>
Focal vs. faecal: Seasonal variation in the diet of wild vervet monkeys from observational and DNA metabarcoding data
<p>1. Assessing the diet of wild animals reveals valuable information about their ecology and trophic relationships that may help elucidate dynamic interactions in ecosystems and forecast responses to environmental changes.</p> <p>2. Advances in molecular biology provide valuable research tools in this field. However, comparative empirical research is still required to highlight strengths and potential biases of different approaches. Therefore, this study compares environmental DNA and observational methods for the same study population and sampling duration.</p> <p>3. We employed DNA metabarcoding assays targeting plant and arthropod diet items in 823 faecal samples collected over 12 months in a wild population of an omnivorous primate, the vervet monkey (<em>Chlorocebus pygerythrus</em>). DNA metabarcoding data were subsequently compared to direct observations.</p> <p>4. We observed the same seasonal patterns of plant consumption with both methods, however, DNA metabarcoding showed considerably greater taxonomic coverage and resolution compared to observations, mostly due to the construction of a local plant DNA database. We found a strong effect of season on variation in plant consumption largely shaped by the dry and wet seasons. The seasonal effect on arthropod consumption was weaker but feeding on arthropods was more frequent in spring and summer, showing overall that vervets adapt their diet according to available resources. The DNA metabarcoding assay outperformed also direct observations of arthropod consumption in both taxonomic coverage and resolution.</p> <p>5. Combining traditional techniques and DNA metabarcoding data can therefore not only provide enhanced assessments of complex diets or reveal trophic interactions to the benefit of wildlife conservationists and managers but also opens new perspectives for behavioural ecologists studying whether diet variation in social species is induced by environmental differences or might reflect selective foraging behaviours.</p>
Data from: Development and validation of targeted environmental DNA (eDNA) metabarcoding for early detection of 69 invasive fishes and aquatic invertebrates
<p>Invasive species are of concern due to their impacts on ecosystems and economies, but they pose significant control challenges. Environmental DNA (eDNA) is a powerful tool in the detection of aquatic organisms at low densities due to high sensitivity and ease of collection. Aquatic eDNA analyses have increased worldwide and are generally either applied to a few target species (quantitative PCR) or for broad taxonomic applications (metabarcoding). Here we describe the development and testing of a hybrid approach that utilized high sensitivity PCR primer sets and high-throughput sequencing (HTS), referred to as <em>targeted metabarcoding</em>, to detect 69 fishes and invertebrates. We identified target species based on reports of globally important invasive species and developed two independent PCR primers for each species (CO1 and a second mtDNA region). We assessed sensitivity and eDNA interference for all 138 primers (2 per species, 69 species) using standard end-point PCR and tested them on 10 eDNA samples spiked with various amounts of one or more of the target species' DNA. The sensitivity of the 138 primer sets ranged between 1.5×10<sup>-5</sup> and 2.64 ng template DNA (mean = 0.069 ng). Primers were also tested for interference effects using plankton eDNA to simulate field conditions. The inclusion of interfering plankton DNA reduced the sensitivity for most primer sets by one or more orders of magnitude (range 0 to 3). Overall, our targeted metabarcoding resulted in the detection of ~ 98% of species in the DNA spiked samples, and, perhaps more importantly, the HTS read count was positively related to the quantity of spiked DNA (P < 0.002). We envision this technique being particularly useful for the early detection of species at low population densities; however, there are diverse applications of targeted metabarcoding for monitoring aquatic community composition and quantifying ecosystem change and health.</p>
Metabarcoding for biodiversity inventory blind spots: A test case using the beetle fauna of an insular cloud forest
<p>Soils harbour a rich arthropod fauna, but many species are still not formally described (Linnaean shortfall), and the distribution of those already described is poorly understood (Wallacean shortfall). Metabarcoding holds much promise to fill this gap, however, nuclear copies of mitochondrial genes, and other artefacts lead to taxonomic inflation, which compromises the reliability of biodiversity inventories. Here we explore the potential of a bioinformatic approach to jointly "denoise" and filter non-authentic mitochondrial sequences from metabarcode reads to obtain reliable soil beetle inventories and address open questions in soil biodiversity research, such as the scale of dispersal constraints in different soil layers. We sampled cloud forest arthropod communities from 49 sites in the Anaga peninsula of Tenerife (Canary Islands). We performed whole organism community DNA (wocDNA) metabarcoding, and built a local reference database with COI barcode sequences of 310 species of Coleoptera for filtering reads and the identification of metabarcoded species. This resulted in reliable haplotype data after considerably reducing nuclear mitochondrial copies and other artefacts. Comparing our results with previous beetle inventories, we found: (i) new species records, potentially representing undescribed species; (ii) new distribution records, and; (iii) validated phylogeographic structure when compared with traditional sequencing approaches. Analyses also revealed evidence for higher dispersal constraint within deeper soil beetle communities, compared to those closer to the surface. The combined power of barcoding and metabarcoding contribute to mitigate the important shortfalls associated with soil arthropod diversity data, and thus address unresolved questions for this vast biodiversity fraction.</p>
rDNA 18S V4 metabarcoding tables (Swarm) for Tara Oceans Expedition (2009-2013), including Tara Polar Circle Expedition (2013)
<p>Reads were grouped into OTUs using the following swarm-based pipeline: paired-end reads were merged with vsearch’s --fastq_mergepairs command (version 2.15.1, allowing for staggered reads; Rognes et al., 2016), and trimmed with cutadapt (version 3.0; Martin, 2011), keeping only reads containing both forward and reverse primers. After trimming, the expected error per read was estimated with vsearch’s command --fastq_filter and the option --eeout. Each sample was then de-replicated, i.e. strictly identical reads were merged, using vsearch’s command --derep_fulllength, and converted into fasta format. Clustering was performed at the sample level with swarm 3.0 using default parameters (Mahé et al., 2015). Prior to global clustering, individual fasta files (one per sample) were pooled and further dereplicated with vsearch. Files containing per-read expected error values were also dereplicated to retain only the lowest expected error for each unique sequence. Global clustering was performed with swarm (using the fastidious option). Cluster representative sequences were then searched for chimeras with vsearch’s command --uchime_denovo using default parameters (Edgar et al., 2011).<br> Clustering results, expected error values, taxonomic assignments, and chimera detection results were used to build a “raw” occurrence table. Reads without primers, reads shorter than 32 nucleotides and reads with uncalled bases (“N”) were discarded. For a “filtered” occurrence table, non-chimeric sequences, sequences with an expected error per nucleotide below 0.0002, and clusters containing at least 2 reads were retained. Since primer trimming is not perfect, some sequences can still contain primer fragments or be excessively trimmed. These sub- or super-sequences were identified using vsearch and merged with their closest, most abundant perfectly trimmed sequence. Finally, occurrence patterns throughout our sample collection were used to further refine the occurrence table. Clusters that contain sub-clusters with only a single-nucleotide difference but with different ecological patterns (defined here as uncorrelated abundance values in at least 5% of the samples) were turned into distinct clusters (https://github.com/frederic-mahe/fred-metabarcoding-pipeline). On the other hand, clusters with similar sequences that had correlated abundance values in at least 95% of the samples, were merged using a re-implementation of lulu's method (Frøslev et al. 2017; https://github.com/frederic-mahe/mumu).</p>
Data supplementing the article "Avoiding quantification bias in metabarcoding: application of a cell biovolume correction factor in diatom molecular biomonitoring" V. Vasselon, A. Bouchez, F. Rimet, S. Jacquet, R. Trobajo, M. Corniquel, K. Tapolczai, I. Domaizon submitted to Methods in Ecology and Evolution journal
<p>These data supplement the article "Avoiding quantification bias in metabarcoding: application of a cell biovolume correction factor in diatom molecular biomonitoring" V. Vasselon, A. Bouchez, F. Rimet, S. Jacquet, R. Trobajo, M. Corniquel, K. Tapolczai, I. Domaizon submitted to Methods in Ecology and Evolution journal</p> <p>The directory contains the following files:</p> <p>1<strong>5 fastq files raw reads (5 mock communities, 3 replicates)</strong><strong>.rar </strong>- contains the 15 fastq files provided by the sequencing platform with demultiplexed DNA reads (raw data prior any bioinformatics treatments).</p> <p><strong>15 fastq files information.xlsx</strong> :</p> <p>- contains the information relative to the 15 fastq files corresponding to the PGM raw data of the 5 mock communities (sequenced with 3 replicates), including: the ID of the fastq files, the mock community name, the replicate number, the final sample Id and the number of raw reads per fastq file.</p> <p>- contains the information of the proportion of the 8 diatoms species (%) used to create the 5 mock communities (estimated from microscopy).</p>
Experimental evaluation of genetic variability based on DNA metabarcoding from the aquatic environment: Insights from the Leray COI fragment
<p>Intraspecific genetic variation is important for the assessment of organisms' resistance to changing environments and anthropogenic pressures. Aquatic DNA metabarcoding provides a non-invasive method in biodiversity research, including investigations at the within-species level. Through the analysis of eDNA samples collected from the Peter the Great Gulf of the Japan Sea, in this study we aimed to evaluate the identification of Amplicon Sequence Variants (ASVs) in marine eDNA among abundant species of the <em>Zostera</em> sp. community: <em>Hexagrammos octogrammus</em>, <em>Pholidapus dybowskii</em> (Teleostei: Perciformes), and <em>Pandalus latirostris</em> (Arthropoda: Decapoda). These species were collected from two distant locations to produce mock communities and gather aquatic eDNA both on the community and individual level. Our approach highlights the efficacy of eDNA metabarcoding in capturing haplotypic diversity and the potential for this methodology to track genetic diversity accurately, contributing to conservation efforts and ecosystem management. Additionally, our results elucidate the impact of nuclear mitochondrial DNA segments (NUMTs) on the reliability of metabarcoding data, indicating the necessity for cautious interpretation of such data in ecological studies. Moreover, we analyzed 83 publicly available <em>COI</em> sequence datasets from common groups of multicellular organisms (Mollusca, Echinodermata, Crustacea, Polychaeta, and Actinopterygii). The results reflect the decrease in population diversity that arises from using the metabarcode compared to the <em>COI</em> barcode.</p>
Data from: Environmental DNA metabarcoding reflects spatiotemporal patterns of fish community shifts in the Scheldt estuary
<p>Estuarine ecosystems face increasing anthropogenic pressures, necessitating effective monitoring methods to mitigate their impacts on the biodiversity they harbour. The use of environmental DNA (eDNA) based detection methods is increasingly recognized as a promising tool to complement other, potentially invasive monitoring techniques. Integrating such eDNA analyses into monitoring frameworks for large spatial ecosystems is still challenging and requires a deeper understanding of the scale and resolution at which eDNA patterns may offer insights in species presence and community composition space and time. The Scheldt estuary, characterized by its diverse habitats and complex currents, is one of the largest Western European tidal river systems. Until now, it remains challenging to obtain accurate information on fish communities living in and migrating through this large ecosystem, consequently confining our knowledge to specific locations. To explore the potential of eDNA-based monitoring, we simultaneously combine stow net fishing with eDNA metabarcoding, to assess the Scheldt estuary's fish communities in space and time. In total, we detected 71 fish species in the estuary using eDNA metabarcoding, partly overlapping with historic fish community data gathered at the different study locations and in contrast to only 42 species using stow net fishing during the same survey period. Community compositions found by both detection methods varied amongst sampling locations, driven by a clear correlation to the salinity gradient. Limited effects of sampling depth and tide were observed on the eDNA metabarcoding data, allowing a significant reduction of the eDNA sampling effort for future eDNA fish monitoring campaigns in this study system. Our results further demonstrate that seasonal shifts in fish species occurrence can be detected using eDNA metabarcoding. Combining eDNA metabarcoding and stow net fishing further enhances our understanding of this vital waterway's diverse fish populations, allowing a higher resolution and more efficient monitoring strategy.</p>
Data from: Community ecology in a bottle: Leveraging eDNA metabarcoding data to predict occupancy of co-occurring species
<p>Detecting environmental DNA (eDNA) of numerous organisms from the same samples has been revolutionized by metabarcoding. However, utilizing the vast amounts of data generated from metabarcoding to predict occupancy probabilities for co-occurring species is currently rare. Here, we demonstrate how metabarcoding data can be used to advance community ecology research through a case study using replicate stream water samples and Bayesian occupancy models to test hypotheses of eDNA occurrence for a native fish (brook trout, Salvelinus fontinalis), its major ectoparasite (gill lice, Salmincola edwardsii), and an introduced potential competitor (brown trout, Salmo trutta). Gill lice DNA occupancy was positively associated with brook trout biomass determined via electrofishing conducted near eDNA sampling sites, suggesting gill lice occupancy is dependent on host density. Leveraging site-specific molecular operational taxonomic units identified from metabarcoding, DNA occupancy of trout and gill lice was often positively predicted by species richness of aquatic insect orders trout commonly feed on, which are also environmental quality indicators. Thus, high-quality habitat that environmentally sensitive salmonids and their primary prey rely on may promote higher fish occupancy rates, further facilitating the spread of fish parasites. An increasing amount of community-level data is being generated from global metabarcoding efforts, and we suggest our framework could be broadly implemented to enhance understanding of factors impacting distributions of co-occurring species, reveal new ecological phenomena, and support management and conservation efforts.</p>
Data from: Primer sets evaluation and sampling method assessment for the monitoring of fish communities in the North-western part of the Mediterranean Sea through eDNA metabarcoding
<p>Environmental DNA (eDNA) metabarcoding appears to be a promising tool for surveying fish communities. However, the effectiveness of this method relies on primer set performance and on a robust sampling strategy. While some studies have evaluated the efficiency of several primers for fish detection, it has not yet been assessed <em>in situ </em>for the Mediterranean Sea. In addition, mainly surface waters were sampled and no filter porosity testing was performed. In this pilot study, our aim was to evaluate the ability of six primer sets, targeting 12S rRNA (AcMDB07; MiFish; Tele04) or 16S rRNA (Fish16S; Fish16SFD; Vert16S) loci, to detect fish species in the Mediterranean Sea using a metabarcoding approach. We also assessed the influence of sampling depth and filter pore size (0.45 µm <em>versus</em> 5 µm filters). To achieve this, we developed a novel sampling strategy allowing simultaneous surface and bottom filtration of large water volumes along on-site the same transect. We found that 16S rRNA primer sets enabled more fish taxa to be detected across each taxonomic level. The best combination was Fish16S/Vert16S/AcMDB07, which recovered 95% of the 97 fish species detected in our study. There were highly significant differences in species composition between surface and bottom samples. Filters of 0.45 µm led to the detection of significantly more fish species. Therefore, to maximize fish detection in the studied area, we recommend to filter both surface and bottom waters through 0.45 µm filters and to use a combination of these three primer sets.</p>
Elk food habits using DNA metabarcoding for plant identification
<p>North American elk (<em>Cervus canadensis</em>) inhabited portions of the Eastern United States until extirpation in the mid-1800s. From 2000 to 2008, 201 elk were reintroduced to the North Cumberland Wildlife Management Area (NCWMA), Tennessee. The stocking source was Elk Island National Park, Alberta Canada where there are two distinct genetic populations isolated from the north and south. This genetic structure has largely persisted in the population after translocation. Food habits were evaluated in the early stages of restoration, but the population has had approximately 20 years to adapt to the landscape, and current food habits are unknown. To assess diet composition using DNA metabarcoding, we collected fecal pellets of elk from 65 openings within the 79,318-ha NCWMA weekly from February to April of 2019. We targeted the ITS2 region of the nuclear ribosomal DNA to amplify vegetation sequences found in the internal portion of the elk feces. DNA metabarcoding of feces was linked to results from an accompanying elk population genetics study to investigate food habits between sexes from the two different genetic groups. The majority (80.298%) of sequences matched plants from 22 genera. The top genera (>5.000%) represented were <em>Vaccinium</em> (15.216%), <em>Festuca</em> (8.446%), <em>Rosa</em> (6.358%), <em>Robinia</em> (5.793%), and <em>Eleagnus</em> (5.186%). Elk heavily used woody plants before and after spring green-up (>50% of diet). However, the quantity of forbs in their diet more than doubled after emergence in the spring. The sex-genetic groups consumed similar vegetation in approximately proportionate amounts. Diversity analyses revealed a significant difference in plant genera sequence detection between males from the two genetic groups, although this finding is likely explained by limited sample size. NCWMA elk used a variety of forage in the winter and DNA metabarcoding analysis allows for a comprehensive analysis of food habits useful for monitoring how elk respond dietarily to habitat management.</p>
Data supplementing the article "Boosting DNA metabarcoding for biomonitoring with phylogenetic estimation of OTUs' ecological profiles" F. Keck, V. Vasselon, F. Rimet, A. Bouchez, and M. Kahlert submitted to Molecular Ecology Resources journal
<p>These data supplement the article "Enhancing DNA metabarcoding for biomonitoring with phylogenetic estimation of OTUs' ecological profiles" F. Keck, V. Vasselon, F. Rimet, A. Bouchez, and M. Kahlert submitted to Molecular Ecology Resources journal</p> <p>The directory contains the following files:</p> <p><strong>278 (139 x 2 replicates) samples fastq files.rar </strong>- contains the 278 fastq files provided by the sequencing platform with demultiplexed and contig DNA reads corresponding to the 139 samples with 2 sequencing replicates (A and B).</p> <p><strong>Counts_diatoms.xlsx </strong>- contains the morphological inventories with species list (Omnidia code) and valve abundances for the 139 samples.</p> <p><strong>Sites_list.xlsx </strong>- contains information regarding the 139 samples, including: River name, GPS coordinates, code used for molecular analysis and corresponding to sequencing fastq names.</p>
Data supplementing the article "Benthic diatom communities in an Alpine river impacted by waste water treatment effluents as revealed using DNA metabarcoding" , submitted to Frontiers in Microbiology
<p>These data supplement the article"Benthic diatom communities in an Alpine river impacted by waste water treatment effluents as revealed using DNA metabarcoding" submitted to Frontiers in Microbiology: </p> <p>The directory contains the following files:</p> <p><strong>64 PGM sequencing files (raw data, fastq files) </strong>- one file is provided for each sample by the sequencing platform with demultiplexed DNA reads (raw data prior any bioinformatics treatments).</p> <p><strong>Sample_Names.xlsx</strong> - contains the information relative to the 64 samples including: the ID used in Mothur analyses (corresponding to the name of the fastq files), the sample name and the raw reads number for each sample.</p>
Raw data supporting metabarcoding unsorted kick-samples research
<p>Raw data supporting: Metabarcoding unsorted kick-samples facilitates macroinvertebrate-based bioassessment with increased taxonomic resolution, while outperforming environmental DNA.</p> <p>Two-step PCRs were performed on each sample replicate using the Qiagen Multiplex PCR Plus Kit.</p> <p>The pool was loaded onto an Illumina MiSeq at 9pM, with 5% Phi-X, using a 600 cycle V3 kit with 300 bp paired end sequencing (index read steps skipped).<br> <br> Also available in the NCBI SRA: <a href="https://www.ncbi.nlm.nih.gov/sra/PRJNA629361">https://www.ncbi.nlm.nih.gov/sra/PRJNA629361</a></p>
Data supplementing the article "Diatom DNA metabarcoding for biomonitoring : strategies to avoid major taxonomical and bioinformatical biases limiting molecular indices capacities" K. Tapolczai, F. Keck, A. Bouchez, F. Rimet, M. Kahlert and V. Vasselon submitted to "Frontiers in Ecology and Evolution" journal
<p>These data supplement the article "Diatom DNA metabarcoding for biomonitoring : strategies to avoid major taxonomical and bioinformatical biases limiting molecular indices capacities" K. Tapolczai, F. Keck, A. Bouchez, F. Rimet, M. Kahlert and V. Vasselon submitted to "Frontiers in Ecology and Evolution" journal.</p> <p>The directory contains the following files:</p> <p><strong>464_samples_fastq_files_(mothur).rar </strong>- contains the 464 fastq files proceed together during the Mothur bioinformatics treatments to produce the OTUs and ISUs tables. As the contig and the demultiplexing steps were performed by the sequencing platform, there is 1 fastq file per sample. From this 464 samples OTU/ISU tables, only information regarding 76 samples were used in this study and are listed in the "<strong>76_samples_list_(mothur).xlsx" </strong>file<strong>.</strong></p> <p><strong>76_samples_list_(mothur).xlsx </strong>- contains the information regarding the 76 samples used to create the OTUs and ISUs tables presented in the paper.</p> <p><strong>76_samples_R1_R2_fastq_files(DADA2).rar - </strong>contains the raw demultiplexed fastq files (R1.fastq and R2.fastq) for each of the 76 samples used in this study to produce the ESVs table using the DADA2 bioinformatics pipeline.</p>
Data from: Performance of DNA metabarcoding, standard barcoding and morphological approaches in the identification of insect biodiversity
<p>A repository containing the data and R code for the associated manuscript titled "Performance of DNA metabarcoding, standard barcoding and morphological approaches in the identification of insect biodiversity"</p> <p>Created by Romana K. Salis</p> <div> <h2>Content Overview</h2> <a href="https://github.com/rksalis/LNU_insectbarcoding/tree/main#content-overview"></a></div> <div> <h3>Rcode:</h3> <a href="https://github.com/rksalis/LNU_insectbarcoding/tree/main#rcode"></a></div> <p>analysis_070923.R</p> <div> <h3>Data files:</h3> <a href="https://github.com/rksalis/LNU_insectbarcoding/tree/main#data-files"></a></div> <p>literature_search_210624.csv barcoding_datasets.csv - barcoding datasets match category counts for the butterflies, bees and wasps</p> <p>barcoding_VB_thresholds.csv - barcoding datasets match category counts for the vietnamese butterfly with different similarity thresholds</p> <p>barcoding_VB_families.csv - barcoding datasets match category counts for the vietnamese butterfly families</p> <p>vietkordfinal_x.csv - intraspecific similarity and geographic distance</p> <p>vietnam_area.csv - intraspecific similarity and sampling area</p> <p>BINsvSp.csv - number of BINs and morphospecies</p> <p>Metabarcoding_Results_CommonSpecies.csv - metabarcoding results including the species only those also identified by morphology</p> <p>Metabarcoding_Results_BINSpecies.csv - metabarcoding results BIN species</p> <p>Metabarcoding_Results_BINs.csv - metabarcoding results all BINs</p>
Mismanagement and poor transparency in the European processed seafood supply revealed by DNA metabarcoding
<p>Raw Illumina MiSeq data in zipped FASTQ format related to the manuscript "Mismanagement and poor transparency in the European processed seafood supply revealed by DNA metabarcoding" (DOI: <a href="https://doi.org/10.1016/j.foodres.2024.114901">https://doi.org/10.1016/j.foodres.2024.114901</a>).</p> <p>Lorusso, L., Shum, P., Piredda, R., Mottola, A., Maiello, G., Cartledge, E. L., Neave E. F., & Di Pinto A., Mariani, S. (2024). Mismanagement and poor transparency in the European processed seafood supply revealed by DNA metabarcoding. Food Research International, <a href="https://doi.org/10.1016/j.foodres.2024.114901">https://doi.org/10.1016/j.foodres.2024.114901</a>.</p>
Script and data from: The best of two worlds: toward large-scale monitoring of biodiversity combining metabarcoding and optimised parataxonomic validation.
<h2>Description</h2> <div> <p>Zenodo linked to : Penel, B., Meynard, C.N., Benoit, L., Bourdonné, A., Clamens, A., Soldati, L., Migeon, A., Chapuis, M.-P., Piry, S., Kergoat, G. and Haran, J. (2025), The best of two worlds: toward large-scale monitoring of biodiversity combining COI metabarcoding and optimized parataxonomic validation. Ecography, 2025: e07699. <a href="https://doi.org/10.1111/ecog.07699">https://doi.org/10.1111/ecog.07699</a></p> <div> <div> <div> <div> <p><strong>Publication abstract </strong></p> </div> </div> </div> <p>In a context of unprecedented insect decline, it is critical to have reliable monitoring tools to measure species diversity and their dynamic at large-scales. High-throughput DNA-based identification methods, and particularly metabarcoding, were proposed as an effective way to reach this aim. However, these identification methods are subject to multiple technical limitations, resulting in unavoidable false-positive and false-negative species detection. Moreover, metabarcoding does not allow a reliable estimation of species abundance in a given sample, which is key to document and detect population declines or range shifts at large scales. To overcome these obstacles, we propose here a Human-Assisted Molecular Identification (HAMI) approach, a framework based on a combination of metabarcoding and image-based parataxonomic validation of outputs and recording of abundance. We assessed the advantages of using HAMI over the exclusive use of a metabarcoding approach by examining 492 mixed beetle samples from a biodiversity monitoring initiative conducted throughout France. On average, 23% of the species are missed when relying exclusively on metabarcoding, this percent being consistently higher in species-rich samples. Importantly, on average, 20% of the species identified by molecular-only approaches correspond to false positives linked to cross-sample contaminations or mis-identified barcode sequences in databases. The combination of molecular methodologies and parataxonomic validation in HAMI significantly reduces the intrinsic biases of metabarcoding and recovers reliable abundance data. This approach also enables users to engage in a virtuous circle of database improvement through the identification of specimens associated with missing or incorrectly assigned barcodes. As such, HAMI fills an important gap in the toolbox available for fast and reliable biodiversity monitoring at large scales.</p> <div> <h4><strong>File description: </strong></h4> <h4>MiSeq raw sequences of the COI barcode from 492 Coleoptera field samples :</h4> <div>The Raw_sequencage_data ZIP directory contains the FASTQ files of the paired-end reads (R1: reads 1; R2: reads 2) produced for each Coleoptera field samples in duplicate using the MiSeq platform GenSeq (ISEM - University of Montpellier)</div> <div> </div> <div>The HAMI_data_script_results_R zip directory contains Rmarkdown script files (.Rmd and .html) and associated data used to analyse the systemic errors of the metabarcoding approach (N= 492 Coleoptera field samples).</div> <div> </div> <div>The HAMI_pipeline zip directory contains all the codes associated with the HAMI pipeline, as well as a ReadMe file and a test data set.</div> <div> </div> <div>The Residual_chimera.zip directory contains lists of MOTUs associated to residual chimeric sequences that were not filtered using FROGS pipeline but secondarily detected with the <em>de novo</em> approach implemented in HAMI pipeline with ‘isBimeraDenovo’ R function from DADA2 v1.28.0. It contains two distinct files according to the two sequencing runs.</div> <div> </div> <div>The NUMTS_filtered.zip directory contains lists of MOTUs that were excluded of the final dataset according to the NUMTS filtering. File xxx_pseudogene_f1_deteled.csv corresponds to MOTUs that were excluded according to the first filtrering step based on DNA sequencing. File xxx_pseudogene_f2_deteled.csv corresponds to merged MOTUs that were excluded according to the second filter based on occurrence and percentage of identity. This folder contains files for the two sequencing runs.</div> </div> </div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.