Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
348
datasets available to search
ShareScore release 0.7.1
Dataset results
348 results for “eDNA”
Linked collectors and determiners for: eDNA‑based detection of the invasive crayfish Pacifastacus leniusculus in streams with a LAMP assay using dependent replicates to gain higher sensitivity.
Natural history specimen data linked to collectors and determiners held within, "eDNA‑based detection of the invasive crayfish Pacifastacus leniusculus in streams with a LAMP assay using dependent replicates to gain higher sensitivity". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/ea16e238-4e23-41fb-9eee-c1ea4f0caa63">https://bionomia.net/dataset/ea16e238-4e23-41fb-9eee-c1ea4f0caa63</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/ea16e238-4e23-41fb-9eee-c1ea4f0caa63">https://gbif.org/dataset/ea16e238-4e23-41fb-9eee-c1ea4f0caa63</a>. Formatted as a Frictionless Data package.
Linked collectors and determiners for: Savu Sea-Indonesia eDNA Dataset.
Natural history specimen data linked to collectors and determiners held within, "Savu Sea-Indonesia eDNA Dataset". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/bbb81807-fc4c-4e4e-bc1c-2caee84a4a06">https://bionomia.net/dataset/bbb81807-fc4c-4e4e-bc1c-2caee84a4a06</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/bbb81807-fc4c-4e4e-bc1c-2caee84a4a06">https://gbif.org/dataset/bbb81807-fc4c-4e4e-bc1c-2caee84a4a06</a>. Formatted as a Frictionless Data package.
Tourmaline files for Sterivex eDNA testing
<p><strong>Tourmaline and QIIME 2 files for 12S (fish) and 16S (microbes) amplicon metabarcoding associated with Sterivex eDNA methods development in Biscayne Bay, Florida </strong>(<a href="https://github.com/aomlomics/sterivex">https://github.com/aomlomics/sterivex</a>).</p> <p>Files are associated with sequence processing using either stand-alone QIIME 2 (12S folder) or <a href="https://github.com/aomlomics/tourmaline">Tourmaline</a> (16S folder), which is a recently developed program that wraps QIIME 2 and Snakemake. Taxonomy was assigned to sequences using QIIME 2-formatted reference database for 16S (SILVA; v.138.1), 16S chloroplasts (PR2; v.4.12), and 12S (<a href="https://zenodo.org/record/4589660#.YONRCRNKg_U">Mitohelper; March 2021 release</a>).</p> <p>Folders with intermediate and final output files:</p> <ul> <li>16S folder - file outputs from Tourmaline</li> <li>16S chloroplast folder - files from QIIME 2</li> <li>12S folder - files from QIIME 2</li> </ul> <p>Reference database artifact files (.qza) used to assign taxonomy:</p> <ul> <li>silva-138-99-tax.qza</li> <li>silva-138-99-seqs.qza</li> <li>pr2_plastid_seqs.qza</li> <li>pr2_plastid_tax.qza</li> <li>12S-16S-18S-seqs_March2021.qza </li> <li>12S-16S-18S-tax_March2021.qza </li> </ul> <p> </p> <p>These files are associated with the following project: <a href="https://github.com/aomlomics/sterivex">https://github.com/aomlomics/sterivex</a>. Raw FASTQ sequence data are available in NCBI SRA under BioProject ID PRJNA728349.</p>
Plant Communities at the Eden Project, UK, derived from soil eDNA
<p>The project seeks to understand the potential for the use of eDNA collected from soil to characterise plant communities. To do so, soils were sampled at the Eden Project in Cornwall UK, within the two covered biomes where we have a good understanding of the structure and composition of plant communities (further quantified with above ground plant coverage inventories). 32 plots were established across 10 different plant assemblages, each of which experiences subtle differences in soil chemistry and microclimate. Each plot consists of a 2 x 2 m quadrat, with four soil aggregates collected at each corner. </p> <p>eDNA was then extracted and amplified following the methods detailed in Zinger et al. (2016) and Donald et al. (2021). The primers used targeted the P6 loop of thechloroplastic trnL intron [primer_fwd: GGGCAATCCTGAGCCAA, primer_rev: CCATTGAGTCTCTGCACCTATC] (Taberlet et al. 2007). 16 Extraction, 54 Sequencing, and 16 PCR controls are included so as to account for potential errors generated during the processing of samples, with a mock community (4 positive controls) of 10 known plant sequences also included to guide filtering thresholds. PCR products were pooled and sequencing libraries were constructed using the Illumina TruSeq NanoPCRFree kit following the supplier’s instructions (Illumina Inc., San Diego, California, USA), except that the ligation product was not PCR amplified to limit tag-jump biases (Taberlet et al 2018). The libraries were then sequenced on an Illumina Hiseq platform (San Diego, CA, USA).</p> <p>Sequencing was conducted by the GenoToul bioinformatics platform (Toulouse, France), with the OBITOOLS package (Boyer et al. 2016). Here, the produced sequence data was processed using the following steps. First, ‘illuminapairedend’ was used to assemble paired-end reads. This algorithm is based on an exact alignment algorithm that considers the quality scores at all positions during the assembly process. Subsequently, we used the ‘ngsfilter’ command to identify and remove the primers and tags on each read, and assign reads to their respective samples (NGS filter file provided: <strong>ngsfilter_TRNL_PLANTS_EDEN_PROJECTb.txt</strong>). This program was used with its default parameters tolerating two mismatches for each of the two primers and no mismatch for the tags. Following this, sequencing reads were dereplicated using the ‘obiuniq’ command. The produced <strong>data.uniq.fasta</strong> file is supplied here. Sequences were then further filtered to remove sequences of low quality (containing Ns or with paired-end alignment scores below 50), and sequences represented by only one read (singletons) using the ‘obigrep’ command. To remove PCR/sequencing errors as well as intraspecific variability, we built OTUs (Operational Taxonomic Units) using the ‘sumaclust’ clustering algorithm (Mercier et al. 2013), which considers the most abundant sequence of each cluster as the cluster representative. OTUs were set at a sequence similarity threshold of 95%. To assign a taxon to plant OTUs, we built a reference sequence database using the ecoPCR programme (Ficetola et al. 2010) on the European Molecular Biology Laboratory (EMBL; release 141). OTUs were then assigned a taxonomy, using OBITOOL’s ecotag programme (Boyer et al. 2016), which performs a global alignment of each OTU sequence (the query) against each reference. The reference taxon assigned to each OTU corresponds to the Last Common Ancestor of all the best-match sequences for the query. </p> <p><br> Datasets were subsequently filtered to remove contaminants as well as artefacts such as PCR chimeras and remaining sequencing errors, using routines implemented in the metabaR R package (Zinger et al 2021), in R version 3.6.1 (R Development Core Team, 2013). The filtering process consisted of four steps: (i) a negative control-based filtering. OTUs whose maximum abundance was found in extraction/PCR negative controls were removed from the dataset, as they were likely to be reagent/aerosol contaminants, better amplified in the absence of competing DNA fragments as it is the case in biological samples. (ii) a reference-based filtering. OTUs which are too dissimilar from sequences available in reference databases are potential chimeras generated during sequencing and amplification. In this study, we chose to set similarity thresholds at 100%. (iii) an abundance-based filtering. This procedure targets incorrect assignment of a few numbers of sequences corresponding to true OTUs occurring to the wrong sample, a phenomenon called “tag-switching”. It consists in setting OTUs abundances to 0 in samples where their abundance represents < 0.03% of the total OTU abundance in the entire dataset. (iv) Finally, we conducted a PCR-based filtering by considering any PCR reaction that yielded less than 1000 reads as non-functional, and removed them from the dataset. The script used for implementing this is provided (<strong>metabaR_Eden_Plants_100sim.html)</strong>, with sequence data processed to remove contaminants, OTUs of low taxonomic resolution, and PCRs with too low a read count. The clean data is provided (<strong>eden_plant_postclean_100sim.rds</strong>).</p> <p>References:</p> <p>Boyer, F. <em>et al.</em> (2016) ‘obitools: a unix-inspired software package for DNA metabarcoding’, <em>Molecular Ecology Resources</em>, 16(1), pp. 176–182. doi:<a href="https://doi.org/10.1111/1755-0998.12428">10.1111/1755-0998.12428</a>.</p> <p>Donald, J. <em>et al. (2021) '</em>‘Multi-taxa environmental DNA inventories reveal distinct taxonomic and functional diversity in urban tropical forest fragments.‘ <em>Global Ecology and Conservation</em> 29 (2021): e01724.</p> <p>Mercier, C. <em>et al.</em> (2013) ‘SUMATRA and SUMACLUST: fast and exact comparison and clustering of sequences’, in <em>Programs and Abstracts of the SeqBio 2013 workshop. Abstract</em>. Citeseer, pp. 27–29.</p> <p>Taberlet, P. <em>et al.</em> (2007) ‘Power and limitations of the chloroplast trn L (UAA) intron for plant DNA barcoding’, <em>Nucleic Acids Research</em>, 35(3), pp. e14–e14. doi:<a href="https://doi.org/10.1093/nar/gkl938">10.1093/nar/gkl938</a>.</p> <p>Taberlet, P. <em>et al.</em> (2018) <em>Environmental DNA: For Biodiversity Research and Monitoring</em>. Oxford University Press.</p> <p>Team, R.C. (2013) <em>R: A language and environment for statistical computing</em>. Vienna, Austria.</p> <p>Zinger, L. <em>et al.</em> (2016) ‘Extracellular DNA extraction is a fast, cheap and reliable alternative for multi-taxa surveys based on soil DNA’, <em>Soil Biology and Biochemistry</em>, 96, pp. 16–19.</p> <p>Zinger, L. et al. (2021) ‘metabaR: An r package for the evaluation and improvement of DNA metabarcoding data quality’, Methods in Ecology and Evolution. DOI: <a href="https://doi.org/10.1111/2041-210X.13552">https://doi.org/10.1111/2041-210X.13552</a></p>
PNBA - Mauritania sedimentary carbon and eDNA raw data
<p>This dataset relates to the publication <PLACEHOLDER>. It contains sediment properties relating to carbon content and granulometry, eDNA metabarcoding and radionuclides required for the analysis.</p> <p> </p> <p><strong>Dataset and variable description --------------------------------------------------------------------------------</strong></p> <p><strong> core-extraction.csv</strong><br> Contains properties related to the sediment cores.</p> <ul> <li>core_id - Unique identifier used for sediment cores</li> <li>habitat - Habitat where core was sampled. One of "Sand" (bare sediment, no vegetation), "Cymodocea" (vegetated by Cymodocea nodosa) or "Zostera" (vegetated by Zostera noltei)</li> <li>date - Sampling date (YYYY/MM/DD)</li> <li>lat - Latitude of sampling site (decimal degrees, WGS84)</li> <li>lon - Longitude of sampling site (decimal degrees, WGS84)</li> <li>sampling_depth - Depth reached by core sampler (cm)</li> <li>core_length - Length of sediment core, as measured in the lab after halving (cm)</li> <li>compaction - Compaction factor of sediment core</li> </ul> <p><strong>loi-and-granulometry.csv</strong><br> Contains results from loss on ignition and dry sieving, used to study organic matter contents and granulometry.</p> <ul> <li>sample_id - Unique sample identifier. Created by concatenating core_id and the sample's middle depth</li> <li>core_id - Identifies which core the sample was retrieved from</li> <li>depth - Middle depth of the sample. All samples have a height of 2 cm. (cm)</li> <li>depth_corr - Middle depth of the sample, corrected for compaction</li> <li>volume - Sample volume (mL or cm3)</li> <li>volume_corr - Sample volume, corrected for compaction (mL or cm3)</li> <li>wet_sample - Wet sample weight (g)</li> <li>dry_sample - Dry sample weight (g)</li> <li>ground_sample - Weight of sample fraction used for loss on ignition (g)</li> <li>burned_sample - Weight of sample fraction used for loss on ignition, after submitting it to burning (g)</li> <li>water_content - Estimated water content of sample (fraction)</li> <li>om_content - Estimated organic matter content of sample (fraction)</li> <li>columns 90000-63000 μm to <1 μm - Weight of sediment particles with size within range described in column name (g)</li> <li>mean_phi - Mean phi of sample (Krumbein phi, φ)</li> <li>percent_gravel - Sediment fraction (in weight) classified as gravel (fraction)</li> <li>percent_sand - Sediment fraction (in weight) classified as sand (fraction)</li> <li>percent_mud - Sediment fraction (in weight) classified as mud(fraction)</li> </ul> <p><strong>radionuclides.csv</strong><br> Contains the results from gamma-ray spectroscopy</p> <ul> <li>core_id - Identifies which core the sample was retrieved from</li> <li>sample_id - Unique sample identifier. Created by concatenating core_id and the sample's middle depth</li> <li>depth - Middle depth of the sample. All samples have a height of 2 cm. (cm)</li> <li>depth_corr - Middle depth of the sample, corrected for compaction</li> <li>remaining columns - Element specific activity for element named in column (Bq/g). Columns with the suffix "_error" represent the measurement error (sigma = 2)</li> </ul>
18S sequences from eDNA surveys of a coral reef in the Maldives
<p>18S sequence data for eDNA samples collected in the Maldives and the associated metadata and tag codes.</p> <p>The title of the study/publication is: Field collections and environmental DNA surveys reveal topographic complexity of coral reefs as a predictor of cryptobenthic biodiversity across small spatial scales</p>
Data from: Koalas, friends, and foes – the application of airborne eDNA for the biomonitoring of threatened species
Open the record for dataset details and reuse information.
Data from: Reduced sampling intensity through key sampling site selection for optimal characterization of riverine fish communities by eDNA metabarcoding
Open the record for dataset details and reuse information.
Unlocking natural history collections to improve eDNA reference databases and biodiversity monitoring
Open the record for dataset details and reuse information.
Data from: Testing multiple substrates for terrestrial biodiversity monitoring using environmental DNA (eDNA) metabarcoding
Open the record for dataset details and reuse information.
Data from: A sedimentary eDNA record of the Atacama Trench reveals biodiversity changes in the most productive marine ecosystem
Open the record for dataset details and reuse information.
Data from: Large-scale eDNA sampling and hierarchical modeling elucidates the importance of stream habitat for eastern hellbender (<em>Cryptobranchus a. alleganiensis</em>) occupancy and eDNA detection
Open the record for dataset details and reuse information.
Data from: Ecological forensic testing: Using multiple primers for eDNA detection of marine vertebrates in an estuarine lagoon subject to anthropogenic influences
Open the record for dataset details and reuse information.
Data from: Treated like dirt: Robust forensic and ecological inferences from soil eDNA after challenging sample storage
Open the record for dataset details and reuse information.
Data from: Leveraging environmental DNA (eDNA) to optimize targeted removal of invasive fishes
Open the record for dataset details and reuse information.
Environmental nucleic acids: a field-based comparison for monitoring freshwater habitats using eDNA and eRNA
Open the record for dataset details and reuse information.
Data from: eDNA metabarcoding of log hollow sediments and soils highlights the importance of substrate type, frequency of sampling and animal size, for vertebrate species detection
Open the record for dataset details and reuse information.
Data from: Sorting states of environmental DNA: Effects of isolation method and water matrix on recovery of membrane-bound, dissolved, and adsorbed states of eDNA
Open the record for dataset details and reuse information.
Validating eDNA Measurements of the Richness and Abundance of Anurans at a Large Scale
Open the record for dataset details and reuse information.
A manager’s guide to using eDNA metabarcoding in marine ecosystems
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.