Skip to main content
zenodoopen

DNA reference database for macroinvertebrates

<p>This repository hosts a ready-to-use database for the study of macroinvertebrate diversity using DNA metabarcoding with the fwhF2/EPTDr2n primer set (Leese et al. 2022).</p> <p>The database was developed from the MIDORI database and completed with some data from the BOLD database.</p> <p>The process of preparation and curation of the database is described below:</p> <p>1. The MIDORI CO1 RAW v247 database was downloaded from the official servers.</p> <p>2. Reference sequences were matched against the fwhF2 forward primer sequences. Non-matching sequences were removed. Bases preceding the forward primer were removed.</p> <p>3. Reference sequences with a length of less than 142 bp were removed. Bases following the position 142 were removed.</p> <p>4. Taxonomy was simplified to keep only 8 ranks: superkingdom, kingdom, phylum, class, order, family, genus and species.</p> <p>5. Taxonomic nomenclature was harmonized using `refdb::refdb_clean_tax_harmonize_nomenclature`</p> <p>6. Extra words were removed from taxonomic names using `refdb::refdb_clean_tax_remove_extra`</p> <p>7. Subspecific information were removed from taxonomic names using `refdb::refdb_clean_tax_remove_subsp`</p> <p>8. Missing taxonomic names were identified and normalized using `refdb::refdb_clean_tax_NA`.</p> <p>9. Hybrids and taxonomic names with qualifiers of uncertainty were converted to NA.</p> <p>10. Duplicates (identical sequences and taxonomy) were removed.</p> <p>11. Sequences with more than ambiguous nucleotides (N) were removed).</p> <p>12. Sequences with low taxonomic precision (above phylum) were removed.</p> <p>13. A random subset of 10 sequences was retained for each taxa.</p> <p>14. For a given taxonomic rank, records with NA values were removed if they were not the only representative of the upper clade.</p> <p>15. For Switzerland: Missing EPT species and IBCH families were searched in BOLD and merged with the reference database. The same filters and cleaning steps (2-14) were then applied.</p> <p>&nbsp;</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4