DNA reference database for macroinvertebrates
<p>This repository hosts a ready-to-use database for the study of macroinvertebrate diversity using DNA metabarcoding with the fwhF2/EPTDr2n primer set (Leese et al. 2022).</p> <p>The database was developed from the MIDORI database and completed with some data from the BOLD database.</p> <p>The process of preparation and curation of the database is described below:</p> <p>1. The MIDORI CO1 RAW v247 database was downloaded from the official servers.</p> <p>2. Reference sequences were matched against the fwhF2 forward primer sequences. Non-matching sequences were removed. Bases preceding the forward primer were removed.</p> <p>3. Reference sequences with a length of less than 142 bp were removed. Bases following the position 142 were removed.</p> <p>4. Taxonomy was simplified to keep only 8 ranks: superkingdom, kingdom, phylum, class, order, family, genus and species.</p> <p>5. Taxonomic nomenclature was harmonized using `refdb::refdb_clean_tax_harmonize_nomenclature`</p> <p>6. Extra words were removed from taxonomic names using `refdb::refdb_clean_tax_remove_extra`</p> <p>7. Subspecific information were removed from taxonomic names using `refdb::refdb_clean_tax_remove_subsp`</p> <p>8. Missing taxonomic names were identified and normalized using `refdb::refdb_clean_tax_NA`.</p> <p>9. Hybrids and taxonomic names with qualifiers of uncertainty were converted to NA.</p> <p>10. Duplicates (identical sequences and taxonomy) were removed.</p> <p>11. Sequences with more than ambiguous nucleotides (N) were removed).</p> <p>12. Sequences with low taxonomic precision (above phylum) were removed.</p> <p>13. A random subset of 10 sequences was retained for each taxa.</p> <p>14. For a given taxonomic rank, records with NA values were removed if they were not the only representative of the upper clade.</p> <p>15. For Switzerland: Missing EPT species and IBCH families were searched in BOLD and merged with the reference database. The same filters and cleaning steps (2-14) were then applied.</p> <p> </p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4