Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
26
datasets available to search
ShareScore release 0.9.0
Dataset results
26 results for “taxonomic bias”
State of biodiversity documentation in the Philippines: Metadata gaps, taxonomic biases, and spatial biases in the DNA barcode data of animal and plant taxa in the context of species occurrence data
<p>These files can be categorized into three groups: (1) raw datasets obtained from public databases (i.e., GBIF, BOLD, and GenBank), (2) manually edited files needed for parsing and analysis, and (3) supplementary files for spatial analysis. All are used in the examination of gaps and biases present in Philippine biodiversity data, which can direct research on the taxa and spatial regions that need more sampling.</p>
What's in a name? Taxonomic and gender biases in the etymology of new species names
<p>As our inventory of Earth's biodiversity progresses, the number of species given a Latin binomial name is also growing. While the coining of species names is bound by rules, the sources of inspiration used by taxonomists are an eclectic mix. We investigated naming trends for nearly 2900 new species of parasitic helminths described in the past two decades. Our analysis indicates that the likelihood of new species being given names that convey some information about them (name derived from morphology, host, or locality of origin) or not (named after an eminent scientist, or for something else) depends on the higher taxonomic group to which the parasite or its host belongs. We also found a consistent gender bias among species named after eminent scientists, with male scientists being immortalised disproportionately more frequently than female scientists. Finally, we found that the tendency for taxonomists to name new species after a family member or close friend has increased over the past twenty years. We end by formulating recommendations for future species naming, aimed at honouring the diverse scientific community regardless of gender or ethnicity and avoiding etymological nepotism and cronyism, while still allowing for creativity in crafting new Latin species names.</p>
Data supplementing the article "Diatom DNA metabarcoding for biomonitoring : strategies to avoid major taxonomical and bioinformatical biases limiting molecular indices capacities" K. Tapolczai, F. Keck, A. Bouchez, F. Rimet, M. Kahlert and V. Vasselon submitted to "Frontiers in Ecology and Evolution" journal
<p>These data supplement the article "Diatom DNA metabarcoding for biomonitoring : strategies to avoid major taxonomical and bioinformatical biases limiting molecular indices capacities" K. Tapolczai, F. Keck, A. Bouchez, F. Rimet, M. Kahlert and V. Vasselon submitted to "Frontiers in Ecology and Evolution" journal.</p> <p>The directory contains the following files:</p> <p><strong>464_samples_fastq_files_(mothur).rar </strong>- contains the 464 fastq files proceed together during the Mothur bioinformatics treatments to produce the OTUs and ISUs tables. As the contig and the demultiplexing steps were performed by the sequencing platform, there is 1 fastq file per sample. From this 464 samples OTU/ISU tables, only information regarding 76 samples were used in this study and are listed in the "<strong>76_samples_list_(mothur).xlsx" </strong>file<strong>.</strong></p> <p><strong>76_samples_list_(mothur).xlsx </strong>- contains the information regarding the 76 samples used to create the OTUs and ISUs tables presented in the paper.</p> <p><strong>76_samples_R1_R2_fastq_files(DADA2).rar - </strong>contains the raw demultiplexed fastq files (R1.fastq and R2.fastq) for each of the 76 samples used in this study to produce the ESVs table using the DADA2 bioinformatics pipeline.</p>
It’s not fur: newspaper article reporting of abandonment and relinquishment of pets exhibit taxonomic biases in framing and language use
Open the record for dataset details and reuse information.
What’s in a name? Taxonomic and gender biases in the etymology of new species names
Open the record for dataset details and reuse information.
Data from: Environmental metabarcodes for insects: in silico PCR reveals potential for taxonomic bias
Studies of insect assemblages are suited to the simultaneous DNA-based identification of multiple taxa known as metabarcoding. To obtain accurate estimates of diversity, metabarcoding markers ideally possess appropriate taxonomic coverage to avoid PCR-amplification bias, as well as sufficient sequence divergence to resolve species. We used in silico PCR to compare the taxonomic coverage and resolution of newly designed insect metabarcodes (targeting 16S) with that of existing markers (16S and COI) and then compared their efficiency in vitro. Existing metabarcoding primers amplified in silico less than 75% of insect species with complete mitochondrial genomes available, whereas new primers targeting 16S provided greater than 90% coverage. Furthermore, metabarcodes targeting COI appeared to introduce taxonomic PCR-amplification bias, typically amplifying a greater percentage of Lepidoptera and Diptera species, while failing to amplify certain orders in silico. To test whether bias predicted in silico was observed in vitro, we created an artificial DNA blend containing equal amounts of DNA from 14 species, representing 11 different insect orders and one arachnid. We PCR-amplified the blend using five primers sets, targeting either COI or 16S, with high-throughput amplicon sequencing yielding more than 6 million reads. In vitro results typically corresponded to in silico PCR predictions, with newly designed 16S primers detecting 11 insect taxa present, thus providing equivalent or better taxonomic coverage than COI metabarcodes. Our results demonstrate that in silico PCR is a useful tool for predicting taxonomic bias in mixed template PCR, and that researchers should be wary of potential bias when selecting metabarcoding markers.
Figure 5 in Taxonomic bias in biodiversity data and societal preferences
Figure 5. Relation between age, origin and quality of the occurrence data for 24 taxonomic classes. Graph showing the first two axes of a Multiple Correspondence Analysis (MCA) performed on 5 million random occurrences. Labels in black represent the categories considered for all occurrences. Classes' names (in green) are placed at the average position of the class occurrences. Occurrence age contains eight time intervals and an Unknown Year category; data origin contains three categories: Specimen for specimen-based occurrences, Observation for observation-based occurrences, and Unknown for unknown origins; data quality contains four categories: Temporal issue for the lack of year or month, Spatial issues for the lack of coordinates, Both issues and No issue.
Figure 4 in Taxonomic bias in biodiversity data and societal preferences
Figure 4. Taxonomic heterogeneity in sampling, occurrence data origin and quality for 24 taxonomic classes. Top: Proportion of species per class recorded in GBIF with at least one occurrence (light green: p>1), with more than 20 occurrences (green: p>20), and with more than 20 spatially distinct occurrences (i.e. "decently" sampled – dark green: p>20d). For all classes, except Aves, less than 1/3 of all species are "decently" sampled. Classes are ranked according to their proportion of "decently" sampled species. Middle: Occurrence origin (basisOfRecord) for each class. Some classes like Amphibia have a high proportion of occurrences based on specimens (blue: living or preserved specimen, material samples or fossils), whereas others like Aves have a majority of occurrences based on observation (orange: machine or human observation, literature). Grey bars show occurrences where the record basis is unknown. Classes are ranked according to their proportion of specimenbased occurrences. Bottom: Data incompleteness. Proportion of occurrences with spatial (purple) or temporal (yellow) inaccuracies for each class. Spatial inaccuracy corresponds to an occurrence lacking coordinates or tagged has having geospatial issues by GBIF. Temporal inaccuracy corresponds to a sampling event with no specified month or year. Classes are ranked according to their proportion of occurrences with spatial issues.
Figure 3 in Taxonomic bias in biodiversity data and societal preferences
Figure 3. Biodiversity occurrences recorded in GBIF between 1900 and 2006. For each curve, the number of occurrences was plotted yearly. Top: black = all 24 classes considered together, yellow = Aves; Middle: yellow = Magnoliopsida, blue = Insecta, green = Liliopsida; Bottom: green = Actinopterygii, yellow = Mammalia, light blue = Reptilia, dark blue = Amphibia, orange = Florideophyceae, purple = Globothalamea.
Figure 2 in Taxonomic bias in biodiversity data and societal preferences
Figure 2. Evolution over time of the taxonomic bias for each class. The larger the circle, the higher the deviation from I, the 'ideal' number of occurrences per class if no taxonomic bias is observed. Red dots indicate negative deviations (i.e. shortfall in occurrences = under-represented classes); green dots indicate positive deviations (i.e. excess of occurrences = over-represented classes).
Figure 1 in Taxonomic bias in biodiversity data and societal preferences
Figure 1. Taxonomic bias in biodiversity occurrence data. The vertical line at x = 0 depicts the 'ideal' number of occurrences per class, where each class is sampled proportionally to its number of known species. Green and red bars show the classes that are over- and under-represented in the GBIF mediated database compared to this 'ideal' sampling, respectively. Insects lack>200 millions occurrences and birds have an excess of>200 millions occurrences compared to an unbiased taxonomic sampling. Because birds and insects are greatly over- and
Figure 4 in Removal of historical taxonomic bias and its impact on biogeographic analyses: a case study of Neotropical tardigrade fauna
Figure 4. PCoA of the dissimilarity values for species compositions for each province from biogeographic regions and transitions zones (Andean region (AR), South American transition zone (SATZ), Neotropical region (NR), and Mexican transition zone (MTZ)) considering a. all data ('false cosmopolitan' and 'indigenous species'; n = 51 provinces) and b. only 'indigenous data' (n = 43 provinces). Points represent provinces. All provinces are connected to the centroid (larger points with black outline), representing the mean of ordination values from all provinces from that area. Points furthest from the rest are interconnected, forming a polygon representing the space occupied by that area in ordination space. The AR is represented by blue points, lines and polygon; SATZ by light green points, lines and polygon; NR by pink points, lines and polygon, and MTZ by purple points, lines and polygon.
Figure 3 in Removal of historical taxonomic bias and its impact on biogeographic analyses: a case study of Neotropical tardigrade fauna
Figure 3. Species accumulation curves for all data ('false cosmopolitan' and 'indigenous species'), and only 'indigenous data' for a. the Neotropical region (NR), b. the Andean region (AR), c. the South American transition zone (SATZ), and d. the Mexican transition zone (MTZ). Orange curves represent all data, violet curves represent only 'indigenous data', and shaded areas around them represent their 95% confidence interval. Each publication containing species records was considered a survey.
Figure 2 in Removal of historical taxonomic bias and its impact on biogeographic analyses: a case study of Neotropical tardigrade fauna
Figure 2. Map of all incidence records of freshwater and limnoterrestrial tardigrades, from 1908 to 2023, present in the Andean and Neotropical regions proposed by Morrone (2015) and Morrone et al. (2022), respectively. Orange circles with black outline represent records. The hierarchy of compartmentalisation is from highest to lowest level: region, transition zone, subregion, dominion, and province. Each level can be subdivided into multiple lower levels under their name (e.g., a region consisting of several subregions). Each biogeographic province is coloured according to its transition zone, subregion or dominion. The Neotropical region (NR), in this study, is composed of the Antillean subregion (ASR), Brazilian subregion (BSR) and Chacoan subregion (CSR). The BSR is composed of the Mesoamerican dominion (MD), Pacific dominion (PD), Boreal Brazilian dominion (BBD) and South Brazilian dominion (BBD). The CSR is composed of the Southeastern Amazonian dominion (SAD), Chacoan dominion (CD) and Paraná dominion (PD). The Andean region (AR), in this study, is composed of the Central Chilean subregion (CCSR), Subantarctic subregion (SSR) and Patagonian subregion (PSR). The acronym for each region, transition zone, subregion or dominion is presented next to its name.
Figure 1 in Removal of historical taxonomic bias and its impact on biogeographic analyses: a case study of Neotropical tardigrade fauna
Figure 1. Biogeographic compartmentalisation of provinces from the Andean and Neotropical regions proposed by Morrone (2015) and Morrone et al. (2022), respectively. The hierarchy of compartmentalisation is from highest to lowest level: region, transition zone, subregion, dominion, and province. Each level can be subdivided into multiple lower levels under their name (e.g., a region consisting of several subregions). Each biogeographic province is coloured according to its transition zone, subregion or dominion. The Neotropical region (NR), in this study, is composed of the Antillean subregion (ASR), Brazilian subregion (BSR) and Chacoan subregion (CSR). The BSR is composed of the Mesoamerican dominion (MD), Pacific dominion (PD), Boreal Brazilian dominion (BBD) and South Brazilian dominion (BBD). The CSR is composed of the Southeastern Amazonian dominion (SAD), Chacoan dominion (CD) and Paraná dominion (PD). The Andean region (AR), in this study, is composed of the Central Chilean subregion (CCSR), Subantarctic subregion (SSR) and Patagonian subregion (PSR). The acronym for each region, transition zone, subregion or dominion is presented next to its name. Under each subregion or dominion, the names of the comprising provinces are listed. The nature of the records present in each province is indicated by an icon of a coloured tardigrade next to its name: provinces without records of tardigrade species (red tardigrade), provinces with only records of 'false cosmopolitan species' (blue tardigrade), provinces with records of 'false cosmopolitan' and 'indigenous species' (orange tardigrade), and provinces with only records of 'indigenous species' (violet tardigrade). The tardigrade icon is in the public domain and was obtained from Phylopic (https://www.phylopic.org).
Figure 5 in Removal of historical taxonomic bias and its impact on biogeographic analyses: a case study of Neotropical tardigrade fauna
Figure 5. Tanglegram of consensus tree exhibiting the similarity between all biogeographic provinces from biogeographic regions and transitions zones (Andean region (AR), South American transition zone (SATZ), Neotropical region (NR), and Mexican transition zone (MTZ)) for all data ('false cosmopolitan' and 'indigenous species', on the left) and only 'indigenous data' (on the right). Each province is connected to itself by a dark grey line. The consensus tree was obtained by resampling (1000×) the row order with the 'recluster' package (Dapporto et al. 2020). Each province is coloured according to their area of origin: provinces from the AR are coloured blue, provinces from the SATZ are coloured light green, provinces from the NR are coloured pink, and provinces from the MTZ are coloured purple. Branches of provinces that appear unrelated to any other are bold and red-coloured. Provinces without any records of 'indigenous species' are not present in the tanglegram. Provinces without any records of tardigrade species were excluded from the clustering analysis.
Table 2 in Removal of historical taxonomic bias and its impact on biogeographic analyses: a case study of Neotropical tardigrade fauna
<p><b>Table 2.</b> Extrapolated richness of limnoterrestrial and freshwater tardigrade species and standard error for each region and transition zone. Estimations were made with all data (‘false cosmopolitan’ and ‘indigenous species’) and only ‘indigenous data for comparison. Estimations were made using the Jackknife1 estimator (based on: Burnham and Overton 1978, 1979).</p><table><tbody><tr><th><b>Area</b></th><th><b>Dataset</b></th><th><b>Number of provinces with at least one record</b></th><th><b>Total number of provinces</b></th><th><b>Observed richness</b></th><th><b>Estimated richness</b></th><th><b>Standard error</b></th></tr></tbody><tbody><tr><th>Neotropical region</th><td>All data (‘false cosmopolitan’ and ‘indigenous species’)</td><td>33</td><td>50</td><td>186</td><td>282.969</td><td>29.094</td></tr><tr><td>‘Indigenous data’</td><td>29</td><td>50</td><td>96</td><td>157.793</td><td>19.499</td></tr><tr><th>Andean region</th><td>All data (‘false cosmopolitan’ and ‘indigenous species’)</td><td>7</td><td>9</td><td>105</td><td>159.857</td><td>31.489</td></tr><tr><td>‘Indigenous data’</td><td>7</td><td>9</td><td>43</td><td>68.714</td><td>14.154</td></tr><tr><th>South America transition zone</th><td>All data (‘false cosmopolitan’ and ‘indigenous species’)</td><td>6</td><td>7</td><td>66</td><td>102.666</td><td>25.011</td></tr><tr><td>‘Indigenous data’</td><td>3</td><td>7</td><td>26</td><td>41.330</td><td>12.995</td></tr><tr><th>Mexican transition zone</th><td>All data (‘false cosmopolitan’ and ‘indigenous species’)</td><td>5</td><td>5</td><td>41</td><td>65.800</td><td>14.881</td></tr><tr><td>‘Indigenous data’</td><td>4</td><td>5</td><td>11</td><td>19.250</td><td>5.068</td></tr></tbody></table>
Data from: Taxonomic identification bias does not drive patterns of abundance and diversity in theropod dinosaurs
<p>The ability of palaeontologists to correctly diagnose and classify new fossil species from incomplete morphological data is fundamental to our understanding of evolution. Different parts of the vertebrate skeleton have different likelihoods of fossil preservation and varying amounts of taxonomic information, which could bias our interpretations of fossil material. Substantial previous research has focused on the diversity and macroevolution of non-avian theropod dinosaurs. Theropods provide a rich dataset for analysis of the interactions between taxonomic diagnosability and fossil preservation. We use specimen data and formal taxonomic diagnoses to create a new metric, the Likelihood of Diagnosis (LoD), which quantifies the diagnostic likelihood of fossil species in relation to bone preservation potential. We use this to assess whether a taxonomic identification bias impacts the non-avian theropod fossil record. We find the patterns of differential species abundance and clade diversity are not a consequence of their relative diagnosability. Although there are other factors that bias the theropod fossil record, our results suggest patterns of relative abundance and diversity for theropods might be more representative of Mesozoic ecology than often considered.</p>
Data from: Taxonomic structure of the fossil record is shaped by sampling bias
Open the record for dataset details and reuse information.
Data from: Environmental metabarcodes for insects: in silico PCR reveals potential for taxonomic bias
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.