Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
41,503
datasets available to search
ShareScore release 0.9.0
Dataset results
41,503 results for “taxonom”
Environmental, Taxonomic, and Stable Isotope Data from Aquatic Insects sampled from Beaver-Engineered Headwater Streams (Adirondack Park, NY; 2024).
This data package contains environmental and biological data from a field study examining aquatic insect assemblage composition and basal resource use in beaver-engineered headwater streams in Adirondack Park, New York. Data was collected from six streams across two watersheds; the Oswegatchie River Watershed and Upper Hudson River Watershed. Three streams were sampled within the Oswegatchie River Watershed; East Creek, Sucker Brook, and Chair Rock Creek located near the Cranberry Lake Biological Station in St. Lawrence County. Three streams were sampled from the Upper Hudson River Watershed; Big Sucker Brook, Little Sucker Brook, and Panther Brook located near SUNY ESF’s Newcomb Campus in Essex County. Site conditions were characterized using densiometer measurements of canopy cover, visual assessments of substrate composition, and river discharge measurements collected with an OTT MF Pro flow meter. Aquatic insect assemblages were sampled using multihabitat active sampling and Hester–Dendy and leaf-bag passive samplers, with specimens identified to genus and assigned to functional feeding groups. Carbon and nitrogen stable isotopes were analyzed for a subset of insect taxa and three basal resource pools; coarse particulate organic matter (CPOM), fine particulate organic matter (FPOM), and periphytic algae. The Bayesian mixing model MixSIAR was used to estimate the proportional contribution of these primary sources to aquatic insect biomass. All data was collected between June and August 2024.
Microbial Community Composition in Lakes - Taxonomic characteristics of the clones at North Temperate Lakes LTER 2000 - 2007
Microbial community composition is inferred by a combination of automated ribosomal intergenic spacer analysis (ARISA) and PCR-generated clone library analysis. Clone libraries include both the 16S rRNA gene and the 16S-23S ribosomal intergenic spacer fragment. Phylogenetic assignments for individual ARISA fragments are obtained by comparing the ARISA fragment length from each clone to all of the profiles stored in our database. We have analyzed over 3900 clones obtained from 41 lakes that represent the range of trophic types found in temperate landscapes. Querying by taxonomic characteristics of the clone allows the user to retrieve clone IDs, sequence data, and characteristics of the sequence (length, chimera status, accession number, taxonomic affiliation). The data can be filtered by clone ID, ARISA fragment length (raw or binned), and/or taxonomic characteristics (Phylum and Phylum-Class). The output includes links to individual clone records, which contain more detailed information about how the clone was generated (researcher, library ID, project ID, primer sets used, etc.).
Taxonomic and ecological database of trees of Western Ghats - TreeGhatsData
<p><em>TreeGhatsData</em> is a compilation of lists of tree taxa found in Western Ghats, South India:</p> <ul> <li>taxa for which the word "tree" appears in habit description in the book <em>Flowering plants of the Western Ghats</em> edited by the Tropical Botanic Garden Research Institute (TBGRI), including planted or cultivated taxa (Nayar, Beegam, and Sibi. 2014);</li> <li>tree taxa described after 2014 in journal articles;</li> <li>taxon names used in forest surveys published by the French Institute of Pondicherry (IFP), in journal articles from 2000, and in the Atlas of endemics of the Western Ghats (Ramesh and Pascal 1997);</li> <li>taxon names reported with "tree" habit in Indian Biodiversity Portal (http://indiabiodiversity.org/).</li> </ul> <p>For each plant name, <em>TreeGhatsData</em> includes the following taxonomic information: family, genus epithet, species epithet, infrataxon rank, infrataxon epithet, authority. Both the family name used in TBGRI book and the corresponding family name according to Angiosperm Phylogeny Group system III (APGIII; Bremer et al. 2009) are provided.</p> <p><em>TreeGhatsData</em> includes the taxonomic status, the reference name and the authority according to TBGRI flora, along with taxonomic status from The Plant List version 1.1 (http://www.theplantlist.org/). From these two sources, a taxonomic status is suggested for each taxon name, with corresponding reference names and authorities.</p> <p><em>TreeGhatsData</em> also includes ecological and biogeographic information from TBGRI and completed by the botanists of French Institute of Pondicherry (IFP).</p> <p>Because most vegetation surveys do not provide taxon names at infraspecific level, <em>TreeGhatsData</em> includes both the infraspecific taxa mentioned in Western Ghats and the corresponding specific binomial names.</p> <p><em>TreeGhatsData</em> is provided as a CSV file with comma separator.</p> <p><strong>Related references</strong></p> <p>Bremer, B., Bremer, K., Chase, M. W., Fay, M. F., Reveal, J. L., Soltis, D. E., Soltis, P. S., Stevens, P. F., Anderberg, A. A., Moore, M. J., Olmstead, R. G., Rudall, P. J., Sytsma, K. J., Tank, D. C., Wurdack, K., Xiang, J. Q. Y. & Zmarzty, S. (2009) An update of the Angiosperm Phylogeny Group classification for the orders and families of flowering plants: APG III. Botanical Journal of the Linnean Society, 161, 105-121.</p> <p>Nayar, T., Rasiya Beegam, A. & Sibi, M. (2014) Flowering plants of the Western Ghats, India, Volume 1 Dicots; Volume 2 Monocots. Jawaharlal Nehru Tropical Botanic Garden and Research Institute.</p> <p>Ramesh, B. & Pascal, J.-P. (1997) Atlas of endemics of the Western Ghats (India): distribution of tree species in the evergreen and semi-evergreen forests. French Institute of Pondicherry, Pondicherry, India.</p>
Silva SSU taxonomic training data formatted for DADA2 (Silva version 138)
<p>These DADA2-formatted training fasta files were derived from the Silva Project's version 138 release. See https://www.arb-silva.de/documentation/release-138/ for database and citation information. The Silva 138 database is licensed under Creative Commons Attribution 4.0 (CC-BY 4.0); see file "SILVA_LICENSE.txt". The fasta files were generated and checked for consistency with version 132 using the R code in the R-markdown document "silva-v138.Rmd".</p> <p>Version 2 removes the dependence on preprocessed files from mothur, which results in a greater number of bacterial and archeal sequences. It also includes a new version of the assignTaxonomy training set that goes through the species level for use with longer amplicons obtained from long-read amplicon sequencing.</p> <p>If you use these files, please cite one or both of the Silva references below (or at the above link) and the DADA2 paper (reference below). I also recommend citing or linking to the Zenodo record for this specific version in your Methods or published source code to record the specific taxonomic database files used in your analysis.</p> <p><strong>NOTE:</strong><strong> </strong>These Version 2 files are intended for use in classifying prokaryotic 16S sequencing data and are not appropriate for classifying eukaryotic ASVs. The new method implemented within DADA2 for constructing these files only includes 100 eukaryotic sequences for use as an outgroup.</p> <p><strong>NOTE:</strong> These Version 2 files have a known problem in 10/883 families and 114/3838 genera. See https://github.com/mikemc/dada2-reference-databases/blob/main/silva-138/v2/bad-taxa.csv for a list of affected taxa and https://github.com/benjjneb/dada2/issues/1293 for more information.</p>
A Standardized Review of Bat Names Across Multiple Taxonomic Authorities
<p>The Bat Eco-Interactions Working Group, in collaboration with GBatNet and the international Bat Taxonomy Group, developed the <strong>Bat Taxonomic Alignment (BTA)</strong> to reconcile taxonomic discrepancies across currently recognized bat species. As knowledge of bat population structure and evolutionary history advances, taxonomic boundaries and species names are frequently revised. To address these changes, the BTA integrates data from ten leading taxonomic authorities and consolidates relationships, synonyms, and historic combinations for over <strong>1,480 valid bat species</strong> across <strong>1,680 taxonomic treatments</strong>.</p> <p>This open-access, searchable tool provides a time-calibrated inventory of Chiroptera taxonomy, documenting valid names, alternative names, subspecies, and synonymies. By aligning these classifications, the BTA enables users to identify unharmonized binomials and trace nomenclatural changes over time. It promotes taxonomic clarity critical for research, biodiversity assessments, and conservation planning, where misidentified or misaligned taxa can lead to gaps in knowledge, resource misallocation, or overlooked species. The BTA thus represents a foundational advancement in bat biodiversity informatics, emphasizing transparency, data provenance, and interoperability across digital taxonomic frameworks.</p> <p> </p>
Introduction to Ancient Metagenomics Textbook (Edition 2025): Taxonomic Profiling, OTU Tables, and Visualisation
<p>Data and conda software environment file for the chapter 'Taxonomic Profiling, OTU Tables, and Visualisation' of the SPAAM Community's textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>
Taxonomic list of Brazilian fruit-bearing plants for human use
<h3>Lista taxonômica de plantas frutíferas para consumo humano, com curadoria da equipe do projeto <a href="https://www.inaturalist.org/projects/pomar-urbano">Pomar Urbano</a>. </h3> <p><em>[see English description below]</em></p> <p><br>As planilhas estão organizadas da seguinte forma:</p> <p><strong>PT_lista_especies_aceitas_v.3.0</strong>: contém os nomes de todas as espécies atualmente indexadas na base de dados do <a href="https://www.inaturalist.org/projects/pomar-urbano">Pomar Urbano</a>.</p> <p><strong>PT_lista_especies_adicionadas_v.3.0</strong>: contém os nomes das novas espécies que passam a integrar a base de dados do Pomar Urbano a partir da versão 3.0.</p> <p><strong>PT_lista_especies_removidas_v3.0</strong>: contém os nomes das espécies removidas da versão 3.0 da lista, e que portanto não fazem mais parte do banco de dados do projeto. </p> <p> </p> <p><strong>Metadados usados nas planilhas:</strong></p> <ul> <li><em>Nome científico</em>: O nome científico completo, com autoria e data, se conhecidos.</li> <li><em>Família</em>: O nome científico completo da família.</li> <li><em>Nome vernacular</em>: nome comum, popular.</li> <li><em>Origem</em><strong>: </strong>Declaração sobre se um organismo foi introduzido em um local e tempo específicos por meio da atividade direta ou indireta dos seres humanos modernos.</li> <li><em>Distribuição geográfica</em>: área geográfica ou região onde uma espécie ocorre no Brasil. Foram considerados como valores válidos para este campo apenas as macrorregiões do Brazil, a saber: S = Sul, SE = Sudeste, CO = Centro-Oeste, NE = Nordeste, N = Norte.</li> <li><em>Última atualização</em>: A data mais recente em que a entrada no catálogo foi alterada, atualizada ou modificada.</li> </ul> <h3>--------------------------------------------------------------------------------------------------------------------------------------<br><br>Taxonomic list of fruit-bearing plants for human consumption, curated by the <a href="https://www.inaturalist.org/projects/pomar-urbano">Pomar Urbano project</a></h3> <p><em>[Vernacular names are presented only in Portuguese; for properly processing in data management tools, downloading a Portuguese language package might be necessary]</em></p> <p>The spreadsheets are organized as follows:</p> <p>EN_list_accepted_species_v.3.0: contains the names of all species currently indexed in the <a href="https://www.inaturalist.org/projects/pomar-urbano">Pomar Urbano</a> database.</p> <p>EN_new_added_species_v.3.0: contains the names of new species that are included in the Pomar Urbano database starting from version 3.0.</p> <p>EN_removed_species_v.3.0: contains the names of species that were present in the version 2.0 of the list and are therefore no longer part of the version 3.</p> <p> </p> <p><strong>Metadata used in the spreadsheets</strong>:</p> <p><em>Scientific Name</em>: The complete scientific name, including authorship and date, if known. <em>ExactMatch</em>: <a href="http://rs.tdwg.org/dwc/terms/scientificName">dwc:scientificName</a>. </p> <p><em>Family</em>: The full scientific name of the family. <em>ExactMatch</em>: <a href="http://rs.tdwg.org/dwc/terms/family">dwc:family.</a></p> <p><em>Vernacular Name</em>: Common or popular name. <em>ExactMatch</em>: <a href="http://rs.tdwg.org/dwc/terms/vernacularName">dwc:vernacularName</a></p> <p><em>Establishment Means</em>: Statement about whether an organism has been introduced to a specific place and time through the direct or indirect activity of modern humans. <em>ExactMatch</em>: <a href="http://rs.tdwg.org/dwc/terms/establishmentMeans">dwc:establishmentMeans</a></p> <p><em>Higher geography</em>: The geographical area or region where a species occurs in Brazil. Only the macroregions of Brazil are considered valid values for this field within this dataset, namely: S = South, SE = Southeast, CO = Central-West, NE = Northeast, N = North. <em>CloseMacth</em>: <a href="http://rs.tdwg.org/dwc/terms/higherGeography">dwc:higherGeography</a></p> <p><em>Last Update</em>: The most recent date on which the catalog entry was changed, updated, or modified. <em>ExactMatch</em>: <a href="http://purl.org/dc/terms/modified">dct:modified</a></p> <p> </p>
Metabarcoding data (number of reads per operational taxonomic unit) from a monitoring study of sandy beach meiofauna before and after sand nourishment (Ahrenshoop, Baltic Sea)
<p>We provide metabarcoding data (number of reads per operational taxonomic unit, OTU) determined from sediment samples collected on the sandy-beach water line of Ahrenshoop (Baltic Sea). Five sampling stations lay within the zone impacted by the sand nourishment between the boundary of the nature reserve in the north east and a site just north of the breakwater (AH01–AH05). An unaffected reference station was located south of Ahrenshoop (close to Niehagen) at the end of the road Pappelallee (PAP). Samples were collected at four dates. The first sampling was carried out before the sand nourishment took place (T0: 14 and 16 September 2021). Three samplings were realised after the impact: T1 (23 March 2022), T2 (27 September 2022), and T3 (28 March 2023). Latitude and longitude of each sampling location per station were recorded at each sampling date using a hand-held GPS application on a mobile phone. At the stations sampling locations varied over time. Prior to the sand nourishment the beach was narrow due to sand erosion in previous years. After the nourishment the additional extent of the beach was approximately 40 m at sampling date T1. Subsequently, progressive sand erosion forced the sampling locations (situated at the water line) further inland at T2 and T3.<br>Samples were taken from the beach-water interface (water line) in the middle of the area between two groynes. Plexiglass cores (inner core diameter 5.4 cm) were inserted vertically into the sediment down to 15 cm depth. Each core was sliced in 5 cm-layers (0–5, 5–10 and 10–15 cm). Sediment horizons were preserved in 96–99% ethanol. <br>Three cores (2 cores at T0) per sampling date were taken for metabarcoding analyses. The organisms were extracted by decantation over a 32-μm sieve. Genomic DNA was extracted from the filters using the DNeasy PowerSoil pro kit (Qiagen). Realtime-PCR was performed to amplify V1&V2, two hypervariable regions of 18S rDNA gene. The sequencing run was performed using the MiSeq Reagent Nanokit v2 (250 cycles paired end) on an Illumina MiSeq platform at the DZMB Metabarcoding lab in Wilhelmshaven, Germany. High-resolution amplicon sequence variants (ASVs) were obtained and compared to the NCBI database to assign taxonomic information to each ASV. The target meiofauna ASVs were further classified into operational taxonomic units (OTUs) with a 3% cut-off threshold using the statistical software R.</p> <p>Here, we present two Tables (as xlsx and tab-delimited files):<br>(1) the taxonomic description of the 843 OTUs and their assigned ID number;<br>(2) the number of reads per OTU per sample (including metadata for each sample: event; date; latitude; longitude; station, core and sample ID; sediment depth).</p> <p>The metabarcoding data are part of a larger ecological study on the influence of sand nourishment on meiofauna communities, which included grain-size and meiofauna abundances (see “related works”).</p> <p><strong>Comment: </strong>Our study is related to but not funded by the project ECAS Baltic: Strategies of ecosystem-friendly coastal protection and ecosystem-supporting coastal adaptation for the German Baltic Sea Coast <a href="https://deutsche-kuestenforschung.de/ecas-baltic.html">https://deutsche-kuestenforschung.de/ecas-baltic.html</a></p>
Metagenome quality metrics and taxonomical annotation visualization through the integration of MAGFlow and BIgMAG (Sup. Material)
<p>Dataset encompassing:</p> <ul> <li>The recovered MAGs by 6 different metagenomics pipelines (ATLAS, DATMA, MetaWRAP, MUFFIN, nf-core/mag and SnakeMAGs) using a mock community as input (SRR8359173 and SRR9328980), complemented with the output from MAGFlow (v1.0.0) using these MAGs as input for their quality assessment and taxonomical annotation. </li> <li>The MAGs produced by nf-core/mag using rice/rhizosphere sequenced libraries (PRJNA663614, PRJNA448773 and PRJNA645385) in either single assembly/single binning or co-assembly/co-binning mode, complemented with the output from MAGFlow (v1.0.0) using these MAGs as input for their quality assessment and taxonomical annotation.</li> <li>Scripts, commands and configuration files to run the different pipelines (ATLAS, DATMA, MetaWRAP, MUFFIN, nf-core/mag and SnakeMAGs) and reproduce the experimental conditions.</li> <li>Outputs, commands and scripts to run Metabinner and Semibin in their default configuration using the rice soil samples co-assembly, along with the MAGFlow (v1.1.0) output to compare these binners against MetaBAT2.</li> </ul>
Graphic Illustration of Kendra Phelp's Talk: A harmonized taxonomic resource is critical for accurately interpreting host-pathogen interactions
<p><a href="https://lib.ku.edu/people/courtney-foat" target="_blank" rel="noopener">Courtney Foat</a>, Advisor for Strategic Initiatives & Organizational Engagement at the University of Kansas, graphically recorded this invited talk by Kendra Phelps at an NSF-supported Workshop: Digital Collections Data and Tracking Disease.</p>
Global taxonomic occurrence grids using GBIF data for species distribution models.
<p>To achieve large geographic coverage, species occurrence databases that are composed of ad hoc species data collections such as that provided by the Global Biodiversity Information Facility (GBIF) are often used. A drawback to using these data is their geographic sampling bias, in which some regions are more intensively sampled than others, while other areas have very little to none reported sampling effort. Uneven sampling effort can mislead conclusions about biodiversity patterns and species distributions (Gotelli & Colwell, 2001; Lobo, 2008).</p> <p>Here we provide taxonomic occurrence grids to help mitigate the effects of sampling bias in species distribution modeling. These grids can be used to exclude areas of (a custom-defined) low sampling effort from the background when sampling for pseudo-absences’ (Phillips et al., 2009; Barbet-Massin et al.,2012). The occurrence grids have a 1 degree spatial resolution using WGS 84 as the geographic coordinate system. Each 1 degree grid cell contains the number of records present in GBIF corresponding to a specific taxonomic group: plants, mammals, reptiles, amphibians, birds and molluscs.</p> <p>To construct the occurrence grids, we used the 1- by 1-degree world latitude and longitude vector grid provided by ESRI (Redlands, California). It has a custom license which permits it reuse as long as ESRI is cited. It was downloaded from : <a href="https://www.arcgis.com/home/item.html?id=f11bcdc5d484400fa926dcce68de3df7">https://www.arcgis.com/home/item.html?id=f11bcdc5d484400fa926dcce68de3df7</a></p> <p>To map spatial sampling effort, the number of georeferenced occurrences corresponding to each taxonomic group contained by each 1- by 1-degree grid cell were counted. The grids were then converted to GeoTIFFs. The raster values correspond to the number of occurrences reported for the grid cells. For the purposes of the <a href="https://osf.io/7dpgr/">TrIAS project</a>, grid cells with fewer than 5 occurrences were removed. The TrIAS taxonomic occurrence grids are used as inputs to the TrIAS risk modelling and mapping workflow: https://github.com/trias-project/risk-modelling-and-mapping. Full (with all grid cells containing at least one occurrence) taxonomic occurrence grids are also provided.</p> <p>GBIF data for each taxonomic group were downloaded using the following criteria: “Basis of Record”: Observation, Machine Observation, Human Observation, Specimen, Material sample, Literature Occurrence, Unknown evidence., "HasCoordinate is true", "HasGeospatialIssue is false", "TaxonKey is Amphibia", "Year 1975-2005".</p> <p><strong>Raster Attributes</strong></p> <table> <tbody> <tr> <td> <p>Attribute</p> </td> <td> <p>Description</p> </td> </tr> <tr> <td> <p>OID</p> </td> <td> <p>numeric row ID</p> </td> </tr> <tr> <td> <p>Value</p> </td> <td> <p>the number of records contained in the grid cell</p> </td> </tr> <tr> <td> <p>Count</p> </td> <td> <p>the number of times the value appears in the raster</p> </td> </tr> </tbody> </table> <p> </p> <p> </p> <p>The extent of each taxonomic occurrence grid:</p> <ul> <li> <p>longitude -180.0; latitude -90.0 (southwest corner)</p> </li> <li> <p>longitude 180.0; latitude 90.0 (northeast corner)</p> </li> </ul> <p> </p> <p><strong>Files:</strong></p> <p>TrIAS taxonomic occurrence grids</p> <p>amphib_1deg_min5.tif</p> <p>birds_1deg_min5.tif</p> <p>mammals_1deg_min5.tif</p> <p>molluscs_1deg_min5.tif</p> <p>reptiles_1deg_min5.tif</p> <p> </p> <p>Raw taxonomic occurrence grids</p> <p>amphib_1deg_grid.tif</p> <p>birds_1deg_grid.tif</p> <p>mammals_1deg_grid.tif</p> <p>molluscs_1deg_grid.tif</p> <p>reptiles_1deg_grid.tif</p> <p><br> </p> <p> </p> <p> </p>
Molecular and Taxonomic Reevaluation of the Digitaria filiformis Complex (Poaceae) including a Globally Extinct, Single Site Endemic from New Hampshire, USA, and a New Species from Mexico
<p>We examine the <em>Digitaria filiformis </em>complex, to determine the proper taxonomic rank and rarity of each taxon. The taxonomy of the <em>D. filiformis </em>complex is highly debated and includes two widespread species, <em>D. filiformis </em>and <em>D. villosa</em>; a possibly extinct species endemic to a single-site in New Hampshire, <em>D. laeviglumis</em>; and a rare species of southern Florida and the West Indies, <em>D.</em><em> dolichophylla. </em>We conducted morphologic comparisons and molecular analysis of the four members of the <em>D. filiformis</em> complex, together with specimens from Mexico and Venezuela purportedly identified as <em>D. laeviglumis</em> (morphology only). Based on results of phylogenetic analyses of plastid and nuclear ITS sequences and morphologic comparisons, we recognize five species in the <em>D. filiformis </em>complex, including a newly described Mexican endemic <em>D. glabrifloris. </em>After field investigation we have moved the global rank of <em>D. laeviglumis </em>from globally historical (GH) to extinct (GX), as there is virtually no likelihood of rediscovery. <em>Digitaria</em><em> dolichophylla </em>is much rarer than previously recognized, moving from secure (T5) to imperiled with extinction (G2).</p>
The tpm metabarcoding DNA sequence database for taxonomic allocations using RDP classifier implemented in DADA2.
<p><strong>The </strong><em>tpm</em><strong> metabarcoding DNA sequence database for taxonomic allocations using the Mothur and DADA2 bio-informatic tools</strong></p> <p>A.C.M. Pozzi<sup>1</sup>, R. Bouchali<sup>1</sup>, L. Marjolet<sup>1</sup>, B. Cournoyer<sup>1</sup></p> <p><sup>1 </sup><em>University of Lyon, UMR Ecologie Microbienne Lyon (LEM), CNRS 5557, INRAE 1418, Université Claude Bernard Lyon 1, VetAgro Sup, Research Team “Bacterial Opportunistic Pathogens and Environment” (BPOE), 69280 Marcy L’Etoile, France.</em></p> <p><strong>Corresponding authors: </strong></p> <ul> <li>A.C.M. Pozzi, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L’Etoile, France. Tel. (+33) 478 87 39 47. Fax. (+33) 472 43 12 23. Email: <a href="mailto:adrien.meynier_pozzi@vetagro-sup.fr">adrien.meynier_pozzi@vetagro-sup.fr</a></li> <li>B. Cournoyer, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L’Etoile, France. Tel. (+33) 478 87 56 47. Fax. (+33) 472 43 12 23. Email: and <a href="mailto:benoit.cournoyer@vetagro-sup.fr">benoit.cournoyer@vetagro-sup.fr</a></li> </ul> <p><strong>Keywords:</strong></p> <p>BACtpm, Bacteria, <em>tpm</em>, thiopurine-<em>S</em>-methyltransferase EC:2.1.1.67, Nucleotide sequences, PCR products, Next-Generation-Sequencing, OTHU</p> <p><strong>Description:</strong></p> <ul> <li>The <em>tpm</em> gene codes for the thiopurine-<em>S</em>-methyltransferase (TPMT), an enzyme that can detoxify metalloid-containing oxyanions and xenobiotics (Cournoyer et al., 1998). Bacterial TPMTs radiated apart from human and animal TPMTs, and showed a vertical evolution in line with the 16S rRNA gene molecular phylogeny (Favre‐Bonté et al., 2005).</li> <li>The <em>tpm</em> database, named BACtpm, was designed to apply the <em>tpm</em>-metabarcoding analytical scheme published in Aigle et al. (2021). It includes the full <em>tpm</em> identifiers, GenBank accession numbers, complete taxonomic records (domain down to strain code) of about 215 nucleotide-long <em>tpm</em> sequences of 840 unique taxa belonging to 139 genera.</li> <li>Nucleotide sequences of <em>tpm</em> (range: 190-233 nucleotides) were either retrieved from public repositories (GenBank) or made available by B. Cournoyer’s research group. Colin et al. (2020) described the PCR and high throughput Illumina Miseq DNA sequencing procedures used to produce <em>tpm</em> sequences.</li> <li>BACtpm v.2.0.1 (June 2021 release) is made available under the Creative Commons Attribution 4.0 International Licence. It can be used for the taxonomic allocations of <em>tpm </em>sequences down to the species and strain levels. Data is stored in the csv format enabling future user to reformat it to fit their specific needs.</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>We thank the worldwide community of microbiologists who made contributions to public databases in the past decades, and made possible the elaboration of the BACtpm database. We also thank the Field Observatory in Urban Hydrology (OTHU, <a href="http://www.graie.org/othu/">www.graie.org/othu/</a>), Labex IMU (Intelligence des Mondes Urbains), the Greater Lyon Urban Community, the School of Integrated Watershed Sciences H2O'LYON, and the Lyon Urban School for their support in the development of this database. This work was funded by the French national research program for environmental and occupational health of ANSES under the terms of project “Iouqmer” EST 2016/1/120, l'Agence Nationale de la Recherche through ANR-16-CE32-0006, ANR-17-CE04-0010, ANR-17-EURE-0018 and ANR-17-CONV-0004, by the MITI CNRS project named Urbamic, and the French water agency for the Rhône, Mediterranean and Corsica areas through the Desir and DOmic projects. We thank former BPOE lab members who contributed to start and expand the BACtpm database: Céline COLINON, Romain MARTI, Emilie BOURGEOIS, Sébastien RIBUN and Yannick COLIN.</p> <p><strong>References:</strong></p> <p>Aigle, A., Colin, Y., Bouchali, R., Bourgeois, E., Marti, R., Ribun, S., Marjolet, L., Pozzi, A.C.M., Misery, B., Colinon, C., Bernardin-Souibgui, C., Wiest, L., Blaha, D., Galia, W., Cournoyer, B., 2021. Spatio-temporal variations in chemical pollutants found among urban deposits match changes in thiopurine S-methyltransferase-harboring bacteria tracked by the tpm metabarcoding approach. Sci. Total Environ. 767, 145425. https://doi.org/10.1016/j.scitotenv.2021.145425</p> <p>Colin, Y., Bouchali, R., Marjolet, L., Marti, R., Vautrin, F., Voisin, J., Bourgeois, E., Rodriguez-Nava, V., Blaha, D., Winiarski, T., Mermillod-Blondin, F., Cournoyer, B., 2020. Coalescence of bacterial groups originating from urban runoffs and artificial infiltration systems among aquifer microbiomes. Hydrol. Earth Syst. Sci. 24, 4257–4273. https://doi.org/10.5194/hess-24-4257-2020</p> <p>Cournoyer, B., Watanabe, S., Vivian, A., 1998. A tellurite-resistance genetic determinant from phytopathogenic pseudomonads encodes a thiopurine methyltransferase: evidence of a widely-conserved family of methyltransferases1The International Collaboration (IC) accession number of the DNA sequence is L49178.1. Biochim. Biophys. Acta BBA - Gene Struct. Expr. 1397, 161–168. https://doi.org/10.1016/S0167-4781(98)00020-7</p> <p>Favre‐Bonté, S., Ranjard, L., Colinon, C., Prigent‐Combaret, C., Nazaret, S., Cournoyer, B., 2005. Freshwater selenium-methylating bacterial thiopurine methyltransferases: diversity and molecular phylogeny. Environ. Microbiol. 7, 153–164. https://doi.org/10.1111/j.1462-2920.2004.00670.x</p>
Pond environmental and taxonomic data for Niwot Ridge and Green Lakes Valley, 2021 - ongoing.
This is a summary of basic environmental data and benthic macroinvertebrates from water in ponds in the vicinity of the Niwot Ridge LTER. Ponds were selected across a range of elevations, sizes, and positions relative to glacial, stream, and lake water sources. Ponds sampled occurred on Niwot Ridge and throughout the Green Lakes Valley.
Mesozooplankton taxonomic density collected using a 1-m diameter ring net with 200-μm mesh at Palmer Station, Antarctica during Palmer LTER field seasons, 2017-2020
Zooplankton are a morphologically and taxonomically diverse group of animals. Many zooplankton feed on phytoplankton in surface waters and thus provide a link between primary producers and higher trophic levels. The numerical density of common mesozooplankton taxa was determined at Palmer LTER Stations B and E. Samples were collected with a 1-m diameter, 200-μm mesh ring net towed obliquely from the surface to a target depth of 50 m and back. Tows were conducted during daytime, and sampling frequency was nominally twice weekly while personnel were at Palmer Station between the months of November and March. The preserved samples were size-fractionated with nested sieves into five size classes (0.2−0.5, 0.5−1, 1−2, 2−5, and >5 mm) prior to microscopic enumeration. Data are provided for the following taxa: copepods Oithona spp., Calanoides acutus (>1 mm only), Calanus propinquus (>1 mm only), Rhincalanus gigas (>1 mm only), and small calanoids (0.2−1 mm), chaetognaths, asteroid larvae, nemertean larvae, and foraminifera (not quantified in all years). Individual size fractions were split and subsampled such that at least 100 individuals of the most abundant taxon were present. Density varies across taxa, seasonally, among years, and between sampling stations. Units of density are individuals per cubic meter.
Macrozooplankton taxonomic density collected using a 1 x 1 m square net with 700-μm mesh at Palmer Station, Antarctica during Palmer LTER field seasons, 2017-2020
Zooplankton are a morphologically and taxonomically diverse group of animals. Many zooplankton feed on phytoplankton in surface waters and thus provide a link between primary producers and higher trophic levels. The numerical density of common macrozooplankton taxa was determined at Palmer LTER Stations B and E. Samples were collected with a 1 x 1 m square, 700-μm mesh Metro net towed obliquely from the surface to a target depth of 50 m and back. Duplicate tows typically were conducted at each sampling site. Tows were conducted during daytime, and sampling frequency was nominally twice weekly while personnel were at Palmer Station between the months of November and March. The catch was sorted and counted live. Data are provided for the following taxa, which dominated biomass: the euphausiids Euphausia superba and Thysanoessa macrura, the thecosome pteropod Limacina rangii, gymnosome pteropods, the salp Salpa thompsoni, amphipods, and larval fishes. Density varies across taxa, seasonally, among years, and between sampling stations. Units of density are individuals per cubic meter.
Phenology data set of plants and birds and other taxonomic groups, as well as agrarian activities and abiotic phenomena from Latvia, 1970-2018
<p>A data set of phenological observations of plants, birds, as well as agrarian activities and abiotic phenomena from Latvia, 1970-2018 is presented. The data include limited number of observations of insects, amphibians, mammals, mushrooms, mollusks and fishes as well. The data was collected by voluntary observers (citizen scientists) and published as paper based yearly bulletins. It includes almost 48 000 individual observations of 159 different phenological phases from 103 locations in Latvia. Each entry is comprised of following fields:</p> <ol> <li>Station: name of the observation station</li> <li>Year: year of observation</li> <li>Season: season of observation as indicated in the primary publication</li> <li>Species: English name of the species observed or description of phenomena observed in case of abiotic occurrences</li> <li>Species Latin: Latine name of the species observed</li> <li>Taxonomic_group: taxonomic group of the species observed or grouping of non-biological phases (“Abiotic” for meteorological phenomena and “Agrarian” for agrarian activities)</li> <li>Phenophase: description of phenological phase observed</li> <li>BBCH: attributed BBCH code for phenological phase observed, where applicable</li> <li>Date: date of the first observation of the phase</li> <li>DoY: day of the year of the first observation of the phase</li> <li>Implausible: flag indicating of the reported date of phenological phase is highly implausible (TRUE) or realistic (FALSE)</li> <li>Wrong_order: flag indicating if the order of the reported phases at a given station and year is not realistic (TRUE) or realistic (FALSE)</li> </ol>
Data release for "OrchID: a Generalized Framework for Taxonomic Classification of Images Using Evolved Artificial Neural Networks"
<p><strong>Abstract</strong></p> <p>Taxonomic expertise for the identification of species is rare and costly. On-going advances in computer vision and machine learning have led to the development of numerous semi- and fully automated species identification systems. However, these systems are rarely agnostic to specific morphology, rarely can perform taxonomic “approximation” (by which we mean partial identification at least to higher taxonomic level if not to species), and frequently rely on costly scientific imaging technologies.</p> <p>We present a generic, hierarchical identification system for automated taxonomic approximation of organisms from images. We assessed the effectiveness of this system using photographs of slipper orchids (Cypripedioideae), for which we implemented image pre-processing, segmentation, and colour and shape feature extraction algorithms to obtain digital phenotypes for 116 species. The identification system trained on these digital phenotypes uses a nested hierarchy of artificial neural networks for pattern recognition and automated classification that mirrors the Linnean taxonomy, such that user-submitted photos can be assigned a genus, section, and species classification by traversing this hierarchy.</p> <p>Performance of the identification system varied depending on photo quality, number of species included for training, and desired taxonomic level for identification. High quality photos were scarce for some taxa and were under-represented in the training set, resulting in imbalanced network training. The image features used for training were sufficient to reliably identify photos to the correct genus but less so to the correct section and species.</p> <p>The outcomes of this project include a library of feature extraction algorithms called <em>ImgPheno</em>, a collection of scripts for neural network training called <em>NBClassify</em>, a library for evolutionary optimization of artificial neural network construction called <em>AI::FANN::Evolving</em> and a planned web application called <em>OrchID</em> for identification of user-submitted images. All project outcomes are open source and freely available.</p> <p><strong>About this release</strong></p> <p>This release corresponds belongs with our response to the reviewers of PLoS One. At this stage of the review cycle the manuscript is assessed as 'minor revision'. Consequently, we don't anticipate making more releases until publication.</p>
Supplementary material 3: World Spider Catalog Bibliographic Data: Treatments from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063
List of journal/publisher by ranked by treatment count exported from the World Spider Catalog 14 October 2014 with total treatments by source, cumulative treatments, and cumulative proportion of treatments.
Supplementary material 2: World Spider Catalog Bibliographic Data: Publications from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063
Ranked list of journal/publisher exported from the World Spider Catalog 14 October 2014 with total articles by source, cumulative articles, and cucmulative proportion of articles.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.