Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
23
datasets available to search
ShareScore release 0.9.0
Dataset results
23 results for “database curation”
AusTraits: a curated plant trait database for the Australian flora
<p>AusTraits is a transformative database, containing measurements on the traits of Australia's plant taxa, standardised from hundreds of disconnected primary sources. So far, data have been assembled from > 300 distinct sources, describing > 500 plant traits and > 34,000 taxa.</p> <p>To handle the harmonising of diverse data sources, we use a reproducible workflow to implement the various changes required for each source to reformat it suitable for incorporation in AusTraits. Such changes include restructuring datasets, renaming variables, changing variable units, changing taxon names. While this repository contains the harmonised data, the raw data and code used to build the resource are also available on the project's GitHub repository, <a href="https://github.com/traitecoevo/austraits.build/">https://github.com/traitecoevo/austraits.build/</a>.</p> <p>Further information on the project is available at the project website <a href="https://austraits.org">austraits.org</a> and in the associated publication (see below).</p> <p><strong>CONTRIBUTORS</strong></p> <p>The project is jointly led by Dr Daniel Falster (UNSW Sydney), Dr Rachael Gallagher (Western Sydney University), Dr Elizabeth Wenk (UNSW Sydney), and Dr Hervé Sauquet (Royal Botanic Gardens and Domain Trust Sydney), with input from > 300 contributors from over > 100 institutions (see full list above). The project was initiated by Dr Rachael Gallagher and Prof Ian Wright while at Macquarie University.</p> <p>We are grateful to the following institutions for contributing data Australian National Botanic Garden, Brisbane Rainforest Action and Information Network, Kew Botanic Gardens, National Herbarium of NSW, Northern Territory Herbarium, Queensland Herbarium, Western Australian Herbarium, South Australian Herbarium, State Herbarium of South Australia, Tasmanian Herbarium, Department of Environment Land Water and Planning Victoria and the Royal Botanic Gardens Victoria.</p> <p>AusTraits has been supported by investment from the Australian Research Data Commons (ARDC), via their "Transformative data collections" (https://doi.org/10.47486/TD044) and "Data Partnerships" (https://doi.org/10.47486/DP720, https://doi.org/10.47486/DP720A) programs; and grants from the Australian Research Council (FT160100113, DE170100208, FT100100910) and Macquarie University, The ARDC is enabled by National Collaborative Research Investment Strategy (NCRIS).</p> <p><strong>ACCESSING AND USE OF DATA</strong></p> <p>The compiled AusTraits database is released under an open source licence (CC-BY), enabling re-use by the community.</p> <p>A requirement of use is that users cite the AusTraits resource paper, which includes all contributors as co-authors:</p> <blockquote> <p>Falster, Gallagher et al (2021) <em>AusTraits, a curated plant trait database for the Australian flora</em>. Scientific Data 8: 254, <a href="https://doi.org/10.1038/s41597-021-01006-6">https://doi.org/10.1038/s41597-021-01006-6</a></p> </blockquote> <p>In addition, we encourage users you to cite the original data sources, wherever possible.</p> <p>Note that under the license data may be redistributed, provided the attribution is maintained.</p> <p>The downloads below provide the data in two formats:</p> <ul> <li>austraits-X.X.X.zip: data in plain text format (.csv, .bib, .yml files). Suitable for anyone, including those using Python.</li> <li>austraits-X.X.X.rds: data as compressed R object. Suitable for users of R (see below).</li> <li> <div>austraits-X.X.X-flattened.rds: contains a flattened version of the dataset for direct loading in R; all data tables are joined into a wider format</div> </li> <li> <div>austraits-X.X.X-flattened.parquet: contains a flattened version of the dataset in parquet format; all data tables are joined into a wider format </div> </li> </ul> <p>For R users, access and manipulation of data is assisted with the <a href="http://github.com/traitecoevo/austraits">austraits R package</a>. The package can both download data and provides examples and functions for running queries.<br><br><strong>STRUCTURE OF AUSTRAITS</strong></p> <p>The compiled AusTraits database contains a series of relational tables and files. These elements include all the data, contextual information submitted with each contributed datasets, database schema, and trait definitions. The file dictionary.html provides the same information in textual format. Similar information is available at <a href="https://traitecoevo.github.io/traits.build-book/">https://traitecoevo.github.io/traits.build-book/</a>.</p> <p><strong>CONTRIBUTING</strong></p> <p>We envision AusTraits as an on-going collaborative community resource that:</p> <ol> <li>Increases our collective understanding the Australian flora;</li> <li>Facilitates accumulation and sharing of trait data;</li> <li>Builds a sense of community among contributors and users; and</li> <li>Aspires to fully transparent and reproducible research of the highest standard.</li> </ol> <p>As a community resource, we are very keen for people to contribute. Assembly of the database is managed on GitHub at <a href="https://github.com/traitecoevo/austraits.build/">https://github.com/traitecoevo/austraits.build/</a>.</p> <p>Here are some of the ways you can contribute:</p> <p><strong>Reporting Errors</strong>: If you notice a possible error in AusTraits, please <a href="https://github.com/traitecoevo/austraits.build/issues">post an issue on GitHub</a>.</p> <p><strong>Refining documentation:</strong> We welcome additions and edits that make using the existing data or adding new data easier for the community.</p> <p><strong>Contributing new data</strong>: We gladly accept new data contributions to AusTraits. See full instructions on how to contribute at <a href="https://github.com/traitecoevo/austraits.build/">https://github.com/traitecoevo/austraits.build/</a>.</p>
pofatu/pofatu-data: Pofatu, a curated and open-access database for geochemical sourcing of archaeological materials
<p>Geochemical fingerprinting of artefacts and sources has proven to be the most effective way to use material evidence in order to reconstruct strategies of raw material procurement, exchange systems, and mobility patterns among past societies. In order to facilitate access to this growing body of data and to promote comparability and reproducibility in provenance studies, we designed Pofatu, the first online and open-access database presenting geochemical compositions and contextual information for archaeological sources and artefacts.</p> <p>The data repository includes a compilation of geochemical data and supporting analytical metadata, as well as the archaeological provenance and context for each sample. All information on Samples related to sources and artefacts can be accessed on this platform or downloaded from Zenodo or GitHub.</p> <p>While most prehistoric quarries and surface procurement sources used in the past have yet to be identified, provenance studies must also integrate wide and reliable geological data. For this reason, we advise Pofatu users to also consult other open-access repositories focusing specifically on geological samples, such as GeoRoc and EarthChem.</p>
ChemTastesDB: A Curated Database of Molecular Tastants
<p><em><strong>ChemTastesDB</strong></em> is a database that includes curated information of 4075 molecular tastants. <strong><em>ChemTastesDB</em></strong> is distributed to the scientific community to expand the information of molecular tastants, which could assist the analysis of the relationships between molecular structure and taste, as well as <em>in</em> <em>silico</em> (QSAR/QSPR) studies for taste prediction. Examples of QSPR approaches for the prediction of molecular taste are given in the following publication: <em>Rojas, C., Abril-González, M., Ballabio, D. & García, F. (2025). ChemTastesPredictor: An ensemble of machine learning classifiers to predict the taste of molecular tastants. Chemometrics and Intelligent Laboratory Systems. 261, 105380. <a href="https://doi.org/10.1016/j.chemolab.2025.105380">https://doi.org/10.1016/j.chemolab.2025.105380</a>.</em></p> <p>The 4075 molecular tastants are categorized into one of the five basic tastes (sweet, bitter, umami sour and salty), as well as to other classes related to non-basic tastes (tasteless, non-sweet, non-bitter, multitaste and miscellaneous). The molecules are categorized into following ten classes: sweet (1313), bitter (1615), umami (220), sour (49), salty (16), multitaste (179), tasteless (232), non-sweet (304), non-bitter (28), and miscellaneous (119).</p> <p><strong><em>ChemTastesDB</em></strong> provides the following information for each molecule: name, PubChem CID, CAS registry number, canonical SMILES string, class taste and the reference to the scientific sources from where data were retrieved. In addition, the molecular structure in the HyperChem (<em>.hin</em>) format of each compound is provided.</p> <p>This is version 2.1 of the <em><strong>ChemTastesDB</strong></em>. In this new version, 1131 newly curated compounds were added. These new molecules were retrieved from 52 new bibliographic references.</p>
Sharkipedia: A Curated Open Access Database of Shark and Ray Life History Traits and Abundance Time-series
<p>This dataset represent the intial launch of Sharkipedia: a curated open access database of shark and ray life history traits and abundance time-series. A curated database of shark and ray biological data is increasingly necessary both to support fisheries management and conservation efforts, and to test the generality of hypotheses of vertebrate macroecology and macroevolution. Sharks and rays are one of the most charismatic, evolutionary distinct, and threatened lineages of vertebrates, comprising around 1,250 species. To accelerate shark and ray conservation and science, we developed Sharkipedia as a curated open-source database and research initiative to make all published biological traits and population trends accessible to everyone. Sharkipedia hosts information on 58 life history traits from 264 sources, for 170 species, from 39 families, and 12 orders related to length (n=9 traits), age (8), growth (12), reproduction (19), demography (5), and allometric relationships (5), as well as 871 population time-series from 202 species. Sharkipedia relies on the backbone taxonomy of the IUCN Red List and the bibliography of Shark-References. Sharkipedia has profound potential to support the rapidly growing data demands of fisheries management, international trade regulation as well as anchoring vertebrate macroecology and macroevolution.</p>
EukRibo: a manually curated eukaryotic 18S rDNA reference database
<p>EukRibo is a manually curated database of reference small-subunit ribosomal RNA gene (18S rDNA) sequences of eukaryotes, specifically aimed at taxonomic annotation of high-throughput metabarcoding datasets. Unlike other reference databases of ribosomal genes, it is not meant to exhaustively capture all publicly available 18S rDNA sequences from the INSDC repositories, but to represent a subset of highly trustable sequences covering the whole known diversity of eukaryotes, with a focus on protists, manually verified taxonomic identifications, and relatively low genetic redundancy.</p> <p>EukRibo is part of a suite of public resources generated by the UniEuk project (www.unieuk.org), which are all designed to follow a common taxonomic framework for maximal interoperability. The high level of taxonomic accuracy of EukRibo, together with a newly designed, phylogenetically-informed annotation approach, allow high confidence in the taxonomic annotation of environmental metabarcodes, as well as identification of new eukaryotic diversity at various taxonomic levels using a connected components approach.</p> <p>* * *</p> <p>Accompanying preprint available at <a href="https://doi.org/10.1101/2022.11.03.515105">https://doi.org/10.1101/2022.11.03.515105</a>.</p> <p>* * *</p> <p><strong>EukRibo ReadMe file, versions 1 and 2</strong></p> <p>Each EukRibo release consists of <strong>4 files</strong>:<br> - a <strong>tsv table </strong>containing the taxonomic and other information about the 18S rDNA sequences included in the release<br> - a <strong>fasta file </strong>containing the <strong>full sequences </strong>as retrieved from the INSDC repositories (NCBI, EMBL-EBI/ENA, DDBJ)<br> - a <strong>fasta file </strong>containing the <strong>variable region V4 </strong>extracted from all these sequences (based on the fragment amplified with the Tara-Oceans V4 primers)<br> - a <strong>fasta file </strong>containing the <strong>variable region V9 </strong>extracted from the subset of sequences where it is present (based on the fragment amplified with the Tara-Oceans V9 primers)</p> <p>The primary goal of EukRibo was to be used to annotate the EukBank meta-dataset of available V4 metabarcoding datasets, and therefore all sequences included in EukRibo contain the variable region V4.<br> Only a subset of these sequences (about 75%) also contain the variable region V9; this is because many 18S rDNA sequences in the INSDC repositories stop before the V9 fragment.</p> <p>Sequences with slightly incomplete V4 or V9 fragments were kept if phylogenetically useful - i.e. if they are the only available representatives of a certain taxonomic lineage.<br> <strong>V4 </strong>We allowed up to 50 missing positions in the relatively conserved area at the 5' end of the V4 fragment (for an average fragment length of about 380 bp); no sequence incomplete at the 3' end of the V4 fragment is included.<br> <strong>V9 </strong>We allowed up to 30 missing positions in the relatively conserved area at the 3' end of the V9 fragment (for an average length of about 135 bp); no sequence incomplete at the 5' end of the V9 fragment is included.<br> We allowed a higher proportion of missing positions for the V9 region because being more conservative would imply losing too many sequences, including entire taxonomic lineages.</p> <p><strong>Version 1 of EukRibo</strong><br> This is the starting version of EukRibo that was used for the taxonomic annotation of the EukBank dataset, with taxonomy strings that were fixed as of October 2020.<br> - Contains 46,345 sequences with a sufficiently complete V4 region; 46,299 with the actual complete V4 region and 46 (about 0.1%) with missing positions at the 5' end.<br> - Of these, 34,438 also include a sufficiently complete V9 region; 23,226 with the actual complete V9 region and 11,206 (about 33%) with missing positions at the 3' end.</p> <p><strong>Version 2 of EukRibo</strong><br> This is a version of EukRibo that was made taxonomically compatible with version 3 of the EukProt database (<a href="https://doi.org/10.1101/2020.06.30.180687">https://doi.org/10.1101/2020.06.30.180687</a>), with taxonomic revisions as of July 2022 as well as additional information on the included selection of sequences that was not provided in the tsv file of version 1.<br> - Contains the exact same selection of sequences as in version 1, with the addition of genus <em>Meteora</em>, the last remaining known supergroup-level eukaryotic lineage for which an 18S rDNA was not previously available. (The <em>Meteora </em>sequence contains the full V4 fragment but does not include a sufficiently complete V9 fragment.)<br> - Only 34,432 sequences with a sufficiently complete V9 region are now retained because of 6 previously unrecognised chimeric sequences where the V9 fragment does not originate from the same organism as the V4 fragment.</p> <p><strong>Files in EukRibo version 1</strong>:<br> 46345_EukRibo.tsv.gz<br> 46345_EukRibo_full_seqs.fas.gz<br> 46345_EukRibo_V4.fas.gz<br> 34438_EukRibo_V9.fas.gz</p> <p>The tsv file contains 6 columns:<br> <strong>gb_accession </strong>- INSDC accession number of the sequence<br> <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2 </strong>- binning of the taxa into strictly monophyletic clades of evolutionary and/or ecological significance<br> <strong>UniEuk_taxonomy_string </strong>- full UniEuk-compatible taxonomic annotation of the sequence<br> - an unlimited number of levels is allowed (going down to strain for isolated organisms or to clone for environmental sequences)<br> - informal names are used for phylogenetically supported clades without formal name<br> <strong>V9 </strong>- presence ('Y') or absence ('N') of a sufficiently complete V9 fragment in the sequence</p> <p><strong>Files in EukRibo version 2</strong>:<br> 46346_EukRibo-02.tsv.gz<br> 46346_EukRibo-02_full_seqs.fas.gz<br> 46346_EukRibo-02_V4.fas.gz<br> 34432_EukRibo-02_V9.fas.gz</p> <p>The tsv file now contains 12 columns:<br> <strong>gb_accession</strong>, <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2</strong>, <strong>UniEuk_taxonomy_string</strong><br> - same columns as in version 1<br> <strong>alternative_strain_names </strong>(new) - provides alternative strain/isolate names when known to help cross-linking genetic data coming from the same organism<br> <strong>V4 </strong>(new) - indicates whether the V4 fragment is complete ('yes - complete') or missing positions at the 5' end ('yes - partial')<br> <strong>V9 </strong>(emended content) - now contains more precise information than in version 1 about whether it is complete ('yes - complete'), missing positions at the 3' end ('yes - partial'), or was excluded, and the 6 possible reasons why ('no - missing', 'no - too incomplete', 'no - chimera', 'no - bad quality', 'no - deletion in V9', 'no - Ns in V9')<br> <strong>EukProt_ID_same_strain </strong>(new) - accession of EukProt datasets from the same isolate<br> <strong>EukProt_ID_different_strain </strong>(new) - accession of EukProt datasets from a different isolate of the same species<br> <strong>columns_modified_since_previous_version </strong>(new) - lists all of the 6 pre-existing columns that have a modified content compared to version 1<br> <strong>remarks </strong>(new) - additional information such as presence of an intron in the V9 fragment, taxonomic identity of the two parts of chimeric sequences, or the presence of Ns or a deletion in the V4 or the V9 fragment (but insufficient to warrant exclusion)</p>
TF-Marker: A comprehensive manually curated database for transcription factors and related markers in specific cell and tissue types in human.
<p>Here, we developed the TF-Marker database (TF-Marker, http://bio.liclab.net/TF-Marker/) which is committed to a comprehensive manual curation of TFs and related markers with experimental evidence in specific cell and tissue types in human. Currently, through reviewing <strong>2,091</strong> published literature, we have manually classified TFs and related markers into five types according to their functions: 1) <strong>TF</strong>: TFs, which regulate the expression of markers; 2) <strong>T Marker</strong>: markers, which are regulated by TFs (TF and T Marker pairs can identify cell types more specifically); 3) <strong>I Marker</strong>: markers, which influence the activity of TFs (I Markers can also influence the development of specific cells and tissues); 4) <strong>TFMarker</strong>: TFs, which play roles as markers (TFMarkers are cell/tissue-specific TFs used as cell markers in biology experiments); and 5) <strong>TF Pmarker</strong>: TFs, which play roles as potential markers. By curating thousands of published literature, <strong>5,905</strong> entries including <strong>1,316</strong> TFs, <strong>1,092</strong> T Markers, <strong>473</strong> I Markers, <strong>1,600</strong> TFMarkers and <strong>1,424</strong> TF Pmarkers, were annotated in <strong>383</strong> cell types and <strong>95</strong> tissue types in human. Moreover, TF-Marker divided markers into disease markers and tissue/cell-specific markers. TF-Marker is an elaborate database, which provides TFs and related markers supported by experimental evidence. We believe TF-Marker will provide strong support for research into cell/tissue-specific TFs and related markers.</p>
A curated database of fungal pathogens and their host range
<p>This database contains a manually curated set of human, animal and plant pathogens, annotated with their confirmed host range and relevant sources. In addition to that, we include additional sets of plant-associated fungi (which may include non-pathogens), as well as fungi with an automatically assigned, putative human, animal or plant host. The labelled fungal species are linked to their representative GenBank genomes wherever possible. Genomes that were screened, but no label was found, are also included.</p> <p><strong>[Last update on: 11 Dec 2022]</strong><br> [Home page: <a href="https://dacs-hpi.gitlab.io/pathogenic-fungi/">https://dacs-hpi.gitlab.io/pathogenic-fungi/</a>]<br> <br> The database is stored in a flat-file format. All metadata are stored in all_data_[date].csv, and all_data_[date].rds contains the same data in a compressed format that can be easily loaded in R. The database was first compiled on 9 Oct 2021 (v1.0), and then updated on 2 Jan 2022 (v1.1) and 11 Dec 2022 (v1.2).</p> <p>The core database is limited to manually confirmed human, animal and plant pathogens with available genomes as of 9 Oct 2021. Those data are a subset of all_data, and are stored in core_fungal_pathogens.csv and core_fungal_pathogens.rds.</p> <p>The temporal-test subset contains confirmed pathogens with genomes added to GenBank between 9 Oct 2021 and 2 Jan 2022.</p> <p>You may also be interested in trained neural network models predicting pathogenic potentials of novel fungi from DNA sequences (<a href="https://zenodo.org/record/5711877">https://zenodo.org/record/5711877</a>) and simulated Illumina read sets used to train them (<a href="https://zenodo.org/record/5846397">https://zenodo.org/record/5846397</a>).<br> <br> See also the preprint: <a href="https://www.biorxiv.org/content/10.1101/2021.11.30.470625">https://www.biorxiv.org/content/10.1101/2021.11.30.470625</a> and <strong>the paper</strong> presented at ECCB '22 and published in <em>Bioinformatics:</em> <a href="https://doi.org/10.1093/bioinformatics/btac495">https://doi.org/10.1093/bioinformatics/btac495.</a></p>
ActDES – a Curated Actinobacterial Database for Evolutionary Studies
<p>ActDES constitutes a novel resource for the community of Actinobacterial researchers that will be useful primarily for two types of analyses: (i) comparative genomic studies - facilitated by reliable orthologs identification across a set of defined, phylogenetically representative genomes, and (ii) phylogenomic studies which will be improved by identification of gene subsets at specified taxonomic level. These studies can then act as a springboard for the study of the evolution of virulence genes, studying the evolution of metabolism and metabolic engineering target identification.</p>
Curated new database list from the 2024 NAR Database issue - Table 1
<p>This is a curated list of 90 new databases recently published in the <a href="https://doi.org/10.1093/nar/gkad1173">2024 Nucleic Acids Research (NAR) Database issue</a>. These databases were specifically categorized as the "new databases" (not previously published) in Table 1 of the <a href="https://doi.org/10.1093/nar/gkad1173">2024 compilation paper</a>. In this dataset, we curated a set of data columns to facilitate knowledge modeling and database integration efforts:</p> <ul> <li>"has API": does the database provide API access, Yes or No.</li> <li>"Primary Entity Type": the primary biomedical entity types captured in the database.</li> <li>"Downloadable": does the database provide direct file downloads, Yes or No.</li> <li>"license specified?": does the database specify a data license explicitly, Yes or No.</li> <li>"knowledge level": the level of knowledge expressed in the database, example values are specified in <a href="https://biolink.github.io/biolink-model/knowledge_level/">the BioLink model</a>.</li> <li>"agent type": the high-level category of agent who originally generated a statement of knowledge or other type of information, example values are specified in <a href="https://biolink.github.io/biolink-model/agent_type/">the BioLink model</a>.</li> <li>"notes": optional extra notes</li> </ul>
Figure 1 in EphemBrazil: a curated online database and dashboard to explore the distribution of mayflies (Insecta: Ephemeroptera) from Brazil
Figure 1 General view of the website and interactive map view tab showing filters on the top and family subtitles in the right corner. Note that no filter is applied and all records are shown.
BFR2: a curated benthic foraminifera ribosomal reference database
<p>The present data set provides a fasta file, a tab-separated text file, and an Excel file. The fasta file includes 5,324 18S rDNA reference sequences for benthic foraminifera. The tab-separated text file includes the following fields: BFR2 number = unique internal sequence accession number; length = sequence length; class = class to which each sequence is assigned; order/suborder/clade = order/suborder/clade to which each sequence is assigned; family = family to which each sequence is assigned; genus = genus to which each sequence is assigned; species = species to which each sequence is assigned; isolate number = unique DNA extraction number; clone/direct: indicates whether it has been directly sequenced or cloned; genbank_accession = NCBI sequence accession number; sampling site = biogeographic region where the specimen has been collected; latitude in decimal degrees, longitude in decimal degrees; collection date = year in which specimen was collected; collector = person who collected the specimen; publication = publication associated to the sequence; journal = journal associated to publication; first author = first author associated to publication/sequence; taxonomic remarks = additional taxonomic information/comment; sampling remarks = addional sampling information/comment.<br><br>List of added or updated:<br>- Clone/direct: indicates whether it has been directly sequenced or cloned.<br>- Latitude and longitude in decimal degrees.<br>- Genbank accession number of 1700 18S rRNA sequences added.</p>
Advancing source tracking: systematic review and source-specific genome database curation of fecally shed prokaryotes
<p>This repository contains several files describing the data discussed in Lindner et al's "Advancing source tracking: systematic review and source-specific genome database curation of fecally shed prokaryotes". </p> <ol> <li>"database.fna" = concatenation of the (draft or complete) genome sequences described in the paper (n=12,730 source associated prokaryotic genomes) which passed quality checks and are species-level representatives (i.e., dereplicated at 95% ANI).</li> <li>"gdef.txt" = a manifest describing which sequences belong to which genomes.</li> <li>"sources.txt" = a manifest describing which genomes belong to which sources.</li> </ol> <p>The sources this database covers:</p> <table> <tbody> <tr> <td>Source Category</td> <td>Species-level <br>genome count</td> <td>Source-specific<br>species-level <br>genome count</td> </tr> <tr> <td>Bird</td> <td>56</td> <td>40</td> </tr> <tr> <td>Cat</td> <td>86</td> <td>6</td> </tr> <tr> <td>Chicken</td> <td>1314</td> <td>887</td> </tr> <tr> <td>Cow</td> <td>39</td> <td>15</td> </tr> <tr> <td>Dog</td> <td>139</td> <td>56</td> </tr> <tr> <td>Pig</td> <td>2764</td> <td>2035</td> </tr> <tr> <td>Ruminant</td> <td>740</td> <td>714</td> </tr> <tr> <td>Human</td> <td>4484</td> <td>3350</td> </tr> <tr> <td>Wastewater</td> <td>3108</td> <td>3097</td> </tr> </tbody> </table> <p> </p> <p>See publication for further details. </p>
Curated Phage Database (CPD) fasta file
<p>This is the fasta file including phage genomes used to generate the Curated Phage Database (CPD) utilized in our in review manuscript "The circulating phageome reflects bacterial infections". The corresponding phage characteristic data will be present in the manuscript as a supplemental file, and can be used to connect a Genbank ID to identified bacteriophage host and phage taxonomic information if known.</p> <p>Please note that this database is built from phage sequences in the NCBI nucleotide repository. Due to field bias towards sequencing human disease-related bacteria and their phage, this database is reflective of this bias and is most representative of bacteriophage associated with human pathogens and as such underrepresents environmental phages in comparison - a limitation to keep in mind when utilizing to interpret potential phage sequences.</p>
Database of 16S sequences from SILVA (r114), filtered, curated and annotated to be used easily by programs of taxonomic assignments
<p>The database used for the taxonomic assignment of reads generally comes from the SILVA database (http://www.arb-silva.de/). The logic behind this database is to use the information from the best one to the worst one. This is why the curated database was splitted in two parts : the [C] sequences for Complete sequences in terms of taxonomy, and the [I] and [E] sequences, for Incomplete and Environmental sequences.</p> <p>Each sequence included into the database must have a specific format summarizing all needed information (example below):<br> >[I]AACY020336309;Archaea(superkingdom);Euryarchaeota(phylum);Thermoplasmata(class);Thermoplasmatales(order);Marine_Group_II(no_rank);;marine_metagenome</p> <p>This sequence is an incomplete one ([I]), with a specific accession number from NCBI or SILVA, or another database (AACY020336309). Then, all taxonomic data is separated using ';' characters, for each considered level (superkingdom, phylum, <br> class, order, family, and genus). The species name is the last one and separated by two ';' characters from the rest of the descriptive line. Finally, the descriptive line must not contain specific characters like spaces. If one or several levels are unknown, this is indicated by 'no_rank'.</p> <p>Another example here for [C] sequences:<br> >[C]AAAK03000010;Bacteria(superkingdom);Firmicutes(phylum);Bacilli(class);Lactobacillales(order);Enterococcaceae(family);Enterococcus(genus);;Enterococcus_faecium_DO<br> This sequence is a complete one ([C]), with a specific accession number from NCBI or SILVA, or another database (AACY020187844). Then, all taxonomic data is separated using ';' characters, for each considered level (superkingdom, phylum, <br> class, order, family, and genus). The species is the last one and separated by two ';' characters from the rest of the descriptive line. Complete sequences must have six levels of information (superkingdom, phylum, class, order, family, and genus). If it is not the case, the sequence will be considered as Incomplete ([I]) (between three and five levels), or Environmental ([E]) (with only the superkingdom and the phylum levels).</p> <p>Another example here for [E] sequences:<br> >[E]U59968;Archaea(superkingdom);Thaumarchaeota(phylum);Soil_Crenarchaeotic_Group(SCG)(no_rank);;uncultured_crenarchaeote<br> This sequence is a environmental one ([E]), with a specific accession number from NCBI or SILVA, or another database (U59968). Then, all taxonomic data is separated using ';' characters, for each considered level (superkingdom, phylum, class, order, family, and genus). The species is the last one and separated by two ';' characters from the rest of the descriptive line. Complete sequences <br> must have six levels of information (superkingdom, phylum, class, order, family, and genus). If it is not the case, the sequence will be considered as Incomplete ([I]) (between three and five levels), or Environmental ([E]) (with only the superkingdom and the phylum levels).</p> <p>More details on the steps defined to clean and define this new database can be available on demand (sebastien.terrat@inra.fr).</p>
Creating, curating, and evaluating a mitogenomic reference database to improve regional species identification using environmental DNA
<p><span>Species detection using eDNA is revolutionizing global capacity to monitor biodiversity. However, the lack of regional, vouchered, genomic sequence information—especially sequence information that includes intraspecific variation—creates a bottleneck for management agencies wanting to harness the complete power of eDNA to monitor taxa and implement eDNA analyses. eDNA studies depend upon regional databases of mitogenomic sequence information to evaluate the effectiveness of such data to detect and identify taxa. We created the Oregon Biodiversity Genome Project to create a database of complete, nearly error-free mitogenomic sequences for all of Oregon's fishes. We have successfully assembled the complete mitogenomes of 313 specimens of freshwater, anadromous, and estuarine fishes representing 24 families, 55 genera, and 129 </span><span>species and lineages. Comparative analyses of these sequences illustrate that many regions of the mitogenome are taxonomically informative, that the short (~150 bp) mitochondrial "barcode" regions typically used for eDNA assays do not consistently diagnose for species, and that complete single or multiple genes of the mitogenome are preferable for identifying Oregon's fishes. This project provides a blueprint for other researchers to follow as they build regional databases, illustrates the taxonomic value and limits of complete mitogenomic sequences, and offers clues as to how current eDNA assays and environmental genomics methods of the future can best leverage this information.</span></p>
A curated database of milky sea observations from 1600 to present
Open the record for dataset details and reuse information.
Creating, curating, and evaluating a mitogenomic reference database to improve regional species identification using environmental DNA
Open the record for dataset details and reuse information.
Integrating QSAR models predicting acute contact toxicity and mode of action profiling in honey bees (A. mellifera): Data curation using open source databases, performance testing and validation
<p>This excel file (DOI: <a href="https://doi.org/10.5281/zenodo.3755675">https://doi.org/10.5281/zenodo.3755675</a>) provides the collection of raw data used for developing the first integrative Quantitative Structure-Activity Relationship (QSAR) model using EFSA's OpenFoodTox, US-EPA ECOTOX and Pesticide Properties DataBase i) to predict acute contact toxicity (LD<sub>50</sub>) and ii) to profile the Mode of Action (MoA) of pesticides active substances in honey bees (<em>Apis mellifera</em>)<em>. </em>Chemical identifiers (e.g. SMILES, CAS n., InChI) and acute contact toxicity data (LD<sub>50</sub>) on honey bees were used to develop and validate i) a two-category QSAR model (toxic/non-toxic; n=411) (sensitivity =0.93), specificity =0.85), balanced accuracy =0.90), Matthews correlation coefficient MCC=0.78), and ii) a regression-based model (n=113) (R2=0.74; MAE=0.52). Similarly, current study proposes the first MoA profiling for 113 pesticides active substances and the first harmonised MoA classification scheme for acute contact toxicity in honey bees, including LD<sub>50s</sub> data points from three different databases such as EFSA's OpenFoodTox, US-EPA ECOTOX and Pesticide Properties DataBase. Such classification allows to further define MoAs and the target site of Plant Protection Products (PPPs) active substances, thus enabling regulators and scientists to refine chemical grouping and toxicity extrapolations for single chemicals and component-based mixture risk assessment of multiple chemicals.</p> <p>The full data collection and analysis of QSAR models, toxicity data (LD<sub>50</sub>) and Mode of Action (Moa) data are described in Carnesecchi et al., 2020 (DOI: doi.org/10.1016/j.scitotenv.2020.139243).</p> <p>This work was supported by the European Food Safety Authority (EFSA) [contract number: OC/EFSA/SCER/2018/01 and NP/EFSA/AFSCO/2016/02 (Edoardo Carnesecchi)].</p>
World Ocean Database Dataset Curated for Analyzing ENSO-Driven Tropical Pacific Oxygen Variability
<p><strong>Description of folders</strong></p> <p><em>- ENSO index time series</em></p> <p>A file containing time series of the Oceanic Niño Index (ONI) is located in the following folder: ENSOtimeseries. ONI time series data was downloaded from http://origin.cpc.ncep.noaa.gov/products/analysis_monitoring/ensostuff/ONI_v5.php on June 4, 2018.</p> <p><em>- Exclusive economic zones (EEZs)</em></p> <p>A shapefile of the world's EEZs, downloaded from http://marineregions.org/downloads.php, is located in the following folder: EEZs.</p> <p><em>- World Ocean Database</em></p> <p><em>In situ</em> profiles of near-surface (0-700 m) temperature, salinity, and O<sub>2</sub> collected between January 1955 and May 2017 were downloaded from the World Ocean Database (WOD) at<a href="https://www.nodc.noaa.gov/OC5/SELECT/dbsearch/dbsearch.html"> </a><a href="https://www.nodc.noaa.gov/OC5/SELECT/dbsearch/dbsearch.html">https://www.nodc.noaa.gov/OC5/SELECT/dbsearch/dbsearch.html</a> on June 25th, 2017 and binned onto a monthly mean, 5-by-5 degree horizontal grid or grouped by EEZ.</p> <p>The binned monthly mean, 5-by-5 degree horizontal gridded data are in the following folders: WODglobalgridded, WODtropicalpacificonlygridded.</p> <p>The EEZ-grouped profiles are in the following folder: WODrawprofs.</p> <p>----------------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p><strong>Associated publication</strong></p> <p>This dataset was used to generate analyses and figures in the following publication:</p> <p>Leung, S., Thompson, L., McPhaden, M. J., & Mislan, K. A. S. (2019). ENSO drives near-surface oxygen and vertical habitat variability in the tropical Pacific. <em>Environmental Research Letters</em>.</p> <p>----------------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p><strong>Associated code</strong></p> <p>After downloading this dataset, run the associated MATLAB code at the following link to generate the figures and analyses in the above publication:</p> <p>http://doi.org/10.5281/zenodo.2648131</p>
AusInverTraits: a curated trait database for Australian invertebrates
<p>The AusInverTraits database is a curated and standardised database of traits data for Australia's invertebrate taxa gathered from primary sources. It is developed in collaboration with <a href="https://austraits.org/">AusTraits</a>.</p> <p> </p> <p><strong>Contributors</strong></p> <p>The database development is jointly led by Dr Payal Bal (University of Melbourne) and Dr Jessica R. Marsh (Charles Darwin University).</p> <p> </p> <p><strong>Data contributors (individuals)</strong>: xxxx</p> <p><strong>Data contributors (institutions)</strong>: xxx</p> <p><strong>Data processing</strong>: Payal Bal, Jane Ogilvie, Junn Kitt Foon, Jess Marsh, Elizabeth Wenk, Sophie Yang</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.