Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
12
datasets available to search
ShareScore release 0.7.1
Dataset results
12 results for “RDP”
The tpm metabarcoding DNA sequence database for taxonomic allocations using RDP classifier implemented in DADA2.
<p><strong>The </strong><em>tpm</em><strong> metabarcoding DNA sequence database for taxonomic allocations using the Mothur and DADA2 bio-informatic tools</strong></p> <p>A.C.M. Pozzi<sup>1</sup>, R. Bouchali<sup>1</sup>, L. Marjolet<sup>1</sup>, B. Cournoyer<sup>1</sup></p> <p><sup>1 </sup><em>University of Lyon, UMR Ecologie Microbienne Lyon (LEM), CNRS 5557, INRAE 1418, Université Claude Bernard Lyon 1, VetAgro Sup, Research Team “Bacterial Opportunistic Pathogens and Environment” (BPOE), 69280 Marcy L’Etoile, France.</em></p> <p><strong>Corresponding authors: </strong></p> <ul> <li>A.C.M. Pozzi, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L’Etoile, France. Tel. (+33) 478 87 39 47. Fax. (+33) 472 43 12 23. Email: <a href="mailto:adrien.meynier_pozzi@vetagro-sup.fr">adrien.meynier_pozzi@vetagro-sup.fr</a></li> <li>B. Cournoyer, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L’Etoile, France. Tel. (+33) 478 87 56 47. Fax. (+33) 472 43 12 23. Email: and <a href="mailto:benoit.cournoyer@vetagro-sup.fr">benoit.cournoyer@vetagro-sup.fr</a></li> </ul> <p><strong>Keywords:</strong></p> <p>BACtpm, Bacteria, <em>tpm</em>, thiopurine-<em>S</em>-methyltransferase EC:2.1.1.67, Nucleotide sequences, PCR products, Next-Generation-Sequencing, OTHU</p> <p><strong>Description:</strong></p> <ul> <li>The <em>tpm</em> gene codes for the thiopurine-<em>S</em>-methyltransferase (TPMT), an enzyme that can detoxify metalloid-containing oxyanions and xenobiotics (Cournoyer et al., 1998). Bacterial TPMTs radiated apart from human and animal TPMTs, and showed a vertical evolution in line with the 16S rRNA gene molecular phylogeny (Favre‐Bonté et al., 2005).</li> <li>The <em>tpm</em> database, named BACtpm, was designed to apply the <em>tpm</em>-metabarcoding analytical scheme published in Aigle et al. (2021). It includes the full <em>tpm</em> identifiers, GenBank accession numbers, complete taxonomic records (domain down to strain code) of about 215 nucleotide-long <em>tpm</em> sequences of 840 unique taxa belonging to 139 genera.</li> <li>Nucleotide sequences of <em>tpm</em> (range: 190-233 nucleotides) were either retrieved from public repositories (GenBank) or made available by B. Cournoyer’s research group. Colin et al. (2020) described the PCR and high throughput Illumina Miseq DNA sequencing procedures used to produce <em>tpm</em> sequences.</li> <li>BACtpm v.2.0.1 (June 2021 release) is made available under the Creative Commons Attribution 4.0 International Licence. It can be used for the taxonomic allocations of <em>tpm </em>sequences down to the species and strain levels. Data is stored in the csv format enabling future user to reformat it to fit their specific needs.</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>We thank the worldwide community of microbiologists who made contributions to public databases in the past decades, and made possible the elaboration of the BACtpm database. We also thank the Field Observatory in Urban Hydrology (OTHU, <a href="http://www.graie.org/othu/">www.graie.org/othu/</a>), Labex IMU (Intelligence des Mondes Urbains), the Greater Lyon Urban Community, the School of Integrated Watershed Sciences H2O'LYON, and the Lyon Urban School for their support in the development of this database. This work was funded by the French national research program for environmental and occupational health of ANSES under the terms of project “Iouqmer” EST 2016/1/120, l'Agence Nationale de la Recherche through ANR-16-CE32-0006, ANR-17-CE04-0010, ANR-17-EURE-0018 and ANR-17-CONV-0004, by the MITI CNRS project named Urbamic, and the French water agency for the Rhône, Mediterranean and Corsica areas through the Desir and DOmic projects. We thank former BPOE lab members who contributed to start and expand the BACtpm database: Céline COLINON, Romain MARTI, Emilie BOURGEOIS, Sébastien RIBUN and Yannick COLIN.</p> <p><strong>References:</strong></p> <p>Aigle, A., Colin, Y., Bouchali, R., Bourgeois, E., Marti, R., Ribun, S., Marjolet, L., Pozzi, A.C.M., Misery, B., Colinon, C., Bernardin-Souibgui, C., Wiest, L., Blaha, D., Galia, W., Cournoyer, B., 2021. Spatio-temporal variations in chemical pollutants found among urban deposits match changes in thiopurine S-methyltransferase-harboring bacteria tracked by the tpm metabarcoding approach. Sci. Total Environ. 767, 145425. https://doi.org/10.1016/j.scitotenv.2021.145425</p> <p>Colin, Y., Bouchali, R., Marjolet, L., Marti, R., Vautrin, F., Voisin, J., Bourgeois, E., Rodriguez-Nava, V., Blaha, D., Winiarski, T., Mermillod-Blondin, F., Cournoyer, B., 2020. Coalescence of bacterial groups originating from urban runoffs and artificial infiltration systems among aquifer microbiomes. Hydrol. Earth Syst. Sci. 24, 4257–4273. https://doi.org/10.5194/hess-24-4257-2020</p> <p>Cournoyer, B., Watanabe, S., Vivian, A., 1998. A tellurite-resistance genetic determinant from phytopathogenic pseudomonads encodes a thiopurine methyltransferase: evidence of a widely-conserved family of methyltransferases1The International Collaboration (IC) accession number of the DNA sequence is L49178.1. Biochim. Biophys. Acta BBA - Gene Struct. Expr. 1397, 161–168. https://doi.org/10.1016/S0167-4781(98)00020-7</p> <p>Favre‐Bonté, S., Ranjard, L., Colinon, C., Prigent‐Combaret, C., Nazaret, S., Cournoyer, B., 2005. Freshwater selenium-methylating bacterial thiopurine methyltransferases: diversity and molecular phylogeny. Environ. Microbiol. 7, 153–164. https://doi.org/10.1111/j.1462-2920.2004.00670.x</p>
RDP Classifier 2.14 and the RDP bacterial and archaeal taxonomy training set No. 19
<p>RDP Classifier 2.14 (August 2023) Release Note:</p><p>The Bacteria and Archaea hierarchy model used by RDP Classifier has been updated to training set No. 19. The new version has over 600 genera and 2500 species added since last version No. 18 released in July 2020. The information that is used to update the RDP taxonomy to training set version No. 19, and RDP Classifier version 2.14 came from publicly available scientific articles and public sequence repository, mostly from International Journal of Systematic and Evolutionary Microbiology (IJSEM), the All-Species Living Tree Project (LTP) and GenBank. </p><p>It is worth noting that most of the phyla have new names, according to article "</p><p>Oren A, Garrity GM. Valid publication of the names of forty-two phyla of prokaryotes. Int J Syst Evol Microbiol. 2021 Oct;71(10). doi: 10.1099/ijsem.0.005056. PMID: 34694987."</p><p>In addition to the files to train and run the RDP Classifier, new file formats are made available to accommodate the needs of users:</p><p>1. A new file trainset19_072023_speciesrank.fa has been added to the release in RDPClassifier_16S_trainsetNo19_rawtrainingdata.zip. This file is NOT needed to train the classifier. In addition to sequences, it contains genus, species, strain, type status and taxonomy rank, which are useful for closest species identification using third-party tools (e.g. BLAST).</p><p>2. Two new files in RDPClassifier_16S_trainsetNo19_QiimeFormat.zip to retrain the RDP Classifier included in Qiime2 package.</p>
RDP taxonomic training data formatted for DADA2 (RDP release 19 - update 2023-08-23)
<p>These DADA2-formatted training fasta files were derived from the Ribosomal Database Project's Training Set 19 and the 2023-08-23 release of the RDP database. https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/</p> <p>These fastas were generated by the following commands using the dada2 R package version 1.35.4:</p> <blockquote> <p>## RDP data: https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/</p> <p>path <- "~/tax/rdp/v19"<br>fn.out.rdp <- "~/Desktop/rdp_19_toGenus_trainset.fa.gz"<br>dada2:::makeTaxonomyFasta_RDP(file.path(path, "trainset19_072023_speciesrank.fa"), <br> file.path(path, "trainset19_db_taxid.txt"), <br> fn.out.rdp, include.species=FALSE,<br> compress=TRUE)</p> <p>fn.out.spc.rdp <- "~/Desktop/rdp_19_toSpecies_trainset.fa.gz"<br>dada2:::makeTaxonomyFasta_RDP(file.path(path, "trainset19_072023_speciesrank.fa"), <br> file.path(path, "trainset19_db_taxid.txt"), <br> fn.out.spc.rdp, include.species=TRUE,<br> compress=TRUE)</p> </blockquote>
RDP Classifier training files for 16S rRNA sequences from GTDB
<p>16S rRNA gene sequences from the <a href="https://gtdb.ecogenomic.org/">Genome Taxonomy Database</a> (GTDB release 220) were used to retrain the <a href="https://github.com/rdpstaff/classifier">RDP Classifier</a> (version 2.13). Two sets of training files are provided:</p> <ul> <li><code>genus.zip</code> - Genus level</li> <li><code>species.zip</code> - Species level</li> </ul> <p>The code in <code>prepare_files.R</code> was used to prepare the GTDB sequence and taxonomy files for retraining the RDP Classifier. Notes:</p> <ul> <li>Steps to retrain the RDP Classifier are adapted from <a href="https://john-quensen.com/tutorials/training-the-rdp-classifier/">https://john-quensen.com/tutorials/training-the-rdp-classifier/</a></li> <li>Python scripts (lineage2taxTrain.py and addFullLineage.py) are available at <a href="https://github.com/rdpstaff/classifier/issues/18">https://github.com/rdpstaff/classifier/issues/18</a></li> <li>The first 1000 training sequences (<code>train_nodups_1000.fasta</code>) are used for benchmarking the classification accuracy (see results at end of <code>prepare_files.R</code>).</li> </ul>
The RDP and GreenGenes taxonomic training sets formatted for DADA2
<p>These DADA2-formatted training fastas were derived from the Ribosomal Database Project (RDP Training Set 14) and the GreenGenes Database Consortium (Version 13.8).</p>
RDP taxonomic training data formatted for DADA2 (RDP trainset 16/release 11.5)
<p>These DADA2-formatted training fasta files were derived from the Ribosomal Database Project's Training Set 16 and the 11.5 release of the RDP database.</p> <p>These fastas were generated by the following commands (using the dada2 R package version 1.5.1):</p> <blockquote> <p>path <- "~/Desktop/RDP/RDPClassifier_16S_trainsetNo16_rawtrainingdata"<br> dada2:::makeTaxonomyFasta_RDP(file.path(path, "trainset16_022016.fa"), file.path(path, "trainset16_db_taxid.txt"), "~/tax/rdp_train_set_16.fa.gz")</p> <p>dada2:::makeSpeciesFasta_RDP("~/Desktop/RDP/current_Bacteria_unaligned.fa", "~/tax/rdp_species_assignment_16.fa.gz")</p> </blockquote>
RDP LSU taxonomic training data formatted for DADA2 (trainingset 11)
<p>#Format RDP taxonomic training set for DADA2<br> #1 Wrangle the RDP trainingsets and unaligned data into the downloads folder by executing this from a terminal and move the file somewhere with >30GB free<br> wget https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata.zip/download<br> wget http://rdp.cme.msu.edu/download/current_Fungi_unaligned.fa.gz<br> #2 Unzip the trainingset file and replace Us with Ts in the fasta by executing in terminal<br> awk 'NR%2==0 {gsub(/[uU]/,"T"); print} NR%2==1' /media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014.fa > /media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014_lsu_fixed_v2.fa<br> #3 Summon the dada2 pkg<br> library(dada2);packageVersion("dada2")<br> #4 Transform the DADA2 formatted training fastas<br> path<-"/media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata"<br> dada2:::makeTaxonomyFasta_RDP(file.path(path, "fungiLSU_train_012014_lsu_fixed_v2.fa"), file.path(path, "fungiLSU_taxid_012014.txt"),"/media/lauren/96BA-19E6/Upload/RDP_LSU_fixed_train_set_v2.fa",compress=FALSE)<br> dada2:::makeSpeciesFasta_RDP("/media/lauren/96BA-19E6/RDPClassifierLSU/current_Fungi_unaligned.fa", "/media/lauren/96BA-19E6/Upload/rdp_species_assignment_LSU_v2.fa", compress=FALSE)</p> <p>#5 Make the compressed DADA2 formatted training fastas in gz and zip format<br> dada2:::makeTaxonomyFasta_RDP("/media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014_lsu_fixed_v2.fa", file.path(path, "fungiLSU_taxid_012014.txt"),"/media/lauren/96BA-19E6/Upload/RDP_LSU_fixed_train_set_v2.fa.gz",compress=TRUE)<br> dada2:::makeTaxonomyFasta_RDP("/media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014_lsu_fixed_v2.fa", file.path(path, "fungiLSU_taxid_012014.txt"),"/media/lauren/96BA-19E6/Upload/RDP_LSU_fixed_train_set_v2.fa.zip",compress=TRUE)<br> dada2:::makeSpeciesFasta_RDP("/media/lauren/96BA-19E6/RDPClassifierLSU/current_Fungi_unaligned.fa", "/media/lauren/96BA-19E6/Upload/rdp_species_assignment_LSU_v2.fa.gz", compress=TRUE)<br> dada2:::makeSpeciesFasta_RDP("/media/lauren/96BA-19E6/RDPClassifierLSU/current_Fungi_unaligned.fa", "/media/lauren/96BA-19E6/Upload/rdp_species_assignment_LSU_v2.fa.zip", compress=TRUE)</p> <p> </p>
GTDB and RefSeq-RDP databases parsed for species assignment
<p>GTDB and RefSeq-RDP databases parsed for the <em>addSpecies()</em> dada2 function. Both databases were parsed with `<a href="https://github.com/antonioggsousa/GTDB-RefSeq-RDP-assign-spp-dada2/blob/master/script/parsing_DB_dada2_spp_assign.py">parsing_DB_dada2_spp_assign.py</a>` python script using as input both databases downloaded at: <a href="https://zenodo.org/record/2541239#.XM2UgCOZPOQ">https://zenodo.org/record/2541239#.XM2UgCOZPOQ</a>. The python script and a detailed description explaining the code and databases versions can be found at: <a href="https://github.com/antonioggsousa/GTDB-RefSeq-RDP-assign-spp-dada2">https://github.com/antonioggsousa/GTDB-RefSeq-RDP-assign-spp-dada2</a></p> <p> </p>
STILNOVO project (Sub.measure 16.2 RDP Tuscany Region 2014-2020)
<p>Dataset of pasture production and quality and milk production and quality of a rotational grazing trial carried out in 2018 in south Tuscany.</p>
RDP taxonomic training data formatted for DADA2 (RDP trainset 18/release 11.5)
<p>These DADA2-formatted training fasta files were derived from the Ribosomal Database Project's Training Set 18 and the 11.5 release of the RDP database.</p> <p>These fastas were generated by the following commands (using the dada2 R package version 1.19.1):</p> <blockquote> <p>## The RDP trainset data was downloaded from: https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/<br> path <- "~/Desktop/RDP/RDPClassifier_16S_trainsetNo18_rawtrainingdata"<br> dada2:::makeTaxonomyFasta_RDP(file.path(path, "trainset18_062020.fa"), <br> file.path(path, "trainset18_db_taxid.txt"), <br> "~/tax/rdp_train_set_18.fa.gz")<br> ## Download "ten_16s.100.fa" from Robert Edgar's taxonomy testing page: https://drive5.com/taxxi/doc/fasta_index.html<br> dada2:::tax.check("~/tax/rdp_train_set_18.fa.gz", "~/Desktop/ten_16s.100.fa")</p> <p>## This function creates the dada2 assignSpecies fasta file for the RDP from the RDP's _Bacteria_unaligned.fa file.<br> dada2:::makeSpeciesFasta_RDP("~/Desktop/RDP/current_Bacteria_unaligned.fa", "~/tax/rdp_species_assignment_18.fa.gz")<br> dada2:::tax.check("~/tax/rdp_species_assignment_18.fa.gz", "~/Desktop/ten_16s.100.fa", mode="species")</p> </blockquote>
Dataset for a Novel Hybrid Cryptography Algorithm for RDP Security
<p>The dataset consists of 95 rows and 8 columns. Each row contains information about an alphanumeric character (e.g. (A-Z), (a-z), (0-9), punctuation, and all special characters). The columns were created based on the steps of the proposed hybrid cryptography algorithm. The features include ASCII_Equivalent, Hex_Values, S_box_Conversion, Decimal_Values_of_S_box, RSA_Encrypted_Values, Replaced_Values, Group, and Class, with Class being the dependent variable. The characters in the Replaced_Values column are assigned a class of 0 or 1 depending on the presence of uppercase or lowercase letters. The dataset was generated using small RSA keys for convenience.</p>
RDP Reliability Study
ClinicalTrials.gov study NCT06057493. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.