Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

12

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

12 results for “RDP”

Learn how ShareScore rates datasets ↗
zenodo48/100

The tpm metabarcoding DNA sequence database for taxonomic allocations using RDP classifier implemented in DADA2.

<p><strong>The </strong><em>tpm</em><strong> metabarcoding DNA sequence database for taxonomic allocations using the Mothur and DADA2 bio-informatic tools</strong></p> <p>A.C.M. Pozzi<sup>1</sup>, R. Bouchali<sup>1</sup>, L. Marjolet<sup>1</sup>, B. Cournoyer<sup>1</sup></p> <p><sup>1 </sup><em>University of Lyon, UMR Ecologie Microbienne Lyon (LEM), CNRS 5557, INRAE 1418, Universit&eacute; Claude Bernard Lyon 1, VetAgro Sup, Research Team &ldquo;Bacterial Opportunistic Pathogens and Environment&rdquo; (BPOE), 69280 Marcy L&rsquo;Etoile, France.</em></p> <p><strong>Corresponding authors: </strong></p> <ul> <li>A.C.M. Pozzi, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L&rsquo;Etoile, France. Tel. (+33) 478 87 39 47. Fax. (+33) 472 43 12 23. Email: <a href="mailto:adrien.meynier_pozzi@vetagro-sup.fr">adrien.meynier_pozzi@vetagro-sup.fr</a></li> <li>B. Cournoyer, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L&rsquo;Etoile, France. Tel. (+33) 478 87 56 47. Fax. (+33) 472 43 12 23. Email: and <a href="mailto:benoit.cournoyer@vetagro-sup.fr">benoit.cournoyer@vetagro-sup.fr</a></li> </ul> <p><strong>Keywords:</strong></p> <p>BACtpm, Bacteria, <em>tpm</em>, thiopurine-<em>S</em>-methyltransferase EC:2.1.1.67, Nucleotide sequences, PCR products, Next-Generation-Sequencing, OTHU</p> <p><strong>Description:</strong></p> <ul> <li>The <em>tpm</em> gene codes for the thiopurine-<em>S</em>-methyltransferase (TPMT), an enzyme that can detoxify metalloid-containing oxyanions and xenobiotics (Cournoyer et al., 1998). Bacterial TPMTs radiated apart from human and animal TPMTs, and showed a vertical evolution in line with the 16S rRNA gene molecular phylogeny (Favre‐Bont&eacute; et al., 2005).</li> <li>The <em>tpm</em> database, named BACtpm, was designed to apply the <em>tpm</em>-metabarcoding analytical scheme published in Aigle et al. (2021). It includes the full <em>tpm</em> identifiers, GenBank accession numbers, complete taxonomic records (domain down to strain code) of about 215 nucleotide-long <em>tpm</em> sequences of 840 unique taxa belonging to 139 genera.</li> <li>Nucleotide sequences of <em>tpm</em> (range: 190-233 nucleotides) were either retrieved from public repositories (GenBank) or made available by B. Cournoyer&rsquo;s research group. Colin et al. (2020) described the PCR and high throughput Illumina Miseq DNA sequencing procedures used to produce <em>tpm</em> sequences.</li> <li>BACtpm v.2.0.1 (June 2021 release) is made available under the Creative Commons Attribution 4.0 International Licence. It can be used for the taxonomic allocations of <em>tpm </em>sequences down to the species and strain levels. Data is stored in the csv format enabling future user to reformat it to fit their specific needs.</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>We thank the worldwide community of microbiologists who made contributions to public databases in the past decades, and made possible the elaboration of the BACtpm database. We also thank the Field Observatory in Urban Hydrology (OTHU, <a href="http://www.graie.org/othu/">www.graie.org/othu/</a>), Labex IMU (Intelligence des Mondes Urbains), the Greater Lyon Urban Community, the School of Integrated Watershed Sciences H2O&#39;LYON, and the Lyon Urban School for their support in the development of this database. This work was funded by the French national research program for environmental and occupational health of ANSES under the terms of project &ldquo;Iouqmer&rdquo; EST 2016/1/120, l&#39;Agence Nationale de la Recherche through ANR-16-CE32-0006, ANR-17-CE04-0010, ANR-17-EURE-0018 and ANR-17-CONV-0004, by the MITI CNRS project named Urbamic, and the French water agency for the Rh&ocirc;ne, Mediterranean and Corsica areas through the Desir and DOmic projects. We thank former BPOE lab members who contributed to start and expand the BACtpm database: C&eacute;line COLINON, Romain MARTI, Emilie BOURGEOIS, S&eacute;bastien RIBUN and Yannick COLIN.</p> <p><strong>References:</strong></p> <p>Aigle, A., Colin, Y., Bouchali, R., Bourgeois, E., Marti, R., Ribun, S., Marjolet, L., Pozzi, A.C.M., Misery, B., Colinon, C., Bernardin-Souibgui, C., Wiest, L., Blaha, D., Galia, W., Cournoyer, B., 2021. Spatio-temporal variations in chemical pollutants found among urban deposits match changes in thiopurine S-methyltransferase-harboring bacteria tracked by the tpm metabarcoding approach. Sci. Total Environ. 767, 145425. https://doi.org/10.1016/j.scitotenv.2021.145425</p> <p>Colin, Y., Bouchali, R., Marjolet, L., Marti, R., Vautrin, F., Voisin, J., Bourgeois, E., Rodriguez-Nava, V., Blaha, D., Winiarski, T., Mermillod-Blondin, F., Cournoyer, B., 2020. Coalescence of bacterial groups originating from urban runoffs and artificial infiltration systems among aquifer microbiomes. Hydrol. Earth Syst. Sci. 24, 4257&ndash;4273. https://doi.org/10.5194/hess-24-4257-2020</p> <p>Cournoyer, B., Watanabe, S., Vivian, A., 1998. A tellurite-resistance genetic determinant from phytopathogenic pseudomonads encodes a thiopurine methyltransferase: evidence of a widely-conserved family of methyltransferases1The International Collaboration (IC) accession number of the DNA sequence is L49178.1. Biochim. Biophys. Acta BBA - Gene Struct. Expr. 1397, 161&ndash;168. https://doi.org/10.1016/S0167-4781(98)00020-7</p> <p>Favre‐Bont&eacute;, S., Ranjard, L., Colinon, C., Prigent‐Combaret, C., Nazaret, S., Cournoyer, B., 2005. Freshwater selenium-methylating bacterial thiopurine methyltransferases: diversity and molecular phylogeny. Environ. Microbiol. 7, 153&ndash;164. https://doi.org/10.1111/j.1462-2920.2004.00670.x</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

RDP Classifier 2.14 and the RDP bacterial and archaeal taxonomy training set No. 19

<p>RDP Classifier 2.14 (August 2023) Release Note:</p><p>The Bacteria and Archaea hierarchy model used by RDP Classifier has been updated to training set No. 19. The new version has over 600 genera and 2500 species added since last version No. 18 released in July 2020. The information that is used to update the RDP taxonomy to training set version No. 19, and RDP Classifier version 2.14 came from publicly available scientific articles and public sequence repository, mostly from International Journal of Systematic and Evolutionary Microbiology (IJSEM), the All-Species Living Tree Project (LTP) and GenBank. &nbsp;</p><p>It is worth noting that most of the phyla have new names, according to article "</p><p>Oren A, Garrity GM. Valid publication of the names of forty-two phyla of prokaryotes. Int J Syst Evol Microbiol. 2021 Oct;71(10). doi: 10.1099/ijsem.0.005056. PMID: 34694987."</p><p>In addition to the files to train and run the RDP Classifier, new file formats are made available to accommodate the needs of users:</p><p>1. A new file trainset19_072023_speciesrank.fa has been added to the release in RDPClassifier_16S_trainsetNo19_rawtrainingdata.zip. This file is NOT needed to train the classifier. In addition to sequences, it contains genus, species, strain, type status and taxonomy rank, which are useful for closest species identification using third-party tools (e.g. BLAST).</p><p>2. Two new files in RDPClassifier_16S_trainsetNo19_QiimeFormat.zip to retrain the RDP Classifier included in Qiime2 package.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

RDP taxonomic training data formatted for DADA2 (RDP release 19 - update 2023-08-23)

<p>These DADA2-formatted training fasta files were derived from the Ribosomal Database Project's Training Set 19 and the 2023-08-23 release of the RDP database. https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/</p> <p>These fastas were generated by the following commands using the dada2 R package version 1.35.4:</p> <blockquote> <p>## RDP data: https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/</p> <p>path &lt;- "~/tax/rdp/v19"<br>fn.out.rdp &lt;- "~/Desktop/rdp_19_toGenus_trainset.fa.gz"<br>dada2:::makeTaxonomyFasta_RDP(file.path(path, "trainset19_072023_speciesrank.fa"),&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;file.path(path, "trainset19_db_taxid.txt"),&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;fn.out.rdp, include.species=FALSE,<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;compress=TRUE)</p> <p>fn.out.spc.rdp &lt;- "~/Desktop/rdp_19_toSpecies_trainset.fa.gz"<br>dada2:::makeTaxonomyFasta_RDP(file.path(path, "trainset19_072023_speciesrank.fa"),&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;file.path(path, "trainset19_db_taxid.txt"),&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;fn.out.spc.rdp, include.species=TRUE,<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;compress=TRUE)</p> </blockquote>

opencc-by-4.0Nov 2024View details →
zenodo40/100

RDP Classifier training files for 16S rRNA sequences from GTDB

<p>16S rRNA gene sequences from the <a href="https://gtdb.ecogenomic.org/">Genome Taxonomy Database</a> (GTDB release 220) were used to retrain the <a href="https://github.com/rdpstaff/classifier">RDP Classifier</a> (version 2.13). Two sets of training files are provided:</p> <ul> <li><code>genus.zip</code> - Genus level</li> <li><code>species.zip</code> - Species level</li> </ul> <p>The code in <code>prepare_files.R</code> was used to prepare the GTDB sequence and taxonomy files for retraining the RDP Classifier. Notes:</p> <ul> <li>Steps to retrain the RDP Classifier are adapted from <a href="https://john-quensen.com/tutorials/training-the-rdp-classifier/">https://john-quensen.com/tutorials/training-the-rdp-classifier/</a></li> <li>Python scripts (lineage2taxTrain.py and addFullLineage.py) are available at <a href="https://github.com/rdpstaff/classifier/issues/18">https://github.com/rdpstaff/classifier/issues/18</a></li> <li>The first 1000 training sequences (<code>train_nodups_1000.fasta</code>) are used for benchmarking the classification accuracy (see results at end of <code>prepare_files.R</code>).</li> </ul>

opencc-by-4.0Jun 2024View details →
zenodo36/100

The RDP and GreenGenes taxonomic training sets formatted for DADA2

<p>These DADA2-formatted training fastas were derived from the Ribosomal Database Project (RDP Training Set 14) and the GreenGenes Database Consortium (Version 13.8).</p>

opencc-by-sa-4.0Sep 2016View details →
zenodo36/100

RDP taxonomic training data formatted for DADA2 (RDP trainset 16/release 11.5)

<p>These DADA2-formatted training fasta files were derived from the Ribosomal Database Project's Training Set 16 and the 11.5 release of the RDP database.</p> <p>These fastas were generated by the following commands (using the dada2 R package version 1.5.1):</p> <blockquote> <p>path &lt;- "~/Desktop/RDP/RDPClassifier_16S_trainsetNo16_rawtrainingdata"<br> dada2:::makeTaxonomyFasta_RDP(file.path(path, "trainset16_022016.fa"), file.path(path, "trainset16_db_taxid.txt"), "~/tax/rdp_train_set_16.fa.gz")</p> <p>dada2:::makeSpeciesFasta_RDP("~/Desktop/RDP/current_Bacteria_unaligned.fa", "~/tax/rdp_species_assignment_16.fa.gz")</p> </blockquote>

opencc-by-4.0May 2017View details →
zenodo36/100

RDP LSU taxonomic training data formatted for DADA2 (trainingset 11)

<p>#Format RDP taxonomic training set for DADA2<br> #1  Wrangle the RDP trainingsets and unaligned data into the downloads folder by executing this from a terminal and move the file somewhere with &gt;30GB free<br> wget https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata.zip/download<br> wget http://rdp.cme.msu.edu/download/current_Fungi_unaligned.fa.gz<br> #2  Unzip the trainingset file and replace Us with Ts in the fasta by executing in terminal<br> awk 'NR%2==0 {gsub(/[uU]/,"T"); print} NR%2==1' /media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014.fa &gt; /media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014_lsu_fixed_v2.fa<br> #3  Summon the dada2 pkg<br> library(dada2);packageVersion("dada2")<br> #4  Transform the DADA2 formatted training fastas<br> path&lt;-"/media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata"<br> dada2:::makeTaxonomyFasta_RDP(file.path(path, "fungiLSU_train_012014_lsu_fixed_v2.fa"), file.path(path, "fungiLSU_taxid_012014.txt"),"/media/lauren/96BA-19E6/Upload/RDP_LSU_fixed_train_set_v2.fa",compress=FALSE)<br> dada2:::makeSpeciesFasta_RDP("/media/lauren/96BA-19E6/RDPClassifierLSU/current_Fungi_unaligned.fa", "/media/lauren/96BA-19E6/Upload/rdp_species_assignment_LSU_v2.fa", compress=FALSE)</p> <p>#5 Make the compressed DADA2 formatted training fastas in gz and zip format<br> dada2:::makeTaxonomyFasta_RDP("/media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014_lsu_fixed_v2.fa", file.path(path, "fungiLSU_taxid_012014.txt"),"/media/lauren/96BA-19E6/Upload/RDP_LSU_fixed_train_set_v2.fa.gz",compress=TRUE)<br> dada2:::makeTaxonomyFasta_RDP("/media/lauren/96BA-19E6/RDPClassifier_fungiLSU_trainsetNo11_rawtrainingdata/fungiLSU_train_012014_lsu_fixed_v2.fa", file.path(path, "fungiLSU_taxid_012014.txt"),"/media/lauren/96BA-19E6/Upload/RDP_LSU_fixed_train_set_v2.fa.zip",compress=TRUE)<br> dada2:::makeSpeciesFasta_RDP("/media/lauren/96BA-19E6/RDPClassifierLSU/current_Fungi_unaligned.fa", "/media/lauren/96BA-19E6/Upload/rdp_species_assignment_LSU_v2.fa.gz", compress=TRUE)<br> dada2:::makeSpeciesFasta_RDP("/media/lauren/96BA-19E6/RDPClassifierLSU/current_Fungi_unaligned.fa", "/media/lauren/96BA-19E6/Upload/rdp_species_assignment_LSU_v2.fa.zip", compress=TRUE)</p> <p> </p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

GTDB and RefSeq-RDP databases parsed for species assignment

<p>GTDB and RefSeq-RDP databases parsed for the&nbsp;<em>addSpecies()</em>&nbsp;dada2 function. Both databases were parsed with `<a href="https://github.com/antonioggsousa/GTDB-RefSeq-RDP-assign-spp-dada2/blob/master/script/parsing_DB_dada2_spp_assign.py">parsing_DB_dada2_spp_assign.py</a>` python script using as input both databases downloaded at:&nbsp;<a href="https://zenodo.org/record/2541239#.XM2UgCOZPOQ">https://zenodo.org/record/2541239#.XM2UgCOZPOQ</a>. The python script and a detailed description explaining the code and databases versions can be found at:&nbsp;<a href="https://github.com/antonioggsousa/GTDB-RefSeq-RDP-assign-spp-dada2">https://github.com/antonioggsousa/GTDB-RefSeq-RDP-assign-spp-dada2</a></p> <p>&nbsp;</p>

opencc-by-4.0May 2019View details →
zenodo28/100

STILNOVO project (Sub.measure 16.2 RDP Tuscany Region 2014-2020)

<p>Dataset of pasture production and quality and milk production and quality of a rotational grazing trial carried out&nbsp;in 2018 in south Tuscany.</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

RDP taxonomic training data formatted for DADA2 (RDP trainset 18/release 11.5)

<p>These DADA2-formatted training fasta files&nbsp;were derived from the Ribosomal Database Project&#39;s Training Set 18&nbsp;and the 11.5 release of the RDP database.</p> <p>These fastas were generated by the following commands (using the dada2 R package version 1.19.1):</p> <blockquote> <p>## The RDP trainset data was downloaded from: https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/<br> path &lt;- &quot;~/Desktop/RDP/RDPClassifier_16S_trainsetNo18_rawtrainingdata&quot;<br> dada2:::makeTaxonomyFasta_RDP(file.path(path, &quot;trainset18_062020.fa&quot;),&nbsp;<br> &nbsp; &nbsp; &nbsp;file.path(path, &quot;trainset18_db_taxid.txt&quot;),&nbsp;<br> &nbsp; &nbsp; &nbsp;&quot;~/tax/rdp_train_set_18.fa.gz&quot;)<br> ## Download &quot;ten_16s.100.fa&quot; from Robert Edgar&#39;s taxonomy testing page: https://drive5.com/taxxi/doc/fasta_index.html<br> dada2:::tax.check(&quot;~/tax/rdp_train_set_18.fa.gz&quot;, &quot;~/Desktop/ten_16s.100.fa&quot;)</p> <p>##&nbsp;This function creates the dada2 assignSpecies fasta file for the RDP from the RDP&#39;s _Bacteria_unaligned.fa file.<br> dada2:::makeSpeciesFasta_RDP(&quot;~/Desktop/RDP/current_Bacteria_unaligned.fa&quot;, &quot;~/tax/rdp_species_assignment_18.fa.gz&quot;)<br> dada2:::tax.check(&quot;~/tax/rdp_species_assignment_18.fa.gz&quot;, &quot;~/Desktop/ten_16s.100.fa&quot;, mode=&quot;species&quot;)</p> </blockquote>

opencc-by-4.0Dec 2020View details →
zenodo24/100

Dataset for a Novel Hybrid Cryptography Algorithm for RDP Security

<p>The dataset consists of 95 rows and 8 columns. Each row contains information about an alphanumeric character (e.g. (A-Z), (a-z), (0-9), punctuation, and all special characters). The columns were created based on the steps of the proposed hybrid cryptography algorithm. The features include ASCII_Equivalent, Hex_Values, S_box_Conversion, Decimal_Values_of_S_box, RSA_Encrypted_Values, Replaced_Values, Group, and Class, with Class being the dependent variable. The characters in the Replaced_Values column are assigned a class of 0 or 1 depending on the presence of uppercase or lowercase letters. The dataset was generated using small RSA keys for convenience.</p>

opencc-by-4.0May 2024View details →
ClinicalTrials.gov24/100

RDP Reliability Study

ClinicalTrials.gov study NCT06057493. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record