Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.9.0
Dataset results
6 results for “Naïve Bayes classifier”
Figure 2 Bar chart represents the percentage of the correct diagnosis using fuzzy diagnosis, K- nearest neighbor and Naïve Bayes classifiers.-Comparison of Fuzzy Diagnosis with K-Nearest Neighbor and Naïve Bayes Classifiers in Disease Diagnosis
<p>Figure 2 Bar chart represents the percentage of the correct diagnosis using fuzzy<br> diagnosis, K- nearest neighbor and Naïve Bayes classifiers.</p>
Figure 1. The comparison of the Area Under the ROC Curve (AUC) for fuzzy Diagnosis, KNN and NB-Comparison of Fuzzy Diagnosis with K-Nearest Neighbor and Naïve Bayes Classifiers in Disease Diagnosis
<p>The area under the receiver operating characteristic (ROC) curve (AUC) is used to measure<br> the performance of fuzzy diagnosis, KNN and NB. In order to show the difference between AUC<br> for the three methods, a single figure which combines the three AUC for the three methods was<br> used for comparison as shown in figure 1 below.</p>
MNBC: a multithreaded Minimizer-based Naïve Bayes Classifier for improved metagenomic sequence classification
<p>These files provide supplementary data underlying the article <a title="https://doi.org/10.1093/bioinformatics/btae601" href="https://doi.org/10.1093/bioinformatics/btae601" target="_blank" rel="noopener noreferrer nofollow">doi.org/10.1093/bioinformatics/btae601</a> (see Figure 1 in the article):</p> <ul> <li>37345_filtered_training_and_test_genomes_list.txt: Refseq assembly sequence filenames of the 37345 filtered training and test genomes</li> <li>taxonomy_37345_filtered_training_and_test_genomes.txt: Taxonomy file for all 37345 filtered training and test genomes</li> <li>Uniform_reference_database_31991_training_genomes_assemblyID_list.txt: Refseq assembly accessions of the 31991 training genomes in the uniform reference database</li> <li>uniform_reference_database.tar.gz_1 to uniform_reference_database.tar.gz_10: Merge them into a single file using the <em>cat</em> command. The folder produced by decompressing this file is the uniform reference database.</li> <li>taxonomy_uniform_reference_database.txt: Taxonomy file for the uniform reference database (i.e. the 31991 training genomes)</li> <li>5354_test_genomes_assemblyID_list.txt: Refseq assembly accessions of the 5354 test genomes</li> <li>testReads_NextSeq_C0.05.fasta.gz: 6562565 150bp-long positive reads randomly generated from the 5354 test genomes, simulating reads sequenced by NextSeq (0.05 coverge)</li> <li>testReads_MiSeq_C0.05.fasta.gz: 3282728 300bp-long positive reads randomly generated from the 5354 test genomes, simulating reads sequenced by MiSeq (0.05 coverge)</li> <li>testReads_Nanopore_C0.05.fasta.gz: 181912 positive reads of normally distributed 1kb-10kb lengths randomly generated from the 5354 test genomes, simulating reads sequenced by Nanopore (0.05 coverge)</li> <li>negaReads_NextSeq_C0.05.fasta.gz: 10143 150bp-long negative reads randomly generated from Chromosome 1 of the Arabidopsis Thaliana reference genome, simulating reads sequenced by NextSeq (0.05 coverge)</li> <li>negaReads_MiSeq_C0.05.fasta.gz: 5072 300bp-long negative reads randomly generated from Chromosome 1 of the Arabidopsis Thaliana reference genome, simulating reads sequenced by MiSeq (0.05 coverge)</li> <li>negaReads_Nanopore_C0.05.fasta.gz: 277 negative reads of normally distributed 1kb-10kb lengths randomly generated from Chromosome 1 of the Arabidopsis Thaliana reference genome, simulating reads sequenced by Nanopore (0.05 coverge)</li> <li>CAMI2_reference_database_16864_genomes_list.txt: Refseq assembly sequence filenames of the 16864 genomes and chromosomes in the reference database for CAMI2</li> <li>taxonomy_CAMI2_reference_database.txt: Taxonomy file for the CAMI2 reference database</li> </ul> <p><strong>Tip</strong>: To directly use the taxonomy file "taxonomy_uniform_reference_database.txt", please use version v1.1 or earlier of the MNBC tool. If using later versions it needs regenerating with the "MNBC taxonomy" program.</p>
Dataset for the paper "Website Fingerprinting: Attacking Popular Privacy Enhancing Technologies with the Multinomial Naïve-Bayes Classifier"
<p>This dataset contains website fingerprints of 775 websites analyzed in the paper "Website Fingerprinting: Attacking Popular Privacy Enhancing Technologies with the Multinomial Naïve-Bayes Classifier" published in the Proceedings of the 2009 ACM workshop on Cloud computing security (CCSW 2009, DOI: 10.1145/1655008.1655013).</p>
MNBC: a multithreaded Minimizer-based Naïve Bayes Classifier for improved metagenomic sequence classification
<p>These files provide supplementary data underlying the article <a title="https://doi.org/10.1093/bioinformatics/btae601" href="https://doi.org/10.1093/bioinformatics/btae601" target="_blank" rel="noopener noreferrer nofollow">doi.org/10.1093/bioinformatics/btae601</a>:</p> <ul> <li>CAMI2_reference_database.tar.gz_1 to CAMI2_reference_database.tar.gz_6: Merge them into a single file using the <em>cat</em> command. The folder produced by decompressing this file is the reference database for CAMI2 (RefSeq sequence filenames are in the file CAMI2_reference_database_16864_genomes_list.txt at <a href="https://dx.doi.org/10.5281/zenodo.10568965">https://dx.doi.org/10.5281/zenodo.10568965</a>).</li> </ul> <p> </p>
MNBC: a multithreaded Minimizer-based Naïve Bayes Classifier for improved metagenomic sequence classification
<p>These files provide supplementary data underlying the article <a title="https://doi.org/10.1093/bioinformatics/btae601" href="https://doi.org/10.1093/bioinformatics/btae601" target="_blank" rel="noopener noreferrer nofollow">doi.org/10.1093/bioinformatics/btae601</a> (see Figure 1 of the article):</p> <ul> <li>37345_filtered_training_and_test_genomes_sequence_files.tar.gz_1 to 37345_filtered_training_and_test_genomes_sequence_files.tar.gz_6: Merge them into a single file using the <em>cat</em> command. The folder produced by decompressing this file contains the filtered RefSeq sequence files of all 37345 training and test genomes, which were used to build the uniform reference database and generate the test reads (The filenames are in the file 37345_filtered_training_and_test_genomes_list.txt at <a href="https://dx.doi.org/10.5281/zenodo.10568965">https://dx.doi.org/10.5281/zenodo.10568965</a>).</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.