Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
12
datasets available to search
ShareScore release 0.9.0
Dataset results
12 results for “kraken”
Kraken 2.0.8 beta taxonomic binning of the CAMI 2 Mouse Gut Toy data set, gold standard pooled assembly
<p>Taxonomic binning of the gold standard pooled assembly<br> <strong>Software: </strong>Kraken<br> <strong>SoftwareVersion: </strong>2.0.8 beta<br> <strong>DataURL: </strong> https://data.cami-challenge.org/participate<br> <strong>SoftwareURL:</strong> https://ccb.jhu.edu/software/kraken2/<br> <strong>ReferenceDatabase:</strong> built 2019-05-22<br> <strong>Taxonomy:</strong> NCBI 2019-05-22<br> <strong>ShortReadsUsed:</strong> False<br> <strong>LongReadsUsed:</strong> False<br> <strong>CommandUsed:</strong> kraken2-build --standard --db kraken2db_std --use-ftp<br> kraken2 --db kraken2db_std --threads 16 --output 19122017_mousegut_scaffolds.kraken --report 19122017_mousegut_scaffolds.kreport anonymous_gsa_pooled.fasta<br> cat 19122017_mousegut_scaffolds | awk '{print $2 "\t" $3}' > 19122017_mousegut_scaffolds.cami</p>
Minimal dummy Kraken DB
<p>This small <a href="http://ccb.jhu.edu/software/kraken/">Kraken</a> DB allows to rapidly perform continuous integration testing of Kraken commands. It has a small memory and disk footprint and can be downloaded on the fly into CI environments.</p>
kraken database for MetaMeta pipeline - Archaea and Bacteria - version 1
<p>kraken database for MetaMeta pipeline version 1. The database was generated by the kraken-build script and it is based on the NCBI RefSeq Complete Genome sequences of Archaea and Bacteria dating from 2015-06-02</p> <p> </p>
Jeu de données de segmentation et de reconnaissance optique de caractères - Kraken - Incunables sévillans 1494-1500
<p>[<strong>Information importante</strong>: ce jeu de données a été versé dans un jeu plus ample (28.000 lignes) contenant majoritairement des données manuscrites: https://zenodo.org/record/7389195]</p> <p> </p> <p>Ce dépôt contient un modèle fonctionnel de reconnaissance optique de caractères, entraîné grâce au logiciel <a href="http://kraken.re">kraken</a> via eScriptorium. Le modèle a été entraîné sur un des incunables du <em>Regimiento de los Prínçipes</em> (connu aussi sous le titre de: <em>Glosa castellana al Regimiento de prínçipes</em>), l'incunable INC/901 de la Bibliothèque nationale d'Espagne.</p> <p> </p> <p>Il contient de même un modèle de segmentation entraîné de même sur kraken après segmentation manuelle sur eScriptorium.</p> <p> </p> <p><strong>Description du jeu de données: </strong></p> <p>Le jeu de données contient 62 pages et 5556 lignes. Le type utilisé par Estanislao Polono pour cet incunable est le 97G (Martín Abad and Moyano Andrés, 2002, p. 61). Ce type est utilisé entre 1494 et 1500. Pour les autres incunables produits à cette époque, voir <em>op.cit</em>, p.112-121.</p> <p>Les zones du modèle de segmentation sont conformes au vocabulaire partagé SegmOnto (<a href="https://segmonto.github.io/">https://segmonto.github.io/</a>).</p> <p> </p> <p><strong>Qualité du modèle: </strong></p> <p>Le modèle a été entraîné sur 5556 lignes. Son taux d'erreur est d'un peu plus de 3% (96.5%). Les vérités terrain sont fournies au format ALTO et jpeg.</p> <p>Deux modèles de segmentation sont fournis, pour les baselines et pour les régions.</p> <p> </p> <p><strong>Crédits et remerciements:</strong></p> <p>Les données ont successivement été entraînées sur Ocropy et Kraken. Pour entraîner originellement <a href="https://zenodo.org/record/2504784">le modèle Ocropy</a> qui a permis de prédire le jeu de données d'entraînement que j'ai ensuite corrigé et utilisé sur Kraken, je me suis amplement servi du manuel rédigé par Jean-Baptiste Camps (ENC-PSL), qui peut être trouvé <a href="https://graal.hypotheses.org/786">sur son carnet de recherche</a>.</p> <p>Merci à Simon Gabay (U. de Neuchâtel) pour son aide sur kraken et pour tous ses conseils méthodologiques.</p> <p><strong>Bibliographie: </strong></p> <p>Kiessling, Benjamin. « Kraken - an Universal Text Recognizer for the Humanities ». DH2019:Complexity, Utrecht, 2019. <a href="https://dev.clariah.nl/files/dh2019/boa/0673.html">https://dev.clariah.nl/files/dh2019/boa/0673.html</a>.</p> <p>Martín Abad, J. and Moyano Andrés, I. (2002). <em>Estanislao Polono</em>.</p> <p><code>« Homemade manuscript OCR (1): OCRopy », <em>Sacré Gr@@l</em>, 6 février 2017, <a href="https://graal.hypotheses.org/786">https://graal.hypotheses.org/786</a> </code></p> <p> </p>
Kraken analysis from shotgun sequenced museum specimens via krona plot visualization
<p>This dataset contains html files with krona plots from Kraken analysis on shotgun sequenced museum specimens.</p>
kraken database for MetaMeta pipeline - Fungal and Viral - 201709
<p>kraken database for MetaMeta pipeline. The database was generated by the kraken-build script and it is based on the NCBI RefSeq Complete Genome sequences of Fungi and Virus dating from 2017-09</p>
PLAsTiCC kraken 2026 Models
<p>PLAsTiCC models simulated with LSST kraken2026 opsim cadence</p>
Kraken 2 manuscript experiment data
<p>This dataset contains the data used in the strain-exclusion experiments in the Kraken 2 manuscript.</p>
Kraken database (2019-01-26)
<p>Kraken database containing genomes from bacteria, archaea, viruses and the human reference genome as of 2019-01-26.</p> <p>A the domain level the database composition is:</p> <ul> <li>Bacteria (85.13%, 358273085 bp)</li> <li>Eukaryota (11.29%, 47508472 bp)</li> <li>Archaea (2.56%, 10774261 bp)</li> <li>Viruses (0.98%, 4115713 bp)</li> <li>Viroids (0.00%, 277 bp)</li> </ul>
Human Kraken database
<p>Human Kraken database built on 04/02/2020 for <a href="https://github.com/nf-core/viralrecon">nf-core/viralrecon</a> pipeline testing,</p> <p>Kraken version:</p> <pre>Kraken version 2.0.8-beta Copyright 2013-2019, Derrick Wood (<a href="mailto:dwood@cs.jhu.edu">dwood@cs.jhu.edu</a>)</pre> <p>Commands:</p> <pre>kraken2-build --db kraken2_human --threads 2 --download-taxonomy kraken2-build --db kraken2_human --threads 2 --download-library human kraken2-build --db kraken2_human --threads 2 --build</pre>
SARS-Cov2 Kraken database
<p>Viral Kraken database built on 04/04/2020 for <a href="https://github.com/nf-core/viralrecon">nf-core/viralrecon</a> pipeline testing,</p> <p>Kraken version:</p> <pre><code>Kraken version 2.0.8-beta Copyright 2013-2019, Derrick Wood (dwood@cs.jhu.edu)</code></pre> <p>Commands:</p> <pre><code>kraken2-build --db kraken2_viral --threads 2 --use-ftp --download-taxonomy kraken2-build --db kraken2_viral --threads 2 --use-ftp --download-library viral kraken2-build --db kraken2_viral --threads 2 --use-ftp --build</code></pre> <p>Files:</p> <pre><code>## Minimal kraken2-build output files required to run kraken2 ## See kraken2_viral.tree.txt for file listing kraken2_viral.tar.gz kraken2_viral.tar.gz.md5 kraken2_viral.tree.txt ## Additional kraken2-build output files that can be used in conjunction with bracken-build ## See kraken2_viral.bracken.tree.txt for file listing kraken2_viral.bracken.tar.gz kraken2_viral.bracken.tar.gz.md5 kraken2_viral.bracken.tree.txt</code></pre> <p> </p>
PLAsTiCC kraken 2044 Models
<p>PLAsTiCC models simulated with LSST kraken2044 opsim cadence</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.