Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
19
datasets available to search
ShareScore release 0.9.0
Dataset results
19 results for “k-mers”
K-mer databases of plant virus sequences for use with the Kodoja workflow
<p><strong>Details</strong></p> <p>This is a gzipped tar file that includes the plant virus database files required to run the Kodoja workflow (https://github.com/abaizan/kodoja)[1]. Kodoja is a workflow for the detection of plant virus sequences in RNA-seq data files that uses two previoulsy published tools Kraken[2] and Kaiju[3].</p> <p>This file contains databases for Kraken [2] and Kaiju [3]. The file includes the kraken database files: database.idx, database.kdb, nodes.dmp, names.dmp and the kaiju database file kaij_library.fmi.</p> <p>These k-mer databases are based on virus sequences in RefSeq [4] (ttps://www.ncbi.nlm.nih.gov/refseq/) with plant hosts as defined in the Virus-Host Database [5] (https://www.genome.jp/virushostdb/).</p> <p><strong>Version</strong><strong> 1.0</strong></p> <p>kodojaDB_v1.0 is based on RefSeq v89 and the Virus-Host Database (accessed 03/09/2018 which is based on RefSeq 89 and Genbank 226.0). The viral partition of RefSeq v89 genome comprises 7946 viruses (ftp://ftp.ncbi.nlm.nih.gov/genomes/refseq/viral/assembly_summary.txt).</p> <p>kodojaDB_v1.0 was created using kodoja_retrieve.py which is part of the kodoja workflow (v0.05) (https://github.com/abaizan/kodoja).</p> <p><strong>References</strong></p> <p>[1] Baizan-Edge, A, Cock, P, MacFarlane, S, McGavin, W, Torrance, T, Jones, S. Kodoja: A workflow for virus detection in plants using k-mer analysis of RNA-sequencing data (under review Nucleic Acids Research). </p> <p>[2] Wood,D.E. and Salzberg,S.L. (2014) Kraken: ultrafast metagenomic sequence classification using exact alignments. <em>Genome Biol.</em>, <strong>15</strong>, R46</p> <p>[3] Menzel,P., Ng,K.L. and Krogh,A. (2016) Fast and sensitive taxonomic classification for metagenomics with Kaiju. <em>Nat. Commun.</em>, <strong>7</strong>, 1–9.</p> <p>[4] O’Leary,N.A., Wright,M.W., Brister,J.R., Ciufo,S., Haddad,D., McVeigh,R., Rajput,B., Robbertse,B., Smith-White,B., Ako-Adjei,D., <em>et al.</em> (2016) Reference sequence (RefSeq) database at NCBI: Current status, taxonomic expansion, and functional annotation. <em>Nucleic Acids Res.</em>, <strong>44</strong>, D733–D745.</p> <p>[5] Mihara,T., Nishimura,Y., Shimizu,Y., Nishiyama,H., Yoshikawa,G., Uehara,H., Hingamp,P., Goto,S. and Ogata,H. (2016) Linking virus genomes with host taxonomy. <em>Viruses</em>, <strong>8</strong>, 10–15</p>
Virtual NGS data used in the paper "Optimal k-mer values for good-quality triploid genome assemblies" and its assembly results
<p>Despite technological advancements, whole-genome sequencing remains technically challenging for organisms with higher-ploidy genomes. Therefore, the use of short-read sequencing platforms for this purpose has been attempted, but the conditions that result in poor-quality genomes have not been elucidated. Therefore, in the present study, simulated sequences mimicking the accumulation of insertion/deletion mutations were created to clarify the permissible differences between homologous chromosomes and the k-mer sizes for good-quality genome assembly from short-read sequencing data for triploid species. The results illustrated that a narrow range of k-mers permits the generation of high-quality assemblies for any level of difference between homologous chromosomes. <br> This dataset consists of the virtual haploid genome (O.fasta), triploid genome data (reference_sequense), NGS read data (NGS_reads), and assembly results (all_contigs) created in this study.<br> </p>
Data from: Memory-bound k-mer selection for large and evolutionary diverse reference libraries
Open the record for dataset details and reuse information.
k-mer matrix Aegilops tauschii diversity panel (Open wild wheat consortium phase II) Part 1/3
Open the record for dataset details and reuse information.
k-mer matrix Aegilops tauschii diversity panel (Open wild wheat consortium phase) Part 3/3
Open the record for dataset details and reuse information.
k-mer matrix Aegilops tauschii diversity panel (Open wild wheat consortium phase II) Part 2/3
Open the record for dataset details and reuse information.
Quantum mechanical electronic and geometric parameters for DNA k-mers as features for machine learning
<p>With the development of advanced predictive modelling techniques, we are witnessing a steep increase in model development initiatives in genomics that employ high-end machine learning methodologies. Of particular interest are models that predict certain genomic or biological characteristics based solely on DNA sequence information. These models, however, treat the DNA sequence as a mere collection of four, A, T, G and C, letters, thus dismissing the past physico-chemical advancements in science that can enable the use of more intricate information about nucleic acid sequences. Here, we provide a comprehensive database of quantum mechanical and geometric features for all the permutations of 7-meric DNA in their representative B, A and Z conformations. The database is generated by employing the applicable high-cost and time-consuming quantum mechanical methodologies. This can thus make it seamless to associate a wealth of novel molecular features to any DNA sequence, by scanning it with a matching k-meric window and pulling the pre-computed values from our database for further use in modelling. We demonstrate the usefulness of our deposited features through their exclusive use in developing a model for A to C mutation rate constants.</p> <p>The DNA k-mer quantum mechanical parameters can also be found <a href="https://github.com/SahakyanLab/DNAkmerQM" target="_blank" rel="noopener">https://github.com/SahakyanLab/DNAkmerQM</a>, the corresponding research and development code from <a href="https://github.com/SahakyanLab/NucleicAcidsQM" target="_blank" rel="noopener">https://github.com/SahakyanLab/NucleicAcidsQM</a>, and the associated pre-print from <a href="https://doi.org/10.1101/2023.01.25.525597" target="_blank" rel="noopener">https://doi.org/10.1101/2023.01.25.525597</a>.</p>
Taxonomic classification based on k-mers
<p>DNA sequencing provides the possibility to obtain complete genomic DNA from environmental samples without the need for laboratory microbiological cultures. To this end, metagenomics, the direct DNA sequencing from microbial communities, has changed radically the field of microbiology, by unearthing a broad space of the planet’s microbial diversity, much of which remains unknown. Metagenomic approaches have become standard methods for identifying the biodiversity and the gene or metabolomic functionalities of bacterial and archaeal communities, with many applications not only in microbial ecology but also in public health as in clinical diagnostics and detection of pathogens. Yet, the decrease in the cost of high throughput sequencing and the great amount of microbial data produced every day, highlight one of the main biological questions: the taxonomic classification of metagenomic short reads. Which organisms are contained in a sample? Are there any features that can be used to identify them?</p> <p>To this end, many algorithms have been developed that achieve high speed, by counting k-mers, short sequence substrings of fixed-length k. In this way for the provided input sequences, a list of features can be computed that describes each one of them. Subsequently, the question now reforms to how can the produced k-mers be used for the taxonomic classification of the input sequences.</p> <p>Sample processing, sequencing, and core amplicon data analysis were performed by the Earth Microbiome Project (www.earthmicrobiome.org), and all amplicon sequence data and metadata have been made public through the EMP data portal (qiita.microbio.me/emp):</p> <ul> <li>Thompson, L. R., Sanders, J. G., McDonald, D., Amir, A., …, Jansson, J. K., Gilbert, J. A., Knight, R., & The Earth Microbiome Project Consortium. (2017). A communal catalogue reveals Earth’s multiscale microbial diversity. Nature, 551:457-463. doi:10.1038/nature24621.</li> </ul>
Multi-reference genome and K-mer based association mapping in Zymoseptoria tritici
<p>Data tables for a study of multi-reference genome and K-mer based association mapping of the fungal wheat pathogen <em>Zymoseptoria tritici</em></p>
Plasmer database for k-mer and genomic features
<p>This is the inital version v1.0 of Plasmer database for k-mer and genomic features.</p> <p>Download and extract the package, and provide the absolute path to the Plasmer command line.</p> <p> </p> <p>For more information about Plasmer, please refer to our GitHub repository at: <a href="https://github.com/nekokoe/plasmer">https://github.com/nekokoe/plasmer</a></p>
Weighted k-mer datasets
<p>These are the datasets used in the experiments of the paper: <em>On Weighted k-mer Dictionaries</em> - Giulio Ermanno Pibiri. Algorithms for Molecular Biology, <strong>18</strong>, Article number: 3 (2023) DOI: <a href="https://almob.biomedcentral.com/articles/10.1186/s13015-023-00226-2">10.1186/s13015-023-00226-2</a>. (A preliminary version of the paper has been published in WABI 2022: <a href="https://doi.org/10.4230/LIPIcs.WABI.2022.9">10.4230/LIPIcs.WABI.2022.9</a>.)</p>
Leveraging basecaller's move table to generate a lightweight k-mer model
<p>The ONT RNA004 dataset that is used to create a 5-mer model using the basecaller's movetable.</p> <p>The dataset is a sub-sample extracted from a dataset that sequenced Universal Human Reference RNA.</p> <p>The specification of the bio sample is here</p> <p>https://www.agilent.com/cs/library/usermanuals/public/740000.pdf</p>
Genomic datasets used for evaluation of k-mer representations and indexes
<p>This record contains genomic datasets, including subsampled <em>k</em>-mer sets for some datasets (files with names containing _subsampled_). Namely, it provides the following datasets:</p> <ul> <li>Two <em>E. coli</em> pan-genomes, obtained as the union of the <em>E. coli</em> genomes from the 661k collection. One contains <em>all</em> genomes (without quality filtering) and for the other (<em>HQ</em>) we applied high-quality filtering.</li> <li><em>S. pneumoniae</em> pan-genome: 616 genomes, as provided in RASE DB <em>S. pneumoniae</em> <a href="https://github.com/c2-d2/rase-db-spneumoniae-sparc/">https://github.com/c2-d2/rase-db-spneumoniae-sparc/</a></li> <li><em>SARS-CoV-2</em> pan-genome, downloaded from GISAID <a href="https://gisaid.org/">https://gisaid.org/</a> (access upon registration) on Jan 25, 2023 (GISAID version 2023/01/23, 14,682,066 genomes, 430 Gbp).</li> <li>Metagenomic sample SRS063932 (Illumina raw reads) of human microbiome with accession SRX023459, download from <a href="https://www.hmpdacc.org/hmp/HMASM/">https://www.hmpdacc.org/hmp/HMASM/</a>. The fastq files were converted to FASTA files using `seqtk seq -A -C`.</li> <li>Human RNA-seq Illumina raw reads with accession SRX348811, downloaded using the prefetch tool from the SRA toolkit and then converted into the FASTA format by<br>`fastq-dump --split-3 --fasta`.</li> <li>Human genome Illumina raw reads with accession SRX016231, downloaded using the prefetch tool from the SRA toolkit and then converted into the FASTA format by<br>`fastq-dump --split-3 --fasta`.</li> <li>Human genome assembly chm13v2.0 (T2T), downloaded from <a href="https://s3-us-west-2.amazonaws.com/human-pangenomics/T2T/CHM13/assemblies/analysis_set/chm13v2.0.fa.gz">https://s3-us-west-2.amazonaws.com/human-pangenomics/T2T/CHM13/assemblies/analysis_set/chm13v2.0.fa.gz</a>.</li> <li>Two MiniKraken datasets (4GB and 8GB), downloaded from <a href="https://ccb.jhu.edu/software/kraken/">https://ccb.jhu.edu/software/kraken/</a>, with the 31-mers dumped using Jellyfish 1.1.12.</li> </ul> <p>The resulting FASTA files (apart from the human genome assembly chm13v2.0 and MiniKraken datasets) were converted to unitigs by GGCAT v1.1.0<br>by `ggcat build -k {kmer-size} -m 200 -j 5 -s {min-freq} -o {preprocessed_unitigs} {input_FASTA}`, where we used $k=128$ and `{min-freq}`=1 for pan-genomes and $k=32$ and `{min-freq}`=2 for dataset from raw reads.</p> <p>Finally, the subsampled files `{dataset}_subsampled_k{$k$}_r0.1.fa.xz` contain 10% randomly chosen distinct canonical $k$-mers from the whole $k$-mer set of the given dataset. The FASTA file contains one subsampled <em>k</em>-mer per sequence.</p>
General and Human Specific (sK2-sK6) Universal k-mer Sets
<p>Data sets produced for the publication<strong> Practical universal k-mer sets for minimizer schemes</strong><strong>. </strong>Dan DeBlasio, Fiyinfoluwa Gbosibo, Carl Kingsford , and Guillaume Marçais presented at ACM-BCB 2019. Intended to be used with the ftrie library at https://github.com/Kingsford-Group/remuval</p>
Human Specific (sK7-sK10) Universal k-mer Sets
<p>Data sets produced for the publication<strong>Practical universal k-mer sets for minimizer schemes</strong><strong>. </strong>Dan DeBlasio, Fiyinfoluwa Gbosibo, Carl Kingsford , and Guillaume Marçais presented at ACM-BCB 2019. Intended to be used with the ftrie library at https://github.com/Kingsford-Group/remuval</p>
KTU: K-mer Taxonomic Units improve the biological relevance of amplicon sequence variant microbiota data
<p>Testing datasets and files for the KTU algorithm</p>
K-mer collision statistics (BLEND: A Fast, Memory-Efficient, and Accurate Mechanism to Find Fuzzy Seed Matches in Genome Analysis)
<p>This dataset contains 1,077 FASTA files and CSV files. Each FASTA file includes 25-character long sequences similar to each other.</p> <p>We have a CSV file for each tool (i.e., minimap2 and BLEND) and configuration (i.e., different number of neighbors in BLEND). CSV files include the non-identical k-mer pairs (16-mers) that generate the same hash value (i.e., collisions). These k-mers are extracted from sequences that are similar to each other. In each line, we show the hash value of the k-mers, the actual sequene pairs that the k-mers are extracted from, k-mer pairs that generate the same hash value, and the edit distance between these k-mers.</p> <p> </p>
Yak k-mer dumps for partition human chrX/Y in de novo assemblies
<p>See <a href="https://github.com/lh3/yak">yak</a> for details.</p>
K-mer matrix for R-gene enrichment data of wheat Watkins diversity panel.
<p>K-mer matrix for RenSeq data of wheat lines including 300 wheat landraces from Watkins collection. The matrix is divided into 40 parts. The file “watkins_matrix_header.txt” contains the accession names and also specifies the order in which presence/absence of a k-mer is scored in the presence/absence matrix. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.