Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

39

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

39 results for “genomic databases”

Learn how ShareScore rates datasets ↗
zenodo44/100

1000 Genomes Project Transposable Element database

<p>Multi-sample VCF with transposable elements across individuals in the 1KGP dataset. Transposable elements were called using RetroSeq&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Databases for exploratory mode of RRE-Finder: A Genome-Mining Tool for Class-Independent RiPP Discovery

<p>RREFinder is a bioinformatic tool for the detection of RiPP Recognition Elements (RREs). See &quot;RRE-Finder: A Genome-Mining Tool for Class-Independent RiPP Discovery&quot;.</p> <p>This database contains the required databases to run exploratory mode of the tool.</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

COMBAT TB Tuberculosis genome annotation database

<p>A Neo4j (version 2.3.3) format graph database containing annotation related to the M. tuberculosis H37Rv genome, created as part of the COMBAT TB project at the South African National Bioinformatics Institute.</p>

opencc-by-4.0May 2016View details →
zenodo40/100

The mOTUs online database provides web-accessible genomic context to taxonomic profiling of microbial communities - Supplementary Tables

<p><strong>Supplementary Table 1:</strong></p> <p>A map between each of the genomes in mOTUs-db (3&rsquo;747&rsquo;151), the&nbsp;associated study and its metagenomic sample (in case of MAGs).</p> <p>Columns:</p> <p><code>&nbsp; &nbsp; GENOME &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &rarr; Unique mOTUs-db name of the genome</code><br><code>&nbsp; &nbsp; STUDY &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&rarr; Unique mOTUs-db name of the study</code><br><code>&nbsp; &nbsp; IS_MAG &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &rarr; True if genome is a MAG, otherwise False&nbsp;</code><br><code>&nbsp; &nbsp; METAGENOMIC_SAMPLE &rarr; Unique name of the metagenomic sample or NA in case of non-MAG genome</code></p> <p>Example:</p> <p><code>&nbsp; &nbsp; GENOME&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;STUDY &nbsp; &nbsp; &nbsp; &nbsp;IS_MAG &nbsp; &nbsp;METAGENOMIC_SAMPLE</code><br><code>&nbsp; &nbsp; ---------------------------------------------------------------------------------------------</code><br><code>&nbsp; &nbsp; ACIN21-1_SAMN05421555_MAG_00000001&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;ACIN21-1&nbsp; &nbsp; &nbsp;True&nbsp; &nbsp; &nbsp; ACIN21-1_SAMN05421555_METAG</code><br><code>&nbsp; &nbsp; RSGB23-1_GCA-006096615-V1_GENO_10000001 &nbsp; &nbsp;RSGB23-1&nbsp; &nbsp; &nbsp;False&nbsp; &nbsp; &nbsp;NA</code></p> <p><strong>Supplementary Table 2:</strong></p> <p>A map between all non-MAG genomes (919&rsquo;090) and their source&nbsp;(e.g. Refseq or JGI).</p> <p>Columns:</p> <p><code>&nbsp; &nbsp; GENOME &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &rarr; Unique mOTUs-db name of the genome</code><br><code>&nbsp; &nbsp; SOURCE_SAMPLE_LINK &rarr; Link to the original location of this genome</code></p> <p>Example:</p> <p><code>&nbsp; &nbsp; #GENOME &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; SOURCE_SAMPLE_LINK</code><br><code>&nbsp; &nbsp; --------------------------------------------------------------------------------------------------------</code><br><code>&nbsp; &nbsp; JGIG23-1_GA0055041_GENO_10000001&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; https://gold.jgi.doe.gov/analysis_project?id=Ga0055041</code><br><code>&nbsp; &nbsp; RSGB23-1_GCA-006717865-V1_GENO_10000001 &nbsp; &nbsp; https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_006717865.1</code></p> <p><strong>Supplementary Table 3:</strong></p> <p>A list of all metagenomic studies processed for the mOTUs-db, their&nbsp;number of samples, the number of reconstructed MAGs and the associated&nbsp;publication.</p> <p>Columns:</p> <p><code>&nbsp; &nbsp; STUDY &nbsp; &nbsp; &nbsp; --&gt; Unique mOTUs-db study identifier</code><br><code>&nbsp; &nbsp; BIOPROJECT &nbsp;--&gt; Public identifier (NCBI/JGI) of metagenomic sequencing project</code><br><code>&nbsp; &nbsp; SAMPLES &nbsp; &nbsp; --&gt; Number of metagenomic samples</code><br><code>&nbsp; &nbsp; MAGs &nbsp; &nbsp; &nbsp; &nbsp;--&gt; Number of reconstructed MAGs</code><br><code>&nbsp; &nbsp; PUBLICATION --&gt; Link to publication</code></p> <p>Example:</p> <p><code>&nbsp; &nbsp; STUDY &nbsp; &nbsp; &nbsp; &nbsp;BIOPROJECT &nbsp; &nbsp;SAMPLES &nbsp; &nbsp;MAGs&nbsp; &nbsp; &nbsp;PUBLICATION</code><br><code>&nbsp; &nbsp; -------------------------------------------------------------------------------------------------</code><br><code>&nbsp; &nbsp; ACIN21-1&nbsp; &nbsp; &nbsp;PRJEB44456 &nbsp; &nbsp;58&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;1,110 &nbsp; &nbsp;https://www.nature.com/articles/s42003-021-02112-2</code></p> <p><strong>Supplementary Table 4:</strong></p> <p>Mapping between mOTUs-db sample identifier, the associated biosample and&nbsp;the environment.</p> <p>Columns:</p> <p><code>&nbsp; &nbsp; SAMPLE &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; --&gt; Unique mOTUS-db sample identifier</code><br><code>&nbsp; &nbsp; BIOSAMPLE &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;--&gt; Public identifier (NCBI/JGI) of metagenomic sample</code><br><code>&nbsp; &nbsp; STUDY &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;--&gt; Unique mOTUs-db study identifier</code><br><code>&nbsp; &nbsp; ENVIRONMENT &nbsp; &nbsp; &nbsp; &nbsp;--&gt; Environment of metagenomic sample</code><br><code>&nbsp; &nbsp; SOURCE_SAMPLE_LINK --&gt; Link to the original location of this sample</code></p> <p>Example:</p> <p><code>&nbsp; &nbsp; #SAMPLE&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; BIOSAMPLE&nbsp; &nbsp; &nbsp;STUDY&nbsp; &nbsp; ENVIRONMENT&nbsp; SOURCE_SAMPLE_LINK</code><br><code>&nbsp; &nbsp; ---------------------------------------------------------------------------------------------------------------------</code><br><code>&nbsp; &nbsp; ACIN21-1_SAMN05421555_METAG&nbsp; SAMN05421555&nbsp; ACIN21-1 marine&nbsp; &nbsp; &nbsp; &nbsp;https://www.ncbi.nlm.nih.gov/biosample/SAMN05421555/</code></p> <p><strong>Supplementary Table 5:</strong></p> <p>A list of environments covered in the mOTUs-db mapped to the respective&nbsp;NCBI taxonomy (if possible)</p> <p>Columns:</p> <p><code>&nbsp; &nbsp; TERM &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; --&gt; Unique environment name</code><br><code>&nbsp; &nbsp; NCBI TAXONOMY ID --&gt; Link to the NCBI taxonomy</code></p> <p>Example:</p> <p><code>&nbsp; &nbsp; TERM&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; NCBI TAXONOMY ID</code><br><code>&nbsp; &nbsp; ----------------------------------------------</code><br><code>&nbsp; &nbsp; activated sludge metagenome&nbsp; &nbsp;NCBI:txid942017</code><br><code>&nbsp; &nbsp; air metagenome &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;NCBI:txid655179</code></p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Metagenome assembled genome database of a human cohort and fecal reactors

<p><strong>HumanCohort_annotations.tsv.zip:</strong> This is the custom MAG database (n=2447 MAGs)&nbsp;and&nbsp;corresponding annotations that&nbsp;were&nbsp;used&nbsp;in&nbsp;Borton 2022: &quot;Targeted curation of the gut microbial gene content modulating human cardiovascular disease&quot;. The citation will be updated upon publication of the manuscript. Metagenome assembled genomes were generated from fecal metagenomes derived from a 54 person cohort and anoxic methylated amine enrichments.&nbsp;</p> <p><strong>HumanCohortmetabolism_summary.xlsx.zip:&nbsp;</strong> This is the annotation summary for 2447 MAGs in the cohort database.&nbsp;</p> <p><strong>Quality_Abundance_CohortMAGs.xlsx: </strong>This is a genome inventory of the&nbsp;2447 MAGs in the cohort database including genome statistics and relative abundance.&nbsp;</p> <p><strong>orig_1D_NMR_fids.zip:&nbsp;</strong>NMR data derived from anoxic methylated amine enrichments.&nbsp;</p>

opencc-by-4.0Apr 2021View details →
dryad40/100

Data from: Improved genome assembly of the whiteleg shrimp Penaeus (Litopenaeus) vannamei using long- and short-read sequences from public databases

Open the record for dataset details and reuse information.

publicMar 2024View details →
zenodo36/100

UHGG v1 database for inStrain genome resolved metagenomic analysis

<p>A series of files that are useful for profiling metagenomic communities with the program inStrain.</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

COMBAT TB Tuberculosis genome annotation database

<p>This is a Neo4j format *M. tuberculosis* reference annotation database. To view it you need&nbsp;a Neo4j (v 2.3)&nbsp;instance. There is a script (<strong>run_db.sh</strong>) that will start a Neo4j instance for&nbsp;you based on this data, using docker. Run that (<strong>bash run_db.sh</strong>) and connect to http://localhost:7474.</p> <p>The database was created by the COMBAT TB project&nbsp;(http://christoffels.sanbi.ac.za/index.php/projects/combat-tb) at the South African National Bioinformatics Institute (SANBI).<br /> <br /> Authors: Thoba Lose, Peter van Heusden, Ziphozakhe Mashologu, Alan Christoffels .</p> <p>The COMBAT TB project is funded by the South African Medical Research Council (MRC) and was supported by the South African<br /> Research Chairs Initiative of the Department of Science and Technology and National Research Foundation of South Africa.</p>

opencc-by-sa-4.0Jun 2016View details →
zenodo36/100

Genome Database: Turnover of strain-level diversity modulates functional traits in the honeybee gut microbiome between nurses and foragers

<p>This repository contains the dataset used in the publication "Turnover of strain-level diversity modulates functional traits in the honeybee gut microbiome between nurses and foragers," which is currently under revision. A pre-print can be found <a href="https://doi.org/10.1101/2022.12.29.522137">here</a>. The database is based on previously published work to create a genomic database of honeybee gut microbes by Kirsten Ellegaard (2021), found <a href="https://zenodo.org/records/4661061">here.</a></p><p>The zipped folder deposited here after unzipping, should contain the following files and directories:</p><ul><li>honeybee_genome.fasta : fasta file containing the host (<i>Apis mellifera</i>) genome sequence</li><li>beebiome_db : fasta file of 198 concatenated genomes with one genome per entry (multi-line fasta) where the headers represent the genome identifier</li><li>beebiome_red_db : fasta file of 39 species representative genomes with one genome per entry (multi-line fasta) where the headers represent the genome identifier to be used for the analysis of intra-specific variation</li><li>fna_files : directory containing genome sequence files and concatenated files where the concatenated files contain one fasta entry renamed to the genome identifier and all contigs concatenated into one entry</li><li>ffn_files : directory containing one file per genome listing the nucleotide sequence of all the predicted genes</li><li>faa_files : directory containing one file per genome listing the amino acid sequence of all the predicted genes</li><li>bed_files : directory containing bed files where the location of each of the predicted genes are indicated based on their position in the concatenated genome file</li><li>single_ortho : directory containing one file per phylotype listing all the single-copy orthogroups (OGs) identified by orthofinder where each line represents an OG id followed by a list of genes from each of the genomes of that phylotype that belong to that OG and the corresponding sequences of these genes can be found in the ffn file belonging to the respective genome</li><li>red_bed_files : directory containing bed files for species representative genomes that only list the positions genes that belong to the core orthogroups of their phylotype</li></ul><p>Further information about how this genome database was used to analyze strain-level diversity can be found in the publication and accompanying code repository.</p>

opengpl-3.0-or-laterSep 2023View details →
zenodo36/100

Database of giant viruses, Mirusviruses genomes, and marker genes

<p>Nucleotide sequences of 1,629 viral genomes (1,518&nbsp;<em>Nucleoviricota</em>&nbsp;and 111&nbsp;<em>Mirusviricota</em>), and marker genes found in these genomes.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Advancing source tracking: systematic review and source-specific genome database curation of fecally shed prokaryotes

<p>This repository contains several files describing the data discussed in Lindner et al's "Advancing source tracking: systematic review and source-specific genome database curation of fecally shed prokaryotes".&nbsp;</p> <ol> <li>"database.fna" = concatenation of the (draft or complete) genome sequences described in the paper (n=12,730 source associated prokaryotic genomes) which passed quality checks and are species-level representatives (i.e., dereplicated at 95% ANI).</li> <li>"gdef.txt" = a manifest describing which sequences belong to which genomes.</li> <li>"sources.txt" = a manifest describing which genomes belong to which sources.</li> </ol> <p>The sources this database covers:</p> <table> <tbody> <tr> <td>Source Category</td> <td>Species-level <br>genome count</td> <td>Source-specific<br>species-level&nbsp;<br>genome count</td> </tr> <tr> <td>Bird</td> <td>56</td> <td>40</td> </tr> <tr> <td>Cat</td> <td>86</td> <td>6</td> </tr> <tr> <td>Chicken</td> <td>1314</td> <td>887</td> </tr> <tr> <td>Cow</td> <td>39</td> <td>15</td> </tr> <tr> <td>Dog</td> <td>139</td> <td>56</td> </tr> <tr> <td>Pig</td> <td>2764</td> <td>2035</td> </tr> <tr> <td>Ruminant</td> <td>740</td> <td>714</td> </tr> <tr> <td>Human</td> <td>4484</td> <td>3350</td> </tr> <tr> <td>Wastewater</td> <td>3108</td> <td>3097</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>See publication for further details.&nbsp;</p>

opencc-by-4.0Feb 2024View details →
dryad36/100

Transitioning from environmental genetics to genomics using mitogenome reference databases

<p><span>Species detection using eDNA is revolutionizing the global capacity to monitor biodiversity. However, the lack of regional, vouchered, genomic sequence information—especially sequence information that includes intraspecific variation—creates a bottleneck for management agencies wanting to harness the complete power of eDNA to monitor taxa and implement eDNA analyses. eDNA studies depend upon regional databases of complete mitogenomic sequence information to evaluate the effectiveness of such data to differentiate, identify and detect taxa. We created the Oregon Biodiversity Genome Project working group to utilize recent advances in sequencing technology to create a database of complete, near error-free mitogenomic sequences for all of Oregon's resident freshwater fishes. So far, we have successfully assembled the complete mitogenomes of 313 specimens of freshwater fish representing 7 families, 55 genera, and 129 (88%) of the 146 resident species and lineages. Our comparative analyses of these sequences illustrate that the short (~150 bp) mitochondrial "barcode" regions typically used for eDNA assays are not consistently diagnostic for species-level identification and that no single region is best for metabarcoding Oregon's fishes. However, often-overlooked intergenic regions of the mitogenome such as the D-loop have the potential to reliably diagnose and differentiate species. This project provides a blueprint for other researchers to follow as they build regional databases. It also illustrates the taxonomic value and limits of complete mitogenomic sequences, and how current eDNA assays and the "PCR-free" environmental genomics methods of the future can best leverage this information.</span></p>

opencc-zeroApr 2022View details →
zenodo36/100

EvoMining genomic and enzyme databases for Actinobacteria, Cyanobacteria, Pseudomonas and Archaea

<p>Databases for EvoMining 2.0</p> <p>Genomic DB is a collection of genomes of a certain taxonomical group, functionally annotated by RAST.</p> <p>Enzyme-DB</p> <p>Actinobacteria</p> <p>Cyanobacteria</p> <p>Pseudomonas</p> <p>Archaea</p> <p>SampleData</p>

opencc-by-4.0Apr 2018View details →
zenodo36/100

Lite Kraken/Bracken databases built using UHGG genomes

<p>Kraken/Braken databases for UHGG genomes.</p> <p>HUMAN_3006.tar.gz: Kraken/Bracken database for 3006 high quality species clusters of the UHGG (Beresford-Jones et al., 2022). Database was built from the single highest quality genome for each species cluster (n=3006).&nbsp;Uses&nbsp;the original&nbsp;GTDB v1.3&nbsp;taxonomy.</p> <p>UHGG_5987_KRAKEN.tar.gz:&nbsp;Kraken/Bracken database for 3006 high quality species clusters of the UHGG (Beresford-Jones et al., 2022).&nbsp;Species clusters are represented by a variable number of high&nbsp;quality genomes (n=5987 in total), selected to maximise represented&nbsp;taxonomic diversity. Uses a custom taxonomy modified from GTDB v2.1 with&nbsp;each species cluster being represented by a&nbsp;species level taxonomic annotation.&nbsp;</p> <p>&nbsp;</p> <p>Methods:</p> <p>Databases built using Kraken&nbsp;v2.1.2 and Bracken v2.6.2. Commands used to build the databases are included below.</p> <p>kraken2-build --build --db Kraken --threads 12</p> <p>bracken-build -d Kraken -k 35 -l 150 -t 12</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Genomically predicted theoretical protein mass database for mass spectrometry (GPMsDB) evaluation datasets

<p>These are datasets obtained for the evaluation of&nbsp;GPMsDB (genomically predicted protein mass database) and its toolkits (GPMsDB-tk/GPMsDB-dbtk). The following datasets are deposited.</p> <ul> <li>The genome sequences of the strains newly sequenced and added using GPMsDB-dbtk (genomes_added.zip)</li> <li>MALDI-TOF-MS peak lists obtained from reference bacterial and archaeal strains (MALDI_peaklists.zip)</li> <li>16S rRNA gene sequences of the faecal isolates (mice_isolates_16S_nanopore.zip)</li> <li>Metagenome-assembled genomes from mouse faeces&nbsp;(mice_MAGs.zip)</li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Plasmer database for k-mer and genomic features

<p>This is the inital version v1.0 of Plasmer database for k-mer and genomic features.</p> <p>Download and extract the package, and provide the absolute path to the Plasmer command line.</p> <p>&nbsp;</p> <p>For more information about Plasmer, please refer to our GitHub repository at: <a href="https://github.com/nekokoe/plasmer">https://github.com/nekokoe/plasmer</a></p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Exposing New Taxonomic Variation with Inflammation – A Model-Specific Genome Database for Microbiome Researchers

<p>Data deposit for CBAJ-DB v1.2</p>

opencc-by-4.0Oct 2022View details →
dryad36/100

Transitioning from environmental genetics to genomics using mitogenome reference databases

Open the record for dataset details and reuse information.

publicApr 2022View details →
zenodo32/100

microbetag : building a thorough database of genome-scale KO annotations

<p>In this repository we keep internal data for the <em><a href="https://hariszaf.github.io/microbetag/">microbetag</a> </em>microbial co-occurrence network annotator.</p> <p><em>microbetag</em>&nbsp;makes use of 2-column files for each genome, indicating the KO term found and a KEGG module in which this terms takes part into. <br>As a single KO term might participates in more than one KEGG modules, the same KO might be more than once in an annotation file.&nbsp;</p> <table> <tbody> <tr> <td> <div>chem_xref.tar.gz</div> </td> <td> <p>The MNXref namespace</p> <ol> <li>The identifier of a chemical compound in an external resource [<a href="https://www.metanetx.org/mnxdoc/mnxref.html#XREF">XREF</a>]</li> <li>The corresponding identifier in the MNXref namespace [<a href="https://www.metanetx.org/mnxdoc/mnxref.html#MNX_ID">MNX_ID</a>]</li> <li>The description given by the external resource [<a href="https://www.metanetx.org/mnxdoc/mnxref.html#STRING">STRING</a>]</li> </ol> <p>MNXref 4.0 release notes: - The third column (evidence tag for the mapping) was suppressed - The descriptions were completed - Deprecated identifiers were moved into they own table below</p> </td> </tr> <tr> <td> <div>gtdb_modelseed_gems.zip</div> </td> <td> <p>for all the GTDB genomes their corresponding <a href="https://patricbrc.org/">PATRIC</a> annotations were gathered. Then, using <a href="https://github.com/ModelSEED/ModelSEEDpy">modelseedpy</a> we constructed their genome scale metabolic reconstructions</p> </td> </tr> <tr> <td> <div>gtdb_kofam_scan_per_module.tar.gz</div> </td> <td> <p>all representative genomes of <a href="https://gtdb.ecogenomic.org/">GTDB</a>&nbsp;(v.202) were parsed and their corresponding `.faa` files were retrieved from the <a href="https://ftp.ncbi.nlm.nih.gov/genomes/all/">NCBI FTP</a>. Then the <a href="https://github.com/takaram/kofam_scan">kofam_scan</a> tool was used to annotate them and finally a <a href="https://github.com/hariszaf/microbetag/blob/clean/mappings/gtdb_mappings/gtdb_annotations_per_module.py">manual script </a>was used to keep KOs of each genome per module.&nbsp;</p> </td> </tr> <tr> <td>SeedSet.pkl.gz</td> <td> <p>A pickle file with the seeds of each GEM included in the&nbsp;<em>gtdb_modelseed_gems.zip </em>file and related to the KEGG MODULES based on the&nbsp;<em>seedId_keggId_module.tsv </em>file you can find on microbetag's GitHub page. &nbsp;Example:</p> <p>PATRIC&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; SeedSet<br>373.172 &nbsp; &nbsp;[cpd00891, cpd00136, cpd00199, cpd01772, cpd00...<br>397278.5 &nbsp; [cpd00891, cpd00136, cpd01772, cpd02698, cpd08...</p> </td> </tr> <tr> <td>NonSeedSet.pkl.gz</td> <td> <p>A pickle file with the non seeds of each GEM included in the <em>gtdb_modelseed_gems.zip </em>file and related to the KEGG MODULES based on the&nbsp;<em>seedId_keggId_module.tsv </em>file you can find on microbetag's GitHub page. &nbsp;Example:</p> <p>PATRIC&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; NonSeedSet<br>64187.548 &nbsp; [cpd00508, cpd00869, cpd00774, cpd03830, cpd00...<br>74426.1719 &nbsp;[cpd00204, cpd00447, cpd20171, cpd03470, cpd00...</p> </td> </tr> <tr> <td>seeds_per_genome.pkl.gz</td> <td> <p>A pickle file with a binary representation of the seeds per genome . &nbsp;Example:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;cpd00493 &nbsp;cpd00296 &nbsp;cpd11431 &nbsp;cpd00063 &nbsp;cpd15717&nbsp; &nbsp;...&nbsp;<br>2162051.4 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;1 &nbsp; &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; &nbsp; 1 &nbsp; &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; &nbsp; 0 &nbsp; &nbsp; &nbsp; &nbsp; 0 &nbsp;...&nbsp; &nbsp; &nbsp;</p> </td> </tr> <tr> <td> <div>nonseeds_per_genome.pkl.gz</div> </td> <td>Like above for non-seeds.</td> </tr> <tr> <td> <div>phen_classes.zip</div> </td> <td>A list of pickle files with the re-trained classes of phenDB for the prediction of functional traits on a genome.</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

SCoV2-VAR: A light-weighted, customizable, and open-source database of 12 million SARS-CoV-2 genomes

<pre>Explosive accumulation of SARS-CoV-2 variants is posing a challenge to monitoring virus mutation and other data dealing, particularly based on centralized databases. The present study aimed to establish a light-weighted, customizable, and open-source database for SARS-CoV-2 genomes and annotations, without any access limit. The database, named SCoV2-VAR, was constructed, based on the variations (VAR) of the full-length SARS-CoV-2 (SCoV2) data uploaded on websites. All sequence samples were subject to quality control, single nucleotide polymorphism (SNP) annotation, format conversion, and final compression before appending to SCoV2-VAR. The final version of SCoV2-VAR (up to Feb 2024) contained more than 12 million SARS-CoV-2 records, with full genome and annotations. SCoV2-VAR was extremely light-weighted, with a storage size of 937 Mb for all 12 million sequences, post a 1: 596 compression. SCoV2-VAR is capable of timely updating, quickly querying, and customizable outputting SARS-CoV-2 sequences and their annotations. Additionally, the present study provided an overview of all 12 million SARS-CoV-2 samples, for both sequences and annotations.<br> <br><br></pre>

opencc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record