Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
39
datasets available to search
ShareScore release 0.9.0
Dataset results
39 results for “genomic databases”
1000 Genomes Project Transposable Element database
<p>Multi-sample VCF with transposable elements across individuals in the 1KGP dataset. Transposable elements were called using RetroSeq </p>
Databases for exploratory mode of RRE-Finder: A Genome-Mining Tool for Class-Independent RiPP Discovery
<p>RREFinder is a bioinformatic tool for the detection of RiPP Recognition Elements (RREs). See "RRE-Finder: A Genome-Mining Tool for Class-Independent RiPP Discovery".</p> <p>This database contains the required databases to run exploratory mode of the tool.</p>
COMBAT TB Tuberculosis genome annotation database
<p>A Neo4j (version 2.3.3) format graph database containing annotation related to the M. tuberculosis H37Rv genome, created as part of the COMBAT TB project at the South African National Bioinformatics Institute.</p>
The mOTUs online database provides web-accessible genomic context to taxonomic profiling of microbial communities - Supplementary Tables
<p><strong>Supplementary Table 1:</strong></p> <p>A map between each of the genomes in mOTUs-db (3’747’151), the associated study and its metagenomic sample (in case of MAGs).</p> <p>Columns:</p> <p><code> GENOME → Unique mOTUs-db name of the genome</code><br><code> STUDY → Unique mOTUs-db name of the study</code><br><code> IS_MAG → True if genome is a MAG, otherwise False </code><br><code> METAGENOMIC_SAMPLE → Unique name of the metagenomic sample or NA in case of non-MAG genome</code></p> <p>Example:</p> <p><code> GENOME STUDY IS_MAG METAGENOMIC_SAMPLE</code><br><code> ---------------------------------------------------------------------------------------------</code><br><code> ACIN21-1_SAMN05421555_MAG_00000001 ACIN21-1 True ACIN21-1_SAMN05421555_METAG</code><br><code> RSGB23-1_GCA-006096615-V1_GENO_10000001 RSGB23-1 False NA</code></p> <p><strong>Supplementary Table 2:</strong></p> <p>A map between all non-MAG genomes (919’090) and their source (e.g. Refseq or JGI).</p> <p>Columns:</p> <p><code> GENOME → Unique mOTUs-db name of the genome</code><br><code> SOURCE_SAMPLE_LINK → Link to the original location of this genome</code></p> <p>Example:</p> <p><code> #GENOME SOURCE_SAMPLE_LINK</code><br><code> --------------------------------------------------------------------------------------------------------</code><br><code> JGIG23-1_GA0055041_GENO_10000001 https://gold.jgi.doe.gov/analysis_project?id=Ga0055041</code><br><code> RSGB23-1_GCA-006717865-V1_GENO_10000001 https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_006717865.1</code></p> <p><strong>Supplementary Table 3:</strong></p> <p>A list of all metagenomic studies processed for the mOTUs-db, their number of samples, the number of reconstructed MAGs and the associated publication.</p> <p>Columns:</p> <p><code> STUDY --> Unique mOTUs-db study identifier</code><br><code> BIOPROJECT --> Public identifier (NCBI/JGI) of metagenomic sequencing project</code><br><code> SAMPLES --> Number of metagenomic samples</code><br><code> MAGs --> Number of reconstructed MAGs</code><br><code> PUBLICATION --> Link to publication</code></p> <p>Example:</p> <p><code> STUDY BIOPROJECT SAMPLES MAGs PUBLICATION</code><br><code> -------------------------------------------------------------------------------------------------</code><br><code> ACIN21-1 PRJEB44456 58 1,110 https://www.nature.com/articles/s42003-021-02112-2</code></p> <p><strong>Supplementary Table 4:</strong></p> <p>Mapping between mOTUs-db sample identifier, the associated biosample and the environment.</p> <p>Columns:</p> <p><code> SAMPLE --> Unique mOTUS-db sample identifier</code><br><code> BIOSAMPLE --> Public identifier (NCBI/JGI) of metagenomic sample</code><br><code> STUDY --> Unique mOTUs-db study identifier</code><br><code> ENVIRONMENT --> Environment of metagenomic sample</code><br><code> SOURCE_SAMPLE_LINK --> Link to the original location of this sample</code></p> <p>Example:</p> <p><code> #SAMPLE BIOSAMPLE STUDY ENVIRONMENT SOURCE_SAMPLE_LINK</code><br><code> ---------------------------------------------------------------------------------------------------------------------</code><br><code> ACIN21-1_SAMN05421555_METAG SAMN05421555 ACIN21-1 marine https://www.ncbi.nlm.nih.gov/biosample/SAMN05421555/</code></p> <p><strong>Supplementary Table 5:</strong></p> <p>A list of environments covered in the mOTUs-db mapped to the respective NCBI taxonomy (if possible)</p> <p>Columns:</p> <p><code> TERM --> Unique environment name</code><br><code> NCBI TAXONOMY ID --> Link to the NCBI taxonomy</code></p> <p>Example:</p> <p><code> TERM NCBI TAXONOMY ID</code><br><code> ----------------------------------------------</code><br><code> activated sludge metagenome NCBI:txid942017</code><br><code> air metagenome NCBI:txid655179</code></p>
Metagenome assembled genome database of a human cohort and fecal reactors
<p><strong>HumanCohort_annotations.tsv.zip:</strong> This is the custom MAG database (n=2447 MAGs) and corresponding annotations that were used in Borton 2022: "Targeted curation of the gut microbial gene content modulating human cardiovascular disease". The citation will be updated upon publication of the manuscript. Metagenome assembled genomes were generated from fecal metagenomes derived from a 54 person cohort and anoxic methylated amine enrichments. </p> <p><strong>HumanCohortmetabolism_summary.xlsx.zip: </strong> This is the annotation summary for 2447 MAGs in the cohort database. </p> <p><strong>Quality_Abundance_CohortMAGs.xlsx: </strong>This is a genome inventory of the 2447 MAGs in the cohort database including genome statistics and relative abundance. </p> <p><strong>orig_1D_NMR_fids.zip: </strong>NMR data derived from anoxic methylated amine enrichments. </p>
Data from: Improved genome assembly of the whiteleg shrimp Penaeus (Litopenaeus) vannamei using long- and short-read sequences from public databases
Open the record for dataset details and reuse information.
UHGG v1 database for inStrain genome resolved metagenomic analysis
<p>A series of files that are useful for profiling metagenomic communities with the program inStrain.</p>
COMBAT TB Tuberculosis genome annotation database
<p>This is a Neo4j format *M. tuberculosis* reference annotation database. To view it you need a Neo4j (v 2.3) instance. There is a script (<strong>run_db.sh</strong>) that will start a Neo4j instance for you based on this data, using docker. Run that (<strong>bash run_db.sh</strong>) and connect to http://localhost:7474.</p> <p>The database was created by the COMBAT TB project (http://christoffels.sanbi.ac.za/index.php/projects/combat-tb) at the South African National Bioinformatics Institute (SANBI).<br /> <br /> Authors: Thoba Lose, Peter van Heusden, Ziphozakhe Mashologu, Alan Christoffels .</p> <p>The COMBAT TB project is funded by the South African Medical Research Council (MRC) and was supported by the South African<br /> Research Chairs Initiative of the Department of Science and Technology and National Research Foundation of South Africa.</p>
Genome Database: Turnover of strain-level diversity modulates functional traits in the honeybee gut microbiome between nurses and foragers
<p>This repository contains the dataset used in the publication "Turnover of strain-level diversity modulates functional traits in the honeybee gut microbiome between nurses and foragers," which is currently under revision. A pre-print can be found <a href="https://doi.org/10.1101/2022.12.29.522137">here</a>. The database is based on previously published work to create a genomic database of honeybee gut microbes by Kirsten Ellegaard (2021), found <a href="https://zenodo.org/records/4661061">here.</a></p><p>The zipped folder deposited here after unzipping, should contain the following files and directories:</p><ul><li>honeybee_genome.fasta : fasta file containing the host (<i>Apis mellifera</i>) genome sequence</li><li>beebiome_db : fasta file of 198 concatenated genomes with one genome per entry (multi-line fasta) where the headers represent the genome identifier</li><li>beebiome_red_db : fasta file of 39 species representative genomes with one genome per entry (multi-line fasta) where the headers represent the genome identifier to be used for the analysis of intra-specific variation</li><li>fna_files : directory containing genome sequence files and concatenated files where the concatenated files contain one fasta entry renamed to the genome identifier and all contigs concatenated into one entry</li><li>ffn_files : directory containing one file per genome listing the nucleotide sequence of all the predicted genes</li><li>faa_files : directory containing one file per genome listing the amino acid sequence of all the predicted genes</li><li>bed_files : directory containing bed files where the location of each of the predicted genes are indicated based on their position in the concatenated genome file</li><li>single_ortho : directory containing one file per phylotype listing all the single-copy orthogroups (OGs) identified by orthofinder where each line represents an OG id followed by a list of genes from each of the genomes of that phylotype that belong to that OG and the corresponding sequences of these genes can be found in the ffn file belonging to the respective genome</li><li>red_bed_files : directory containing bed files for species representative genomes that only list the positions genes that belong to the core orthogroups of their phylotype</li></ul><p>Further information about how this genome database was used to analyze strain-level diversity can be found in the publication and accompanying code repository.</p>
Database of giant viruses, Mirusviruses genomes, and marker genes
<p>Nucleotide sequences of 1,629 viral genomes (1,518 <em>Nucleoviricota</em> and 111 <em>Mirusviricota</em>), and marker genes found in these genomes.</p>
Advancing source tracking: systematic review and source-specific genome database curation of fecally shed prokaryotes
<p>This repository contains several files describing the data discussed in Lindner et al's "Advancing source tracking: systematic review and source-specific genome database curation of fecally shed prokaryotes". </p> <ol> <li>"database.fna" = concatenation of the (draft or complete) genome sequences described in the paper (n=12,730 source associated prokaryotic genomes) which passed quality checks and are species-level representatives (i.e., dereplicated at 95% ANI).</li> <li>"gdef.txt" = a manifest describing which sequences belong to which genomes.</li> <li>"sources.txt" = a manifest describing which genomes belong to which sources.</li> </ol> <p>The sources this database covers:</p> <table> <tbody> <tr> <td>Source Category</td> <td>Species-level <br>genome count</td> <td>Source-specific<br>species-level <br>genome count</td> </tr> <tr> <td>Bird</td> <td>56</td> <td>40</td> </tr> <tr> <td>Cat</td> <td>86</td> <td>6</td> </tr> <tr> <td>Chicken</td> <td>1314</td> <td>887</td> </tr> <tr> <td>Cow</td> <td>39</td> <td>15</td> </tr> <tr> <td>Dog</td> <td>139</td> <td>56</td> </tr> <tr> <td>Pig</td> <td>2764</td> <td>2035</td> </tr> <tr> <td>Ruminant</td> <td>740</td> <td>714</td> </tr> <tr> <td>Human</td> <td>4484</td> <td>3350</td> </tr> <tr> <td>Wastewater</td> <td>3108</td> <td>3097</td> </tr> </tbody> </table> <p> </p> <p>See publication for further details. </p>
Transitioning from environmental genetics to genomics using mitogenome reference databases
<p><span>Species detection using eDNA is revolutionizing the global capacity to monitor biodiversity. However, the lack of regional, vouchered, genomic sequence information—especially sequence information that includes intraspecific variation—creates a bottleneck for management agencies wanting to harness the complete power of eDNA to monitor taxa and implement eDNA analyses. eDNA studies depend upon regional databases of complete mitogenomic sequence information to evaluate the effectiveness of such data to differentiate, identify and detect taxa. We created the Oregon Biodiversity Genome Project working group to utilize recent advances in sequencing technology to create a database of complete, near error-free mitogenomic sequences for all of Oregon's resident freshwater fishes. So far, we have successfully assembled the complete mitogenomes of 313 specimens of freshwater fish representing 7 families, 55 genera, and 129 (88%) of the 146 resident species and lineages. Our comparative analyses of these sequences illustrate that the short (~150 bp) mitochondrial "barcode" regions typically used for eDNA assays are not consistently diagnostic for species-level identification and that no single region is best for metabarcoding Oregon's fishes. However, often-overlooked intergenic regions of the mitogenome such as the D-loop have the potential to reliably diagnose and differentiate species. This project provides a blueprint for other researchers to follow as they build regional databases. It also illustrates the taxonomic value and limits of complete mitogenomic sequences, and how current eDNA assays and the "PCR-free" environmental genomics methods of the future can best leverage this information.</span></p>
EvoMining genomic and enzyme databases for Actinobacteria, Cyanobacteria, Pseudomonas and Archaea
<p>Databases for EvoMining 2.0</p> <p>Genomic DB is a collection of genomes of a certain taxonomical group, functionally annotated by RAST.</p> <p>Enzyme-DB</p> <p>Actinobacteria</p> <p>Cyanobacteria</p> <p>Pseudomonas</p> <p>Archaea</p> <p>SampleData</p>
Lite Kraken/Bracken databases built using UHGG genomes
<p>Kraken/Braken databases for UHGG genomes.</p> <p>HUMAN_3006.tar.gz: Kraken/Bracken database for 3006 high quality species clusters of the UHGG (Beresford-Jones et al., 2022). Database was built from the single highest quality genome for each species cluster (n=3006). Uses the original GTDB v1.3 taxonomy.</p> <p>UHGG_5987_KRAKEN.tar.gz: Kraken/Bracken database for 3006 high quality species clusters of the UHGG (Beresford-Jones et al., 2022). Species clusters are represented by a variable number of high quality genomes (n=5987 in total), selected to maximise represented taxonomic diversity. Uses a custom taxonomy modified from GTDB v2.1 with each species cluster being represented by a species level taxonomic annotation. </p> <p> </p> <p>Methods:</p> <p>Databases built using Kraken v2.1.2 and Bracken v2.6.2. Commands used to build the databases are included below.</p> <p>kraken2-build --build --db Kraken --threads 12</p> <p>bracken-build -d Kraken -k 35 -l 150 -t 12</p>
Genomically predicted theoretical protein mass database for mass spectrometry (GPMsDB) evaluation datasets
<p>These are datasets obtained for the evaluation of GPMsDB (genomically predicted protein mass database) and its toolkits (GPMsDB-tk/GPMsDB-dbtk). The following datasets are deposited.</p> <ul> <li>The genome sequences of the strains newly sequenced and added using GPMsDB-dbtk (genomes_added.zip)</li> <li>MALDI-TOF-MS peak lists obtained from reference bacterial and archaeal strains (MALDI_peaklists.zip)</li> <li>16S rRNA gene sequences of the faecal isolates (mice_isolates_16S_nanopore.zip)</li> <li>Metagenome-assembled genomes from mouse faeces (mice_MAGs.zip)</li> </ul>
Plasmer database for k-mer and genomic features
<p>This is the inital version v1.0 of Plasmer database for k-mer and genomic features.</p> <p>Download and extract the package, and provide the absolute path to the Plasmer command line.</p> <p> </p> <p>For more information about Plasmer, please refer to our GitHub repository at: <a href="https://github.com/nekokoe/plasmer">https://github.com/nekokoe/plasmer</a></p>
Exposing New Taxonomic Variation with Inflammation – A Model-Specific Genome Database for Microbiome Researchers
<p>Data deposit for CBAJ-DB v1.2</p>
Transitioning from environmental genetics to genomics using mitogenome reference databases
Open the record for dataset details and reuse information.
microbetag : building a thorough database of genome-scale KO annotations
<p>In this repository we keep internal data for the <em><a href="https://hariszaf.github.io/microbetag/">microbetag</a> </em>microbial co-occurrence network annotator.</p> <p><em>microbetag</em> makes use of 2-column files for each genome, indicating the KO term found and a KEGG module in which this terms takes part into. <br>As a single KO term might participates in more than one KEGG modules, the same KO might be more than once in an annotation file. </p> <table> <tbody> <tr> <td> <div>chem_xref.tar.gz</div> </td> <td> <p>The MNXref namespace</p> <ol> <li>The identifier of a chemical compound in an external resource [<a href="https://www.metanetx.org/mnxdoc/mnxref.html#XREF">XREF</a>]</li> <li>The corresponding identifier in the MNXref namespace [<a href="https://www.metanetx.org/mnxdoc/mnxref.html#MNX_ID">MNX_ID</a>]</li> <li>The description given by the external resource [<a href="https://www.metanetx.org/mnxdoc/mnxref.html#STRING">STRING</a>]</li> </ol> <p>MNXref 4.0 release notes: - The third column (evidence tag for the mapping) was suppressed - The descriptions were completed - Deprecated identifiers were moved into they own table below</p> </td> </tr> <tr> <td> <div>gtdb_modelseed_gems.zip</div> </td> <td> <p>for all the GTDB genomes their corresponding <a href="https://patricbrc.org/">PATRIC</a> annotations were gathered. Then, using <a href="https://github.com/ModelSEED/ModelSEEDpy">modelseedpy</a> we constructed their genome scale metabolic reconstructions</p> </td> </tr> <tr> <td> <div>gtdb_kofam_scan_per_module.tar.gz</div> </td> <td> <p>all representative genomes of <a href="https://gtdb.ecogenomic.org/">GTDB</a> (v.202) were parsed and their corresponding `.faa` files were retrieved from the <a href="https://ftp.ncbi.nlm.nih.gov/genomes/all/">NCBI FTP</a>. Then the <a href="https://github.com/takaram/kofam_scan">kofam_scan</a> tool was used to annotate them and finally a <a href="https://github.com/hariszaf/microbetag/blob/clean/mappings/gtdb_mappings/gtdb_annotations_per_module.py">manual script </a>was used to keep KOs of each genome per module. </p> </td> </tr> <tr> <td>SeedSet.pkl.gz</td> <td> <p>A pickle file with the seeds of each GEM included in the <em>gtdb_modelseed_gems.zip </em>file and related to the KEGG MODULES based on the <em>seedId_keggId_module.tsv </em>file you can find on microbetag's GitHub page. Example:</p> <p>PATRIC SeedSet<br>373.172 [cpd00891, cpd00136, cpd00199, cpd01772, cpd00...<br>397278.5 [cpd00891, cpd00136, cpd01772, cpd02698, cpd08...</p> </td> </tr> <tr> <td>NonSeedSet.pkl.gz</td> <td> <p>A pickle file with the non seeds of each GEM included in the <em>gtdb_modelseed_gems.zip </em>file and related to the KEGG MODULES based on the <em>seedId_keggId_module.tsv </em>file you can find on microbetag's GitHub page. Example:</p> <p>PATRIC NonSeedSet<br>64187.548 [cpd00508, cpd00869, cpd00774, cpd03830, cpd00...<br>74426.1719 [cpd00204, cpd00447, cpd20171, cpd03470, cpd00...</p> </td> </tr> <tr> <td>seeds_per_genome.pkl.gz</td> <td> <p>A pickle file with a binary representation of the seeds per genome . Example:</p> <p> cpd00493 cpd00296 cpd11431 cpd00063 cpd15717 ... <br>2162051.4 1 0 0 1 0 0 0 0 0 0 ... </p> </td> </tr> <tr> <td> <div>nonseeds_per_genome.pkl.gz</div> </td> <td>Like above for non-seeds.</td> </tr> <tr> <td> <div>phen_classes.zip</div> </td> <td>A list of pickle files with the re-trained classes of phenDB for the prediction of functional traits on a genome.</td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p> <p> </p>
SCoV2-VAR: A light-weighted, customizable, and open-source database of 12 million SARS-CoV-2 genomes
<pre>Explosive accumulation of SARS-CoV-2 variants is posing a challenge to monitoring virus mutation and other data dealing, particularly based on centralized databases. The present study aimed to establish a light-weighted, customizable, and open-source database for SARS-CoV-2 genomes and annotations, without any access limit. The database, named SCoV2-VAR, was constructed, based on the variations (VAR) of the full-length SARS-CoV-2 (SCoV2) data uploaded on websites. All sequence samples were subject to quality control, single nucleotide polymorphism (SNP) annotation, format conversion, and final compression before appending to SCoV2-VAR. The final version of SCoV2-VAR (up to Feb 2024) contained more than 12 million SARS-CoV-2 records, with full genome and annotations. SCoV2-VAR was extremely light-weighted, with a storage size of 937 Mb for all 12 million sequences, post a 1: 596 compression. SCoV2-VAR is capable of timely updating, quickly querying, and customizable outputting SARS-CoV-2 sequences and their annotations. Additionally, the present study provided an overview of all 12 million SARS-CoV-2 samples, for both sequences and annotations.<br> <br><br></pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.