Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
345
datasets available to search
ShareScore release 0.7.1
Dataset results
345 results for “genome annotation”
Training data for 'Genome annotation with Apollo' tutorial (Galaxy Training Material)
<p>Published scaffolds from the Apis mellifera assembly Amel_4.5 and Official Gene Set 3.2.</p> <p>Source: <a href="http://hymenopteragenome.org/beebase/?q=download_sequences">http://hymenopteragenome.org/beebase/?q=download_sequences</a></p>
Peromyscus genome annotation
<p>Genome annotations and predicted proteins and transcripts for Peromyscus attwateri, nudipes, aztecus, melanophrys. </p>
The Caedibacter taeniospiralis genome sequence and annotation
<p>Interest in host-symbiont interactions is continuously increasing, not only due to the relevance and<br> prevalence of microbiomes. Started with the detection and description of novel symbionts, attention<br> moves to the molecular consequences and innovations of symbioses. However, molecular analysis requires<br> genome data which is dicult to obtain from obligate intracellular and uncultivated bacteria.We describe<br> here the Caedibacter taeniospiralis genome and transcriptome, identified by DNA and RNA Dual-Seq<br> of infected paramecia.</p> <p>Data inlcudes a genome FASTA file, annotations of genes and operons in gff format, and the compelte annotation in Geneious format.</p>
Genome annotations of Solanaceae species
<p>This archive contains genome annotations of <em>Solanaceae</em> species (i.e., <em>S. lycopersicum</em>, <em>S. pennellii</em> and <em>S. tuberosum</em>). The annotation files include gene modes and genetic markers from the <a href="https://solgenomics.net/">Sol Genomics Network</a> (SGN) resource that were converted to semantically interoperable format using the <a href="http://10.5281/zenodo.1076437">SIGA.py</a> command-line tool. The data are (re)distributed in:</p> <ul> <li><a href="https://github.com/The-Sequence-Ontology/Specifications/blob/master/gff3.md">Generic Feature Format</a> files (.gff)</li> <li><a href="https://sqlite.org/">SQLite</a> database files (.db)</li> <li><a href="https://www.w3.org/TR/turtle/">RDF/</a><a href="https://www.w3.org/TR/turtle/">Turle</a> files (gzip-ed .ttl)</li> </ul>
Improved genome assembly and annotation of the soybean aphid (Aphis glycines Matsumura)
<p>Updated genome assembly and annotation of <em>Aphis glycines</em> biotype 4.</p> <p><strong>Overview of files included in this release:</strong></p> <p><strong>Frozen release:</strong></p> <p>Updated <em>A. glycines </em>biotype 4 genome assembly: Aphis_glycines_4.v2.1.scaffolds.fa.gz </p> <p>BRAKER2 gene models for updated <em>A. glycines </em>biotype 4 genome assembly: Aphis_glycines_4.v2.1.scaffolds.fa.gff</p> <p>BRAKER2 protein sequences: Aphis_glycines_4.v2.1.scaffolds.fa.gff.aa.fa</p> <p>BRAKER2 nucleotide coding sequences: Aphis_glycines_4.v2.1.scaffolds.fa.gff.CDS.fa</p> <p><strong>Unfiltered raw intermediate genome assemblies:</strong></p> <p>Canu assembly of biotype 4 PacBio data from Wenger et. al. (2017): canu.fa.gz</p> <p>DBG2OLC hybrid assembly of selected biotype 4 MiSeq data and biotype 4 PacBio data from Wenger et. al. (2017): DBG2OLC.fa.gz</p> <p>Merged Canu and DBG2OLC assembly created with quickmerge: quickmerge.fa.gz</p> <p>Pilon polished (2 rounds) quickmerge assembly: quickmerge.pilon_r2.fa.gz</p> <p><strong>Mitochondrial and endosymbiont contigs extracted from the pilon polished quickmerge assembly: </strong></p> <p><em>A. glycines </em>biotype 4 mitochondrial genome: Aphis_glycines_4_Buchnera_v1.fa</p> <p><em>A. glycines </em>biotype 4 <em>Buchnera aphidicola</em> contigs: Aphis_glycines_4_Buchnera_v1.fa</p> <p><em>A. glycines </em>biotype 4 <em>Wolbachia</em> contigs: Aphis_glycines_4_Buchnera_v1.fa</p> <p><strong>Other files:</strong></p> <p>MUSCLE alignment of <em>A. glycines </em>v1, <em>A. glycines </em>biotype 4 v2.1 and <em>Drosophila melanogaster</em> R6.22 Osiris proteins in fasta format: D_mel_v1_v2_osiris.prots.muscle.fasta</p> <p>FastTree Maximum Likelihood phylogeny based on the MUSCLE alignment of Osiris genes in newick format: D_mel_v1_v2_osiris.prots.muscle.FastTree.nwk</p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
HIV-1 genome and annotation
<p>HIV-1 genome and annotation datasets for use with Galaxy Training Materials (<a href="https://training.galaxyproject.org/">https://training.galaxyproject.org</a>). This repository contains three datasets:</p> <ol> <li>hbx2.fa - genomic sequence of HIV-1 derived from GenBank entry K03455.1 </li> <li>hxb2.bed - coordinates of genomic features and drug resistance mutations </li> <li>hxb2.dr.bed - a subset of the annotation data containing drug resistance mutations only.</li> </ol> <p>Coordinates of drug resistance mutations are derived from Los Alamos National Lab HIV <a href="https://www.hiv.lanl.gov/content/sequence/HIV/MAP/hxb2.xls">database data</a>.</p>
Genome assembly and annotation of an apple variety 'RubyMac'
<p>In this dataset, we provided the contig-level genome assembly and annotation of an apple tree called 'RubyMac', which is growing in Michigan, USA (43°04'53.1"N 85°43'13.5"W). In this tree, the upper branches carried a sport mutation as compared with lower branches.</p>
Genome Annotation Xff Temecula1 (NCBI accession GCF_000007245.1) with Bakta
<p>The genome of <em>Xff </em>strain Temecula1 (NCBI accession GCF_000007245.1) was re-annotated using the Bakta pipeline (Schwengers et al., 2021).</p>
A human gut Faecalibacterium prausnitzii fatty acid amide hydrolase- Genome Annotation
<p>Genome annotation for <em>F. prausnitzii </em>Bg7063</p> <p><em>Science </em><strong>386</strong>, eado6828 (2024)</p> <p>DOI: 10.1126/science.ado6828</p> <p>Undernutrition in Bangladeshi children is associated with disruption of postnatal gut microbiota assembly; compared with standard therapy, a microbiota-directed complementary food (MDCF) substantially improved their ponderal and linear growth. Here, we characterize a fatty acid amide hydrolase (FAAH) from a growth-associated intestinal strain of Faecalibacterium prausnitzii cultured from these children. This enzyme, expressed and purified from Escherichia coli, hydrolyzes a variety of N-acylamides, including oleoylethanolamide (OEA), neurotransmitters, and quorum sensing N-acyl homoserine lactones; it also synthesizes a range of N-acylamides, notably N-acyl amino acids. Treating germ-free mice with N-oleoylarginine and N-oleolyhistidine, major products of FAAH OEA metabolism, markedly affected expression of intestinal immune function pathways. Administering MDCF to Bangladeshi children considerably reduced fecal OEA, a satiety factor whose levels were negatively correlated with abundance and expression of their F. prausnitzii FAAH. This enzyme, structurally and catalytically distinct from mammalian FAAH, is positioned to regulate levels of a variety of bioactive molecules.</p> <p> </p>
Supplementary Data for AGouTI - flexible Annotation of Genomic and Transcriptomic Intervals
<p>The data allows to replicate the use-case scenario described in the manuscript "AGouTI - flexible Annotation of Genomic and Transcriptomic Intervals" by Jan G. Kosiński and Marek Żywicki.</p>
Lathyrus sativus LS007 genome assembly and annotation Rbp1.0
<p>Genome assembly of grass pea (<em>Lathyrus sativus</em> L.) genotype LS007, assembled from PromethION nanopore data and polished using Illumina HiSeq PE data. The assembly was annotated using the mikado-minos pipeline developed by the Earlham Institute. Also included is a separate annotation track for repeat sequences produced using the DANTE pipleline. </p> <p> </p> <p>For any questions regarding this dataset, contact peter.emmrich@jic.ac.uk</p> <p> </p> <p>Note: ctg14433 has been manually corrected based on sequenced amplicon data. Files have been updated accordingly.</p> <p> </p> <p><strong>Assembly files:</strong></p> <p>Lsativus_LS007_Rbp1.0.7z - compressed complete assembly without scaffolding. The annotation refers to this assembly</p> <p>Rbp_9 largest HiC scaffolds.7z - compressed fasta file of the largest 9 scaffolds following HiC scaffolding</p> <p>Lsat_LS007_Rbp_chloroplast.fasta - fasta file of the complete LS007 chloroplast genome</p> <p>Lsat_LS007_Rbp_mitochondrion.fasta - fasta file of the complete LS007 mitochondrial genome</p> <p> </p> <p><strong>Annotation tracks:</strong></p> <p>LATSA3860_EIv1.0.annotation.gff3</p> <p>DANTE_transposable_element_protein_domains.gff3</p> <p>Full_length_LTR_retrotransposons.gff3</p> <p>Repeat_annotation_classI_classII_satellites.gff3</p> <p> </p> <p><strong>Annotation FASTA files:</strong></p> <p>LATSA3860_EIv1.0.annotation.gff3.cds.fasta</p> <p>LATSA3860_EIv1.0.annotation.gff3.cdna.fasta</p> <p>LATSA3860_EIv1.0.annotation.gff3.pep.fasta</p> <p> </p> <p><strong>Summaries and statistics:</strong></p> <p>LATSA3860_EIv1.0.annotation.gff3.final_table.tsv</p> <p>LATSA3860_EIv1.0.annotation.gff3.mikado_stats.txt</p> <p>LATSA3860_EIv1.0.annotation.gff3.biotype_conf.summary</p> <p>LATSA3860_EIv1.0.annotation.gff3.final_table.tsv</p> <p>LATSA3860_EIv1.0.annotation.gff3.pep.fasta.functional_annotation.tsv</p> <p>NOT_UPDATED_LATSA3860_EIv1.0.annotation.gff3.metrics.tsv *</p> <p>Blobtools_passed_contigs.txt - list of all contigs of the assembly that pass the BlobTools filter (Streptophyta, 20-100x coverage, >50 kbp) </p> <p> </p> <p>*this file has not been updated to reflect the correction to ctg14433</p>
Micractinium rhizosphaerae NFX-FRZ genome annotation, fasta file, amino acid
<p>Micractinium rhizosphaerae NFX-FRZ annotation. FASTA file, amino acid.</p>
Metagenome-Assembled Genome DRAM Annotations (EMERGE 97% dereplicated MAGs)
<p>This is the combined DRAM annotation outputs for the 1,864 97% dereplicated metagenome-assembled genomes from Stordalen Mire, Sweden. </p> <ul> <li>1864_97percentmags_annotations_combined.tsv.gz</li> <li>1864_97percentmags_metabolism_summary.xlsx</li> <li>product_0.html</li> <li>product_1.html</li> </ul> <p>METHODS:</p> <p>MAGs were annotated and distilled using DRAM (v1.4.0).</p> <p>FUNDING:<br> This research is a contribution of the EMERGE Biology Integration Institute ((https://emerge-bii.github.io/), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.<br> We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.<br> This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.<br> A portion of this research was performed under the Facilities Integrating Collaborations for User Science (FICUS) program (proposal: 10.46936/fics.proj.2017.49950/60006215 and 10.46936/10.25585/60001148) and used resources at the DOE Joint Genome Institute (<a href="https://www.google.com/url?q=https://ror.org/04xm1d337&sa=D&source=docs&ust=1674859614742521&usg=AOvVaw2XgXYw9eI4JIXRMKn3S9Se">https://ror.org/04xm1d337</a>) and the Environmental Molecular Sciences Laboratory (<a href="https://www.google.com/url?q=https://ror.org/04rc0xn13&sa=D&source=docs&ust=1674859614742655&usg=AOvVaw3UXdoHIFmVjc-mXUhDXYQt">https://ror.org/04rc0xn13</a>), which are DOE Office of Science User Facilities operated under Contract Nos. DE-AC02-05CH11231 (JGI) and DE-AC05-76RL01830 (EMSL).</p>
Inferring and comparing metabolisms across heterogeneous sets of annotated genomes using AuCoMe
<p>CONTENT OF THIS ARCHIVE</p> <p>The Zenodo archive is composed of one file and four main directories:<br> * <strong>analyses</strong> gathers three subdirectories: algae, bacteria, and fungi. It includes all files used to create the figures, supplemental figures, and results of the paper.</p> <p>* <strong>code</strong> contains all AuCoMe and PADMET codes.</p> <p> – <strong>aucome_v0.5.1</strong> this directory gathers the code of AuCoMe used to run the three datasets.</p> <p> – <strong>padmet_v5.0.1</strong> this directory contains the code of PADMET used to run AuCoMe.</p> <p>* <strong>datasets</strong> this directory gathers all datasets on which AuCoMe was run: the bacterial, fungal, and algal datasets, and the 32 synthetic datasets, which contain an <em>E. coli</em> K–12 MG1655 genome to which various degradations were applied, together with 28 other bacterial genomes. It also encompasses the version 23.5 of MetaCyc database.</p> <p>* <strong>scripts_analyses</strong> this directory contains several scripts to generate the figures, supplemental figures and a script to degrade the <em>E. coli</em> K–12 MG1655 genome.</p> <p> </p> <p> </p> <p>1/ Content of the <strong>analyses</strong> repertory<br> It is composed of three subdirectories: <strong>algae</strong>, <strong>bacteria</strong>, and <strong>fungi</strong>.</p> <p>1.1/ Content of the <strong>algae</strong> subdirectory<br> It encompasses 9 files.</p> <p>* <strong>Figure_2_algal_nb_reactions.tsv</strong> for each species of the algal dataset, this file gives the number of reactions at each AuCoMe step. It was used to create figure 2D.</p> <p>* <strong>Figure_S10_Deepec_algal.tsv</strong> for each species of the algal dataset, at each AuCoMe step (robust orthology, non-robust orthology, and annotation or orthology), several measures were computed, i.e.: the number of reactions, the number of ECs, the number of ECs validated by DeepEC, and the ratio number of ECs valided by DeepEC / number of ECs. It was used to design figure S10(b).</p> <p>* <strong>Table_S6_50_random_reactions_found.xlsx</strong> contains manual validation of 50 randomly chosen reactions found in any of the species (is the Supplemental Table S6).</p> <p>* <strong>Table_S7_50_random reactions absent.xlsx</strong> includes manual validation of 50 reactions absent from a species and randomly chosen (is the Supplemental Table S7).</p> <p>* <strong>Table_S8_reactions_common_only_Cokamuranus_Sjaponica.xlsx</strong> encompasses reactions common to <em>Saccharina japonica</em> and <em>Cladosiphon okamuranus</em> but not found in other brown algae (is the Supplemental Table S8).</p> <p>* <strong>Table_S9_homologues_Esiliculosus_Sjaponica.xlsx</strong> contains additional homologs in <em>E. siliculosus</em> found by BLASTP searches for sequences inferred to be present only in <em>C. okamuranus</em> and <em>S. japonica</em> (is the Supplemental Table S9).</p> <p>* <strong>Table_S10_o-aminophenol_Esiliculosus_holomogues.xlsx</strong> includes additional o-aminophenol oxidases from <em>E. siliculosus</em> and their homologs in other stramenopiles. It is the Supplemental Table S10 with more detail (like the amino acid sequences).</p> <p>* <strong>Table_S11_reactions_cryptophytes_haptophytes_stramenopiles_archeplastida.xlsx</strong> encompasses reactions distinguishing the cryptophyte, haptophyte, stramenopile, and archeplastida groups (is the Supplemental Table S11).</p> <p>* <strong>Table_S12_pathways_cryptophytes_haptophytes_stramenopiles_archeplastida.xlsx</strong> contains shared metabolic pathways as well as the absence of pathways between chryptophytes, haptophytes, stramenopiles, and archaeplastida (is the Supplemental Table S12).</p> <p> </p> <p>1.2/ Content of the <strong>bacteria</strong> subdirectory<br> It gathers 12 files and 9 repertories.</p> <p>* <strong>aucome_final.tsv</strong> output file of the figure S4 comparison bacteria.py script, for each of the 29 bacterial metabolic networks produced with AuCoMe, this table contains the number of ECs, the number of unique ECs, the number of total reactions, the number of enzymatic reactions with genes, the number of enzymatic reactions without genes, and the number of spontaneous reactions.</p> <p>* <strong>carveme_stat.tsv</strong> output file of the figure S4 comparison bacteria.py script, for each of the 29 bacterial metabolic networks produced with CarveMe, this table contains the number of ECs, the number of unique ECs, the number of total reactions, the number of enzymatic reactions with genes, the number of enzymatic reactions without genes, and the number of spontaneous reactions.</p> <p>* <strong>ecocyc.padmet</strong> contains the EcoCyc database version 23.5 at the PADMet, is used to generate the Supplemental Fig. S5.</p> <p>* <strong>Figure_2_bacterial_nb_reactions.tsv</strong> for each species of the bacterial dataset, this file gives the number of reactions at each AuCoMe step. It was used to create figure 2B.</p> <p>* <strong>Figure_3_nb_reactions_step.tsv</strong> for each dataset of the 32 synthetic bacterial datasets, this file enumerates the number of reactions at each AuCoMe step. It was used to create figure 3A.</p> <p>* <strong>Figure_3_fmeasure_steps.tsv</strong> for each dataset of the 32 synthetic bacterial datasets, this file indicates the values of the F-measures resulting of the comparison of the GSMNs recovered for each <em>E. coli</em> K–12 MG1655 genome replicate with the gold-standard network EcoCyc. It was used to create figure 3B.</p> <p>* <strong>Figure_S4_output</strong> contains 3 output files of the figure_S4_comparison_bacteria.py script:</p> <p> – <strong>Figure_S4_boxplot_networks.svg</strong> is the Supplemental Figure S4 in high resolution.</p> <p> – <strong>Figure_S4_boxplot_networks.tsv</strong> contains the number of reactions, the type of reactions (All, Reactions with genes, ...), and the used software. Thes data were produced and used in the figure_S4_comparison_bacteria.py script.</p> <p> – <strong>Figure_S4_barplot_time_networks.svg</strong> for each software, shows the required time in seconds used to reconstruct these bacterial metabolic networks.</p> <p>* <strong>Figure_S5_output</strong> encompasses 3 output files of figure S5 reference catalog.py script:</p> <p> – <strong>Figure_S5_ec_union.svg</strong> is the Supplemental Figure S5 in high resolution.</p> <p> – <strong>Figure_S5_ec_union_venn.svg</strong> another visualisation of presenting the results of the Supplemental Fig. S5.</p> <p> – <strong>Figure_S5_refence_ec_catalog_K12MG1655.tsv</strong> contains an EC catalog to <em>E. coli</em> K-12 MG1655 from the BIGG, EcoCyc, KEGG, and ModelSEED databases. This file is used to produce the Supplemental Figure S5.</p> <p>* <strong>Figure_S6_output</strong> includes 2 output files of the figure_S6.py script:</p> <p> – <strong>Figure_S6_comparison_all.svg</strong> is the Supplemental Figure S6 in high resolution.</p> <p> – <strong>Figure_S6_comparison_all.tsv</strong> contains data used to produce the Supplemental Figure S6.</p> <p>* <strong>gapseq_stat.tsv</strong> output file of the figure_S4_comparison_bacteria.py script, for each of the 29 bacterial metabolic networks produced with gapseq, this table contains the number of ECs, the number of unique ECs, the number of total reactions, the number of enzymatic reactions with genes, the number of enzymatic reactions without genes, and the number of spontaneous reactions.</p> <p>* <strong>jsons_bigg</strong> todate contains the five metabolic networks of <em>E. coli</em> K–12 MG1655 that can find in BIGG at JSON format. These files correspond to the BIGG reference metabolic network on the Supplemental Figure S5.</p> <p>* <strong>jsons_modelseed</strong> todate includes the metabolic network of <em>E. coli</em> K–12 MG1655 that can find in ModelSEED at JSON format. It is the ModelSEED reference metabolic network on the Supplemental Figure S5.</p> <p>* <strong>kegg_ecs.txt</strong> input file of the figure_S5_reference_catalog.py script, it contains matches between EC numbers and all the entries of <em>E. coli</em> K–12 MG1655 in the KEGG database.</p> <p>* <strong>mapping_modelseed_ec.tsv</strong> input file of the figure_S4_comparison_bacteria.py script, it encompasses matches between ModelSEED reactions and EC numbers.</p> <p>* <strong>modelseed_stat.tsv</strong> output file of the figure_S4_comparison_bacteria.py script, for each of the 29 bacterial metabolic networks produced with ModelSEED, this table contains the number of ECs, the number of unique ECs, the number of total reactions, the number of enzymatic reactions with genes, the number of enzymatic reactions without genes, and the number of spontaneous reactions.</p> <p>* <strong>networks_aucome</strong> for each of the 29 bacteria, contains a metabolic networks at the PADMet format obtained with AuCoMe.</p> <p>* <strong>networks_carveme</strong> for each of the 29 bacteria, contains a metabolic networks at the SBML format got to CarveMe.</p> <p>* <strong>networks_gapseq</strong> composes of 29 subdirectories (one for each bacterium). All these subdirectories contain 10 files about the metabolic networks a obtained with gapseq:</p> <p> – <strong>species-all-Pathways.tbl</strong> encompasses data on pathways at TBL format.</p> <p> – <strong>species-all-Reactions.tbl</strong> includes data on reactions at TBL format.</p> <p> – <strong>species-draft.RDS</strong> is a draft metabolic network at RDS (R Data Format).</p> <p> – <strong>species-draft.xml</strong> is a draft metabolic network at SBML format.</p> <p> – <strong>species-medium.csv</strong> encompasses all the metabolites allow the default medium.</p> <p> – <strong>species.RDS</strong> is the final metabolic network at RDS (R Data Format).</p> <p> – <strong>species-rxnWeights.RDS</strong> is a temporary file nedeed to gapseq fill at RDS (R Data Format).</p> <p> – <strong>species-rxnXgenes.RDS</strong> is a temporary file nedeed to gapseq fill at RDS (R Data Format).</p> <p> – <strong>species-Transporter.tbl</strong> includes data on transporters at TBL format.</p> <p> – <strong>species.xml</strong> is the final metabolic network at SBML format.</p> <p>* <strong>networks_modelseed</strong> includes two subdirectories:</p> <p> – <strong>sbml</strong> for each of the 29 bacteria, encompasses a metabolic networks at the SBML format got to ModelSEED.</p> <p> – <strong>tsv</strong> for each of the 29 bacteria, contains two TSV files:</p> <p> – <strong>genomeset__species.gbk_genome.fbamodel-compounds.tsv</strong> includes data on compounds at TSV format.</p> <p> – <strong>genomeset__species.gbk_genome.fbamodel-reactions.tsv</strong> encompasses data on reactions at TSV format.</p> <p>* <strong>time_carveme.txt</strong> input file of the figure S4 comparison bacteria.py script, for each of the 29 bacteria it stores the running time of CarveMe (in seconds) to reconstruct a metabolic network.</p> <p>* <strong>time_gapseq.txt</strong> input file of the figure S4 comparison bacteria.py script, for each of the 29 bacteria it stores the running time of gapseq (in seconds) to reconstruct a metabolic network.</p> <p> </p> <p>1.3/ Content of the <strong>fungi</strong> repertory<br> It contains three files and five directories.</p> <p>* <strong>All-pathways-of-S.-cerevisiae-S288c.txt</strong> encompasses all the YeastCyc pathways.</p> <p>* <strong>Figure_2_fungal_nb_reactions.tsv</strong> for each species of the fungal dataset, this file gives the number of reactions at each AuCoMe step. It was used to create figure 2C.</p> <p>* <strong>Figure_S7_output</strong> contains 11 output files of the figure S7 comparison pathway fungi.py script:</p> <p> – <strong>completion_pathway_species.svg</strong> for each of the 5 fungi (<em>L. bicolor</em>, <em>N. crassa</em>, <em>R. oryzae</em>, <em>S. cerevisiae</em> S288C, and <em>S. pombe</em>), contains a subfigure of the Supplemental Fig. S7.</p> <p> – <strong>fungi_stats.tsv</strong> is the Supplemental Table S5.</p> <p> – <strong>pathway_venn_species.png</strong> for each of the 5 fungi (<em>L. bicolor</em>, <em>N. crassa</em>, <em>R. oryzae</em>, <em>S. cerevisiae</em> S288C, and <em>S. pombe</em>), includes a Venn diagram about all the pathways found with the 3 software (AuCoMe, gapseq, and ModelSEED).</p> <p>* <strong>Figures_S8_S9_output</strong> contains 11 files, in all these files, a comparison of all pathways of metabolic networks of <em>S. cerevisiae</em> S288C obtained with AuCoMe and gapseq to those of YeastCyc was released.</p> <p> – <strong>comparison_yeastcyc.png</strong> is a picture about number of pathways true positive, false positive, and false negative are found, according the used method (AuCoMe and gapseq).</p> <p> – <strong>completion_pathway_gapseq.svg</strong> includes the number of pathways common or specific to YeastCyc and gapseq with their completeness ratio predicted by gapseq.</p> <p> – <strong>Figure_S8_completion_pathway_aucome.svg</strong> contanis the number of pathways common or specific to YeastCyc and AuCoMe with their completeness ratio predicted by AuCoMe, is the Supplemental Figure S8.</p> <p> – <strong>Figure_S9_venn_diagram_70_100.svg</strong> is the Supplemental Figure S9. All pathways of AuCoMe, gapseq and YeastCyc with a completion rate between 50% and 70% are compared.</p> <p> – <strong>venn_diagram.svg</strong> in this picture, all pathways are compared.</p> <p> – <strong>venn_diagram 50.svg</strong> all pathways of AuCoMe, gapseq and YeastCyc with a completion rate less than 50% are compared.</p> <p> – <strong>venn_diagram_50_gapseq.svg</strong> all pathways of gapseq whatever their completion rate are compared to the AuCoMe and YeastCyc pathways with a completion rate less than 50%.</p> <p> – <strong>venn_diagram_50_70.svg</strong> all pathways of AuCoMe, gapseq, and YeastCyc with a completion rate between 50% and 70% are compared.</p> <p> – <strong>venn_diagram_50_70_gapseq.svg</strong> all pathways of gapseq whatever their completion rate are compared to the AuCoMe and YeastCyc pathways with a completion rate between 50% and 70%.</p> <p> – <strong>venn diagram_70_100_gapseq.svg</strong> all pathways of gapseq whatever their completion rate are compared to the AuCoMe and YeastCyc pathways with a completion rate between 70% and 100%.</p> <p> – <strong>yeast_cyc_comparison.tsv</strong> contains the number of pathways true positive, false positive, and false negative are found, according the used method (AuCoMe and gapseq).</p> <p>* <strong>Figure_S10_Deepec_fungal.tsv</strong> for each species of the fungal dataset, at each AuCoMe step (robust orthology, non-robust orthology, and annotation or orthology), several measures were computed, i.e.: the number of reactions, the number of ECs, the number of ECs valided by DeepEC, and ratio number of ECs validated by DeepEC / number of ECs. It was used to design figure S10(a).</p> <p>* <strong>networks_aucome</strong> for each of the 5 fungi (<em>L. bicolor</em>, <em>N. crassa</em>, <em>R. oryzae</em>, <em>S. cerevisiae</em> S288C, and <em>S. pombe</em>), contains a metabolic networks at the PADMet format obtained with AuCoMe.</p> <p>* <strong>networks_gapseq</strong> is composed of 5 subdirectories (one for each fungus). All these subdirectories contain two files about the metabolic networks a obtained with gapseq:</p> <p> – <strong>species-all-Pathways.tbl</strong> encompasses data on pathways at TBL format.</p> <p> – <strong>species-all-Reactions.tbl</strong> includes data on reactions at TBL format.</p> <p>* <strong>networks_modelseed</strong> for each of the 5 fungi (<em>L. bicolor</em>, <em>N. crassa</em>, <em>R. oryzae</em>, <em>S. cerevisiae</em> S288C, and <em>S. pombe</em>), contains two TSV files:</p> <p> – <strong>species.gbk_genome.draftModel-compounds.tsv</strong> includes data on compounds at TSV format.</p> <p> – <strong>species.gbk_genome.draftModel-reactions.tsv</strong> encompasses data on reactions at TSV format.</p> <p> </p> <p> </p> <p>2/ Content of the <strong>code</strong> repertory<br> It gathers two directories <strong>aucome v0.5.1</strong> and <strong>padmet_v5.0.1</strong>.</p> <p>2.1/ Content of the <strong>aucome v0.5.1</strong> subdirectory<br> This directory contains a copy of the AuCoMe project on the GitHub site: <a href="https://github.com/AuReMe/aucome">https://github.com/AuReMe/aucome</a> (downloaded the 15/11/2022). It is composed of two subdirectories and five files:<br> * <strong>LICENCE</strong> licence of the AuCoMe software.</p> <p>* <strong>README.rst</strong> README of the AuCoMe software.</p> <p>* <strong>requirements.txt</strong> contains the list of requires Python packages.</p> <p>* <strong>setup.cfg</strong> contains metadata about AuCoMe package and is used with setup.py to distribute AuCoMe.</p> <p>* <strong>setup.py</strong> contains various information relevant to the AuCoMe package including options and metadata. Then, it is used to distribute AuCoMe with PyPI. It is also used to create an entrypoint when installing it with pip.</p> <p>* <strong>recipes</strong> this subdirectory contains two files:<br> – <strong>Dockerfile</strong> contains instructions to run AuCoMe in a Docker environment.<br> <br> – <strong>Singularity</strong> contains instructions to run AuCoMe in a Singularity container.</p> <p>* <strong>aucome</strong> this directory contains 11 Python files:<br> – <strong>__init__.py</strong> indicates the directory as a python module.</p> <p> – <strong>__main__.py</strong> contains the functions implementing the command-line interface of AuCoMe.</p> <p> – <strong>analysis.py</strong> contains the functions to analyse the AuCoMe results.</p> <p> – <strong>check.py</strong> contains the functions to check the input files.</p> <p> – <strong>compare.py</strong> contains the functions to compare the AuCoMe results between two distinct subgroups.</p> <p> – <strong>orthology.py</strong> contains the functions to propagate reaction through orthology.</p> <p> – <strong>reconstruction.py</strong> contains the functions to perform the reconstruction of draft GSMNs by using Pathway Tools in a parallel implementation.<br> <br> – <strong>spontaneous.py</strong> contains the functions to add spontaneous reactions to some GSMNs if it completes MetaCyc metabolic pathway.</p> <p> – <strong>structural.py</strong> contains the functions to check that no reactions are missing due to missing gene structures. A genomic search is performed for all reactions present in one organism but not in another.<br> <br> – <strong>utils.py</strong> contains a function to analyse the configuration file.</p> <p> – <strong>workflow.py</strong> contains functions to run all the steps of AuCoMe.</p> <p> </p> <p>2.2/ Content of the <strong>padmet_v5.0.1</strong> subdirectory<br> This directory contains a copy of the PADMET project on the GitHub site: <a href="https://github.com/AuReMe/padmet/">https://github.com/AuReMe/padmet/</a> (downloaded the 15/11/2022). It is composed of two subdirectories and six files:<br> * <strong>CHANGELOG.md </strong>records of all notable changes made in the PADMET project.</p> <p>* <strong>docs</strong> this repertory contains all the documentation files of PADMET package in the RST format.</p> <p>* <strong>LICENCE</strong> licence of the PADMET package.</p> <p>* <strong>README.md</strong> manual of the PADMET package.</p> <p>* <strong>requirements.txt</strong> contains the list of requires Python packages.</p> <p>* <strong>setup.cfg</strong> contains metadata about PADMET package and is used with setup.py to distribute PADMET.</p> <p>* <strong>setup.py</strong> contains various information relevant to the PADMET package including options and metadata. Then, it is used to distribute PADMET with PyPI. It is also used to create an entrypoint when installing it with pip.<br> <br> * <strong>padmet</strong> this repertory grathers two files and two subdirectories:<br> – <strong>__init__.py</strong> indicates the version of PADMET.</p> <p> – <strong>__main__.py</strong> contains the functions implementing the command-line interface of PADMET.</p> <p> – <strong>classes</strong> contains 7 files.</p> <p> – <strong>utils</strong> contains 4 files and 3 subdirectories.</p> <p><br> 2.2.1/ Content of the <strong>class</strong> subdirectory<br> The class repertory contains 7 files.<br> * <strong>__init__.py </strong>indicates the directory as a python module.</p> <p>* <strong>instantiation.py</strong> contains a function to instantiate padmet object.</p> <p>* <strong>node.py</strong> contains a class defining a Node object which is representing an element in a metabolic network (e.g: compound, reaction).</p> <p>* <strong>padmetRef.py</strong> contains a class defining a PadmetRef object which is representing a database of metabolic network.</p> <p>* <strong>padmetSpec.py</strong> creates a PadmetSpec object which is representing the metabolic network of a species/organism based on a reference database PadmetRef.</p> <p>* <strong>policy.py</strong> contains a class defining a Policy object that is defining the types of Relations and Nodes of a network.</p> <p>* <strong>relation.py</strong> contains a class defining a Relation object which is representing a link between two elements (Node) in a metabolic network.</p> <p><br> 2.2.2/ Content of the <strong>utils</strong> subdirectory<br> The utils directory contains 4 files and 3 subdirectories.<br> * <strong>__init__.py</strong> indicates the directory as a python module.</p> <p>* <strong>gbr.py</strong> implements a lexical analysis to handle genes relationship associated with a reaction, either a complex (with and relation between genes) or isozyme (with or relation between genes).</p> <p>* <strong>sbmlPlugin.py</strong> contains functions to handle SBML element (ex: species or reaction), then it returns all the sections named notes in a dictionary.<br> <br> * <strong>utils.py</strong> contains a function that checks paths of file.</p> <p>* <strong>connection</strong> this subdirectory contains 22 files:<br> - <strong>__init__.py</strong> indicates the directory as a python module.</p> <p> – <strong>biggAPI_to_padmet.py</strong> allows to extract the BIGG database from the API to create a padmet. An Internet access is required.</p> <p> – <strong>check_orthology_input.py</strong> is written to check if the metabolic network and the proteome of the model organism use the same identifiers for genes (or at least more than a given cutoff), before running orthology based reconstruction.</p> <p> – <strong>enhanced_meneco_output.py</strong> extracts the results from Meneco gap-filling to add more information to the gap-filled reactions. Then it returns a PADMET file with more information for each reaction.</p> <p> – <strong>extract_orthofinder.py</strong> after running Orthofinder on n FASTA files, it reads the output file ’Orthogroups.tsv’ to identify the orthologous genes. It is used by AuCoMe to extract the orthologous genes.<br> <br> – <strong>extract_rxn_with_gene_assoc.py</strong> from a given SBML file, it creates a SBML with only the reactions associated to a gene.<br> <br> – <strong>gbk_to_faa.py</strong> extracts protein sequence from a GenBank into a FASTA file with Biopython package.</p> <p> – <strong>gene_to_targets.py</strong> from a list of genes, it gets the products associated with the reactions linked to the genes. For example: R1 is linked to G1, R1 produces M1 and M2, this script outputs: M1, M2.</p> <p> – <strong>get_metacyc_ontology.py</strong> from the PadmetRef of MetaCyc, it creates the MetaCyc ontology.</p> <p> – <strong>metexploreviz_export.py</strong> converts a PADMET object representing a metabolic network into a JSON compatible with MetExplore.<br> <br> – <strong>modelSeed_to_padmet.py</strong> from ModelSEED reactions and pathways files, it creates a PADMET.<br> <br> – <strong>network_to_gnn.py</strong> creates input for GNN (Graph Neural Networks) from PADMET or SBML.</p> <p> – <strong>padmet_to_asp.py</strong> converts PADMET to Answer Set Programming.</p> <p> – <strong>padmet_to_matrix.py</strong> creates a stoichiometry matrix from a PADMET file, in which the columns represent the reactions and rows represent metabolites.</p> <p> – <strong>padmet_to_padmet.py</strong> allows to merge 1-n PADMET.<br> <br> – <strong>padmet_to_tsv.py</strong> converts a PADMET representing a database (PadmetRef) and/or a PADMET representing a model (PadmetSpec) to TSV files.</p> <p> – <strong>pgdb_to_padmet.py</strong> reads a PGDB folder (from BIOCYC/Pathway Tools) and creates a PADMET. It is used by AuCoMe to create PADMET files from PGDB in the annotation-based step.</p> <p> – <strong>sbmlGenerator.py</strong> contains functions to generate SBML files from PADMET and TXT files usign the libsbml package. It is used by AuCoMe to create SBML files at the annotation-based, orthology and final steps.</p> <p> – <strong>sbml_to_curation_form.py</strong> extracts one or several reactions from a SBML file to the form used in AuReMe for curation.</p> <p> – <strong>sbml_to_padmet.py</strong> converts a SBML file into a PADMET file (with or without a reference database).</p> <p> – <strong>sbml_to_sbml.py</strong> creates a SBML file from another one. Use it to change the SBML level.</p> <p> – <strong>wikiGenerator.py</strong> contains all necessary functions to generate wiki pages from a PADMET file and update a wiki online. It requires WikiManager module (with wikiMate, Vendor).</p> <p>* <strong>exploration</strong> this subdirectory contains 15 files:<br> - <strong>__init__.py</strong> indicates the directory as a python module.</p> <p> – <strong>compare_padmet.py</strong> compares 1-n PADMET files, and creates a folder with 4 output files (compounds.tsv, genes.tsv, pathways.tsv and reactions.tsv). It is used by AuCoMe to create these files to analyse the metabolic networks.</p> <p> – <strong>compare_sbml.py</strong> compares 2 or 1-n SBML, then it creates two output files reactions.tsv and metabolites.tsv with the reactions/metabolites in each SBML files.</p> <p> – <strong>compare_sbml_padmet.py</strong> compares reaction identifiers in SBML versus PADMET, then returns the number of reactions in both, and reaction identifiers not in SBML or not in PADMET.</p> <p> – <strong>convert_sbml_db.py</strong> uses the MetaNetX database to check or convert a SBML. Flat files from MetaNetx are required to run this script. They can be found in the AuReMe workflow or from the MetaNetx website.</p> <p> – <strong>dendrogram_reactions_distance.py</strong> uses the reactions.tsv file from compare_padmet.py to create a dendrogram using the R package pvclust. It has been used in the article to create the metabolic dendrogram.</p> <p> – <strong>flux_analysis.py</strong> runs the flux balance analyse with cobra package on an already defined reaction. It needs to set in the SBML the value ’objective_coefficient’ to 1.</p> <p> – <strong>get_pwy_from_rxn.py</strong> from a file containing a list of reaction, it returns the pathways where these reactions are involved.</p> <p> – <strong>padmet_stats.py</strong> creates a PADMET stats file (named padlet_stats.tsv) containing the number of pathways, reactions, genes and compounds inside the one or several PADMET files.</p> <p> – <strong>pathway_production.py</strong> compares 1-n PADMET objects to show the pathway input/output for them.</p> <p> – <strong>prot2genome.py</strong> contains function to search a genome using protein sequences and Gene-Protein-Reaction associations. It is used in the structural search step of AuCoMe.</p> <p> – <strong>report_network.py</strong> creates reports of a PADMET file, and it writes three TSV files (all metabolites.tsv, all_pathways.tsv, and all_reactions.tsv).</p> <p> – <strong>visu_network.py</strong> allows to visualize a metabolic network on a compounds perspectives.</p> <p> – <strong>visu_path.py</strong> allows to visualize a pathway in PADMET network.<br> <br> – <strong>visu_similarity_gsmn.py</strong> visualize similarity between metabolic networks using MDS.</p> <p>* <strong>management</strong> this subdirectory contains 5 files:<br> – <strong>__init__.py</strong> indicates the directory as a python module.</p> <p> – <strong>manual_curation.py</strong> updates a PadmetSpec object by filling specific forms. It either creates new reaction(s) to PADMET file, or it adds/removes reaction(s) from a PadmetRef.</p> <p> – <strong>padmet_compart.py</strong> for a given PADMET file, it checks and updates compartment.</p> <p> – <strong>padmet_medium.py</strong> for a given set of compounds representing the growth medium (or seeds), it creates two reactions in order to maintain consistency of the network for flux analysis.</p> <p> – <strong>relation_curation.py</strong> for a given PADMET file, it adds or removes relations between nodes.</p> <p> </p> <p> </p> <p>3/ Content of the <strong>datasets</strong> subdirectory<br> It contains a files and four repertories.</p> <p>* <strong>metacyc 23.5.padmet</strong> the version 23.5 of the <a href="https://metacyc.org/">MetaCyc</a> database in the PADMET format. It was used by AuCoMe to reconstruct all the metabolic networks. Hence metacyc 23.5.padmet is required to reproduce the article results.</p> <p>3.1/ Content of the <strong>algal</strong>, <strong>bacterial</strong>, and <strong>fungal</strong> directories<br> These three directories are composed of 8 subdirectories and a Supplemental Table (respectively Table S1, Table S2 and Table S3 in bacterial, fungal and algal directories):<br> * <strong>FASTA</strong> contains the proteome of each species as a FASTA file.</p> <p>* <strong>cleaned_GBKs</strong> for each species, it contains the annotated genome, with the protein sequences in a GenBank format file.</p> <p>* <strong>dictionaries</strong> for some species, genes needed to be renamed for compatibility reasons. This folder contains CSV files with the mapping between the old names of genes and the new ones.</p> <p>* <strong>annotated_DATs</strong> contains a subdirectory per species with all the output files from Pathway Tools v23.5, without any post-treatment, in the DAT format.</p> <p>* <strong>annotated_PADMETs</strong> for each species, it contains a metabolic network of the draft reconstruction step of AuCoMe, in the PADMET format.<br> <br> * <strong>final_PADMETs</strong> for each species, it contains a metabolic network generated by the AuCoMe workflow, at the PADMET format.</p> <p>* <strong>final_SBMLs</strong> for each species, it contains a metabolic network generated by the AuCoMe workflow, in the SBML format.</p> <p>* <strong>panmetabolism</strong> is composed of 7 files describing the final metabolic networks:<br> – <strong>genes.tsv</strong> contains, for each organism, the list of genes and the associated reactions.</p> <p> – <strong>metabolites.tsv</strong> contains the list of metabolites present in the panmetabolism. Then, for each metabolite and for each organism, it lists the reactions that produced this compound and the reactions that consumed it.<br> <br> – <strong>pathways.tsv</strong> contains the list of pathways present in the panmetabolism. For each pathway and for each organism, it indicates the number of reactions present in this pathway, and the names of these reactions.<br> <br> – <strong>reactions.tsv</strong> contains the list of reactions present in the panmetabolism. Then for each reaction, it indicates whether or not it belongs to an organism. If a reaction is found in a species, the genes associated with the reaction are also listed.</p> <p> – <strong>pvclust_reaction_dendrogram.png</strong> based on the presence/absence matrix of reactions in different species of the dataset, it computes the Jaccard distances between these species, and it applies a hierarchical clustering on these data with a complete linkage to create a dendrogram. The R package pvclust is used to create the dendrogram, with bootstrap resampling. For each node, a p-value indicates how strong the cluster is supported by data. This dendrogram is provided as a PNG picture.</p> <p><br> 3.2/ Content of the <strong>synthetic_bacterial</strong> repertory<br> The synthetic_bacterial repertory contains the Supplemental Table S4 and 32 subdirectories named Run_00, Run_01, . . . , etc, Run 31. Each subdirectory is composed of 9 files:<br> * <strong>K_12_MG1655.gbk</strong> the annotated genome of <em>E. coli K–12 MG1655</em> to which degradation of the functional and/or structural annotations was applied.</p> <p>* <strong>annotated_K_12_MG1655.sbml</strong> the metabolic network of <em>E. coli K–12 MG1655</em> output of the draft reconstruction step of AuCoMe in the SBML format.</p> <p>* <strong>annotated_K_12_MG1655.padmet</strong> the metabolic network of <em>E. coli K–12 MG1655</em> output of the draft reconstruction step of AuCoMe in the PADMET format.</p> <p>* <strong>orthology_K_12_MG1655.sbml</strong> the metabolic network of <em>E. coli K–12 MG1655</em> output of the orthology propagation step of AuCoMe in the SBML format.</p> <p>* <strong>orthology_K_12_MG1655.padmet</strong> the metabolic network of <em>E. coli K–12 MG1655</em> output of the orthology propagation step of AuCoMe in the PADMET format.<br> <br> * <strong>structural_K_12_MG1655.sbml</strong> the metabolic network of <em>E. coli K–12 MG1655</em> output of the structural verification step of AuCoMe in the SBML format.<br> <br> * <strong>structural_K_12_MG1655.padmet</strong> the metabolic network of <em>E. coli K–12 MG1655</em> output of the structural verification step of AuCoMe in the PADMET format.</p> <p>* <strong>final_K_12_MG1655.sbml</strong> the metabolic network of <em>E. coli K–12 MG1655</em> output of the AuCoMe workflow in the SBML format.</p> <p>* <strong>final_K_12_MG1655.padmet</strong> the metabolic network of <em>E. coli K–12 MG1655</em> output of the AuCoMe worflow in the PADMET format.</p> <p> </p> <p> </p> <p>4/ Content of the <strong>scripts_analyses</strong> subdirectory<br> The scripts repertory contains 12 files:</p> <p>* <strong>bacteria_random_degradation.py</strong> was used to degrade the <em>E. coli</em> K–12 MG1655 genome. The procedure for the genome degradation is described in the algorithm 1.</p> <p>* <strong>figure_2_algal_dataset.py</strong> for each species of the algal dataset, and at each AuCoMe step. This script allows to generate the figure 2D.</p> <p>* <strong>figure_2_bacterial_dataset.py</strong> for each species of the bacterial dataset, and at each AuCoMe step. This script allows to generate the figure 2B.</p> <p>* <strong>figure_2_fungal_dataset.py</strong> for each species of the fungal dataset, and at each AuCoMe step. This script allows to generate the figure 2C.</p> <p>* <strong>figure_3_degradation.py</strong> allows to generate the figure 3B from the figure_3_fmeasure_steps.tsv file (described above).</p> <p>* <strong>figure_5_mds.py</strong> allows to generate the figure 5A from two reactions.tsv files of the algal dataset (annotation-based and final).</p> <p>* <strong>figure_S4_comparison_bacteria.py</strong> computes statistics on all the 29 bacterial metabolic networks reconstructed with AuCoMe, CarveMe, gapseq and ModelSEED, it uses the mapping_modelseed_ec.tsv, soft_stat.tsv files and bacteria/networks soft directories, then it creates the Supplemental Fig. S4 and files inside the analyses/bacteria/Figure_S4_output repertory.</p> <p>* <strong>figure_S5_reference_catalog.py</strong> reads the ecocyc.padmet, kegg ecs.txt files, and the jsons bigg/, jsons modelseed/ directories, then it creates the Figure_S5_refence_ec_catalog_K12MG1655.tsv and the Supplemental Fig. S5.</p> <p>* <strong>figure_S6.py</strong> reads the Figure_S5_refence_ec_catalog_K12MG1655.tsv file and those inside the analyses/bacteria/networks_soft directories about <em>E. coli</em> K–12 MG1655, it generates the Supplemental Fig. S6.</p> <p>* <strong>figure_S7_comparison_pathway_fungi.py</strong> reads metacyc_23.5.padmet file and those inside the analyses/fungi/networks_soft directories, then it create five completion pathway species.svg pictures which composed the Supplemental Fig. S7.</p> <p>* <strong>figures_S8_S9.py</strong> reads the metacyc_23.5.padmet, and metabolic networks of <em>S. cerevisiae</em> S288 reconstructed with AuCoMe, gapseq, and YeastCyc. For all pathways of both AuCoMe and gapseq networks, it also computes their completion rates. Then it compares the results obtained with AuCoMe and gapseq on <em>S. cerevisiae</em> S288C to YeastCyc according to the completion rates of their pathways. It generates Supplemental Fig. S8 and S9.</p> <p>* <strong>figure_S12_supervenn.py</strong> allows to generate the figure S12, it reads the reactions.tsv file of the algal dataset at the final AuCoMe step, and another tabular file that contains abbreviated names of species.</p>
Additional annotation, alignment, and results from Ka/Ks analysis for Chromosomal-level reference genome assembly of the African Spiny Mouse (Acomys cahirinus)
<p><strong>Annotation files, alignments, and results summaries from Chromosomal-level reference genome assembly of the African Spiny Mouse (Acomys cahirinus).</strong></p> <p>Pairwise genome alignments contain the .maf suffix</p> <p>FASTA alignments from stitched gene blocks contain the .fasta suffix</p> <p>CSV file containing the Ka/Ks results</p> <p>RepeatMasker .out file</p>
Annotation of genes encoding enzymes across marine phytoplankton genomes
<p>Phytoplankton cells span a large size range, from picoplankton (<2µm), nanoplankton (2 to 20µm), microplankton (20 to 200µm) to macroplankton (200 to <2000µm). Cell size interacts with multiple selective pressures, including cellular metabolic rate, light absorption, nutrient uptake, cell nutrient quotas, trophic interactions and diffusional exchanges with the environment. Beyond simple size, cells of different shapes differ in surface area to volume ratio. For example, more elongated cells, such as pennate diatoms, have a larger surface area to volume ratio compared to more rounded cells, such as centric diatoms, of equivalent biovolume, which can in turn influence diffusional exchanges between cells and their environment. We assembled metadata on diverse marine phytoplankters, in parallel with genomic or transcriptomic data annotations to identify genes encoding enzymes, to facilitate analyses of genomic patterns of encoded enzymes across diverse taxa, sizes, growth forms and origins of strains.</p>
Genome sequences and gene annotations for two Ophryocystis lineages
<p>Assembly, annotation, and gene sequences for the <em>Ophryocystis </em>lineages sequenced in "Genome sequence of <em>Ophryocystis elektroscirrha</em>, an apicomplexan parasite of monarch butterflies: cryptic diversity and response to host-sequestered plant chemicals." Each of the two lineages has three associated files: a genome sequence file (.fa), an annotation in .gff3 format, and gene sequences in .fna format. Sequences generated for <em>Ophryocystis elektroscirrha </em>come from direct DNA extraction and sequencing effort and are hosted elsewhere on NCBI as well. The other lineage, prefixed Ophryocystis-elektroscirrha_like, was bioinformatically extracted from the genome of an infected host. As such, we are less confident in its completeness and it is not archived elsewhere. </p>
Training data for 'Genome annotation with Funannotate' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for genome annotation with funannotate.</p> <p>Genome was assembled following the GTN Flye assembly tutorial, then masked with RepeatMasker.</p> <p>RNASeq data: SRR8534859 reads were mapped to the genome using STAR (toolshed.g2.bx.psu.edu/repos/iuc/rgrnastar/rna_star/2.7.8a+galaxy0), then the bam was downsampled (10% with toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_DownsampleSam/2.18.2.1) to reduce the size of the dataset. Fastq files were then extracted from the resulting bam file (toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_SamToFastq/2.18.2.1).</p> <p>SwissProt_subset.fasta is a subset of SwissProt proteins that are known to have some similarity with the genome (found using Diamond against the genome, then extracting sequences matching with e-value < 0.0001).</p>
Test Genome Annotation Dataset
<p>This folder contains a set of data which could be used to annotate a drosophila melanogaster genome including the genome, isoseq and rnaseq data. There is also the genbank annotation for QC purposes</p> <p>##Genome<br> GCF_000001215.4_Release_6_plus_ISO1_MT_genomic.fna.gz</p> <p>##Annotation<br> GCF_000001215.4_Release_6_plus_ISO1_MT_genomic.gff.gz</p> <p>##RNAseq<br> fastq files downloaded from SRA with the following accessions:<br> ###Testes<br> SRR21753563</p> <p>SRR21753562</p> <p>###Ovary<br> SRR22757316</p> <p>SRR22757307</p> <p>###Brain<br> SRR24426130</p> <p><br> ##Isoseq<br> fastq files downloaded from SRA with the following accessions:<br> ###Head<br> SRR19344254</p> <p>SRR19344255</p> <p>SRR19344256</p> <p>###Ovary<br> SRR12634521</p> <p>SRR12634520</p> <p>###Testes<br> SRR12634519</p> <p>SRR12634518</p> <p> </p>
Annotation Data: Decoding the chromosome-scale genome of the nutrient-rich Agaricus subrufescens: A Resource for fungal biology and biotechnology
<p><strong>Decoding the chromosome-scale genome of the nutrient-rich Agaricus subrufescens: A Resource for fungal biology and biotechnology</strong></p> <p>Genome annotation data</p> <p><strong>Genome Browser:</strong> <a href="https://plantgenomics.ncc.unesp.br/gen.php?id=Asub">https://plantgenomics.ncc.unesp.br/gen.php?id=Asub</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.