Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
133
datasets available to search
ShareScore release 0.9.0
Dataset results
133 results for “gene annotation”
Gene Annotations of 49 Bacillariophyta Genome Assemblies
<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div> </div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a></p> </div> <h2>Files</h2> <div>The following gff3-files with structural and functional genome annotation are included in the compressed archive Bacillariophyta_annotations.tar.gz:</div> <div> </div> <div>Asterionella_formosa.gff3<br>Asterionellopsis_glacialis.gff3<br>Bacterosira_constricta.gff3<br>Chaetoceros_muellerii.gff3<br>concatenated_output.gff3<br>Conticribra_guillardii.gff3<br>Conticribra_weissflogii.gff3<br>Craspedostauros_australis.gff3<br>Cyclostephanos_invisitatus.gff3<br>Cyclostephanos_tholiformis.gff3<br>Cyclotella_atomus.gff3<br>Cyclotella_baltica.gff3<br>Cyclotella_choctawhatcheeana.gff3<br>Cyclotella_cryptica.gff3<br>Cylindrotheca_fusiformis.gff3<br>Detonula_confervacea.gff3<br>Discostella_pseudostelligera.gff3<br>Discostella_stelligera.gff3<br>Discostella_stelligeroides.gff3<br>Epithemia_pelagica.gff3<br>Fistulifera_pelliculosa.gff3<br>Fistulifera_solaris.gff3<br>Fragilaria_radians.gff3<br>Fragilariopsis_cylindrus.gff3<br>Licmophora_abbreviata.gff3<br>Mediolabrus_comicus.gff3<br>Nitzschia_palea.gff3<br>Nitzschia_putrida.gff3<br>Porosira_glacialis.gff3<br>Psammoneis_japonica.gff3<br>Pseudo-nitzschia_multiseries.gff3<br>Pseudo-nitzschia_pungens.gff3<br>Skeletonema_costatum.gff3<br>Skeletonema_marinoi.gff3<br>Skeletonema_menzelii.gff3<br>Skeletonema_potamos.gff3<br>Skeletonema_tropicum.gff3<br>Stephanocyclus_meneghinianus.gff3<br>Stephanodiscus_minutulus.gff3<br>Stephanodiscus_triporus.gff3<br>Thalassiosira_allenii.gff3<br>Thalassiosira_delicatula.gff3<br>Thalassiosira_exigua.gff3<br>Thalassiosira_gravida.gff3<br>Thalassiosira_livingstoniorum.gff3<br>Thalassiosira_mediterranea.gff3<br>Thalassiosira_oceanica.gff3<br>Thalassiosira_ordinaria.gff3<br>Thalassiosira_pacifica.gff3<br>Thalassiosira_profunda.gff3</div> <div> </div> <div>To extract the dataset, execute the following command:</div> <div> </div> <div><code>tar -xvf Bacillariophyta_annotations.tar.gz</code></div> <h2>Genome Assemblies</h2> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <div> </div> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p> </p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p> </p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p> </p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^>/ s/ .*//' genome.fasta > genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <div> </div> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <div> </div> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>This release contains a gene set where a results of an OrthoFinder run that did not include genes on contigs that are suspected to be contaminants or horizontal gene transfer candidates were used to filter single exon genes. This means the gene and transcript counts changed compared to the previous release.</p> <h2>License</h2> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div>
A Curated Gene and Biological System Annotation of Adverse Outcome Pathways Related to Human Health
<p>Adverse Outcome Pathways (AOPs) are multi-scale models of biological mechanisms connecting molecular initiating events to adverse outcomes through measurable key events. AOPs can guide the use and development of new approach methodologies (NAMs) aimed at reducing animal experimentation in chemical safety assessment. Here, we present a comprehensive molecular annotation of AOPs relevant to human health to embed the AOP framework into molecular data interpretation, which supports the development and application of novel AOP-based approaches in biomedical research.</p> <p>Please cite the following publication alongside this Zenodo entry when using the data:</p> <p>Saarimäki, L.A., Fratello, M., Pavel, A. <em>et al.</em> A curated gene and biological system annotation of adverse outcome pathways related to human health. <em>Sci Data</em> <strong>10</strong>, 409 (2023). https://doi.org/10.1038/s41597-023-02321-w</p>
Multifaceted quality assessment of gene repertoire annotation with OMArk
<p>Dataset associated to the OMArk paper.</p><p>Contain eight archives:</p><p>Supplementary_Tables</p><p>The Supplementary Table files referred to in the paper</p><p>OMAmerDB:</p><p>The OMAmer database constructed using the whole dataset of the OMA database (November 2022 Release) and used in the paper. An OMAmer database is necessary to run OMArk.</p><p>Simulation:<br>Proteomes with artificially introduced errors, contaminants or depleted completeness, used to assess OMArk's performance. The archive contains the generated proteomes (Simulated_Data) and their OMArk quality assessments (omark). They also contains the OMAmer results (OMAmerResults) that were used to run OMArk and BUSCO completeness assessments (BUSCO).</p><p>*Note that for storage efficiency, only the non-redundant part of the data (added errors, added contamination, random fraction of proteomes) are stored there. The full modified proteome can be regenerated from these data and the source proteomes.</p><p>Reference Proteomes:</p><p>The UniProt Reference Proteomes (Proteomes) (2021_04) and their proteome quality assesment results according to OMArk. The archive contains the source proteome FASTA (Source folder), OMAmer results for these proteomes (omamer folder) , OMArk results (omark folder), and BUSCO completeness assesments (BUSCO folder). It also contains a subfolder that contains part of the Contamination detection experiment (Contamination folder).</p><p>Ensembl_Metazoa_AssemblyChange.<br><br>Contains Ensembl Metazoa proteomes with version change between version 52 and 54 as well as their quality assesment resuls for both version. The archive contains the source proteomes FASTA (Source folder), a Splice file that group together all proteins coded by the same gene (Splice folder), omamer results for the proteomes (omamer folder) and the omark results (omark folder)</p><p>MissingGenesBLAST<br><br>Contains sequences of HOGs considered as missing in the Human proteome, that was used to look for sequences in the human genome.</p><p>Ensembl_NCBI_Results</p><p>Contains OMArk and BUSCO results for Ensembl and NCBI proteomes. These results were then used to evaluate OMArk biais due to source of proteomes in the OMA database.</p><p>Notebooks<br>Jupyter Notebooks that were used to perform the analysis described in the paper<br><br> </p>
Metagenomes: gene function and family annotations
<p>Functional annotations of genes for all contigs in 1,782 metagenomes.</p> <p>Genes were annotated to three sources: (1) COGs, (2) Pfams, and (3) <em>de novo</em> families from reference sequences. These gene annotations are used to train and run PlasX.</p>
Gene annotation files for Fraxinus excelsior (European ash) genome assembly BATG-0.5
<p>Gene annotation files for <em>Fraxinus excelsior</em> genome assembly v. BATG-0.5, published in Nature (doi:10.1038/nature20786). These files were previously hosted on the Ash Tree Genomes website (http://www.ashgenome.org/transcriptomes) and first made available for download via that site on 2016-02-08.</p> <p>The following annotation files are available:</p> <p>### GFF file of all gene models (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3</p> <p>### FASTA file of all cDNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.pep.fa</p> <p>### Functional annotation for each gene model (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.functional_annotation.tsv</p> <p>### GFF file of all gene models (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3</p> <p>### FASTA file of all cDNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.pep.fa</p> <p>### GFF file for gene models identified as probable transposable element related sequences (excluded from the other files)<br> Fraxinus_excelsior_38873_TGAC_v2.transposable_elements.gff3</p> <p><br> NB: The annotation files include preliminary annotations for genes within the organellar scaffolds (gene models FRAEX38873_v2_000400370-FRAEX38873_v2_000401330), which were not reported in the publication of the BATG0.5 assembly (doi:10.1038/nature20786).</p>
The terrestrial carnivorous plant Utricularia reniformis sheds light on environmental and life-form genome plasticity: Annotation, Gene Ontology and raw data
<p><strong>Description:</strong> In this work, we deeply sequenced (genome and transcriptome of different organs), assembled, and analyzed the 311-Mbp genome of the terrestrial carnivorous plant <em>U. reniformis</em> (Lentibulariaceae). This project presents great importance to the understanding of genomic, evolutive and functional aspects of<em> U. reniformis</em>, which may, with the next-generation sequencing and computational biology approaches shed light to a better understanding not only for the biology and evolution of <em>Utricularia</em> genus, but also for other genera and lineages of the Lentibulariaceae family. Here we present all the raw data generated, including annotation and gene ontology files.</p> <p><strong>External Information</strong></p> <p><a href="https://genomevolution.org/coge/GenomeInfo.pl?gid=54799">Genome Browser</a> avaliable at CoGe Portal (https://genomevolution.org/coge/GenomeInfo.pl?gid=54799)</p> <p><a href="http://https://www.ncbi.nlm.nih.gov/bioproject/290588">GenBank </a><a href="http://https://www.ncbi.nlm.nih.gov/bioproject/290588">Bioproject</a> (https://www.ncbi.nlm.nih.gov/bioproject/290588) for raw genomic and transcriptomic reads</p> <p><a href="https://bv.fapesp.br/en/auxilios/84264/genomics-and-transcriptomics-of-utricularia-reniformis-lentibulariaceae-an-evolutive-and-function/">FAPESP grant website</a> contaning the project abstract and other information.</p> <p><strong>Papers published related to <em>Utricularia reniformis</em> genome</strong></p> <pre><strong>[1]</strong> Silva SR, Diaz YC, Penha HA, Pinheiro DG, Fernandes CC, Miranda VF, MichaelTP, Varani AM. <strong>The Chloroplast Genome of Utricularia reniformis Sheds Light on the Evolution of the ndh Gene Complex of Terrestrial Carnivorous Plants from the Lentibulariaceae Family</strong>. PLoS One. 2016 Oct 20;11(10):e0165176. doi:<strong><a href="https://www.ncbi.nlm.nih.gov/pubmed/27764252">10.1371/journal.pone.0165176</a></strong>. </pre> <pre><strong>[2] </strong>Silva SR, Alvarenga DO, Aranguren Y, Penha HA, Fernandes CC, Pinheiro DG, Oliveira MT, Michael TP, Miranda VFO, Varani AM. <strong>The mitochondrial genome of the terrestrial carnivorous plant Utricularia reniformis (Lentibulariaceae): Structure, comparative analysis and evolutionary landmarks.</strong> PLoS One. 2017 Jul19;12(7):e0180484. doi: <strong><a href="https://www.ncbi.nlm.nih.gov/pubmed/28723946">10.1371/journal.pone.0180484</a></strong>.</pre> <pre><strong>[3] </strong>Silva SR, Moraes AP, Penha HA, Julião MHM, Domingues DS, Michael TP, Miranda VFO, Varani AM. <strong>The Terrestrial Carnivorous Plant Utricularia reniformis Sheds Light on Environmental and Life-Form Genome Plasticity.</strong> Int J Mol Sci. 2019 Dec 18;21(1). pii: E3. doi: <strong><a href="https://www.ncbi.nlm.nih.gov/pubmed/31861318">10.3390/ijms21010003</a></strong>.</pre> <p><strong>Acknowledgements</strong></p> <p>This work was supported by Sao Paulo Research Foundation FAPESP, Grant ID: [1325164-6]</p> <p> </p> <p><strong>---------------------------------------------------------</strong><br> <strong>FILES DESCRIPTION</strong><br> <strong>---------------------------------------------------------</strong><br> <br> ----------------<br> <strong>ANNOT-vFinal.sql: </strong>MySQL database containing all integrated annotation information of Urenif and Ugibba<br> ----------------<br> <strong>TABLE fields description</strong><br> gene_name gene name generated by EVidence Modeler + PASA<br> length gene lenght<br> status duplicate_gene_classifier status (0:singleton, 1:dispersed, 2:proximal, 3: tandem, 4:WGD)<br> product gene product <br> GOterms Blast2GO/OmicsBox GOterms<br> GO_mapping Blast2GO/OmicsBox GOterms derived from direct mapping (UniProt)<br> GO_annotation Blast2GO/OmicsBox annotated GOterms<br> GO_interpro Blast2GO/OmicsBox derived from InterProScan<br> EC Blast2GO/OmicsBox EC number<br> EC_name Blast2GO/OmicsBox enzyme name<br> NOG_annot EggNOG annotation description<br> NOG_EC EggNOG EC number<br> NOG_GO EggNOG GOterms<br> NOG_class EggNOG COG/KOG classfication<br> KEGG_Pathway EggNOG KEGG pathyways<br> KEGG_ko EggNOG KEGG ko<br> CAZy EggNOG CAZy enzymes<br> TAIR_gene Closest A. thaliana gene name (homologous) TAIR database lasted version<br> TAIR_annot Closest A. thaliana gene product (homologous) TAIR database lasted version <br> ortho MCL clustering among Vvinifera, Athaliana, and Slycopersicum (S:singleton, C: clustered, Y: shared)<br> ortho_two MCL clustering among Urenif and Ugibba (S:singleton, C: clustered, Y: shared)<br> -<br> -<br> ----------------<br> <strong>CEGs.zip </strong> 336 shared and concatenated CEGs from Urenif, U. gibba, Genlisea nigrocaulis, G. hispidula, G. aurea, G. pygmaea, and G. repens.<br> ----------------</p> <p><strong>ProcessRepeats_mod</strong> Modified version of RepeatMasker, ProcessRepeats script for detection of plant evolutionary lineages<br> ----------------</p> <p><strong>----------------------------------------------------------------------------------------------------------------------------------------------<br> <em>Utricularia gibba</em> files<br> ----------------------------------------------------------------------------------------------------------------------------------------------</strong><br> <strong>Ugibba</strong><strong>-no-masked.fa </strong> Ugibba genome excluding organellar genomes (provided by Lan et al., 2017)<br> <strong>Ugibba-softmasked.fa</strong> Ugibba genome RepeatMasker softmasked and excluding organellar genomes (provided by Lan et al., 2017)<br> <strong>Ug.collinearity </strong> MCScanX collinearity file<br> <strong>Ug-duplicates.txt</strong> MCScanX duplicate_gene_classifier short report<br> <strong>Ug.gene_type </strong> MCScanX duplicate_gene_classifier full report<br> <strong>Ug.tandem </strong> Ugibba tandem genes generated by MCScanX tool<br> <strong>Ugibba_annot.annot </strong> Blast2GO/OmicsBox annotation file (eudicotyledons filtered and Viridiplantae GOSlim) <strong>Ugibba_annot-</strong><strong>noclean</strong><strong>.</strong><strong>annot</strong><strong> </strong> Blast2GO/OmicsBox annotation file (not filtered)<br> <strong>Ugibba</strong><strong>.cDNA</strong> Ugibba cDNAs fasta file<br> <strong>Ugibba</strong><strong>.CDS </strong> Ugibba CDSs fasta file<br> <strong>Ugibba</strong><strong>-EVM.all-no-TEs-PASA-ANNOTATED.gff3</strong> Ugibba GFF3 file fully annotated (including gene products and GO terms)</p> <p><strong>Ugibba</strong><strong>-EVM.all-no-TEs-PASA.gff3</strong> Ugibba GFF3 file fully annotated (genes only)<br> <strong>Ugibba_export.txt</strong> Blast2GO/OmicsBox full exported table<br> <strong>Ugibba_fasta.fasta</strong> Blast2GO/OmicsBox Ugibba fasta proteins containg annotation (product and GO terms)<br> <strong>ugibba_frozen_cleaned-validated.box</strong> Full Blast2GO/OmicsBox file</p> <p><strong>ugibba_frozen.box</strong> Full Blast2GO/OmicsBox file (containing TEs genes annotation)</p> <p><strong>ugibba_nogs_emapper_annotations.box</strong> Full Blast2GO/OmicsBox EggNOG file (containing TEs genes annotation)</p> <p><strong>Ugibba_GAF.txt</strong> GAF file<br> <strong>Ugibba</strong><strong>.gene</strong> Ugibba gene fasta file<br> <strong>Ugibba_GOstat.txt </strong> GOstat file<br> <strong>Ugibba</strong><strong>-PASA-assemblies.fasta </strong> Ugibba PASA assemblies<br> <strong>Ugibba</strong><strong>-PASA.stats </strong> Ugibba annotation STATS<br> <strong>Ugibba</strong><strong>.</strong><strong>prot</strong><strong> </strong> Ugibba protein fasta file<br> <strong>Ugibba</strong><strong>-RepeatMasker.gff </strong> Ugibba RepeatMasker gff file<br> <strong>Ugibba</strong><strong>-RepeatMasker.gff3 </strong> Ugibba RepeatMasker gff3 file<br> <strong>Ugibba</strong><strong>-RepeatMasker.tbl </strong> Ugibba RepeatMasker results<br> <strong>Ugibba</strong><strong>-RepeatMasker-v2.gff3</strong> Ugibba RepeatMasker gff3 second version file<br> <strong>Ugibba</strong><strong>-RNAseq-assembled.fasta </strong> Ugibba RNAseq assembled transcriptome (Trinity)<br> <strong>Ugibba_TEs_DANTE_2019.fa </strong> Ugibba TEs library, detected by REPET and annotated by PASTEC and DANTE<br> <strong>Ugibba_WEGO.txt </strong> WEGO file</p> <p><strong>----------------------------------------------------------------------------------------------------------------------------------------------<br> <em>Utricularia reniformis</em> files<br> ----------------------------------------------------------------------------------------------------------------------------------------------</strong><br> <strong>Urenif</strong><strong>-no-masked.fa </strong> Urenif genome excluding organellar genomes<br> <strong>Urenif</strong><strong>-</strong><strong>softmasked</strong><strong>.fa</strong> Urenif genome RepeatMasker softmasked and excluding organellar genomes<br> <strong>Ur.collinearity </strong> MCScanX collinearity file<br> <strong>Ur-duplicates.txt </strong> MCScanX duplicate_gene_classifier short report<br> <strong>Ur.gene_type</strong> MCScanX duplicate_gene_classifier full report<br> <strong>Ur.tandem</strong> Urenif tandem genes generated by MCScanX tool<br> <strong>Urenif_annot.annot</strong> Blast2GO/OmicsBox annotation file (eudicotyledons filtered and Viridiplantae GOSlim)<br> <strong>Urenif_annot-</strong><strong>noclean</strong><strong>.</strong><strong>annot</strong> Blast2GO/OmicsBox annotation file (not filtered)<br> <strong>Urenif</strong><strong>.cDNA</strong> Urenif cDNAs fasta file<br> <strong>Urenif</strong><strong>.CDS </strong> Urenif cDNAs fasta file<br> <strong>Urenif</strong><strong>-EVM.all-no-TEs-PASA-ANNOTATED.gff3</strong> Urenif GFF3 file fully annotated (including gene products and GO terms)</p> <p><strong>Urenif</strong><strong>-EVM.all-no-TEs-PASA.gff3</strong> Urenif GFF3 file fully annotated (genes only)<br> <strong>Urenif_export.txt</strong> Blast2GO/OmicsBox full exported table<br> <strong>Urenif_fasta.fasta</strong> Blast2GO/OmicsBox Urenif fasta proteins containg annotation (product and GO terms)<br> <strong>urenif_frozen_cleaned-validated.box</strong> Full Blast2GO/OmicsBox file</p> <p><strong>urenif_frozen.box</strong> Full Blast2GO/OmicsBox file (containing TEs genes annotation)</p> <p><strong>urenif_nogs_emapper_annotations.box</strong> Full Blast2GO/OmicsBox EggNOG file (containing TEs genes annotation)<br> <strong>Urenif_GAF.txt </strong> GAF file<br> <strong>Urenif</strong><strong>.gene</strong> Urenif gene fasta file<br> <strong>Urenif_GOStat.txt </strong> GOstat file<br> <strong>Urenif</strong><strong>-PASA-assemblies.fasta</strong> Urenif PASA assemblies<br> <strong>Urenif</strong><strong>-PASA.stats </strong> Urenif annotation STATS<br> <strong>Urenif</strong><strong>.</strong><strong>prot</strong><strong> </strong> Urenif protein fasta file<br> <strong>Urenif</strong><strong>-RepeatMasker.gff </strong> Urenif RepeatMasker gff file<br> <strong>Urenif</strong><strong>-RepeatMasker.gff3 </strong> Urenif RepeatMasker gff3 file<br> <strong>Urenif</strong><strong>-RepeatMasker.tbl </strong> Urenif RepeatMasker results<br> <strong>Urenif</strong><strong>-RepeatMasker-v2.gff3 </strong> Urenif RepeatMasker gff3 second version file<br> <strong>Urenif</strong><strong>-RNAseq-assembled.fasta </strong> Urenif RNAseq assembled transcriptome (Trinity)<br> <strong>Urenif_TEs_DANTE_2019.fa </strong> Urenif TEs library, detected by REPET and annotated by PASTEC and DANTE<br> <strong>Urenif_WEGO.txt </strong> WEGO file<br> <strong>----------------------------------------------------------------------------------------------------------------------------------------------<br> ----------------------------------------------------------------------------------------------------------------------------------------------</strong></p>
UnientrezDB: Large-scale Gene Ontology Annotation Dataset and Evaluation Benchmarks with Unified Entrez Gene Identifiers
<p>Our work focuses on providing a comprehensive dataset and benchmarks for evaluating gene ontology annotations using a unified system of Entrez Gene Identifiers.</p>
Gene Annotations of 49 Bacillariophyta Genome Assemblies (Individual gff3 files)
<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div> </div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a> . It is a copy of the data hostet at <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a> , but instead of storing one archive will all gff3 files included, the gff3 files are here hosted, individually. This copy was made upon request from the RDA Working Group "FAIRification of Genomic Annotations – metadata harmonisation at scale".</p> <div> <h2>Files</h2> <div>The following gzip-compressed gff3-files with structural and functional genome annotation are included:</div> <div> </div> <div>Asterionella_formosa.gff3.gz<br>Asterionellopsis_glacialis.gff3.gz<br>Bacterosira_constricta.gff3.gz<br>Chaetoceros_muellerii.gff3.gz<br>concatenated_output.gff3.gz<br>Conticribra_guillardii.gff3.gz<br>Conticribra_weissflogii.gff3.gz<br>Craspedostauros_australis.gff3.gz<br>Cyclostephanos_invisitatus.gff3.gz<br>Cyclostephanos_tholiformis.gff3.gz<br>Cyclotella_atomus.gff3.gz<br>Cyclotella_baltica.gff3.gz<br>Cyclotella_choctawhatcheeana.gff3.gz<br>Cyclotella_cryptica.gff3.gz<br>Cylindrotheca_fusiformis.gff3.gz<br>Detonula_confervacea.gff3.gz<br>Discostella_pseudostelligera.gff3.gz<br>Discostella_stelligera.gff3.gz<br>Discostella_stelligeroides.gff3.gz<br>Epithemia_pelagica.gff3.gz<br>Fistulifera_pelliculosa.gff3.gz<br>Fistulifera_solaris.gff3.gz<br>Fragilaria_radians.gff3.gz<br>Fragilariopsis_cylindrus.gff3.gz<br>Licmophora_abbreviata.gff3.gz<br>Mediolabrus_comicus.gff3.gz<br>Nitzschia_palea.gff3.gz<br>Nitzschia_putrida.gff3.gz<br>Porosira_glacialis.gff3.gz<br>Psammoneis_japonica.gff3.gz<br>Pseudo-nitzschia_multiseries.gff3.gz<br>Pseudo-nitzschia_pungens.gff3.gz<br>Skeletonema_costatum.gff3.gz<br>Skeletonema_marinoi.gff3.gz<br>Skeletonema_menzelii.gff3.gz<br>Skeletonema_potamos.gff3.gz<br>Skeletonema_tropicum.gff3.gz<br>Stephanocyclus_meneghinianus.gff3.gz<br>Stephanodiscus_minutulus.gff3.gz<br>Stephanodiscus_triporus.gff3.gz<br>Thalassiosira_allenii.gff3.gz<br>Thalassiosira_delicatula.gff3.gz<br>Thalassiosira_exigua.gff3.gz<br>Thalassiosira_gravida.gff3.gz<br>Thalassiosira_livingstoniorum.gff3.gz<br>Thalassiosira_mediterranea.gff3.gz<br>Thalassiosira_oceanica.gff3.gz<br>Thalassiosira_ordinaria.gff3.gz<br>Thalassiosira_pacifica.gff3.gz<br>Thalassiosira_profunda.gff3.gz</div> <div> </div> <div>To extract individual files after download execute the following command:</div> <div> </div> <div><code>gunzip *.gff3.gz</code></div> <h2>Genome Assemblies</h2> <p> </p> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <p> </p> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p> </p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p> </p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p> </p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^>/ s/ .*//' genome.fasta > genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <p> </p> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <p> </p> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>The submission and release was made upon request of the RDA working group "FAIRification of Genomic Annotations – metadata harmonisation at scale". The contained data is identical to <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a></p> <h2>License</h2> <p> </p> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div> <p> </p> </div> </div>
Gene annotations of Amphibolurus muricatus (jacky dragon), Intellagama lesueurii (Australian water dragon), Phrynocephalus przewalskii (Przewalski's toadhead agama), and Phrynocephalus vlangalii (Ching Hai toadhead agama)
<p><strong>Annotation file and associated FASTA files for <em>A. muricatus</em> assembly AmpMurF_3.0</strong><br> • AmpMurF3.gff3.tar.gz: EVidenceModeler (EVM) gene model annotation file.<br> • AmpMurF3.cds.tar.gz: EVM gene models coding sequences.<br> • AmpMurF3.pep.tar.gz: EVM gene models coding sequences translated into amino acid sequences.</p> <p><strong>Annotation file and associated FASTA files for <em>A. muricatus</em> assembly AmpMurM_3.0</strong><br> • AmpMurM3.gff3.tar.gz: EVidenceModeler (EVM) gene model annotation file.<br> • AmpMurM3.cds.tar.gz: EVM gene models coding sequences.<br> • AmpMurM3.pep.tar.gz: EVM gene models coding sequences translated into amino acid sequences.</p> <p><strong>Annotation file and associated FASTA files for <em>I. lesueurii</em> (Australian water dragon; assembly EWD_hifiasm_HiC generated as part of the AusARG consortium)</strong><br> • Intellagama_lesueurii.evm.final.add_replace_buscoV5_homolog.final.gff3.tar.gz: EVidenceModeler (EVM) gene model annotation file.<br> • Intellagama_lesueurii.evm.final.add_replace_buscoV5_homolog.final.cds.fa.tar.gz: EVM gene models coding sequences.<br> • Intellagama_lesueurii.evm.final.add_replace_buscoV5_homolog.final.pep.fa.tar.gz: EVM gene models coding sequences translated into amino acid sequences.</p> <p><strong>Annotation file and associated FASTA files for <em>P. przewalskii</em> (Przewalski’s toadhead agama; see PMID ID 30808754 and CNGBdb accession no. CNP0000203) </strong><br> • Phrynocephalus_przewalskii.evm.final.add_replace_buscoV5_homolog.gff3.tar.gz: EVidenceModeler (EVM) gene model annotation file.<br> • Phrynocephalus_przewalskii.evm.final.add_replace_buscoV5_homolog.cds.fa.tar.gz: EVM gene models coding sequences.<br> • Phrynocephalus_przewalskii.evm.final.add_replace_buscoV5_homolog.pep.fa.tar.gz: EVM gene models coding sequences translated into amino acid sequences.</p> <p><strong>Annotation file and associated FASTA files for <em>P. vlangalii</em> (Ching Hai toadhead agama; see PMID ID 30808754 and CNGBdb accession no. CNP0000203)</strong><br> • Phrynocephalus_vlangalii.evm.final.add_replace_busco_homolog.gff3.tar.gz: EVidenceModeler (EVM) gene model annotation file.<br> • Phrynocephalus_vlangalii.evm.final.add_replace_busco_homolog.cds.fa.tar.gz: EVM gene models coding sequences.<br> • Phrynocephalus_vlangalii.evm.final.add_replace_busco_homolog.pep.fa.tar.gz: EVM gene models coding sequences translated into amino acid sequences.</p>
Worldwide Fraxinus Genome Assemblies, Annotations, and Gene Families v0.2
<p>The worldwide <em>Fraxinus</em> genome project was conducted to assess the pathogenic resistance of 34 Ash tree species to Ash Dieback and Emerald Ash Borer. The project is led by Dr. Richard Buggs. v0.1 genomes are available at ashgenome.org and on ENA. As part of Josiah Seaman's PhD thesis, he improved the assembly of 13 genomes included here as v0.2. New de novo annotations, gene families, and all associated files are included for future studies and reproducibility. </p> <p>Annotations are done with GeMoMa using F. excelsior as a reference (Keilwagen et al. 2016). Gene families are defined as genes originating from a single copy at the last common ancestor with Solanum. Orthofinder outputs reconciled gene trees, aligned CDS, and gene families (Emms and Kelly 2015; Tekaia 2016). Species tree was calibrated based on fossil evidence using r8s, RAxML across 25,182,399 sites (SpeciesTreeAlignment.fa). More methods details can be found in the full Chapter two of Josiah Seaman's PhD thesis (2021).</p> <p>I'd be happy to talk with you if you'd like any additional information or help visualizing your genomic data. You can find the tools used to browse this data at https://fluentdna.com/ and http://graphgenome.org/ Contact me at josiah@newline.us</p>
Dataset for "Bacterial genome annotation" and "AMR gene detection" workflows
<p>This dataset is associated with the workflows "Bacterial genome annotation" and "AMR gene detection in an assembled bacterial genome".</p>
MGBC-26640: nucleotide sequences for gene annotations
<p>Nucleotide sequences of annotated genes from the 26,640 high-quality, non-redundant genomes of the MGBC.</p>
Aliarcobacter butzleri gene annotation and transcriptome data
<p>Genomes assembled sequences, functional annotation files (Prokka), logFC table of 3 A. butzleri strains isolated from human (LMG 10828<sup>T</sup>, LMG 11119, 31).</p>
Processed and annotated yeast gene expression data from yeast2 and ygs98 platforms
<p>This dataset contains the following files:</p> <ul> <li><em>yeast2_processed_rds.tar.gz -</em> processed gene expression matrices from the yeast2 platform. The data is stored in binary R format (.rds).</li> <li><em>ygs98_processed_rds.tar.gz </em>- processed gene expression matrices from the yeast2 platform. The data is stored in binary R format (.rds).</li> <li><em>yeast2-curated-annotations.txt</em> - metadata for the yeast2 platform.</li> <li><em>ygs98-curated-annotations.txt</em> - metadata for the ygs98 platform.</li> </ul> <p> </p>
eggNOG Mapper annotations of Mouse, Dog and Pig gut gene catalogs
<p><a href="https://github.com/jhcepas/eggnog-mapper">eggNOG-mapper</a> annotations of <a href="https://doi.org/10.1038/nbt.3353">mouse</a>, <a href="https://doi.org/10.1186/s40168-018-0450-3">dog</a> and <a href="https://doi.org/10.1038/nmicrobiol.2016.161">pig</a> gut, and <a href="https://doi.org/10.1038/nbt.2942">IGC</a> gene catalogs.</p>
Cell specificity of human regulatory annotations and their genetic effects on gene expression
<p>Processed data for the manuscript: Cell specificty of regulatory annotations and their genetics effects on gene expression.</p> <p>The code to generate these data is included in the zip-file: regulatoryAnnotations_comparisons-master.zip</p> <p>and at GitHub: https://github.com/ParkerLab/regulatoryAnnotations_comparisons</p> <p>Individual dataset information:</p> <p>numberOfSegments.dat.gz: Number of segments in each annotation</p> <p>lengthDistributionOfSegments.dat.gz: Length distribution of segments in each annnotation</p> <p>genomeCoverage.dat.gz: Total genome coverage of regions in each annoation</p> <p>overlapFraction.dat.gz: Basepair level overlap between two pairs of annotations</p> <p>annotations_chromatinStateOverlap.dat.gz: Overlap of annotations with chromatin states</p> <p>information.fourcells_enhancer.dat.gz: Information content of the mean enhancer chromatin posterior probability for annotation segments across four cell types</p> <p>information.fourcells_promoter.dat.gz: Information content of the mean promoter chromatin posterior probability for annotation segments across four cell types</p> <p>gwas_enrichment_stats.txt.gz: Enrichment of GWAS SNPs in annotations</p> <p>GTEx_v7.ESI_medianTPM_min0.15.dat.gz : ESI for genes across 50 tissues</p> <p>gtexv7_lcleqtl.enrichment.dat.gz: Enrichment of GTEx v7 LCL eQTL in annotations</p> <p>gtexv7_lcleqtl.binned_lclESI.enrichment.dat.gz: Enrichment of GTEx v7 LCL eQTL binned by lclESI in annotations</p> <p>K562.gtexv7_bloodeqtl.fdr0.1.prune0.8.maf0.2.ld0.99.annotations.dat.gz: GTEx v7 blood eQTL effect sizes in K562 annotations </p> <p>GM12878.gtexv7_lcleqtl.fdr0.1.prune0.8.maf0.2.ld0.99.annotations.dat.gz: GTEx v7 LCL eQTL effect sizes in GM12878 annotations</p> <p>GM12878.dsqtl.prune0.8.maf0.2.ld0.99.annotations.dat.gz : GM12878 DNase QTL effect sizes in GM12878 annotations</p> <p>GM12878.allelicBiasResults.FracRef.downsampled30.annotations.withMAF.dat.gz: GM12878 ATAC-seq allelic bias effect sizes in GM12878 annotations </p> <p> <br> <br> <br> </p>
The North Pacific Eukaryotic Gene Catalog: metatranscriptome assemblies with taxonomy, function and abundance annotations
<p>This data continues with the development of the unprocessed NPEGC Trinity <em>de novo</em> metatranscriptome assemblies, uploaded to this Zenodo repository for raw assemblies: <a href="../records/7332796">The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3</a><br><br>A full description of this data is published in Scientific Data, available here: <a href="https://www.nature.com/articles/s41597-024-04005-5" target="_blank" rel="noopener">The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations</a>. Please cite this publication if your research uses this data:<br><br>Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., & Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. <em>Scientific Data</em>, <em>11</em>(1), 1161.</p> <p><br>Excerpts of key processing steps are sampled below with links to the detailed code on the main github code repository: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog">https://github.com/armbrustlab/NPac_euk_gene_catalog</a></p> <p><br>Processing and annotation of protein-level NPEGC metatranscripts is done in 6 primary steps:<br>1. Six-frame translation into protein sequences<br>2. Frame-selection of protein-coding translation frames<br>3. Clustering of protein sequences at 99% sequence identity<br>4. Taxonomic annotation against MarFERReT v1.1 + MARMICRODB v1.0 multi-kingdom marine reference protein sequence library with DIAMOND<br>5. Functional annotation against Pfam 35.0 protein family HMM profiles using HMMER3<br>6. Functional annotation against KOfam HMM profiles (KEGG release 104.0) using KofamScan v1.3.0<br><br><code># Define local NPEGC base directory here:</code><br><code>NPEGC_DIR="/mnt/nfs/projects/armbrust-metat"</code></p> <p><code># Raw assemblies are located in the /assemblies/raw/ directory</code><br><code># for each of the metatranscriptome projects</code><br><code>PROJECT_LIST="D1PA G1PA G2PA G3PA G3PA_diel"</code></p> <p><code># raw Trinity assemblies:</code><br><code>RAW_ASSEMBLY_DIR="${NPEGC_DIR}/${PROJECT}/assemblies/raw"</code><br><br><strong>Translation</strong><br>We began processing the raw metatranscriptome assemblies by six-frame translation from nucleotide transcripts into three forward and three reverse reading frame translations, using the transeq function in the EMBOSS package. We add a cruise and sample prefix to the sequence IDs to ensure unique identification downstream (ex, `>TRINITY_DN2064353_c0_g1_i1_1` to `>G1PA_S09C1_3um_TRINITY_DN2064353_c0_g1_i1_1` for the S09C1_3um sample in the G1PA assemblies). See <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/aa_data/NPEGC.6tr_frame_selection_clustering.sh">NPEGC.6tr_frame_selection_clustering.sh</a> for full code description.<br><br>Example of six-frame translation using transeq<br><code>transeq -auto -sformat pearson -frame 6 -sequence 6tr/${PREFIX}.Trinity.fasta -outseq 6tr/${PREFIX}.Trinity.6tr.fasta</code><br><br><strong>Frame selection</strong><br>We use a custom frame-selection python script <a href="https://github.com/armbrustlab/marferret/blob/main/scripts/python/keep_longest_frame.py">keep_longest_frame.py</a> to determine the longest coding length in each open reading frame and retain this sequence (or multiple sequences if there is a tie) for downstream analyses. See <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/aa_data/NPEGC.6tr_frame_selection_clustering.sh">NPEGC.6tr_frame_selection_clustering.sh</a> for full code description.<br><br><strong>Clustering by sequence identity</strong><br>To reduce sequence redundancy and near-identical sequences, we cluster protein sequences at the 99% sequence identity level and retain the sequence cluster representative in a reduced-size FASTA output file. See <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/aa_data/NPEGC.6tr_frame_selection_clustering.sh">NPEGC.6tr_frame_selection_clustering.sh</a> for full code description of linclust/mmseqs clustering.<br><br>Sample of linclust clustering script: core mmseqs function<br><code>function NPEGC_linclust {</code><br><code># make an index of the fasta file:</code><br><code>$MMSEQS_DIR/mmseqs createdb $FASTA_PATH/$FASTA_FILE NPac.$STUDY.bf100.db</code><br><code># cluster sequences at $MIN_SEQ_ID</code><br><code>$MMSEQS_DIR/mmseqs linclust NPac.${STUDY}.bf100.db NPac.${STUDY}.clusters.db NPac_tmp --min-seq-id ${MIN_SEQ_ID}</code><br><code># retieve cluster representatives:</code><br><code>$MMSEQS_DIR/mmseqs result2repseq NPac.${STUDY}.bf100.db NPac.${STUDY}.clusters.db NPac.${STUDY}.clusters.rep</code><br><code># generate flat FASTA output with cluster reps</code><br><code>$MMSEQS_DIR/mmseqs result2flat NPac.${STUDY}.bf100.db NPac.${STUDY}.bf100.db NPac.${STUDY}.clusters.rep NPac.${STUDY}.bf100.id99.fasta --use-fasta-header</code><br><code>}</code><br><br>Corresponding files uploaded to this repository: Gzip-compressed FASTA files after translation, frame-selection, and clustering at 99% sequence identity (.bf100.id99.aa.fasta.gz)<br><strong> </strong><em> NPac.G1PA.bf100.id99.aa.fasta.gz</em><br><em> NPac.G2PA.bf100.id99.aa.fasta.gz</em><br><em> NPac.G3PA.bf100.id99.aa.fasta.gz</em><br><em> NPac.G3PA_diel.bf100.id99.aa.fasta.gz</em><br><em> NPac.D1PA.bf100.id99.aa.fasta.gz</em><br><br><strong>MarFERReT + MARMICRODB taxonomic annotation with DIAMOND</strong></p> <p>Taxonomy was inferred for the NPEGC metatranscripts with the DIAMOND fast read alignment software against the <a href="../records/10586950">MarFERReT v1.1 + MARMICRODB v1.0 multi-kingdom marine reference protein sequence library (v1.1)</a>, a combined database of the <a href="https://doi.org/10.1038/s41597-023-02842-4">MarFERReT v1.1 marine microbial eukaryote sequence library</a> and <a href="https://doi.org/10.5281/zenodo.3520509">MARMICRODB v1.0 </a>prokaryote-focused marine genome database. See <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/aa_data/NPEGC.diamond_taxonomy.log.sh">NPEGC.diamond_taxonomy.log.sh</a> for full description of DIAMOND annotation.</p> <p>Excerpt of core DIAMOND function:<br><code>function NPEGC_diamond {</code><br><code># FASTA filename for $STUDY</code><br><code>FASTER_FASTA="NPac.${STUDY}.bf100.id99.aa.fasta"</code><br><code># Output filename for LCA results in lca.tab file:</code><br><code>LCA_TAB="NPac.${STUDY}.MarFERReT_v1.1_MMDB.lca.tab"</code><br><code>echo "Beginning ${STUDY}"</code><br><code>singularity exec --no-home --bind ${DATA_DIR} \</code><br><code> "${CONTAINER_DIR}/diamond.sif" diamond blastp \</code><br><code> -c 4 --threads $N_THREADS \</code><br><code> --db $MFT_MMDB_DMND_DB -e $EVALUE --top 10 -f 102 \</code><br><code> --memory-limit 110 \</code><br><code> --query ${FASTER_FASTA} -o ${LCA_TAB} >> "${STUDY}.MarFERReT_v1.1_MMDB.log" 2>&1</code><br><code>}</code><br><br>Corresponding files uploaded to this repository: Gzip-compressed diamond lowest common ancestor predictions with NCBI Taxonomy against a combined MarFERReT + MARMICRODB taxonomic library (*.Pfam35.domtblout.tab.gz)<br><em> NPac.G1PA.MarFERReT_v1.1_MMDB.lca.tab.gz</em><br><em> NPac.G2PA.MarFERReT_v1.1_MMDB.lca.tab.gz</em><br><em> NPac.G3PA.MarFERReT_v1.1_MMDB.lca.tab.gz</em><br><em> NPac.G3PA_diel.MarFERReT_v1.1_MMDB.lca.tab.gz</em><br><em> NPac.D1PA.MarFERReT_v1.1_MMDB.lca.tab.gz</em><br><br><strong>Pfam 35.0 functional annotation using HMMER3</strong><br>Clustered protein sequences were annotated against the Pfam 35.0 collection of 19,179 protein family Hidden Markov Models (HMMs) using <a href="http://hmmer.org/">HMMER 3.3 </a> with the <a href="https://academic.oup.com/nar/article/49/D1/D412/5943818">Pfam 35.0 protein family database</a>. Pfam annotation code is documented here: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/aa_data/NPEGC.hmmer_function.sh">NPEGC.hmmer_function.sh</a><br><br>Excerpt of core hmmsearch function:<br><br><code>function NPEGC_hmmer {</code><br><code># Define input FASTA</code><br><code>INPUT_FASTA="NPac.${STUDY}.bf100.id99.aa.fasta"</code><br><code># hmmsearch call:</code><br><code>hmmsearch --cut_tc --cpu $NCORES --domtblout $ANNOTATION_DIR/${STUDY}.Pfam35.domtblout.tab $HMM_PROFILE ${INPUT_FASTA}</code><br><code># compress output file:</code><br><code>gzip $ANNOTATION_DIR/${STUDY}.Pfam35.domtblout.tab</code><br><code>}</code><br><br>Corresponding files uploaded to this repository: Gzip-compressed hmmsearch domain table files for Pfam35 queries (*.Pfam35.domtblout.tab.gz)<br><em> G1PA.Pfam35.domtblout.tab.gz</em><br><em> G2PA.Pfam35.domtblout.tab.gz</em><br><em> G3PA.Pfam35.domtblout.tab.gz</em><br><em> G3PA_diel.Pfam35.domtblout.tab.gz</em><br><em> D1PA.Pfam35.domtblout.tab.gz</em><br><br></p> <p><strong>KEGG functional annotation using KofamScan v1.3.0</strong></p> <p>Clustered protein sequences were annotated against the KEGG collection (release 104.0) of 20,819 protein family Hidden Markov Models (HMMs) using <a href="https://github.com/takaram/kofam_scan" target="_blank" rel="noopener">KofamScan </a>and KofamKOALA. Kofam annotation code is documented here: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/aa_data/NPEGC.kofamscan_function.sh">NPEGC.kofamscan_function.sh</a></p> <p>Excerpt of core NPEGC_kofam function:</p> <p><code># Core function to perform KofamScan annotation</code><br><code>function NPEGC_kofam {</code><br><code> # Define input FASTA</code><br><code> local INPUT_FASTA="NPac.${STUDY}.bf100.id99.aa.fasta"</code></p> <p><code> # KofamScan call</code><br><code> ${KOFAM_DIR}/kofam_scan-1.3.0/exec_annotation -f detail-tsv -E ${EVALUE} -o ${ANNOTATION_DIR}/NPac.${STUDY}.bf100.id99.aa.tsv ${FASTA_DIR}/${INPUT_FASTA}</code></p> <p><code> # Keep best hit (data is already sorted by KofamScan)</code><br><code> sort -uk1,1 ${ANNOTATION_DIR}/NPac.${STUDY}.bf100.id99.aa.tsv > ${ANNOTATION_DIR}/NPac.${STUDY}.bf100.id99.aa.best.kofam.tsv</code></p> <p><code> # Compress output file</code><br><code> gzip ${ANNOTATION_DIR}/NPac.${STUDY}.bf100.id99.aa.tsv</code></p> <p><code> # Compress best.kofam output file</code><br><code> gzip ${ANNOTATION_DIR}/NPac.${STUDY}.bf100.id99.aa.best.kofam.tsv</code><br><code>}</code></p> <p><br><code># filter hits with a score > 30 in R</code></p> <p>Corresponding files uploaded to this repository: Gzip-compressed KofamScan domain table files for Kofam queries (*.best.Kofam.incT30.csv.gz):<em><br> NPac.G1PA.bf100.id99.aa.best.Kofam.incT30.csv.gz</em><em><br> NPac.G2PA.bf100.id99.aa.best.Kofam.incT30.csv.gz</em><br><em> NPac.G3PA.UW.bf100.id99.aa.best.Kofam.incT30.csv.gz</em><br><em> NPac.G3PA.diel.bf100.id99.aa.best.kofam.incT30.csv.gz</em><br><em> NPac.D1PA.diel.bf100.id99.aa.best.kofam.incT30.csv.gz<br><br></em>The full kofamscan tables with score >30 are deposited here: <a title="The North Pacific Eukaryotic Gene Catalog: KOfam protein function annotations" href="../records/13743267" target="_blank" rel="noopener">https://zenodo.org/records/13743267</a></p>
The North Pacific Eukaryotic Gene Catalog: KOfam protein function annotations
<p><strong>KEGG functional annotation using KofamScan v1.3.0</strong></p> <p>These tables are larger alternative versions to the KOfam tables included in the North Pacific Eukaryotic Gene Catalog protein data repository here: <a href="../records/12630398">https://zenodo.org/records/12630398</a><br><br>A full description of this data is published in Scientific Data, available here: <a href="https://www.nature.com/articles/s41597-024-04005-5" target="_blank" rel="noopener">The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations</a>. Please cite this publication if your research uses this data:<br><br>Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., & Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. <em>Scientific Data</em>, <em>11</em>(1), 1161.<br><br>Clustered protein sequences were annotated against the KEGG collection (release 104.0) of 20,819 protein family Hidden Markov Models (HMMs) using <a href="https://github.com/takaram/kofam_scan" target="_blank" rel="noopener">KofamScan </a>and KofamKOALA. Kofam annotation code is documented in the project github repository here: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/aa_data/NPEGC.kofamscan_function.sh">NPEGC.kofamscan_function.sh</a></p> <p>Excerpt of core NPEGC_kofam function:</p> <p><code># Define input FASTA</code><br><code>local INPUT_FASTA="NPac.${STUDY}.bf100.id99.aa.fasta"</code></p> <p><code># KofamScan call</code><br><code>${KOFAM_DIR}/kofam_scan-1.3.0/exec_annotation -f detail-tsv -E ${EVALUE} -o ${ANNOTATION_DIR}/NPac.${STUDY}.bf100.id99.aa.tsv ${FASTA_DIR}/${INPUT_FASTA}</code></p> <p>Unprocessed annotation results were filtered with a minimum score of 30 to remove low-scoring matches:<br><br><code>zcat NPac.<em><u>NPacID</u></em>.kofam.tsv.gz | awk -F'\t' '{ gsub(/"/, "", $5); $5 = $5 + 0; if ($5 >= 30) print }' | gzip > NPac.<em><u>NPacID</u></em>.UW.bf100.id99.aa.incT30.tsv.gz</code></p>
Neurogenomic divergence during speciation by reinforcement of mating behaviors in chorus frogs (Pseudacris) – De novo reference transcriptome: Assemblerd contigs and gene annotations
<p>Assembled contigs (Trinity) and gene annotations (Trinotate) of a reference transcriptome for the Upland Chorus Frog, <em>Pseudacris feriarum</em>. Data to assemble the contigs were obtained by sequencing four tissue types: Brain, eyes, testis, and somatic (liver/heart/lung/skin/muscle). Raw reads are stored in the NCBI-SRA database (BioProject PRJNA723357).</p>
Gene and repeat annotation for snowy owl (Bubo scandiacus) and selected species
<p>Here we provide the gene and repeat annotation for snowy owl (<em>Bubo scandiacus</em>), in addition to gene and repeat annotation done for some species this was compared to. It is unfortunately currently not possible to upload repeat annotation tracks to an international nucleotide sequence database such as ENA. While uploading the gene annotation is possible, some of the cross references to different databases in the functional annotation are removed. Further, the names of the entries in the publicly available genome assemblies on ENA have different names than what is found in the annotation tracks here, so we also provide the FASTA files for the snowy owl assemblies (bBubSca1.1.hap1.fasta.gz and bBubSca1.1.hap2.fasta.gz). Ideally, all this should have been available via ENA.</p> <p>We annotated the snowy owl genome assemblies, in addition to downy woodpecker (<em>Dryobates pubescens</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_014839835.1">GCA_014839835.1</a>), Northern Carmine bee-eater (<em>Merops nubicus</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_009819595.1">GCA_009819595.1</a>), Northern goshawk (<em>Accipiter gentilis</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_929443795.2">GCA_929443795.2</a>) and barn owl (<em>Tyto alba</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCF_018691265.1#/st">GCF_018691265.1</a>), since no genome annotation was publicly available for these species. We used a pre-release version of the EBP-Nor genome annotation pipeline (<a href="https://github.com/ebp-nor/GenomeAnnotation">https://github.com/ebp-nor/GenomeAnnotation</a>). First, AGAT (https://zenodo.org/record/7255559) agat_sp_keep_longest_isoform.pl and agat_sp_extract_sequences.pl were used on the GRCg7b (GCA_016699485.1) chicken genome assembly and annotation to generate one protein (the longest isoform) per gene. Miniprot (Li, 2023) was used to align the proteins to the curated assemblies. UniProtKB/Swiss-Prot (Consortium et al., 2022) release 2022_03 in addition to the vertebrata part of OrthoDB v11 (Kuznetsov et al., 2022) were also aligned separately to the assemblies. Red (Girgis, 2015) was run via redmask (<a href="https://github.com/nextgenusfs/redmask">https://github.com/nextgenusfs/redmask</a>) on the snowy owl assemblies to mask repetitive areas (we used the soft-masked genome assemblies available at NCBI for the other species). GALBA (Brůna et al., 2023; Buchfink et al., 2015; Hoff and Stanke, 2018; Li, 2023; Stanke et al., 2006) was run with the chicken proteins using the miniprot mode on the masked assemblies. The funannotate-runEVM.py script from Funannotate was used to run EvidenceModeler (Haas et al., 2008) on the alignments of chicken proteins, UniProtKB/Swiss-Prot proteins, vertebrata proteins and the predicted genes from GALBA. The resulting predicted proteins were compared to the protein repeats that Funannotate distributes using DIAMOND blastp and the predicted genes were filtered based on this comparison using AGAT. The filtered proteins were compared to the UniProtKB/Swiss-Prot release 2022_03 using DIAMOND (Buchfink et al., 2015) blastp to find gene names and InterProScan was used to discover functional domains. AGATs agat_sp_manage_functional_annotation.pl was used to attach the gene names and functional annotations to the predicted genes. EMBLmyGFF3 (Norling et al., 2018) was used to combine the fasta files and GFF3 files into a EMBL format for submission to ENA. These files end in gff.gz (the ones ending in fa.out.gff.gz are repeat annotations), proteins.fa.gz and mrna.fa.gz. </p> <p>All species in this study downy woodpecker (<em>Dryobates pubescens</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_014839835.1">GCA_014839835.1</a>), Northern Carmine bee-eater (<em>Merops nubicus</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_009819595.1">GCA_009819595.1</a>), Northern goshawk (<em>Accipiter gentilis</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_929443795.2">GCA_929443795.2</a>), barn owl (T<em>yto alba</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCF_018691265.1#/st">GCF_018691265.1</a>), chicken (<em>Gallus gallus</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCF_016699485.2/">GCF_016699485.2</a>), zebra finch (<em>Taeniopygia guttat</em>a; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_003957565.4">GCA_003957565.4</a>) and California condor (<em>Gymnogyps californianus</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCF_018139145.2">GCF_018139145.2</a>) in addition to hap1 of snowy owl was repeat masked with a bird-specific library from <a href="https://www.pnas.org/doi/abs/10.1073/pnas.1616702114">https://www.pnas.org/doi/abs/10.1073/pnas.1616702114</a>, provided by Alexander Suh. These files are named such as MerNubi.fa.out.gff.gz, MerNubi.fa.masked.gz and MerNubi.fna.cat.gz. </p> <p>We have also included the species specific repeat library as generated by RepeatModeler running on hap1 of snowy owl. This is called bBubSca1.1.hap1.repeatlibrary.fa.gz, with the files bBubSca1.1.hap1.fasta.masked.gz, bBubSca1.1.hap1.fasta.out.gff.gz, <span>bBubSca1.1.hap1.divsum.gz </span>and bBubSca1.1.hap1.fasta.cat.gz resulting from running RepeatMasker one hap1 using that library.</p> <p>From the Genespace analyses we have included all files including OrthoFinder results. This is found in the file genespace.tgz.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.