Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
345
datasets available to search
ShareScore release 0.7.1
Dataset results
345 results for “genome annotation”
The terrestrial carnivorous plant Utricularia reniformis sheds light on environmental and life-form genome plasticity: Annotation, Gene Ontology and raw data
<p><strong>Description:</strong> In this work, we deeply sequenced (genome and transcriptome of different organs), assembled, and analyzed the 311-Mbp genome of the terrestrial carnivorous plant <em>U. reniformis</em> (Lentibulariaceae). This project presents great importance to the understanding of genomic, evolutive and functional aspects of<em> U. reniformis</em>, which may, with the next-generation sequencing and computational biology approaches shed light to a better understanding not only for the biology and evolution of <em>Utricularia</em> genus, but also for other genera and lineages of the Lentibulariaceae family. Here we present all the raw data generated, including annotation and gene ontology files.</p> <p><strong>External Information</strong></p> <p><a href="https://genomevolution.org/coge/GenomeInfo.pl?gid=54799">Genome Browser</a> avaliable at CoGe Portal (https://genomevolution.org/coge/GenomeInfo.pl?gid=54799)</p> <p><a href="http://https://www.ncbi.nlm.nih.gov/bioproject/290588">GenBank </a><a href="http://https://www.ncbi.nlm.nih.gov/bioproject/290588">Bioproject</a> (https://www.ncbi.nlm.nih.gov/bioproject/290588) for raw genomic and transcriptomic reads</p> <p><a href="https://bv.fapesp.br/en/auxilios/84264/genomics-and-transcriptomics-of-utricularia-reniformis-lentibulariaceae-an-evolutive-and-function/">FAPESP grant website</a> contaning the project abstract and other information.</p> <p><strong>Papers published related to <em>Utricularia reniformis</em> genome</strong></p> <pre><strong>[1]</strong> Silva SR, Diaz YC, Penha HA, Pinheiro DG, Fernandes CC, Miranda VF, MichaelTP, Varani AM. <strong>The Chloroplast Genome of Utricularia reniformis Sheds Light on the Evolution of the ndh Gene Complex of Terrestrial Carnivorous Plants from the Lentibulariaceae Family</strong>. PLoS One. 2016 Oct 20;11(10):e0165176. doi:<strong><a href="https://www.ncbi.nlm.nih.gov/pubmed/27764252">10.1371/journal.pone.0165176</a></strong>. </pre> <pre><strong>[2] </strong>Silva SR, Alvarenga DO, Aranguren Y, Penha HA, Fernandes CC, Pinheiro DG, Oliveira MT, Michael TP, Miranda VFO, Varani AM. <strong>The mitochondrial genome of the terrestrial carnivorous plant Utricularia reniformis (Lentibulariaceae): Structure, comparative analysis and evolutionary landmarks.</strong> PLoS One. 2017 Jul19;12(7):e0180484. doi: <strong><a href="https://www.ncbi.nlm.nih.gov/pubmed/28723946">10.1371/journal.pone.0180484</a></strong>.</pre> <pre><strong>[3] </strong>Silva SR, Moraes AP, Penha HA, Julião MHM, Domingues DS, Michael TP, Miranda VFO, Varani AM. <strong>The Terrestrial Carnivorous Plant Utricularia reniformis Sheds Light on Environmental and Life-Form Genome Plasticity.</strong> Int J Mol Sci. 2019 Dec 18;21(1). pii: E3. doi: <strong><a href="https://www.ncbi.nlm.nih.gov/pubmed/31861318">10.3390/ijms21010003</a></strong>.</pre> <p><strong>Acknowledgements</strong></p> <p>This work was supported by Sao Paulo Research Foundation FAPESP, Grant ID: [1325164-6]</p> <p> </p> <p><strong>---------------------------------------------------------</strong><br> <strong>FILES DESCRIPTION</strong><br> <strong>---------------------------------------------------------</strong><br> <br> ----------------<br> <strong>ANNOT-vFinal.sql: </strong>MySQL database containing all integrated annotation information of Urenif and Ugibba<br> ----------------<br> <strong>TABLE fields description</strong><br> gene_name gene name generated by EVidence Modeler + PASA<br> length gene lenght<br> status duplicate_gene_classifier status (0:singleton, 1:dispersed, 2:proximal, 3: tandem, 4:WGD)<br> product gene product <br> GOterms Blast2GO/OmicsBox GOterms<br> GO_mapping Blast2GO/OmicsBox GOterms derived from direct mapping (UniProt)<br> GO_annotation Blast2GO/OmicsBox annotated GOterms<br> GO_interpro Blast2GO/OmicsBox derived from InterProScan<br> EC Blast2GO/OmicsBox EC number<br> EC_name Blast2GO/OmicsBox enzyme name<br> NOG_annot EggNOG annotation description<br> NOG_EC EggNOG EC number<br> NOG_GO EggNOG GOterms<br> NOG_class EggNOG COG/KOG classfication<br> KEGG_Pathway EggNOG KEGG pathyways<br> KEGG_ko EggNOG KEGG ko<br> CAZy EggNOG CAZy enzymes<br> TAIR_gene Closest A. thaliana gene name (homologous) TAIR database lasted version<br> TAIR_annot Closest A. thaliana gene product (homologous) TAIR database lasted version <br> ortho MCL clustering among Vvinifera, Athaliana, and Slycopersicum (S:singleton, C: clustered, Y: shared)<br> ortho_two MCL clustering among Urenif and Ugibba (S:singleton, C: clustered, Y: shared)<br> -<br> -<br> ----------------<br> <strong>CEGs.zip </strong> 336 shared and concatenated CEGs from Urenif, U. gibba, Genlisea nigrocaulis, G. hispidula, G. aurea, G. pygmaea, and G. repens.<br> ----------------</p> <p><strong>ProcessRepeats_mod</strong> Modified version of RepeatMasker, ProcessRepeats script for detection of plant evolutionary lineages<br> ----------------</p> <p><strong>----------------------------------------------------------------------------------------------------------------------------------------------<br> <em>Utricularia gibba</em> files<br> ----------------------------------------------------------------------------------------------------------------------------------------------</strong><br> <strong>Ugibba</strong><strong>-no-masked.fa </strong> Ugibba genome excluding organellar genomes (provided by Lan et al., 2017)<br> <strong>Ugibba-softmasked.fa</strong> Ugibba genome RepeatMasker softmasked and excluding organellar genomes (provided by Lan et al., 2017)<br> <strong>Ug.collinearity </strong> MCScanX collinearity file<br> <strong>Ug-duplicates.txt</strong> MCScanX duplicate_gene_classifier short report<br> <strong>Ug.gene_type </strong> MCScanX duplicate_gene_classifier full report<br> <strong>Ug.tandem </strong> Ugibba tandem genes generated by MCScanX tool<br> <strong>Ugibba_annot.annot </strong> Blast2GO/OmicsBox annotation file (eudicotyledons filtered and Viridiplantae GOSlim) <strong>Ugibba_annot-</strong><strong>noclean</strong><strong>.</strong><strong>annot</strong><strong> </strong> Blast2GO/OmicsBox annotation file (not filtered)<br> <strong>Ugibba</strong><strong>.cDNA</strong> Ugibba cDNAs fasta file<br> <strong>Ugibba</strong><strong>.CDS </strong> Ugibba CDSs fasta file<br> <strong>Ugibba</strong><strong>-EVM.all-no-TEs-PASA-ANNOTATED.gff3</strong> Ugibba GFF3 file fully annotated (including gene products and GO terms)</p> <p><strong>Ugibba</strong><strong>-EVM.all-no-TEs-PASA.gff3</strong> Ugibba GFF3 file fully annotated (genes only)<br> <strong>Ugibba_export.txt</strong> Blast2GO/OmicsBox full exported table<br> <strong>Ugibba_fasta.fasta</strong> Blast2GO/OmicsBox Ugibba fasta proteins containg annotation (product and GO terms)<br> <strong>ugibba_frozen_cleaned-validated.box</strong> Full Blast2GO/OmicsBox file</p> <p><strong>ugibba_frozen.box</strong> Full Blast2GO/OmicsBox file (containing TEs genes annotation)</p> <p><strong>ugibba_nogs_emapper_annotations.box</strong> Full Blast2GO/OmicsBox EggNOG file (containing TEs genes annotation)</p> <p><strong>Ugibba_GAF.txt</strong> GAF file<br> <strong>Ugibba</strong><strong>.gene</strong> Ugibba gene fasta file<br> <strong>Ugibba_GOstat.txt </strong> GOstat file<br> <strong>Ugibba</strong><strong>-PASA-assemblies.fasta </strong> Ugibba PASA assemblies<br> <strong>Ugibba</strong><strong>-PASA.stats </strong> Ugibba annotation STATS<br> <strong>Ugibba</strong><strong>.</strong><strong>prot</strong><strong> </strong> Ugibba protein fasta file<br> <strong>Ugibba</strong><strong>-RepeatMasker.gff </strong> Ugibba RepeatMasker gff file<br> <strong>Ugibba</strong><strong>-RepeatMasker.gff3 </strong> Ugibba RepeatMasker gff3 file<br> <strong>Ugibba</strong><strong>-RepeatMasker.tbl </strong> Ugibba RepeatMasker results<br> <strong>Ugibba</strong><strong>-RepeatMasker-v2.gff3</strong> Ugibba RepeatMasker gff3 second version file<br> <strong>Ugibba</strong><strong>-RNAseq-assembled.fasta </strong> Ugibba RNAseq assembled transcriptome (Trinity)<br> <strong>Ugibba_TEs_DANTE_2019.fa </strong> Ugibba TEs library, detected by REPET and annotated by PASTEC and DANTE<br> <strong>Ugibba_WEGO.txt </strong> WEGO file</p> <p><strong>----------------------------------------------------------------------------------------------------------------------------------------------<br> <em>Utricularia reniformis</em> files<br> ----------------------------------------------------------------------------------------------------------------------------------------------</strong><br> <strong>Urenif</strong><strong>-no-masked.fa </strong> Urenif genome excluding organellar genomes<br> <strong>Urenif</strong><strong>-</strong><strong>softmasked</strong><strong>.fa</strong> Urenif genome RepeatMasker softmasked and excluding organellar genomes<br> <strong>Ur.collinearity </strong> MCScanX collinearity file<br> <strong>Ur-duplicates.txt </strong> MCScanX duplicate_gene_classifier short report<br> <strong>Ur.gene_type</strong> MCScanX duplicate_gene_classifier full report<br> <strong>Ur.tandem</strong> Urenif tandem genes generated by MCScanX tool<br> <strong>Urenif_annot.annot</strong> Blast2GO/OmicsBox annotation file (eudicotyledons filtered and Viridiplantae GOSlim)<br> <strong>Urenif_annot-</strong><strong>noclean</strong><strong>.</strong><strong>annot</strong> Blast2GO/OmicsBox annotation file (not filtered)<br> <strong>Urenif</strong><strong>.cDNA</strong> Urenif cDNAs fasta file<br> <strong>Urenif</strong><strong>.CDS </strong> Urenif cDNAs fasta file<br> <strong>Urenif</strong><strong>-EVM.all-no-TEs-PASA-ANNOTATED.gff3</strong> Urenif GFF3 file fully annotated (including gene products and GO terms)</p> <p><strong>Urenif</strong><strong>-EVM.all-no-TEs-PASA.gff3</strong> Urenif GFF3 file fully annotated (genes only)<br> <strong>Urenif_export.txt</strong> Blast2GO/OmicsBox full exported table<br> <strong>Urenif_fasta.fasta</strong> Blast2GO/OmicsBox Urenif fasta proteins containg annotation (product and GO terms)<br> <strong>urenif_frozen_cleaned-validated.box</strong> Full Blast2GO/OmicsBox file</p> <p><strong>urenif_frozen.box</strong> Full Blast2GO/OmicsBox file (containing TEs genes annotation)</p> <p><strong>urenif_nogs_emapper_annotations.box</strong> Full Blast2GO/OmicsBox EggNOG file (containing TEs genes annotation)<br> <strong>Urenif_GAF.txt </strong> GAF file<br> <strong>Urenif</strong><strong>.gene</strong> Urenif gene fasta file<br> <strong>Urenif_GOStat.txt </strong> GOstat file<br> <strong>Urenif</strong><strong>-PASA-assemblies.fasta</strong> Urenif PASA assemblies<br> <strong>Urenif</strong><strong>-PASA.stats </strong> Urenif annotation STATS<br> <strong>Urenif</strong><strong>.</strong><strong>prot</strong><strong> </strong> Urenif protein fasta file<br> <strong>Urenif</strong><strong>-RepeatMasker.gff </strong> Urenif RepeatMasker gff file<br> <strong>Urenif</strong><strong>-RepeatMasker.gff3 </strong> Urenif RepeatMasker gff3 file<br> <strong>Urenif</strong><strong>-RepeatMasker.tbl </strong> Urenif RepeatMasker results<br> <strong>Urenif</strong><strong>-RepeatMasker-v2.gff3 </strong> Urenif RepeatMasker gff3 second version file<br> <strong>Urenif</strong><strong>-RNAseq-assembled.fasta </strong> Urenif RNAseq assembled transcriptome (Trinity)<br> <strong>Urenif_TEs_DANTE_2019.fa </strong> Urenif TEs library, detected by REPET and annotated by PASTEC and DANTE<br> <strong>Urenif_WEGO.txt </strong> WEGO file<br> <strong>----------------------------------------------------------------------------------------------------------------------------------------------<br> ----------------------------------------------------------------------------------------------------------------------------------------------</strong></p>
Sardinops sagax genome assemblies and annotations
<p>Files included are the genome assemblies for each haplotype (hap 1 and hap 2) of Sardinops sagax and the corresponding annotation files for each haplotype.</p> <p> </p> <p> </p>
LyBar v2.0 Genome Assembly and Annotation for Lycium barbarum
<p>LyBar v2.0 genome assembly and annotation files for <em>Lycium barbarum</em>.</p>
Gene Annotations of 49 Bacillariophyta Genome Assemblies (Individual gff3 files)
<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div> </div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a> . It is a copy of the data hostet at <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a> , but instead of storing one archive will all gff3 files included, the gff3 files are here hosted, individually. This copy was made upon request from the RDA Working Group "FAIRification of Genomic Annotations – metadata harmonisation at scale".</p> <div> <h2>Files</h2> <div>The following gzip-compressed gff3-files with structural and functional genome annotation are included:</div> <div> </div> <div>Asterionella_formosa.gff3.gz<br>Asterionellopsis_glacialis.gff3.gz<br>Bacterosira_constricta.gff3.gz<br>Chaetoceros_muellerii.gff3.gz<br>concatenated_output.gff3.gz<br>Conticribra_guillardii.gff3.gz<br>Conticribra_weissflogii.gff3.gz<br>Craspedostauros_australis.gff3.gz<br>Cyclostephanos_invisitatus.gff3.gz<br>Cyclostephanos_tholiformis.gff3.gz<br>Cyclotella_atomus.gff3.gz<br>Cyclotella_baltica.gff3.gz<br>Cyclotella_choctawhatcheeana.gff3.gz<br>Cyclotella_cryptica.gff3.gz<br>Cylindrotheca_fusiformis.gff3.gz<br>Detonula_confervacea.gff3.gz<br>Discostella_pseudostelligera.gff3.gz<br>Discostella_stelligera.gff3.gz<br>Discostella_stelligeroides.gff3.gz<br>Epithemia_pelagica.gff3.gz<br>Fistulifera_pelliculosa.gff3.gz<br>Fistulifera_solaris.gff3.gz<br>Fragilaria_radians.gff3.gz<br>Fragilariopsis_cylindrus.gff3.gz<br>Licmophora_abbreviata.gff3.gz<br>Mediolabrus_comicus.gff3.gz<br>Nitzschia_palea.gff3.gz<br>Nitzschia_putrida.gff3.gz<br>Porosira_glacialis.gff3.gz<br>Psammoneis_japonica.gff3.gz<br>Pseudo-nitzschia_multiseries.gff3.gz<br>Pseudo-nitzschia_pungens.gff3.gz<br>Skeletonema_costatum.gff3.gz<br>Skeletonema_marinoi.gff3.gz<br>Skeletonema_menzelii.gff3.gz<br>Skeletonema_potamos.gff3.gz<br>Skeletonema_tropicum.gff3.gz<br>Stephanocyclus_meneghinianus.gff3.gz<br>Stephanodiscus_minutulus.gff3.gz<br>Stephanodiscus_triporus.gff3.gz<br>Thalassiosira_allenii.gff3.gz<br>Thalassiosira_delicatula.gff3.gz<br>Thalassiosira_exigua.gff3.gz<br>Thalassiosira_gravida.gff3.gz<br>Thalassiosira_livingstoniorum.gff3.gz<br>Thalassiosira_mediterranea.gff3.gz<br>Thalassiosira_oceanica.gff3.gz<br>Thalassiosira_ordinaria.gff3.gz<br>Thalassiosira_pacifica.gff3.gz<br>Thalassiosira_profunda.gff3.gz</div> <div> </div> <div>To extract individual files after download execute the following command:</div> <div> </div> <div><code>gunzip *.gff3.gz</code></div> <h2>Genome Assemblies</h2> <p> </p> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <p> </p> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p> </p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p> </p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p> </p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^>/ s/ .*//' genome.fasta > genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <p> </p> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <p> </p> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>The submission and release was made upon request of the RDA working group "FAIRification of Genomic Annotations – metadata harmonisation at scale". The contained data is identical to <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a></p> <h2>License</h2> <p> </p> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div> <p> </p> </div> </div>
Fraxinus pennsylvanica genome assembly and annotation
<p>We report the first chromosome-level assembly for green ash (<em>Fraxinus pennsylvanica</em>) to assist in ash breeding efforts to propagate resistance to the emerald ash borer. The final haploid assembly consists of 23 chromosomes and 87 unplaced scaffolds of 10 kb or more. Over 99% of the bases anchored to the chromosomes. The assembly spans 757 Mb and consists of 49.43% repetitive DNA. Gene annotation yielded 35,470 high-confidence gene models, all located on the chromosomes and assigned to 22,976 Asterid Orthogroups.</p>
Mycobacteroides abscessus subp. bolletii strain associated with a persistent infection (genome assembly and annotation dataset)
<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) of a <strong><em>Mycobacteroides abscessus subp. bolletti </em></strong>strain associated with a persistente infection. (genome anotation was performed using Bakta v1.2.2 https://github.com/oschwengers/bakta)</p> <p>The raw sequence reads were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB57933; Run Accession: ERR10554471).</p>
Genome assembly and annotation files for Corylus americana accessions 'Rush' and 'Winkler'
<p>The native shrub American hazelnut (<em>Corylus americana</em>) is currently used in breeding programs that are aiming to develop commercially viable hazelnut varieties for the U.S. Upper Midwestern U.S. This species provides significant ecological benefits as it is a perennial crop and well-adapted to this region. Breeding cycles for perennial species are long, and may benefit from the use of predictive methods such as genomic selection to reduce cycle time and increase the efficiency of field trials.</p> <p>High-quality reference genome assemblies are very useful for the implementation marker-assisted selection and genomic prediction, and we therefore developed the first chromosome-scale reference assemblies for <em>C. americana</em>, using the accessions 'Rush' and 'Winkler'. Initial draft assemblies were created using HiFi PacBio reads and Arima Hi-C sequencing to assemble genomes into 11 pseudomolecules. We then utilized Oxford Nanopore reads and a high-density genetic map in order to perform error correction. N50 scores were calculated to be 31.9 Mb and 35.3 Mb for 'Rush' and 'Winkler', respectively, while 97.1% (for 'Winkler') and 90.2% (for 'Rush') of the total genome was assembled into the 11 pseudomolecules. Gene prediction was performed using both RNAseq libraries as well as protein homology data. 'Rush' had a BUSCO score of 99.0 for its assembly and 99.0 for its annotation, while 'Winkler' had corresponding scores of 96.9 and 96.5, indicating extremely high-quality assemblies.</p> <p>These two independent, de novo assemblies enable unbiased assessment of structural variation across the genome, as well as patterns of syntenic relationships within C. americana and the <em>Corylus</em> genus. These assemblies are also an important first step in providing a resource for using next-generation sequencing data in the improvement of <em>C. americana</em>. We demonstrate this utility through the generation of high-density SNP marker sets from genotyping-by-sequencing data for 1,343 <em>C. americana</em>, <em>C. avellana</em>, and <em>C. americana</em> x <em>C. avellana</em> hybrids, in order to assess population structure in natural and breeding populations. Finally, the transcriptomes of these assemblies, as well as several other recently published <em>Corylus</em> genomes, were utilized to perform phylogenetic analysis of sporophytic self-incompatibility (SSI) in hazelnut, providing further evidence of unique molecular pathways governing self-incompatibility in Corylus not exhibited in other well-studied SSI systems. We hope these assemblies will aide in the application of modern breeding methods to the development of commercially viable hazelnut varieties for the U.S. Upper Midwest.</p>
Picea mariana isolate 40-10-1 genome annotation
<p>Genome annotation of <em>Picea mariana</em> isolate 40-10-1. Gene models were identified using BRAKER v2.1.6 and functionally annotated with EnTAP v0.10.8. </p> <p>pmariana-v1.gff - Genome annotation in GFF3 format<br> pmariana-v1_proteins.faa - FASTA file of translated protein-coding sequences<br> pmariana-v1_transcripts.fa - FASTA file of transcripts (CDS)</p>
Genome annotations of darkbarbel catfish (Pelteobagrus vachelli)
<p>Genome annotations (Pva.annotation.v1.gff.gz), along with predicted coding sequences (Pva.annotation.v1.cds.fa.gz) and protein sequences (Pva.annotation.v1.pep.fa.gz) of darkbarbel catfish genome assembly.</p>
Annotation of inverted repeats displaying features of pble STIR or IR in the hg38 genome model
<p>Annotation of inverted repeats displaying features of pble STIR or IR in the hg38 genome model. The annotation of <em>pble</em>-like inner inverted repeats was done using Palindrome (EMBOSS package). The output file was then filtered using pal2gff (https://github.com/Leelouh/pal2gff/blob/main/pal2gff.py), using as parameters a repeat size between 5 and 15 nucleotides, a spacer between pairs of inverted repeats (IRs) of 2 to 10 nucleotides, and a number of mismatches within repeats ranging from 0 to 1. These parameters were chosen taking into account those of the inner IRs found at ends of invertebrate pbles.</p>
Training data for 'Genome annotation with Maker' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for genome annotation with Maker.</p> <p>It is based on data used in <a href="http://weatherby.genetics.utah.edu/MAKER/wiki/index.php/MAKER_Tutorial_for_WGS_Assembly_and_Annotation_Winter_School_2018">another Maker tutorial</a>.</p> <p>The full genome was <a href="https://www.ncbi.nlm.nih.gov/genome/?term=Schizosaccharomyces%20pombe[Organism]&cmd=DetailsSearch">downloaded from NCBI</a>, and mitochondria sequence removed from it for simplicity.</p>
Worldwide Fraxinus Genome Assemblies, Annotations, and Gene Families v0.2
<p>The worldwide <em>Fraxinus</em> genome project was conducted to assess the pathogenic resistance of 34 Ash tree species to Ash Dieback and Emerald Ash Borer. The project is led by Dr. Richard Buggs. v0.1 genomes are available at ashgenome.org and on ENA. As part of Josiah Seaman's PhD thesis, he improved the assembly of 13 genomes included here as v0.2. New de novo annotations, gene families, and all associated files are included for future studies and reproducibility. </p> <p>Annotations are done with GeMoMa using F. excelsior as a reference (Keilwagen et al. 2016). Gene families are defined as genes originating from a single copy at the last common ancestor with Solanum. Orthofinder outputs reconciled gene trees, aligned CDS, and gene families (Emms and Kelly 2015; Tekaia 2016). Species tree was calibrated based on fossil evidence using r8s, RAxML across 25,182,399 sites (SpeciesTreeAlignment.fa). More methods details can be found in the full Chapter two of Josiah Seaman's PhD thesis (2021).</p> <p>I'd be happy to talk with you if you'd like any additional information or help visualizing your genomic data. You can find the tools used to browse this data at https://fluentdna.com/ and http://graphgenome.org/ Contact me at josiah@newline.us</p>
COMBAT TB Tuberculosis genome annotation database
<p>A Neo4j (version 2.3.3) format graph database containing annotation related to the M. tuberculosis H37Rv genome, created as part of the COMBAT TB project at the South African National Bioinformatics Institute.</p>
Dataset for "Bacterial genome annotation" and "AMR gene detection" workflows
<p>This dataset is associated with the workflows "Bacterial genome annotation" and "AMR gene detection in an assembled bacterial genome".</p>
Genbank annotation and sequence of the genome of a Rickettsiales symbiont of Reticulomyxa filosa
<p>The genome sequence of a Rickettsiales symbiont was selectively from the genome sequencing reads of its host, Reticulomyxa filosa. Those sequences were obtained from a previous third-party study (doi: 10.1016/j.cub.2013.11.027). The selective assembly procedure was based on GC content and coverage of contigs obtained from the total read sets with SPAdes. Full details are provided in the manuscript file.<br>The obtained symbiont sequence was then annotated with Prokka, and the gbk output file selected.</p><p>This genome was obtained and analysed in the context of a larger genome comparative studies on the Rickettsiales, aimed to investigate the evolutionary patterns within the whole lineage.</p>
Genome annotation workflow for Effrenium voratum rt-383
<p>Scripts of complete genome annotation workflow for Effrenium voratum rt-383, associated with the key genome paper (Shah et al., 2024, Massive genome reduction predates the divergence of Symbiodiniaceae dinoflagellates, under review in <em>ISME Journal</em>). An earlier preprint of this manuscript is available at <em>bioRxiv</em>: <a href="https://doi.org/10.1101/2023.03.24.534093" target="_blank" rel="noopener">https://doi.org/10.1101/2023.03.24.534093</a>.</p> <p>See <strong>README_Evrt383.txt</strong> for more detail.</p>
Genome annotation workflow for Effrenium voratum CCMP421
<p>Scripts of complete genome annotation workflow for Effrenium voratum CCMP421, associated with the key genome paper (Shah et al., 2024, Massive genome reduction predates the divergence of Symbiodiniaceae dinoflagellates, under review in <em>ISME Journal</em>). An earlier preprint of this manuscript is available at <em>bioRxiv</em>: <a href="https://doi.org/10.1101/2023.03.24.534093" target="_blank" rel="noopener">https://doi.org/10.1101/2023.03.24.534093</a>.</p> <p>See <span><strong>README_EvCCMP421.txt</strong> </span>for more detail.</p>
Annotation files related to the Telomere-to-Telomere genome assembly of the clubroot pathogen Plasmodiophora brassicae (GCA_036867785.1)
<p>This repository contains annotation files related to the T2T genome aseembly of <em>Plasmodiophora brassicae</em>. Link to the NCBI genome submission- https://www.ncbi.nlm.nih.gov/bioproject/1071157</p> <p><strong>Description of the files :</strong></p> <p><strong>GCA_036867785.1_ULAVAL_Pb3A_genomic.fna</strong> - Soft-masked genome sequence FASTA file representing 20 chromosomes.</p> <p><strong>sequence_report.jsonl</strong> - Detailed information about individual chromosome seqeunce.</p> <p><strong>PBTT_annotation.gtf</strong> - GTF file corresponding to the genomic FASTA file.The GTF file was generated by BRAKER3 and contains information about all possible transcripts.</p> <p><strong>PBTT_CDS_longest_isoform.fasta</strong> - Contains 10521 FASTA sequences representing the CDS of only the longest isoform of the gene models.</p> <p><strong>PBTT_protein_longest_isoform.fasta</strong> - Contains 10521 FASTA sequences representing the amino acid sequences of only the longest isoform of the gene models.</p>
Genome annotation file for a draft genome assembly for Nucella lapillus
<p><span>A male specimen of wild <em>Nucella lapillus</em>, measuring approximately 1.5–3 cm in length, was collected <span>from a rocky shore (mid-upper shore) at low tide from near the quay at <span>Portnahaven, Isle of Islay, Argyll and Bute, Scotland (National Grid Reference NR 16614 51966)</span> on <span>26th June 2023. Genomic DNA was extracted from the non-shell tissue of the specimen, and sequenced using PacBio HIFI and Oxford Nanopore technolgies (ONT) long read sequencing platforms. </span></span>The genome assembly was derived from 40.6 Gb of PacBio HiFi reads (read <span>N50, 11291; N90, 9246</span>), and 61.1 Gb of ONT data (read <span>N50, </span>3643<span>; N90, </span>1546). </span>Annotation of protein-coding genes in the cleaned and masked genome assembly of <em>Nucella lapillus</em> was performed using GALBA v1.0.11, an automated pipeline that uses proteins from a closely related species to assist in the training of gene prediction using AUGUSTUS. Proteins from <em>Rapana venosa</em> were provided for this purpose, and the miniprot option was used to perform the protein-to-genome alignments. Functional annotation of predicted protein-coding genes was performed using eggNOG-mapper v2.1.12, and additionally annotated with best hit BLAST results (v2.16.0) against the proteomes of the following marine gastropod species: <em>Rapana venosa</em>, <em>Littorina. saxatilis</em>, <em>Pomocea canaliculata, Stramonita haemastoma</em> and <em>Haliotis rufescens</em>.</p>
Syzygium jambos genome assembly and annotation
<p><em>De novo</em> genome assembly and annotation of <em>Syzygium jambos</em>.</p> <p>The following files are available:</p> <ul> <li>sjam.fa.gz: reference genome sequence in fasta format</li> <li>sjam.gff3.gz: gene annotation in GFF3 format</li> <li>sjam.gtf.gz: gene annotation in GTF format</li> <li>sjam.tx.fa.gz: transcript sequences in fasta format</li> <li>sjam.cds.fa.gz: coding sequences in fasta format</li> <li>sjam.prot.fa.gz: protein sequences in fasta format</li> <li>sjam.tsv.gz: gene functional annotation in TSV format</li> <li>sjam.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.