Skip to main content
zenodoopen

Assembly data files Oryzias dopingdopingensis

<p>Identification and masking of&nbsp;repetitive&nbsp;elements in the genome sequence of&nbsp;<em>O. dopingdopingensis&nbsp;</em>was performed with the following bioinformatic tool case. Nucleotides were masked using the DUST algorithm with dustmasker (version 1.0.0, part of blast+ 2.9.0&nbsp;(Altschul et al., 1990; Camacho et al., 2009)&nbsp;(Kuzio et al., unpublished but described in&nbsp;(Morgulis et al., 2006). Tandem Repeats were identified with Tandem Repeat Finder (trf version 4.09)&nbsp;(Benson, 1999).&nbsp;A species-specific&nbsp;<em>de novo</em>&nbsp;repeat library was built with RepeatModeler v1.0.11 (<a href="http://www.repeatmasker.org/RepeatModeler/">http://www.repeatmasker.org/RepeatModeler/</a>). Repeat Elements were located in the genome sequence using RepeatMasker (version 4.1.0) (<a href="http://www.repeatmasker.org/">http://www.repeatmasker.org</a>) with the&nbsp;<em>de&nbsp;</em><em>novo</em>&nbsp;and<em>&nbsp;Danio rerio&nbsp;</em>libraries. The information from all four repeat&nbsp;analyses&nbsp;was merged and the genome was softmasked with bedtools (2.29.2)&nbsp;(Quinlan &amp; Hall, 2010)&nbsp;PMID: 20110278; PMCID: PMC2832824.]. All steps of masking&nbsp;repetitive&nbsp;regions were performed with scripts&nbsp;provided by the sigenae platform, following the workflow from&nbsp;(Feron et al., 2020).</p> <p>For the identification of genes the masked genome was annotated with funannotate&nbsp;(Palmer &amp; Stajich, 2019). The sequences were sorted by length with the&nbsp;&lsquo;funannotate sort&rsquo;&nbsp;function,&nbsp;followed by a gene prediction with&nbsp;&lsquo;funannotate predict&rsquo;. No training&nbsp;based on RNA-Seq data was&nbsp;performed since&nbsp;it was not available for&nbsp;this species. Additional external evidence from transcripts and proteins&nbsp;was&nbsp;added. As transcript evidence, gene predictions from&nbsp;<em>Oryzias latipes</em>&nbsp;(NCBI Bioproject:PRJNA183868; Assembly: GCF_002234675.1)&nbsp;(Kasahara et al., 2007)&nbsp;and&nbsp;<em>Oryzias melastigma</em>&nbsp;(NCBI Bioproject: PRJNA401159 ; Assembly: ASM292280v2)&nbsp;(Kim et al., 2018)&nbsp;were used. As protein evidence, a protein set from Oryzias javanicus (NCBI Bioprject : PRJNA505405 ; Assembly: GCA_003999625.1)&nbsp;(Lee et al., 2020), manually annotated reference sequences from UniProt Knowledgebase (UniProtKB) (Release&nbsp;2020_02 (22-Apr-2020) UniProtKB/Swiss-Prot with&nbsp;562,253 entries )&nbsp;(Apweiler et al., 2004)&nbsp;and a set of&nbsp;orthologous&nbsp;sequences generated in this study.&nbsp;Furthermore, the&nbsp;<em>de novo</em>&nbsp;gene&nbsp;predictors&nbsp;were trained with the Busco dataset of actinopterygii_odb10.&nbsp;Gene prediction&nbsp;resulted in a total of 56658 genes.</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0