Assembly data files Oryzias dopingdopingensis
<p>Identification and masking of repetitive elements in the genome sequence of <em>O. dopingdopingensis </em>was performed with the following bioinformatic tool case. Nucleotides were masked using the DUST algorithm with dustmasker (version 1.0.0, part of blast+ 2.9.0 (Altschul et al., 1990; Camacho et al., 2009) (Kuzio et al., unpublished but described in (Morgulis et al., 2006). Tandem Repeats were identified with Tandem Repeat Finder (trf version 4.09) (Benson, 1999). A species-specific <em>de novo</em> repeat library was built with RepeatModeler v1.0.11 (<a href="http://www.repeatmasker.org/RepeatModeler/">http://www.repeatmasker.org/RepeatModeler/</a>). Repeat Elements were located in the genome sequence using RepeatMasker (version 4.1.0) (<a href="http://www.repeatmasker.org/">http://www.repeatmasker.org</a>) with the <em>de </em><em>novo</em> and<em> Danio rerio </em>libraries. The information from all four repeat analyses was merged and the genome was softmasked with bedtools (2.29.2) (Quinlan & Hall, 2010) PMID: 20110278; PMCID: PMC2832824.]. All steps of masking repetitive regions were performed with scripts provided by the sigenae platform, following the workflow from (Feron et al., 2020).</p> <p>For the identification of genes the masked genome was annotated with funannotate (Palmer & Stajich, 2019). The sequences were sorted by length with the ‘funannotate sort’ function, followed by a gene prediction with ‘funannotate predict’. No training based on RNA-Seq data was performed since it was not available for this species. Additional external evidence from transcripts and proteins was added. As transcript evidence, gene predictions from <em>Oryzias latipes</em> (NCBI Bioproject:PRJNA183868; Assembly: GCF_002234675.1) (Kasahara et al., 2007) and <em>Oryzias melastigma</em> (NCBI Bioproject: PRJNA401159 ; Assembly: ASM292280v2) (Kim et al., 2018) were used. As protein evidence, a protein set from Oryzias javanicus (NCBI Bioprject : PRJNA505405 ; Assembly: GCA_003999625.1) (Lee et al., 2020), manually annotated reference sequences from UniProt Knowledgebase (UniProtKB) (Release 2020_02 (22-Apr-2020) UniProtKB/Swiss-Prot with 562,253 entries ) (Apweiler et al., 2004) and a set of orthologous sequences generated in this study. Furthermore, the <em>de novo</em> gene predictors were trained with the Busco dataset of actinopterygii_odb10. Gene prediction resulted in a total of 56658 genes.</p>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0