Single Nucleotide Polymorphisms (SNPs) identified from the whole genome sequences of hilsa shad (Tenualosa ilisha) of the Bay of Bengal
<p>The data file contains 792,939 isolated SNPs identified by discoSnp++ v2.3.x (Uricaru et al., 2015) from the whole genome sequence of T. ilisha of the Bay of Bengal. The central sequence of length 2k-1 is seen in upper case, while the flanking sequences are seen in lower case. SNP_higher/lower: one of the two alleles. id: id of the SNP (each SNP has a unique id).</p> <p>FOR SNPs:</p> <p>P_i:pos_Alt1/Alt2: Information about a ith SNP (If more than a unique SNP is found, the following format is used: P_1:pos_Alt1/Alt2,P_2:pos_Alt1/Alt2,...</p> <p>pos: position of the SNP with respect to the starting position of the bubble, i.e. the starting of the upper case sequence.</p> <p>Alt1: One of the two alleles</p> <p>Alt2: the other</p> <p>FOR INDELs:</p> <p>P_1:pos_size_repeatSize</p> <p>pos: predicted position of the indel with respect to the starting position of the bubble, i.e. the starting of the upper case sequence.</p> <p>size: predicted size of the indel</p> <p>repeatSize: Size of the longest sequence both prefix of the indel and prefix of the sequence located just after the insertion.</p> <p>high/low: sequence complexity. If the sequence if of low complexity (e.g. ATATATATATATATAT) this variable would be low</p> <p>nb_pol: number of polymorphism.</p> <p>left_unitig_length: size of the full left extension.</p> <p>right_unitig_length: size of the right extension.</p> <p>left_contig_length: size of the full left extension.</p> <p>right_contig_length: size of the right extension.</p> <p>C1: number of reads mapping the central upper case sequence from the first read set.</p> <p>C2: number of reads mapping the central upper case sequence from the second read set.</p> <p>Q1 [if reads were given in fastq]: average phred quality of the central nucleotide from the mapped reads from the first read set.</p> <p>Q2 [if reads were given in fastq]: average phred quality of the central nucleotide from the mapped reads from the second read set.</p> <p>G1: Genotype of the variant in the first read set.</p> <p>G2: Genotype of the variant in the second read set.</p> <p>rank: ranks the predictions according to their read coverage in each condition favoring SNPs that are discriminant between conditions.</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0