Skip to main content
zenodoopen

Supporting data for Novel functional sequences uncovered through a bovine multi-assembly graph

<p><strong>Description of the datasets</strong></p> <p>Data are organized as a folder&nbsp;and compressed with tar.gz.</p> <p>You need to unzip the folder using the command <em>tar -xz</em><em>v</em><em>f</em> data.tar.gz. Unzipping will output a folder named <em>data_tidy</em>, which is organized as follow:</p> <ul> <li>graph.gfa : Graph in GFA format constructed from 6 cattle assemblies</li> <li>nonref.fa : Non-reference sequences extracted from the graph</li> <li>nonref.fa.masked: Hard masked repetitive regions version of nonref.fa</li> <li>nonref_woflanking.fa: Nonref.fa without flanking sequences</li> <li>nonref_woflanking.fa.masked: Masked version of nonref_woflanking.fa</li> <li>augustus_predict.gtf: Annotated gene models of Augustus from non-ref sequences</li> <li>augustus_prot.fa: Protein fasta of the predicted gene models from Augustus</li> <li>breeds_assembled.gtf: Annotation of the StringTie assembled across-breed transcriptome</li> <li>breeds_expressed.tsv: Expression data of breeds_assembled.gtf</li> <li>de_assembled.gtf: Annotation of the StringTie assembled differentially-expressed transcriptome on non-ref sequences</li> <li>de_expression.tsv: Differential expression results from de_assembled.gtf</li> <li>variant_nonref.tsv: Variants called from non-ref sequences (-1, 0, 1, 2 indicates no call, hom ref, het, and hom alt respectively)</li> </ul>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4

Topics