Skip to main content
zenodoopen

Genome and Transcriptome references based on hg19 from UCSC, 2015

<p>rsem.transcripts.nant2015.fa.gz - bgzipped FASTA reference of transcriptomes</p><p>genome.nant2015.fa.gz - bgzipped FASTA human genome reference, with several viral sequences added.</p><p>refseq.txt.gz - Exact sequence accessions and mapping coordinates for a RefSeq transcriptome based off the UCSC genome browser for hg19.</p><p>Coordinates are BED-style, with one row per transcript, and 1+ transcript per gene.</p><p>Column annotation</p><p>1. RefSeq Accession</p><p>2. Chromosome</p><p>3. Strand</p><p>4. thinStart (gene boundary, including UTR)</p><p>5. thinEnd (gene boundary, including UTR)</p><p>6. thickStart (CDS boundary)</p><p>7. thinStart (CDS boundary)</p><p>8. number of exons</p><p>9. comma separated exon starts</p><p>10. comma separate exon ends</p><p>11. common gene name</p><p>12. refseq gene id</p><p>13. 0 if non-primary transcript, 1 if primary transcript</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
20
Reuse readiness
8
Engagement
0