Skip to main content
zenodoopen

Intergenic RNAPII Atlas : output data

<p>This dataset&nbsp;represents the RNAPII (RNAP2)&nbsp;Atlas of potentially transcribed intergenic regions of the human genome by integrating 906 high quality human Chromatin-ImmunoPrecipitation sequencing (ChIP-seq) biosamples targeting the RNA Polymerase II, obtained from public data warehouses.&nbsp;</p> <p><strong>Github Code available here :&nbsp;&nbsp;</strong><a href="https://github.com/benoitballester/Pol2Atlas">https://github.com/benoitballester/Pol2Atlas</a>&nbsp;</p> <p><strong>The dataset consists of 5 zipped&nbsp;folders described&nbsp;below:&nbsp;</strong></p> <p><strong>./pol2_consensuses/:</strong><br> &nbsp; &nbsp; consensuses.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; Location of intergenic RNAP2 consensuses in bed format for the hg38 assembly.&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; First three columns are genomic locations, 4th column is consensus ID,&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; 5th column is the number of datasets with RNAP2 observed at this RNAP2 consensus,<br> &nbsp; &nbsp; &nbsp; &nbsp; 6th column is strand (not used), 7-8th columns is consensus centroid.<br> &nbsp; &nbsp; consensusesHg19.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; Location of intergenic RNAP2 consensuses in hg19 assembly. ~1000 are missing due to liftover.<br> &nbsp; &nbsp; &nbsp; &nbsp; Consensus ID is matching with the hg38 one.<br> &nbsp; &nbsp; matrix.mtx:<br> &nbsp; &nbsp; &nbsp; &nbsp; RNAP2 occupancy consensus-dataset binary matrix in sparse matrix market format.<br> &nbsp; &nbsp; &nbsp; &nbsp; Corresponding row annotation are RNAP2 consensuses.<br> &nbsp; &nbsp; &nbsp; &nbsp; Corresponding column annotation are datasets stored in dataset.txt.<br> &nbsp; &nbsp; datasets.txt:<br> &nbsp; &nbsp; &nbsp; &nbsp; See matrix.mtx<br> &nbsp; &nbsp; clusterConsensuses_Labels.txt:<br> &nbsp; &nbsp; &nbsp; &nbsp; Assigned cluster for each RNAP2 consensus.<br> &nbsp; &nbsp; intersectIntergPol2.tsv:<br> &nbsp; &nbsp; &nbsp; &nbsp; RNAP2 consensuses with cluster ID and intersections with reference databases.<br> &nbsp; &nbsp; cluster_bed/:<br> &nbsp; &nbsp; &nbsp; &nbsp; consensuses.bed splitted per cluster.<br> &nbsp; &nbsp; saf_files/:<br> &nbsp; &nbsp; &nbsp; &nbsp; Files typically used for read counting with featureCounts. Suffixes:<br> &nbsp; &nbsp; &nbsp; &nbsp; _500 : RNAP2 consensuses standardized to 1kbp.<br> &nbsp; &nbsp; &nbsp; &nbsp; _all : All RNAP2 consensuses including genic.<br> &nbsp; &nbsp; &nbsp; &nbsp; Hg19 : Intergenic RNAP2 Lifted to Hg19.</p> <p>&nbsp;</p> <p><strong>./rnap2_all_peaks/:</strong><br> &nbsp; &nbsp; all_peaks.bed.gz:<br> &nbsp; &nbsp; &nbsp; &nbsp; Concatenated bed file with all POLR2A peaks from all experiments, genome wide, for the hg38 assembly.&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; Peaks are filtered with a MACS2 qvalue &gt; 1e-5, datasets with less than 100 peaks in intergenic regions are removed.<br> &nbsp; &nbsp; &nbsp; &nbsp; First three columns are genomic locations, 4th column contains sample of origin of the peak, 5th column is the<br> &nbsp; &nbsp; &nbsp; &nbsp; MACS2 q-value, 6th column is dna strand (not used), 7-8th are peak &quot;summit&quot;. 9th column contains an r,g,b value<br> &nbsp; &nbsp; &nbsp; &nbsp; corresponding to the biotype of origin (Blood / Immune, Brain, Embryo...) for easy visualization in a genome browser.<br> &nbsp; &nbsp; &nbsp; &nbsp; Legend is available in legend.png. Conversion table between rgb values and biotype in palette.csv.<br> &nbsp; &nbsp; &nbsp; &nbsp; Note that singletons are removed when creating consensus peaks.<br> &nbsp; &nbsp; all_peaks_interg.bed.gz:<br> &nbsp; &nbsp; &nbsp; &nbsp; Same as above, but for intergenic regions only (excluding 1kb before TSS and 1kb after TES).</p> <p>&nbsp;</p> <p><strong>./count_tables_rnaseq/:</strong><br> &nbsp; &nbsp; ENCODE/:<br> &nbsp; &nbsp; &nbsp; &nbsp; counts.mtx.gz:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Count table in sparse matrix market format. Row corresponds to samples, columns to Pol II probes (Pol2_500.saf).<br> &nbsp; &nbsp; &nbsp; &nbsp; samples.csv.gz:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Matching row annotation for count matrix.<br> &nbsp; &nbsp; &nbsp; &nbsp; encode_total_rnaseq_annot.tsv.gz:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Sample annotation (not ordered!).<br> &nbsp; &nbsp; GTEx/:<br> &nbsp; &nbsp; &nbsp; &nbsp; counts.mtx.gz:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Count table in sparse matrix market format. Row corresponds to samples, columns to Pol II probes (Pol2_500.saf).<br> &nbsp; &nbsp; &nbsp; &nbsp; samples.csv.gz:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Matching row annotation for count matrix.<br> &nbsp; &nbsp; &nbsp; &nbsp; sample_annot.tsv.gz:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Sample annotation (not ordered!).<br> &nbsp; &nbsp; TCGA/:<br> &nbsp; &nbsp; &nbsp; &nbsp; counts.mtx.gz:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Count table in sparse matrix market format. Row corresponds to samples, columns to Pol II probes (Pol2_500.saf).<br> &nbsp; &nbsp; &nbsp; &nbsp; samples.csv.gz:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Matching row annotation for count matrix.<br> &nbsp; &nbsp; &nbsp; &nbsp; annotation_table.tsv.gz:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Sample annotation (not ordered!).<br> &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</p> <p><br> <strong>./cancer_markers/:</strong><br> &nbsp; &nbsp; bed/:<br> &nbsp; &nbsp; &nbsp; &nbsp; DE_Tumor_vs_Normal/:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; TCGA-*/:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; allWithStats.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; FDR, mean difference in pearson residuals and log2(FC) for each RNAP2 probe. Warning: probes are prefiltered to have &gt; 1 read in 3 samples, make sure to use row index to match&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; with RNAP2 consensuses.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; allDE.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; All DE (cancer vs normal) probes in bed format for this cancer.&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 5th column has been replaced by enrichment p-value.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DE_downreg.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Downregulated (cancer vs normal) probes in bed format for this cancer.&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 5th column has been replaced by enrichment p-value.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DE_upreg.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Upregulated (cancer vs normal) probes in bed format for this cancer.&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 5th column has been replaced by enrichment p-value.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; classifier_TCGA-*:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Performance of a machine learning tumor-normal tissues classifier using Pol II probes as input.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; globally_DE.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Probes DE in 7+ cancers (FPR permutation threshold). Last column indicates the number of cancers this probe is DE in.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; globally_Down regulated.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Probes DE in 6+ cancers (FPR permutation threshold). Last column indicates the number of cancers this probe is Down regulated in.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; globally_Up regulated.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Probes DE in 5+ cancers (FPR permutation threshold). Last column indicates the number of cancers this probe is Up regulated in.<br> &nbsp; &nbsp; &nbsp; &nbsp; subtypes/:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; BRCA/:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; allWithStats_BRCA.*.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; FDR, mean difference in pearson residuals and log2(FC) for each RNAP2 probe for DE test of sample from this subtype against normal samples.&nbsp;<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp;Warning: probes are prefiltered to have &gt; 1 read in 3 samples, make sure to use row index to match&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; with RNAP2 consensuses.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; bed_BRCA.*.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; All DE (subtype vs normal) probes in bed format for this cancer.&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; bed_uniqueDE_BRCA.*.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; All DE (subtype vs normal) probes in bed format for this cancer and not DE in any other subtype.&nbsp;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; TCGA_survival/:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; TCGA-*/:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; prognostic.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; All probes associated with survival for this cancer.&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 5th column has been replaced by p-value.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; stats.csv:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Cox linear model statistics for each Pol II probe.&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Warning: probes are prefiltered to have &gt; 1 read in 3 samples, make sure to use row index to match&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; with RNAP2 consensuses.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; globally_prognostic.bed:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Probes associated with survival in 5+ cancers (FPR permutation threshold). 5th column has been replaced with&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; the number of cancers this probe is associated with survival in.<br> &nbsp; &nbsp; tabular/:<br> &nbsp; &nbsp; &nbsp; &nbsp; Same as above but stored in a tabular binary format for DE and survival.</p> <p>&nbsp;</p> <p><strong>./metacluster_markers/:</strong><br> &nbsp; &nbsp; bed/:<br> &nbsp; &nbsp; &nbsp; &nbsp; allPol2_datasetCount:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; For each tissue, all Pol II consensuses, with 5th column indicating the number of<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; datasets (RNAP2, GTEx, ENCODE, TCGA tumour and normal) in which the RNAP2 consensus is<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; considered a marker.<br> &nbsp; &nbsp; &nbsp; &nbsp; robust_2_datasets_per_tissue:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; For each tissue, Pol II consensuses considered marker in 2+ datasets out of 5&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; (RNAP2, GTEx, ENCODE, TCGA tumour and normal).<br> &nbsp; &nbsp; tabular/:<br> &nbsp; &nbsp; &nbsp; &nbsp; Each Pol II consensus with marker information stored in a binary format.<br> &nbsp;&nbsp;</p>

ShareScore

28/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
0
Engagement
0

Topics