BPNet manuscript data 1
<p>This repository holds the trained BPNet models and all the output files produced using the models including TF-MoDISco files and motif instances.</p> <p><strong>Files:</strong></p> <ul> <li><strong>output.tar.gz</strong> <ul> <li>Full output directory containing trained BPNet models, contribution scores, TF-MoDISco results, motif instances, etc.</li> <li>See <a href="https://github.com/kundajelab/bpnet-manuscript">https://github.com/kundajelab/bpnet-manuscript</a> for more information (archived at <a href="https://zenodo.org/record/4294814">https://zenodo.org/record/4294814</a> with DOI: <a href="https://doi.org/10.5281/zenodo.4294813">https://doi.org/10.5281/zenodo.4294813</a>).</li> </ul> </li> <li><strong>bpnet.model.h5</strong> <ul> <li>BPNet Keras model. This model can be directly used via the model repository Kipoi: <a href="http://kipoi.org/models/BPNet-OSKN">http://kipoi.org/models/BPNet-OSKN</a> . </li> </ul> </li> <li><strong>regions.bed</strong> <ul> <li>ChIP-nexus peak regions (resized to 1 kb width) where BPNet was trained and interpreted on.</li> <li>Third column denotes the TF for which the peak was called.</li> </ul> </li> <li><strong>motif-instances.bed</strong> <ul> <li>BPNet motif instances for 11 representative motifs</li> <li>Stored as BED file with the following columns <ul> <li>Chromosome</li> <li>Start</li> <li>End</li> <li>Motif name</li> <li>'match' score - similarity to the CWM computed using the continuous Jaccard distance metric between the CWM and L1 normalized contribution scores</li> <li>Strand</li> <li>'contrib' score - computed as the L1 norm of the contribution scores at the motif instance position</li> <li>The log odds score with respect to the PWM derived from the PFM (i.e. classical PWM score)</li> </ul> </li> </ul> </li> <li><strong><TF>.preds.<strand>.bw</strong> <ul> <li>Model predictions in `regions.bed` for a particular TF in {Oct4, Sox2, Nanog, Klf4} and strand in {pos, neg}.</li> </ul> </li> <li><strong><TF>.importance.counts.bw</strong> <ul> <li>Total count contribution scores for a particular TF</li> </ul> </li> <li><strong><TF>.importance.profile.bw</strong> <ul> <li>Profile contribution scores. These were used to run TF-MoDISco and downstream analysis.</li> </ul> </li> <li><strong>PWM,CWM,ChIP-nexus-profiles.tar.gz</strong> <ul> <li>Directory containing PFM, CWM and aggregated ChIP-nexus footprint for 11 representative motifs.</li> </ul> </li> <li><strong>Figure-5d-periodicity.csv</strong> <ul> <li>Raw data corresponding to Figure 5d where the 10bp periodicity of contribution scores was computed for each motif.</li> </ul> </li> <li><strong>dfabf.Oct4-Sox2-subset.parq</strong> <ul> <li>Parquet file (read with `pd.read_parquet`) of genomics motif instance pairs and along with the BPNet-predicted interaction score.</li> <li>See <a href="https://github.com/kundajelab/bpnet-manuscript/blob/dd481a04318dbbd7ea1aa9097ec3b5de8be89f93/src/figures/08-motif-interactions.genomic.ipynb">08-motif-interactions.genomic.ipynb</a> for more information on how to generate plots for Figure 5 from this data.</li> </ul> </li> </ul>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0