Skip to main content
zenodoopen

Personalized genomes for DL models supporting data

<p>Archive of models and data associated with our manuscript <a href="https://www.biorxiv.org/content/10.1101/2024.10.15.618510v1">"Training deep learning models on personalized genomic sequences improves variant effect prediction"</a>.</p> <p>Code for training and benchmarking LCL models is available at <a href="https://github.com/Danko-Lab/clipnet_ablation">https://github.com/Danko-Lab/clipnet_ablation</a>, whereas code for training and benchmarking K562 models is available at <a href="https://github.com/Danko-Lab/clipnet_k562/">https://github.com/Danko-Lab/clipnet_k562/</a>.</p> <p><strong>Model files &amp; metadata:</strong></p> <ul> <li><strong>n{i}_run{j}.tar</strong> <ul> <li>CLIPNET LCL models trained on i individuals</li> </ul> </li> <li><strong>subsample_individuals_ids.tar</strong> <ul> <li>text files containing lists of the individuals used to train the above models.</li> </ul> </li> <li><strong>reference_models.tar</strong> <ul> <li>CLIPNET LCL model trained on data from 67 PRO-cap libraries, but using hg38 sequences instead of personal genomes.</li> </ul> </li> <li><strong>clipnet_k562_reference.tar</strong><br> <ul> <li>hg38-trained model described above transfer learned to K562.</li> </ul> </li> </ul> <p><strong>Benchmark data:</strong></p> <ul> <li><strong>across_loci_metrics.tar</strong> <ul> <li>benchmarks of LCL models at predicting transcription initiation at individual CREs within the genome</li> </ul> </li> <li><strong>qtl_metrics.tar</strong> <ul> <li>benchmarks of LCL models at predicting differences in transcription initiation between individuals at initiation QTLs</li> </ul> </li> <li><strong>k562_data.tar</strong><br> <ul> <li>benchmarks of the reference-trained K562 model and <a href="https://zenodo.org/records/11196189">one transferred over from the personalized CLIPNET model</a> on MPRA data from <a href="https://www.biorxiv.org/content/10.1101/2024.05.05.592437v1">https://www.biorxiv.org/content/10.1101/2024.05.05.592437v1</a></li> </ul> </li> </ul>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0