Training and test data for antibody humanness evaluation
<p>### Training and test data for humanness evaluation</p> <p>This data was collected in conjunction with and used for<br>training and testing for Parkinson / Wang et al 2024. The<br>data is organized as follows:</p> <p>- Heavy chain training and multispecies test data (under the heavy chain folder)<br> - The conslidated cAb rep file contains training human sequences<br> - The test sample sequences folder contains fasta files with test sequences for each species<br>- Light chain training and multispecies test data (under the light chain folder)<br> - The conslidated cAb rep file contains training human sequences<br> - The test sample sequences folder contains fasta files with test sequences for each species<br>- Abybank data (under the abybank compiled data folder)<br> - This folder contains separate folders for heavy and light chain<br> - Each subfolder contains test data for a more diverse species set under fasta files for each species<br>- Humanization test data (under the humanization test data folder)<br> - The sequences in the parental.fa file were originally humanized as part of drug discovery programs<br> - The experimental.fa file contains the humanization results<br>- IMGT and ADA data (under the imgt test data folder)<br> - The imgt mab db fa and tsv files contain sequences and species assignments for IMGT mAb DB<br> - The thera ada fa file contains sequences evaluated in the clinic<br> - The Therapeutic ADA txt file contains anti drug antibody results for those antibodies<br>- VDJ statistics (under the vdj_statistics_eval folder)</p> <p>The data was retrieved from the following sources.</p> <p>1. All heavy and light chain training data is from the cAb-Rep database from [Guo et al.](https://pubmed.ncbi.nlm.nih.gov/31649674/)<br>2. All testing data is from the Observed Antibody Space [(OAS) database](https://opig.stats.ox.ac.uk/webapps/oas/)</p> <p>The training and test data show is after filtering for quality. The testing data was additionally randomly sampled to yield a set of 50,000 sequences for each species, then filtered to remove duplicates. The human test data was checked to ensure no overlap with the human training set.</p> <p><br>The IMGT, ADA and humanization test data was retrieved from Prihoda et al. and<br>the associated [Github repo](https://github.com/Merck/BioPhi-2021-publication).</p> <p>See Parkinson et al. 2024 and the associated github repos for more details on how models other than<br>SAM / AntPack were evaluated on this data.</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0