Benchmark datasets for testing AIRR-seq data processing with pyIgMap pipeline
<p>This is a set of raw FASTQ files produced by various AIRR-seq protocols, stored here for benchmark conveinience and for reference purposes.</p> <p>Currently it contains the following datasets:</p> <table> <tbody> <tr> <td><strong>id</strong></td> <td><strong>fastq</strong></td> <td><strong>reference</strong></td> <td><strong>method</strong></td> <td><strong>description</strong></td> </tr> <tr> <td>allergy</td> <td> <p>ERR7425614_1.fastq.gz,</p> <p>ERR7425614_2.fastq.gz</p> </td> <td>https://doi.org/10.7554/eLife.79254</td> <td>5'RACE, UMI, long MiSeq reads, IGH w/ isotype</td> <td>Longitudinal full-length IGH repertoire profiling and clonal lineage dynamics in memory B cells, plasmablasts and plasma cells of human peripheral blood</td> </tr> <tr> <td>covid</td> <td> <p>fmba_TRAB_R1.fastq.gz,</p> <p>fmba_TRAB_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1101/2023.11.08.566227</p> </td> <td>DNA multiplex, UMI, TRA+TRB mix, NextSeq</td> <td>TCR sequencing in COVID-19 convalescent and healthy donors</td> </tr> <tr> <td>natprot</td> <td> <p>PMID27490633_R1.fastq.gz,</p> <p>PMID27490633_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1038/nprot.2016.093</p> </td> <td>5'RACE, UMI, long MiSeq reads, high-quality overlap, IGH no isotype</td> <td>High-quality full-length immunoglobulin profiling with unique molecular barcoding</td> </tr> <tr> <td>brnaseq</td> <td> <p>SRR3743469_R1.fastq.gz,</p> <p>SRR3743469_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1016/j.immuni.2016.08.012</p> </td> <td>RNA-Seq, all chains, B-cells</td> <td>Primary human mature naïve B-cells (IgD+CD38lo; NB) and GCB-cells (CD77+CD38hi; GCB) were purified from tonsils of healthy individuals. RNA-seq libraries were prepared using the Illumina TruSeq RNA sample kits according to the manufacturer.</td> </tr> <tr> <td>uhrr</td> <td> <p>UHRR_full_R1.fastq.gz,</p> <p>UHRR_full_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1038/s41598-021-04583-z</p> </td> <td>RNA-Seq, bulk</td> <td>Universal Human Reference RNA</td> </tr> <tr> <td>10x</td> <td> <p>10x_bcr_R1.fastq.gz,</p> <p>10x_bcr_R2.fastq.gz,</p> <p>10x_tcr_R1.fastq.gz,</p> <p>10x_tcr_R2.fastq.gz</p> </td> <td> <p>https://www.10xgenomics.com/datasets/human-pbmc-from-a-healthy-donor-10-k-cells-v-2-2-standard-5-0-0</p> </td> <td>10x Genomics vdj</td> <td>See reference. AIRR-seq is split into TCR and BCR parts</td> </tr> </tbody> </table> <p> </p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0