Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
51
datasets available to search
ShareScore release 0.9.0
Dataset results
51 results for “repertoire sequencing”
PRJNA638224 - BCR repertoire sequencing from COVID-19 patients
<p><strong>Description</strong></p> <p>These are the processed BCR repertoire sequence data that accompany the following manuscript: “Deep sequencing of B cell receptor repertoires from COVID-19 patients reveals strong convergent immune signatures”. The manuscript preprint is available at doi: <a href="https://doi.org/10.1101/2020.05.20.106294">https://doi.org/10.1101/2020.05.20.106294</a>. The raw sequence data are available on SRA under the BioProject PRJNA638224</p> <p> </p> <p><strong>Sequence processing</strong></p> <p>The Immcantation framework (docker container v3.0.0) was used for sequence processing. Briefly, paired-end reads were joined based on a minimum overlap of 20 nt, and a max error of 0.2, and reads with a mean phred score below 20 were removed. Primer regions, including UMIs and sample barcodes, were then identified within each read, and trimmed. Together, the sample barcode, UMI, and constant region primer were used to assign molecular groupings for each read. Within each grouping, usearch, was used to subdivide the grouping, with a cutoff of 80% nucleotide identity, to account for randomly overlapping UMIs. Each of the resulting groupings is assumed to represent reads arising from a single RNA. Reads within each grouping were then aligned, and a consensus sequence determined. For each processed sequence, IgBlast was used to determine V, D and J gene segments, and locations of the CDRs and FWRs. Isotype was determined based on comparison to germline constant region sequences. Sequences annotated as unproductive by IgBlast were removed.</p> <p> </p> <p><strong>Sequence data column description</strong></p> <ul> <li><strong>sample_id </strong>Unique identifier for each sequencing library</li> <li><strong>sequence_id </strong>Unique identifier for a sequence within a sample_id</li> <li><strong>sequence_alignment </strong>IMGT gapped nucleotide sequence</li> <li><strong>germline_alignment </strong>IMGT gapped germline sequence</li> <li><strong>v_call </strong>IGHV gene segment(s) and allele</li> <li><strong>d_call </strong>IGHD gene segment(s) and allele</li> <li><strong>j_call </strong>IGHJ gene segment(s) and allele</li> <li><strong>c_call </strong>Isotype subclass</li> <li><strong>junction </strong>Junction nucleotide sequence</li> <li><strong>junction_aa </strong>Junction amino acid sequence</li> <li><strong>duplicate_count </strong>UMI count for the given unique sequence</li> <li><strong>consensus_count </strong>Raw read count for the given unique sequence</li> </ul> <p> </p> <p><strong>Sequence metadata column description</strong></p> <ul> <li><strong>sample_id </strong>Unique identifier for each sequencing library</li> <li><strong>bioproject_accession </strong>NCBI BioProject accession number</li> <li><strong>biosample_accession </strong>NCBI BioSample accession number</li> <li><strong>sra_accession </strong>NCBI SRA accession number</li> <li><strong>sex </strong>Sex of patient</li> <li><strong>age </strong>Age of patient at time of sampling</li> <li><strong>ethnicity </strong>Ethnicity of patient</li> <li><strong>health_state </strong>One of worsening, stable, or improving</li> </ul>
Immune repertoire sequencing reveals differences in treatment response to camrelizumab plus platinum-based chemotherapy in advanced ESCC
Open the record for dataset details and reuse information.
Pre-processed B cell receptor repertoire sequencing data from BioProject PRJNA527941
<p><strong>Data Processing</strong></p> <p> </p> <p>Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2). Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2)</p> <p> </p> <p><strong>software_versions</strong> pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8</p> <p><strong>quality_thresholds</strong> FilterSeq.py pRESTO Q>20</p> <p><strong>paired_reads_assembly</strong> AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001</p> <p><strong>primer_match_cutoffs</strong> MaskPrimers.py pRESTO C primer & V primer maxerror 0.2</p> <p><strong>consensus_building</strong> BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5</p> <p><strong>collapsing_method</strong> CollapseSeq.py pRESTO</p> <p><strong>germline_database </strong>IMGT</p> <p> </p> <p><strong>Format</strong></p> <p> </p> <p>Processed sequences are provided in a tab delimited file format, including the following annotations:</p> <p> </p> <p><strong>C_CALL </strong>Isotype subclass</p> <p><strong>SEQUENCE_ID </strong>Sequence identifier</p> <p><strong>V_CALL </strong>V segment gene and allele</p> <p><strong>D_CALL </strong>D segment gene and allele</p> <p><strong>J_CALL </strong>J segment gene and allele</p> <p><strong>JUNCTION_LENGTH </strong>Junction length</p> <p><strong>CONSCOUNT </strong>Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence.</p> <p><strong>DUPCOUNT </strong>UMI count for the given unique sequence</p> <p><strong>ISOTYPE </strong>Constant region primer (isotype)</p> <p><strong>MU_COUNT_CDR_R </strong>Number of replacement mutations in CDR region</p> <p><strong>MU_COUNT_CDR_S </strong>Number of silent mutations in CDR region</p> <p><strong>MU_COUNT_FWR_R </strong>Number of replacement mutations in FWR region</p> <p><strong>MU_COUNT_FWR_S </strong>Number of silent mutations in FWR region</p> <p><strong>MUT_TOTAL </strong>Total number of mutations in V gene </p> <p><strong>SEQUENCE_INPUT </strong>Full length sequence</p> <p><strong>SEQUENCE_IMGT </strong>Gapped IMGT sequence</p> <p><strong>V_GERM_START_VDJ </strong>position of the first nucleotide in ungapped V germline sequence alignment</p> <p><strong>JUNCTION </strong>Junction nucleotide sequence</p> <p><strong>GERMLINE_IMGT_D_MASK </strong>IMGT-gapped germline nucleotide sequence with ns masking the NP1-D-NP2 regions</p> <p><strong>Run </strong>ID of sequencing run</p> <p><strong>Sample_type </strong>The tissue sampled (e.g Peripheral Blood, bone marrow, ..)</p> <p><strong>Sex </strong>Sex of the Subject</p> <p><strong>Age </strong>Age of the subject</p> <p><strong>UNIQUE_ID </strong>Subject identifier </p> <p><strong>SAMPLE_ID </strong>Sample identifier, linking back to raw data</p> <p><strong>Subset </strong>Defined B cell subset </p> <p><strong>Repertoire </strong>Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG)</p> <p><strong>R_SCDR </strong>R/S ratio in CDR region</p> <p><strong>R_SFWR </strong>R/S ratio in FWR region</p> <p><strong>V_FAM </strong>V family gene</p> <p><strong>V_GENE </strong>V segment gene</p> <p><strong>D_GENE </strong>D segment gene</p> <p><strong>J_GENE </strong>J segment gene</p> <p><strong>Clust_Rank </strong>Cluster rank</p> <p><strong>Clust_REPRES </strong>Cluster representative</p> <p><strong>Clust_SIZE </strong>Cluster size</p> <p><strong>Clust_MAXFREQ </strong>Cluster maximum frequency</p> <p><strong>Clust_SHAREDNESS </strong>Cluster sharedness</p> <p><strong>CDR3_AA_GRAVY </strong>CDR3 hydrophobicity index</p> <p><strong>CDR3_AA_CHARGE </strong>CDR3 charge</p> <p><strong>CDRH3PDB </strong>CDRH3 PDB (Structure) code</p> <p><strong>H1Canon </strong>H1 Canonical class</p> <p><strong>H2Canon </strong>H2 Canonical class</p> <p><strong>H1_GERMLINE </strong>H1 Germline Canonical class</p> <p><strong>H2_GERMLINE </strong>H2 Germline Canonical class</p> <p> </p> <p><strong>References</strong></p> <p>1. Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O’Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein. 2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires. <em>Bioinformatics</em>30: 1930–1932.</p> <p>2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein. 2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data. <em>Bioinformatics</em>31: 3356–3358.</p> <p>3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool. <em>Nucleic Acids Res.</em>41.</p> <p>4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads. <em>Genome Res.</em>21: 936–939.</p>
Single-cell repertoire and transcriptome sequencing reveals clonally expanded and transcriptionally distinct lymphocytes in aged CNS
<p>Single-cell repertoire and transcriptome sequencing reveals clonally expanded and transcriptionally distinct lymphocytes in aged CNS. Gene expression and immune receptor repertoire sequencing was performing for both B and T cells. This dataset contains the VDJ sequencing information for the four samples. Each B cell and T cell library was sequenced across four lanes. </p> <p> </p> <p>Files with _WT_ in their name correspond to the young (4-6 week B6 mice) </p> <p>Files with _12_ in their name before the BDJ or VDJ text correspond to the 12-month-old cohort.</p> <p>Files with _18_ in their name before the BDJ or VDJ text correspond to the 18-month-old cohort in which four brains were pooled.</p> <p>Files with 4_18_ in their name before the BDJ or VDJ text correspond to the 18-month-old mouse that was processed and sequenced alone. </p> <p> </p> <p>The L001 - L004 in the file names indicates the sequencing lane. Samples with BDJ correspond to the B cell repertoire library (B cell VDJ). Samples with TDJ correspond to the T cell repertoire library (T cell VDJ). </p>
Single-cell immune repertoire sequencing of two convalescent COVID-19 patients
<p>Single-cell immune repertoire sequencing of two convalescent COVID-19 patients using 10x genomics 5' immune profiling. Resulting output files are from the count and vdj functions from 10x genomic's cellranger v3.1.0. </p>
Pre-processed IgH repertoire sequencing data from BioProject PRJNA748239
<p><strong>Data Processing</strong></p> <p>Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2). Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2)</p> <p> </p> <p>software_versions pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8</p> <p>quality_thresholds FilterSeq.py pRESTO Q>20</p> <p>paired_reads_assembly AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001</p> <p>primer_match_cutoffs MaskPrimers.py pRESTO C primer & V primer maxerror 0.2</p> <p>consensus_building BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5</p> <p>collapsing_method CollapseSeq.py pRESTO</p> <p>germline_database IMGT</p> <p> </p> <p>Format</p> <p> </p> <p>Processed sequences are provided in a tab delimited file format, including the following annotations:</p> <p> </p> <p><strong>C_CALL</strong> Isotype subclass</p> <p><strong>SEQUENCE_ID</strong> Sequence identifier</p> <p><strong>V_CALL</strong> V segment gene and allele</p> <p><strong>D_CALL</strong> D segment gene and allele</p> <p><strong>J_CALL</strong> J segment gene and allele</p> <p><strong>JUNCTION_LENGTH</strong> Junction length</p> <p><strong>CONSCOUNT</strong> Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence.</p> <p><strong>DUPCOUNT</strong> UMI count for the given unique sequence</p> <p><strong>ISOTYPE</strong> Constant region primer (isotype)</p> <p><strong>MU_COUNT_CDR_R</strong> Number of replacement mutations in CDR region</p> <p><strong>MU_COUNT_CDR_S</strong> Number of silent mutations in CDR region</p> <p><strong>MU_COUNT_FWR_R</strong> Number of replacement mutations in FWR region</p> <p><strong>MU_COUNT_FWR_S</strong> Number of silent mutations in FWR region</p> <p><strong>MUT_TOTAL</strong> Total number of mutations in V gene </p> <p><strong>NP_LENGTH</strong> Total number of N and P additions<strong> </strong></p> <p><strong>SEQUENCE_INPUT</strong> Full length sequence</p> <p><strong>SEQUENCE_IMGT</strong> Gapped IMGT sequence</p> <p><strong>V_GERM_START_VDJ</strong> position of the first nucleotide in ungapped V germline sequence alignment</p> <p><strong>JUNCTION</strong> Junction nucleotide sequence</p> <p><strong>GERMLINE_IMGT_D_MASK</strong> IMGT-gapped germline nucleotide sequence with ns masking the NP1-D-NP2 regions</p> <p><strong>CDR3_AA_GRAVY</strong> CDR3 hydrophobicity</p> <p><strong>CDR3_AA_BULK</strong> CDR3 bulkiness</p> <p><strong>CDR3_AA_ALIPHATIC</strong> Normalized aliphatic index</p> <p><strong>CDR3_AA_POLARITY</strong> CDR3 polarity</p> <p><strong>CDR3_AA_CHARGE</strong> normalised net charge</p> <p><strong>CDR3_AA_BASIC</strong> Basic side chain residue content</p> <p><strong>CDR3_AA_ACIDIC</strong> Acidic side chain residue content</p> <p><strong>CDR3_AA_AROMATIC</strong> aromatic side chain conten</p> <p><strong>Subset</strong> Defined B cell subset </p> <p><strong>Repertoire</strong> Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG)</p> <p><strong>R_SCDR</strong> R/S ratio in CDR region</p> <p><strong>R_SFWR</strong> R/S ratio in FWR region</p> <p><strong>V_GENE</strong> V segment gene</p> <p><strong>D_GENE</strong> D segment gene</p> <p><strong>J_GENE</strong> J segment gene</p> <p><strong>V_FAM</strong> V family gene</p> <p><strong>Run</strong> ID of sequencing run</p> <p><strong>Sex</strong> Sex of the Subject</p> <p><strong>Age</strong> Age of the subject</p> <p><strong>UNIQUE_ID</strong> Subject identifier </p> <p><strong>SAMPLE</strong> Sample identifier, linking back to raw data</p> <p><strong>Bcellno</strong> Number of input B cells</p> <p><strong>Cells</strong> Cell type </p> <p> </p> <p><strong>References</strong></p> <p>1. Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O’Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein. 2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires. <em>Bioinformatics</em>30: 1930–1932.</p> <p>2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein. 2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data. <em>Bioinformatics</em>31: 3356–3358.</p> <p>3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool. <em>Nucleic Acids Res.</em>41.</p> <p>4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads. <em>Genome Res.</em>21: 936–939.</p>
Systematic profiling of full-length immunoglobulin and T-cell receptor repertoire diversity in rhesus macaque through long read transcriptome sequencing
<p>Using long read sequencing, we sequenced four Indian-origin rhesus macaque tissues. From raw full-length, non-chimeric circular consensus sequencing (CCS) reads, we obtained high quality, full-length sequences for over 6,000 unique immunoglobulin and T-cell receptor transcripts, without the need for sequence assembly.</p>
Data from: Defining the alloreactive T cell repertoire using high-throughput sequencing of mixed lymphocyte reaction culture
The cellular immune response is the most important mediator of allograft rejection and is a major barrier to transplant tolerance. Delineation of the depth and breadth of the alloreactive T cell repertoire and subsequent application of the technology to the clinic may improve patient outcomes. As a first step toward this, we have used MLR and high-throughput sequencing to characterize the alloreactive T cell repertoire in healthy adults at baseline and 3 months later. Our results demonstrate that thousands of T cell clones proliferate in MLR, and that the alloreactive repertoire is dominated by relatively high-abundance T cell clones. This clonal make up is consistently reproducible across replicates and across a span of three months. These results indicate that our technology is sensitive and that the alloreactive TCR repertoire is broad and stable over time. We anticipate that application of this approach to track donor-reactive clones may positively impact clinical management of transplant patients.
TCR repertoire sequencing related to "Unique roles of coreceptor-bound LCK in helper and cytotoxic T cells"
<p>This archive contains datasets needed for recapitulating the analysis of TCR repertoires for the manuscript <em>“Unique roles of coreceptor-bound LCK in helper and cytotoxic T cells”</em> by Horkova et al., 2022. The code for the analysis can be found on GitHub: https://github.com/Lab-of-Adaptive-Immunity/lck-tcrseq. Raw data are deposited in the SRA (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA872031).</p> <p>The Zenodo archive contains the following files, which are needed to run the analysis script:</p> <ul> <li>merged outputs from MiXCR <code>merged_TCR_repertoires.csv</code></li> <li>metadata file <code>metadata_Lck.csv</code></li> <li>TRA and TRB repertoires prepared for processing with the Immunarch package <code>immdata_tra.rds</code>, <code>immdata_trb.rds</code></li> </ul>
Data from: Defining the alloreactive T cell repertoire using high-throughput sequencing of mixed lymphocyte reaction culture
Open the record for dataset details and reuse information.
Murine TCR-beta repertoire sequencing
The purpose of this project was to determine the effects of chronic exposure to unpredictable mild socio-environmental stressors during gestation on the TCR-beta repertoire of newborn mice.
High throughput sequencing of immunoglobulin (Ig) heavy (IgH) and light (IgL) chain repertoires
GEO Series GSE48870. Mus musculus. 40 samples. Type: Expression profiling by high throughput sequencing.
Parallel T-cell cloning and deep sequencing of the transcripts of human MAIT cells reveal stable oligoclonal TCRβ repertoire
GEO Series GSE56055. Homo sapiens. 19 samples. Type: Other.
Deep Characterization of the Human Antibody Response to Natural Infection Using Longitudinal Immune Repertoire Sequencing
GEO Series GSE123158. Homo sapiens. 210 samples. Type: Other.
Novel TCR sequencing and cloning methods for sensitive and quantitative interrogation of repertoires and rapid isolation of tumor-reactive TCRs
GEO Series GSE225984. Homo sapiens; Mus musculus. 34 samples. Type: Other.
Single cell RNA sequencing and TCR repertoire analysis of MIS-C affected patients versus healthy controls and severe adult COVID-19
GEO Series GSE184330. Homo sapiens. 16 samples. Type: Expression profiling by high throughput sequencing.
High-throughput Ig and TCR repertoire single-cell sequencing analysis in rhesus macaques
GEO Series GSE179722. Macaca mulatta. 5 samples. Type: Expression profiling by high throughput sequencing.
Single Cell Immunophenotyping of Lyme Erythema Migrans (Bulk TCR Repertoire Sequencing)
GEO Series GSE172225. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
Single-Cell Analysis of Transcriptome and TCR Sequencing Reveals Immune Cell Atlas and Functional Heterogeneity of T Cell Repertoire in Murine Heart Transplantation
GEO Series GSE249989. Mus musculus. 6 samples. Type: Expression profiling by high throughput sequencing.
High-throughput sequencing of the IGH repertoire in human peripheral blood and bone marrow
GEO Series GSE181689. Homo sapiens. 2 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.