Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
595
datasets available to search
ShareScore release 0.9.0
Dataset results
595 results for “Repertoire”
Supplementary data for: Chromosome-level genome assembly and circadian gene repertoire of the Patagonia blennie Eleginops maclovinus
<p>This dataset contains the genome assembly and associated annotation of the Patagonian Blennie (<em>Eleginops maclovinus</em>), the closest extant taxon to the Antarctic notothenioid radiation. In addition to the characterization of the <em>E. maclovinus </em>genome, the dataset includes a description of circadian rhythm orthologs for <em>E. maclovinus</em>, other notothenenioid taxa, and teleost outgroups, as well as a copy of the bioinformatic scripts used for the assembly, annotation, and other downstream analysis.</p>
Tissue-specific features of the T cell repertoire following allogeneic hematopoietic cell transplantation in human and mouse
<p>T cells are the central drivers of many inflammatory diseases, but the repertoire of tissue-resident T cells at sites of pathology in human organs remains poorly understood. We examined the site-specificity of T cell receptor (TCR) repertoires across tissues (5-18 tissues per patient) in prospectively collected autopsies of patients with and without graft-versus-host disease (GVHD), a potentially lethal tissue-targeting complication of allogeneic hematopoietic cell transplantation, as well as in mouse models of GVHD. Anatomic similarity between tissues was a key determinant of TCR repertoire composition within patients, independent of disease or transplant status. The T cells recovered from peripheral blood and spleen in patients and mice captured a limited portion of the TCR repertoire detected in tissues. Whereas few T cell clones were shared across patients, motif-based clustering revealed shared repertoire signatures across patients in a tissue-specific fashion. T cells at disease sites had a tissue-resident phenotype and were of donor origin based on single-cell chimerism analysis. These data demonstrate the complex composition of T cell populations that persist in human tissues at the end-stage of an inflammatory disorder following lymphocyte-directed therapy. These findings also underscore the importance of studying T cells in tissues rather than blood for tissue-based pathologies and suggest the tissue-specific nature of both the endogenous and post-transplant T cell landscape.</p>
Immune repertoire sequencing reveals differences in treatment response to camrelizumab plus platinum-based chemotherapy in advanced ESCC
Open the record for dataset details and reuse information.
Data from: A novel approach to quantifying mammal locomotor repertoires using scoring and cluster analysis
Open the record for dataset details and reuse information.
Rapid evolution of host repertoire and geographic range in a young and diverse genus of montane butterflies
Open the record for dataset details and reuse information.
Data from: Plant ammonium sensitivity is associated with the external pH adaptation, repertoire of nitrogen transporters, and nitrogen requirement
Open the record for dataset details and reuse information.
Data from: Vocal repertoire expansion in singing mice by co-opting a conserved midbrain circuit node
Open the record for dataset details and reuse information.
Supplementary data for: Chromosome-level genome assembly and circadian gene repertoire of the Patagonia blennie Eleginops maclovinus
Open the record for dataset details and reuse information.
Tissue-specific features of the T cell repertoire following allogeneic hematopoietic cell transplantation in human and mouse
Open the record for dataset details and reuse information.
Pre-processed B cell receptor repertoire sequencing data from BioProject PRJNA527941
<p><strong>Data Processing</strong></p> <p> </p> <p>Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2). Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2)</p> <p> </p> <p><strong>software_versions</strong> pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8</p> <p><strong>quality_thresholds</strong> FilterSeq.py pRESTO Q>20</p> <p><strong>paired_reads_assembly</strong> AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001</p> <p><strong>primer_match_cutoffs</strong> MaskPrimers.py pRESTO C primer & V primer maxerror 0.2</p> <p><strong>consensus_building</strong> BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5</p> <p><strong>collapsing_method</strong> CollapseSeq.py pRESTO</p> <p><strong>germline_database </strong>IMGT</p> <p> </p> <p><strong>Format</strong></p> <p> </p> <p>Processed sequences are provided in a tab delimited file format, including the following annotations:</p> <p> </p> <p><strong>C_CALL </strong>Isotype subclass</p> <p><strong>SEQUENCE_ID </strong>Sequence identifier</p> <p><strong>V_CALL </strong>V segment gene and allele</p> <p><strong>D_CALL </strong>D segment gene and allele</p> <p><strong>J_CALL </strong>J segment gene and allele</p> <p><strong>JUNCTION_LENGTH </strong>Junction length</p> <p><strong>CONSCOUNT </strong>Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence.</p> <p><strong>DUPCOUNT </strong>UMI count for the given unique sequence</p> <p><strong>ISOTYPE </strong>Constant region primer (isotype)</p> <p><strong>MU_COUNT_CDR_R </strong>Number of replacement mutations in CDR region</p> <p><strong>MU_COUNT_CDR_S </strong>Number of silent mutations in CDR region</p> <p><strong>MU_COUNT_FWR_R </strong>Number of replacement mutations in FWR region</p> <p><strong>MU_COUNT_FWR_S </strong>Number of silent mutations in FWR region</p> <p><strong>MUT_TOTAL </strong>Total number of mutations in V gene </p> <p><strong>SEQUENCE_INPUT </strong>Full length sequence</p> <p><strong>SEQUENCE_IMGT </strong>Gapped IMGT sequence</p> <p><strong>V_GERM_START_VDJ </strong>position of the first nucleotide in ungapped V germline sequence alignment</p> <p><strong>JUNCTION </strong>Junction nucleotide sequence</p> <p><strong>GERMLINE_IMGT_D_MASK </strong>IMGT-gapped germline nucleotide sequence with ns masking the NP1-D-NP2 regions</p> <p><strong>Run </strong>ID of sequencing run</p> <p><strong>Sample_type </strong>The tissue sampled (e.g Peripheral Blood, bone marrow, ..)</p> <p><strong>Sex </strong>Sex of the Subject</p> <p><strong>Age </strong>Age of the subject</p> <p><strong>UNIQUE_ID </strong>Subject identifier </p> <p><strong>SAMPLE_ID </strong>Sample identifier, linking back to raw data</p> <p><strong>Subset </strong>Defined B cell subset </p> <p><strong>Repertoire </strong>Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG)</p> <p><strong>R_SCDR </strong>R/S ratio in CDR region</p> <p><strong>R_SFWR </strong>R/S ratio in FWR region</p> <p><strong>V_FAM </strong>V family gene</p> <p><strong>V_GENE </strong>V segment gene</p> <p><strong>D_GENE </strong>D segment gene</p> <p><strong>J_GENE </strong>J segment gene</p> <p><strong>Clust_Rank </strong>Cluster rank</p> <p><strong>Clust_REPRES </strong>Cluster representative</p> <p><strong>Clust_SIZE </strong>Cluster size</p> <p><strong>Clust_MAXFREQ </strong>Cluster maximum frequency</p> <p><strong>Clust_SHAREDNESS </strong>Cluster sharedness</p> <p><strong>CDR3_AA_GRAVY </strong>CDR3 hydrophobicity index</p> <p><strong>CDR3_AA_CHARGE </strong>CDR3 charge</p> <p><strong>CDRH3PDB </strong>CDRH3 PDB (Structure) code</p> <p><strong>H1Canon </strong>H1 Canonical class</p> <p><strong>H2Canon </strong>H2 Canonical class</p> <p><strong>H1_GERMLINE </strong>H1 Germline Canonical class</p> <p><strong>H2_GERMLINE </strong>H2 Germline Canonical class</p> <p> </p> <p><strong>References</strong></p> <p>1. Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O’Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein. 2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires. <em>Bioinformatics</em>30: 1930–1932.</p> <p>2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein. 2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data. <em>Bioinformatics</em>31: 3356–3358.</p> <p>3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool. <em>Nucleic Acids Res.</em>41.</p> <p>4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads. <em>Genome Res.</em>21: 936–939.</p>
Dataset: The TCR repertoire reconstitution in multiple sclerosis: comparing one-shot and continuous immunosuppressive therapies
<p>This dataset, containing TCRbeta-chain data, is the basis for the following publication in Frontiers in Immunology: The TCR repertoire reconstitution in multiple sclerosis: comparing one-shot and continuous immunosuppressive therapies. The file key can be found in the file: file_key.xlsx. Relevant methodological details maybe found in the corresponding publication.</p>
TCRb repertoires of murine CD8T cells following viral infection
<p>TCRbeta repertoires from Tcf7-GFP transgenic GFP mice were FACS isolated and sequenced following viral infection. </p>
Alignment free identification of clones in Bcell receptor repertoires
<p>The data sets are used for the analysis of the Alignment free [1] clonal identification approach.</p> <p>[1] "Alignment free identification of clones in B cell receptor repertoires", Ofir Lindenbaum, Nima Nouri, Yuval Kluger, Steven H. Kleinstein .</p> <p>Preprint available at: https://www.biorxiv.org/content/10.1101/2020.03.30.017384v1</p>
Single-cell repertoire and transcriptome sequencing reveals clonally expanded and transcriptionally distinct lymphocytes in aged CNS
<p>Single-cell repertoire and transcriptome sequencing reveals clonally expanded and transcriptionally distinct lymphocytes in aged CNS. Gene expression and immune receptor repertoire sequencing was performing for both B and T cells. This dataset contains the VDJ sequencing information for the four samples. Each B cell and T cell library was sequenced across four lanes. </p> <p> </p> <p>Files with _WT_ in their name correspond to the young (4-6 week B6 mice) </p> <p>Files with _12_ in their name before the BDJ or VDJ text correspond to the 12-month-old cohort.</p> <p>Files with _18_ in their name before the BDJ or VDJ text correspond to the 18-month-old cohort in which four brains were pooled.</p> <p>Files with 4_18_ in their name before the BDJ or VDJ text correspond to the 18-month-old mouse that was processed and sequenced alone. </p> <p> </p> <p>The L001 - L004 in the file names indicates the sequencing lane. Samples with BDJ correspond to the B cell repertoire library (B cell VDJ). Samples with TDJ correspond to the T cell repertoire library (T cell VDJ). </p>
Single-cell immune repertoire sequencing of two convalescent COVID-19 patients
<p>Single-cell immune repertoire sequencing of two convalescent COVID-19 patients using 10x genomics 5' immune profiling. Resulting output files are from the count and vdj functions from 10x genomic's cellranger v3.1.0. </p>
Comparative analyses of the Hymenoscyphus fraxineus and Hymenoscyphus albidus genomes reveals potentially adaptive differences in secondary metabolite and transposable element repertoires
<p><strong>Background </strong>The dieback epidemic decimating common ash (<em>Fraxinus excelsior</em>) in Europe is caused by the invasive fungus <em>Hymenoscyphus fraxineus</em>. In this study we analyzed the genomes of <em>H. fraxineus</em> and <em>H. albidus</em>, its native but, now essentially displaced, non-pathogenic sister species, and compared them with several other members of <em>Helotiales</em>. The focus of the analyses was to identify signals in the genome that may explain the rapid establishment of <em>H. fraxineus</em> and displacement of <em>H. albidus</em>.</p> <p><strong>Results</strong> The genomes of <em>H. fraxineus</em> and <em>H. albidus </em>showed a high level of synteny and identity. The assembly of <em>H. fraxineus </em>is 13 Mb longer than that of <em>H. albidus’, </em>most of this difference can be attributed to higher dispersed repeat content ((i.e transposable elements [TEs]) in <em>H. fraxineus</em>. In general, TE families in <em>H. fraxineus</em>showed more signals of repeat-induced point mutations (RIP) than in <em>H. albidus</em>, especially in Long-terminal repeat (LTR)/Copia and LTR/Gypsy elements. Comparing gene family expansions and 1:1 orthologs, relatively few genes show signs of positive selection between species. However, several of those that did appeared to be associated with secondary metabolite genes families, including gene families containing two of the genes in the <em>H. fraxineus-</em>specific, <em>hymenosetin </em>biosynthetic gene cluster (BGC).</p> <p><strong>C</strong><strong>onclusion </strong>The genomes of <em>H. fraxineus</em> and <em>H. albidus</em> show a high degree of synteny, and are rich in both TEs and BGCs, but the genomic signatures also indicated that <em>H. albidus</em> may be less well equipped to adapt and maintain its ecological niche in a rapidly changing environment. </p> <p><strong>Data included</strong></p> <p>This post contains the alternate structural and functional annotations of the genomes of Helotealean fungi used in the study.</p>
Data from: Diversity and evolution of the transposable element repertoire in arthropods with particular reference to insects
Background: Transposable elements (TEs) are a major component of metazoan genomes and are associated with a variety of mechanisms that shape genome architecture and evolution. Despite the ever-growing number of insect genomes sequenced to date, our understanding of the diversity and evolution of insect TEs remains poor. Results: Here, we present a standardized characterization and an order-level comparison of arthropod TE repertoires, encompassing 62 insect and 11 outgroup species. The insect TE repertoire contains TEs of almost every class previously described, and in some cases even TEs previously reported only from vertebrates and plants. Additionally, we identified a large fraction of unclassifiable TEs. We found high variation in TE content, ranging from less than 6 % in the antarctic midge (Diptera), the honey bee and the turnip sawfly (Hymenoptera) to more than 58 % in the malaria mosquito (Diptera) and the migratory locust (Orthoptera), and a possible relationship between the content and diversity of TEs and the genome size. Conclusion: While most insect orders exhibit a characteristic TE composition, we also observed intraordinal differences, e.g., in Diptera, Hymenoptera, and Hemiptera. Our findings shed light on common patterns and reveal lineage-specific differences in content and evolution of TEs in insects. We anticipate our study to provide the basis for future comparative research on the insect TE repertoire.
Bayesian inference of ancestral host-parasite interactions under a phylogenetic model of host repertoire evolution
<p>Intimate ecological interactions, such as those between parasites and their hosts, may persist over long time spans, coupling the evolutionary histories of the lineages involved. Most methods that reconstruct the coevolutionary history of such interactions make the simplifying assumption that parasites have a single host. Many methods also focus on congruence between host and parasite phylogenies, using cospeciation as the null model. However, there is an increasing body of evidence suggesting that the host ranges of parasites are more complex: that host ranges often include more than one host and evolve via gains and losses of hosts rather than through cospeciation alone. Here, we develop a Bayesian approach for inferring coevolutionary history based on a model accommodating these complexities. Specifically, a parasite is assumed to have a host repertoire, which includes both potential hosts and one or more actual hosts. Over time, potential hosts can be added or lost, and potential hosts can develop into actual hosts or vice versa. Thus, host colonization is modeled as a two-step process that may potentially be influenced by host relatedness. We first explore the statistical behavior of our model by simulating evolution of host-parasite interactions under a range of parameter values. We then use our approach, implemented in the program RevBayes, to infer the coevolutionary history between 34 Nymphalini butterfly species and 25 angiosperm families. Our analysis suggests that host relatedness among angiosperm families influences how easily Nymphalini lineages gain new hosts.</p>
Deep repertoire mining uncovers ultra-broad coronavirus neutralizing antibodies targeting multiple spike epitopes
<p><strong>Abstract:</strong> Development of vaccines and therapeutics that are broadly effective against known and emergent coronaviruses is an urgent priority. We screened the circulating B cell repertoires of COVID-19 survivors and vaccinees to isolate over 9,000 SARS-CoV-2-specific monoclonal Abs (<strong>mAbs</strong>), providing an expansive view of the SARS-CoV-2-specific Ab repertoire. Among the recovered antibodies was TXG-0078, an NTD-specific neutralizing mAb that recognizes diverse alpha- and beta-coronaviruses. TXG-0078 achieves its exceptional binding breadth while utilizing the same VH1-24 variable gene signature and heavy chain-dominant binding pattern seen in other NTD supersite-specific neutralizing Abs with much narrower specificity. We also report the discovery of CC24.2, a pan-sarbecovirus neutralizing antibody that targets a novel RBD epitope and shows similar neutralization potency against all tested SARS-CoV-2 variants, including BQ.1.1 and XBB.1.5. A cocktail of TXG-0078 and CC24.2 protects <i>in vivo</i>, suggesting potential use in variant-resistant therapeutic Ab cocktails and as templates for pan-coronavirus vaccine design.</p><p><strong>Datasets: </strong>This repository contains the 10x Genomic cellranger outputs (matrix and vdj contig files) as well as complied functional characterization dataset used to generate figures on the publication "Deep repertoire mining uncovers ultra-broad coronavirus neutralizing antibodies targeting multiple spike epitopes". </p><p>Post-vaccination samples for donors CC10, CC25, CC31, CC66 were processed in single 10x Genomic reactions. The timepoints samples consist of multiplexing donors CC10, CC25, CC31, CC66 into one 10x Genomic reaction. Similarly, donors CC26, CC42, CC62, CC67 were multiplexed into a single 10x Genomic reaction.</p><p><strong>Files:</strong></p><p>feature names.csv - csv file with sort bait/antigen barcode key </p><p>feature_reference.csv - csv file with cell hash and antigen barcode reference</p><p>filtered_contig<i>_</i>annotations.csv - High-level annotations of each high-confidence contigs from cell-associated barcodes. This is a subset of all_contig_annotations.csv.</p><p>filtered_contig.fasta - filtered antibody fasta</p><p>filtered_matrix.mtx.gz - 10x Genomic matrix file for filtered cells. Contains counts data for feature and gene expression library.</p><p>raw_matrix.mtx.gz - 10x Genomic matrix file for unfiltered cells. Contains counts data for feature and gene expression library.</p><p><strong>Code: </strong>All code used to generate analysis and figures is available under the MIT license on Github<br> </p>
Demographics and song repertoire sizes of cirl bunting populations
<p>In order to improve conservation outcomes translocation or reintroduction of individuals may be necessary. When song learning birds are translocated, changes in the cultural diversity of song repertoires, or abnormal vocalisations, in the new population can be a problem. We monitored song production over 8 years in a reintroduced population of the cirl bunting (<em>Emberiza cirlus</em>). Chicks were removed from nests in Devon, UK, between 2006-2011, translocated at six days old to be hand-reared and released in Cornwall, UK. Recordings at the release site in 2011 showed a significantly reduced population repertoire and individuals sang abnormal song types compared to the source populations in Devon. However, recordings in 2019, showed population song repertoire had reached the level of source populations of similar size, and song types were species typical. Our study shows that species can recover from a cultural bottleneck and suggests that, for some song learning birds, if translocation of nestlings is necessary it may not lead to long-term problems for communication and thus population persistence. For future translocations of nestlings, we recommend that efforts are made to provide tutoring to enable song learning. This may be achieved by providing recordings but may also include providing adult song tutors. In addition, playback of 'normal' songs to translocated populations may aid in development of species typical song repertoires, although care must be taken that this is not disturbing the reintroduced birds.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.