Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,650
datasets available to search
ShareScore release 0.9.0
Dataset results
3,650 results for “antibody”
Finding an antibody to detect EZH1 protein expression – Part 4
<p>Follow up to previous posts where we determine that we have an antibody to detect EZH1 expression in AML patient cells.</p>
Dataset Inhibition of membrane-bound BAFF by the anti-BAFF antibody belimumab
<p>This dataset is related to "Inhibition of membrane-bound BAFF by the anti-BAFF antibody belimumab" (Kowalczyk-Quintas C, Chevalley D, Willen L, Jandus C, Vigolo M, Schneider P).</p>
Maternal antibodies provide strain-specific protection against infection with the Lyme disease pathogen in bank voles
<p>Raw data for manuscript titled, "Maternal antibodies provide strain-specific protection against infection with the Lyme disease pathogen in bank voles". This manuscript was submitted to Applied and Environmental Microbiology and was assigned the manuscript ID number AEM01887-19R1.</p>
CXCR4 antibody specific staining
<p>CXCR4 antibody (#PA3305, Invitrogen Inc., USA) is truly reacting to the antigens in question (specific staining). </p> <p>Tissue: Oral cancer (20X magnification).</p> <p> </p>
MMSEQS meets AntiRef90: reference clusters of human antibody sequences
<p>This data set contains pre-computed mmseqs databases for the antiref fasta files created by <em>Briney et al.</em> </p> <p>Please cite the original work if you use any of the databases provided here.</p> <p>Sources:</p> <ul> <li><a href="https://github.com/brineylab/antiref">Antiref GitHub</a></li> <li><a href="../records/7474336">Antiref Zenodo</a></li> <li><a href="https://academic.oup.com/bioinformaticsadvances/article/3/1/vbad109/7247530?login=true">Antiref Paper</a></li> </ul> <p> </p> <p>The mmseqs databases were created as follows:</p> <p> </p> <p>```</p> <p>aria2x -x16 -s16 --input-file antiref_links.txt<br>snakemake -s antiref_mmseqs.smk --jobs 1 --cores 1 --local-cores 250</p> <p>```</p> <p> </p> <p>Please check the summary repository for the fasta files and snakemake files. In this sub repo we only store the antiref files matching the title of the repo.</p> <p> </p>
Appraise the immunogenicity of designed antibody using deep learning
<p>Appraise the immunogencity of designed antibody using deep learning</p> <p>Link to deep learning model: <a href="https://gitlab.developers.cam.ac.uk/ch/sormanni/abnativ">Yusuf Hamied Department of Chemistry / Sormanni Lab / AbNatiV · GitLab (cam.ac.uk)</a></p> <p>Python code: run_abnativ_humanness_score.py</p> <p>Dataset: antibody_affinity_protein_sabdab_vhvl_immunebuilder_outfiles_nomissing_h.fasta</p> <p>Example output: *res_scores.csv and *seq_scores.csv</p>
Design antibody using LLM
<p>Design antibody variants using LLM</p> <p>Link to LLM: <a href="https://github.com/brianhie/efficient-evolution">GitHub - brianhie/efficient-evolution: Efficient evolution from protein language models</a></p> <p>Python code: use_llm_to_design_variants.py</p> <p>Dataset: antibody_affinity_protein_sabdab_vhvl_immunebuilder_outfiles_nomissing.fasta</p> <p>eeout,.txt: example output files</p>
Antibody dataset Kd and LLM embedding
<p>A dataset of ~500 antibodies with binding affinity Kd obtained from SAbDab via Therapeutic Data Commons. LLM embedding Ablang2</p> <p>Python code: get_antibody_llm_embedding.py</p> <p>Dataset sequence: antibody_affinity_protein_sabdab_vhvl.csv</p> <p>Dataset LLM embedding: antibody_affinity_protein_sabdab_vhvl_ablang2seqenc.csv </p>
Streptococcus pyogenes pharyngitis elicits diverse antibody responses to key vaccine antigens influenced by the imprint of past infections.
<p>Here you will find the raw data (RawData.RData) and code (CHIVAS_SEROLOGY_Code.Rmd, an R Markdown file) for generating the analysis and figures for the following publication:</p> <p><strong><em>Streptococcus pyogenes</em> pharyngitis elicits diverse antibody responses to key vaccine antigens influenced by the imprint of past infections.</strong></p> <p>Joshua Osowicki1,2,3 #, Hannah R Frost1 #, Kristy I Azzopardi1, Alana L Whitcombe4, Reuben McGregor4, Lauren H. Carlton4, Ciara Baker1, Loraine Fabri1,5,6, Manisha Pandey7, Michael F Good7, Jonathan R. Carapetis8,9,10, Mark J Walker11,12,13, Pierre R Smeesters1,2,5,6, Paul V Licciardi2,14, Nicole J Moreland4 *, Danika L Hill15 *, Andrew C Steer1,2,3 *</p> <p>Provided in the RData file are the following items: </p> <p><strong>Dataframes: </strong></p> <p>"outcome" : clinical variables associated with human challenge for each participant</p> <p>"data" : ELISA and functional antibody responses for human challenge participants. Each timepoint and isotype for each antigen as seperate column)</p> <p>"data_long": Data equivalent to "data" file but in long format, i.e. One column for each antigen, timepoint and isotype as factors. </p> <p>"data.melt" : Data equivalent to "data" file but in longer format , i.e. timepoint, isotype and antigen as factors, 'value' as ELISA AU. </p> <p>"luminex" : IgG responses to 6 antigens analysed by luminex bead-based assay in human challenge participants.</p> <p>"luminex.children" : IgG responses to 6 antigen analysed by luminex bead-based assay in children</p> <p><strong>Vectors:</strong></p> <p>"pharyngitis" : participant "id" for the 19 individuals that developed pharyngitis. </p> <p>"Antigen.Order" : relates to "Main" antigen classification used in Figure 2</p> <p>'additional" : relates to "Additional </p> <p><strong>Function: </strong></p> <p>"custom_theme" : used as a theme when using ggplot to graph. </p> <p>Adobe Illustrator or Inkscape were used to generate the final image files for publication, with some graph editing to axes labels, font size, adding p-values etc. </p> <p> </p> <p><em><strong>Additional files: </strong></em></p> <p> 3 .csv files have been included for download</p> <p>"ELISA_data_wide_format.csv", a wide format data table of 25 human challenge individuals and 219 variables. Equivalent to the 'data' dataframe in the RData file</p> <p>"CHIVAS_luminex.csv", a long format data table of 25 human challenge participants at 1 week, 1 month, and 3 months. Equivalent to the 'luminex' dataframe in the RData file. </p> <p>"Luminex.children.csv", a datatable of 6 luminex variables for 39 children (healthy and post pharyngitis). Equivalent to the 'luminex.children' dataframe in the RData file. </p> <p> </p>
Data mining antibody sequences for database searching in bottom-up proteomics
<p>Mass spectrometry (MS)-based proteomics is a powerful method for identifying and quantifying antibodies. Among the various MS approaches, bottom-up proteomics is especially effective for analyzing thousands of antibodies in complex mixtures. In this method, proteins are enzymatically digested into smaller peptides, typically using the protease trypsin, which are then analyzed via mass spectrometry. These peptides are matched to sequences in standard databases like UniProt or NCBI-RefSeq for identification.</p> <p>However, a major limitation of this approach is the absence of comprehensive disease-specific antibody databases. Current databases, such as UniProt, include only a fraction of the antibody sequences present in the human body. For instance, as of January 2024, UniProt contains just 38,800 immunoglobulin sequences, far short of the billions of antibodies the human immune system can produce. As a result, relying on such limited databases can lead to under-detection of antibodies, particularly those associated with specific diseases. Expanding antibody databases with disease-specific sequences is crucial for improving the accuracy of MS-based proteomics in identifying antibodies relevant to human health.</p> <p>Recently, through next-generation sequencing of antibody gene repertoires, it has become possible to obtain billions of antibody sequences (in amino acid format) by annotating, translating, and numbering antibody gene sequences. These large numbers of sequences are now available in public databases such as the <a href="https://opig.stats.ox.ac.uk/webapps/oas/" rel="nofollow">Observed Antibody Space</a>. We hypothesize that using these theoretical antibody sequences as new databases for bottom-up proteomics could address the current lack of antibody coverage in standard databases.</p> <p>We developed a workflow to create disease-specific antibody peptide databases for bottom-up proteomics. The workflow details are available on <a href="https://github.com/trinhxt/SDU_Immunoinformatics">GitHub</a>. The database and metadata files generated by this workflow are stored in this Zenodo dataset, and they are used in DAT-DB — a web application that allows researchers to obtain FASTA files of disease-specific antibody peptides for direct use in bottom-up proteomics (see <a href="https://trinhxt.shinyapps.io/DAT-DB/">Demo version</a>).</p> <p>Each database file in this dataset is in <em>.duckdb</em> format and contains tables with 10 columns: Sequence, Filename, Patient, BSource, BType, Isotype, N_patient, N_antibody, Length_aa, and CDR3. The "<strong>Sequence</strong>" column contains tryptic peptides. "<strong>Filename</strong>" is the file where the data was collected. "<strong>Patient</strong>" refers to the patient number as listed in <em>metadata2.csv</em>. "<strong>BSource</strong>" refers to the B-cells' source, and "<strong>BType</strong>" refers to the type of B-cells. "<strong>Isotype</strong>" specifies the antibody isotype (IgA, IgD, IgE, IgG, IgM, or Bulk). "<strong>N_patient</strong>" indicates the number of patients having this peptide, and "<strong>N_antibody</strong>" specifies the number of antibodies containing this peptide. "<strong>Length_aa</strong>" indicates the number of amino acids in the peptide, while "<strong>CDR3</strong>" shows whether the peptide is found in the CDR3 region.</p> <p>The file <em>metadata1.csv</em> contains information about each database file, while <em>metadata2.csv</em> provides details about the sources of the collected antibodies.</p>
p-IgGen Dataset: Cleaned paired and unpaired antibody sequence data for machine learning applications.
<p>This data is released alongside "p-IgGen: A Paired Antibody Generative Language Model", which contains full details on the data processing and cleaning.</p> <p>p-IgGen Paper: https://www.biorxiv.org/content/10.1101/2024.08.06.606780v1 .</p> <p>OAS: https://opig.stats.ox.ac.uk/webapps/oas/</p> <p> </p>
Skin autonomous antibody production regulates host-microbiota interactions
<p>Supplemental Tables containing bulk BCR-sequencing clonotype results for all samples and for sequences with somatic hypermutations. Data, analysis and results are described in more detail in the accompaning publication Gribonika et al., "Host-microbiota interaction is regulated by autonomous skin-intrinsic germinal centers", Nature, 2024</p>
Focused learning by antibody language models using preferential masking of non-templated regions
<p><strong>Motivation.</strong> While existing antibody language models (AbLMs) excel at predicting germline residues, they often struggle with mutated and non-templated residues, which concentrate in the complementarity-determining regions (CDRs) and are crucial for determining antigen-binding specificity. Many of these models are trained using a masked language modeling (MLM) objective with uniform masking probabilities; however, antibody recombination is modular in nature, creating relatively distinct regions of high and low complexity (non-templated and templated, respectively). We sought to determine whether and to what extent AbLMs can improve when trained using an alternative masking strategy based on this observation.</p> <p><strong>Results.</strong> We developed a variation on MLM called <strong><em>Preferential Masking</em></strong>, which alters masking probabilities to amplify training signals from the CDR3. We pre-trained two AbLMs using either uniform or preferential masking and observed that the latter improves pre-training efficiency and residue prediction accuracy in the highly variable CDR3. Preferential masking also improves antibody classification by native chain pairing and binding specificity, suggesting improved CDR3 understanding and indicating that non-random, learnable patterns help govern antibody chain pairing. We further show that specificity classification is largely informed by residues in the CDRs, demonstrating that AbLMs learn meaningful patterns that align with immunological understanding.</p> <p><strong>Files. </strong>The following files are included in this repository:</p> <ul> <li><strong><em>uniform_250k.tar.gz</em></strong>: Model weights for the Uniform-250k model.</li> <li><strong><em>uniform_350k.tar.gz</em></strong>: Model weights for the Uniform-350k model.</li> <li><strong><em>preferential_250k.tar.gz</em></strong>: Model weights for the Preferential-250k model.</li> <li><strong><em>train-eval-test_cdr-mask.tar.gz</em></strong>: Datasets used to train all three models above. Compressed folder containing three files: <em>A_train.csv</em>, <em>A_eval.csv</em>, and <em>B_test.csv</em>. Each row contains a natively paired sequence with its corresponding label-encoded CDR mask, designed to align with the tokenized amino acid sequence. Sequences were obtained from <a href="https://doi.org/10.1038/s41586-022-05371-z">Jaffe et al.</a> and <a href="https://doi.org/10.1016/j.celrep.2024.114307">Hurtado et al</a>. These are referenced in the paper as Dataset A (<em>A_train.csv, A_eval.csv)</em>, and Dataset B (<em>B_test.csv</em>)<em>.</em></li> <li><strong><em>test-set_annotations.tar.gz</em></strong>: Unpaired annotations for all test set (Dataset B) sequences: <em>B_test-set_annotations.csv</em>. Used for Fig. 3 and Fig. 4D. Annotations can be mapped back to the paired sequences using their `sequence_id` and `locus` information.</li> <li><strong><em>pair_classification.tar.gz</em></strong>: Two classification datasets used to train the classifier models in Figure 4: <em>C_native-0_shuffled-1.csv</em> (Dataset C) and <em>D_native-0_shuffled-1.csv</em> (Dataset D). Dataset C sequences were obtained from <a href="https://doi.org/10.1038/s41586-022-05371-z">Jaffe et al.</a> and <a href="https://doi.org/10.1016/j.celrep.2024.114307">Hurtado et al</a> (Dataset B), and Dataset D sequences were obtained from <a href="https://doi.org/10.1038/s41590-022-01230-1">Phad et al</a> and data generated as part of this study.</li> <li><strong><em>CoV_classification.tar.gz</em></strong>: Classification dataset used to train the classifier models in Figure 5: <em>E_hd-0_cov-1.csv </em>(Dataset E). CoV antibody sequences were obtained from <a href="https://doi.org/10.1093/bioinformatics/btaa739">CoV-AbDAb</a>, and healthy donor sequences were obtained from <a href="https://doi.org/10.1038/s41590-022-01230-1">Phad et al</a>.</li> </ul> <p><strong>Code.</strong> All code used for model training, testing, and figure generation is available under the MIT license on <a href="https://github.com/brineylab/preferential-masking-paper">GitHub.</a></p> <p> </p>
Project files provided as supporting information to the manuscript "How Communication Pathways Bridge Local and Global Conformations in an IgG4 Antibody: a Molecular Dynamics Study"
<p>June 23, 2021</p> <p>Thomas Tarenzi, Marta Rigoli and Raffaello Potestio</p> <p>==================================</p> <p>The dataset contains the following folders:</p> <p>- contact_area_binding_site: files with the computed surface area, used for the calculation of the contact area between PD-1 and the antibody Fab (Fig. S41).</p> <p>- hbonds_ab-pd1: number of hydrogen bonds between the antigen and the antibody, for each holo cluster (Fig. S41).</p> <p>- mutual_information: matrices with the computed mutual information, for each pair of residues (Fig. S34, S35, S46). The folder contains also the generalized correlation coefficients (Fig. S43) and the correlation scores (Fig. 4, S44, S45), computed from the mutual informations.</p> <p>- networks: communities - for each cluster, each residue is assigned to a community within the interaction network (Fig. S32, S33). betweenness - the values of edge betweenness for each cluster (S30, S31).</p> <p>- output_clustering: each frame of the apo and holo simulations is assigned a cluster index, on the basis of the structural similarity (Fig. S4).</p> <p>- PAD: per-residue values of PAD parameter, for apo and holo systems (Fig. 3, S39).</p> <p>- PCA: principal component analysis for each conformational cluster (Section S2.1).</p> <p>- representative_structures: representative structures for each conformational cluster (Fig. 2).</p> <p>- r_gyr_antibody: radii of gyration of the antibody, for each cluster (Fig. 2).</p> <p>- r_gyr_hinge: radii of gyration of the sole hinge segment, for each cluster (Fig. S36).</p> <p>- RMSD_antibody: distributions of the root-mean-square deviation of the antibody, for each cluster (Fig. S5). </p> <p>- RMSD_antigen: root-mean-square deviation of the antigen PD-1, for each cluster (Fig. S42).</p> <p>- RMSD_binding_site: distributions of the root-mean-square deviation of the residues belonging to the paratope, for each cluster (Fig. S40, S47).</p> <p>- RMSD_matrix: root-mean-square deviation between structures belonging to different pairs of clusters (Fig. S9).</p> <p>- rmsf_antigen: root-mean-square fluctuation of the antigen PD-1, for each cluster (Fig. S42).</p> <p>- rmsf_hinge: difference between the total root-mean-square fluctuations of the two hinge segments, for each cluster (Fig. S38).</p> <p>- salt_bridge: distribution of distances between residues R979 and D1377 (Fig. 4).</p> <p>- sasa_domains: contact area between Fab and Fc antibody domains, for each cluster (Fig. S8).</p> <p>- sasa_hinge: solvent accessible surface area of each hinge segment, for each cluster (Fig. S37).</p>
Data from: Population-based screening for hepatitis C antibodies and active infection using a point-of-care test in a low prevalence area
<p><span><span><span><span><span><span><span><span><span><span><span><b>Background.</b> Data on the true prevalence of hepatitis C virus (HCV) infection in the </span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>general population is essential to health policies. We evaluated a program implementing free universal HCV screening using a non-invasive point-of-care test (POCT) (OraQuick-HCV rapid test) in oral fluid in an urban area in Valencia, South-Eastern Spain. </span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Methods.</b> A cross-sectional study was performed during 2015-2017. Free HCV screening was offered by regular mail to 11,500 individuals aged 18 and over, randomly selected from all census residents in the Health Department. All responding participants filled in a questionnaire about HCV infection risk factors and were tested in their tertiary Hospital. In those with a positive POCT, results were confirmed by enzyme-immunoassay and HCV-RNA.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Results.</b> 1,206 persons agreed to participate (response rate: 11.16%). HCV antibodies were detected in 19 (1.60%) cases (age-sex standardized rate: 1.31%; 95%CI: 0.82-2.07), but only 8 showed positive HCV-RNA (age-sex standardized rate: 0.56%; 95%CI: 0.28-1.14). The majority (89%) of the cases were born before 1965 and 74% had at least one known risk factor for HCV infection. All anti-HCV positive individuals were already aware of their infection, and no undiagnosed cases were detected. The performance of the POCT was excellent for detecting active infection. </span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Conclusions.</b> These preliminary data suggest that HCV population screening with a POCT is feasible but, in our setting, mailing recruiting is not effective (11% response rate). The low prevalence of HCV antibodies and active infection in the participant population (with no new diagnoses made) suggests that, in our setting, underdiagnosis may be uncommon.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Files uploaded include the study database (Stata v.13) and the do.file of the study.</b></span></span></span></span></span></span></span></span></span></span></span></p>
Antibody structure prediction method benchmark results
<p>Antibody Fv structures generated for benchmarking recent antibody structure prediction methods in "Antibody structure prediction using interpretable deep learning." All successfully predicted targets are provided for each method, although only those where every method succeeded are included in accuracy metrics. Each DeepAb prediction includes 50 generated decoys from which the final structure was selected.</p>
Monoclonal antibodies from humans with Mycobacterium tuberculosis exposure or latent infection recognize distinct arabinomannan epitopes
<p>The surface polysacharide arabinomannan (AM) and related glycolipid lipoarabinomannan (LAM) play critical roles in tuberculosis pathogenesis. Human antibody responses to AM/LAM are heterogenous and knowledge of reactivity to specific glycan epitopes at the monoclonal level is limited, especially in individuals who can control <i><span>M. tuberculosis</span></i> infection<span>. </span>We generated human IgG mAbs to AM/LAM from B cells of two asymptomatic individuals exposed to or latently infected with <i><span>M. tuberculosis</span></i>. We here show that two of these mAbs have high affinity to AM/LAM, are non-competing, and recognize different glycan epitopes <span>distinct from other anti-AM/LAM mAbs reported. Both mAbs recognize virulent </span><i><span>M. tuberculosis</span></i><span> and nontuberculous mycobacteria with marked differences, can be used for the detection of urinary LAM, and can detect </span><i><span>M. tuberculosis</span></i><span> and LAM in infected lungs. These mAbs enhance our understanding of the spectrum of antibodies to AM/LAM epitopes </span>in humans <span>and</span> are valuable for <span>tuberculosis</span> diagnostic and research applications.</p>
HLA epitopic mismatch and risk of developing Donor Specific Antibodies in pediatric kidney transplantation
<p>The aim of this study was to evaluate the correlation between epitopic HLA mismatches and the risk of developing DSA in pediatric kidney recipients. There are very few publications concerning epitope HLA mismatch in pediatric kidney transplantation although this topic has been growing in scientific articles in recent years and is promising in adult recipients. The majority of children with chronic renal failure will require multiple kidney transplants during their lifetime and it is therefore important that we can limit the risk of antibody development in order to give them the best chance of survival for their future kidney transplants.</p>
Dataset and R code: Prior exposure to B. pertussis shapes the mucosal antibody response to acellular pertussis booster vaccination
<p>The R code and dataset for the figures created in the Nature communications manuscript titled "<strong>Prior exposure to <em>B. pertussis </em>shapes the mucosal antibody response to acellular pertussis booster vaccination"</strong>.</p> <p>contains:</p> <p>- excel dataset including the parameters needed for the figures</p> <p>- R code document with the code used to produce the figures and statistical analyses</p>
Adaptive, randomized, non-inferiority trial to evaluate the efficacy of monoclonal antibodies in outpatients with mild or moderate COVID-19
<p><span>The dataset is based on the results of a trial called MANTICO (Clinical trial number NCT05205759). The study is a non-inferiority randomised controlled trial comparing the clinical efficacy of bamlanivimab/etesevimab, casirivimab/imdevimab, and sotrovimab in outpatients aged 50 or older with early COVID-19. The primary outcome was COVID-19 progression (hospitalisation, need of supplemental oxygen therapy, or death through day 14). </span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.