Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

595

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

595 results for “Repertoire”

Learn how ShareScore rates datasets ↗
zenodo36/100

Enhancing comparative T-cell receptor repertoire analysis in small biological samples through pooling homologous cell samples from multiple mice

<p>All data files used to generate the figures in the paper are shared in this project.</p> <p>Scripts are available on <a href="https://github.com/i3-unit/CRM_24" target="_blank" rel="noopener">GitHub</a>.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Code and data to "The shifting of buffer crop repertoires in pre-industrial north-eastern Europe "

<p>This code and data can be used to replicate the plots and figures of the paper and to trace the correlation and tests of climate variability and crop development in the study area.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Immune repertoires of de-identified TCGA tumor samples assembled by TRUST4

<p>We applied TRUST4 to assemble the&nbsp;immune repertoire data from TCGA tumor samples. Because TCGA has restricted access permission, the sample IDs are de-identified and the sequence is at the amino acid level. The data is used in the study of "Comprehensive characterizations of immune receptor repertoire in tumors and cancer immunotherapy studies".&nbsp; The format is:</p> <p>CancerType_RandomID Chain_Type CDR3_AminoAcid VGene JGene ConstantGene Abundance</p> <p>(Note that the deidentified RandomID is different from the previous version (version 1))</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities: validation cohort meta data and parsed TCR repertoire data

<p>Meta data corresponding the the validation cohort for the paper,&nbsp;&quot;Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities&quot;&nbsp;by Magdalena L Russell, Aisha Souquette, David M Levine, Stefan A Schattgen, E Kaitlynn Allen, Guillermina Kuan, Noah Simon, Angel Balmaseda, Aubree Gordon, Paul G Thomas, Frederick A Matsen IV, and Philip Bradley. These meta data include:&nbsp;</p> <p>(1) SNP genotypes for the two SNPs which overlap with the discovery cohort<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;- (nicaragua_snp_genotypes_ints.tsv) -- SNP genotypes as integers<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;- (nicaragua_snp_genotypes_strings.tsv) -- SNP genotypes as allele strings&nbsp;<br> (2) the ancestry PCs for each individual in the validation cohort (nicaragua_snp_ancestry_PCA.tsv)<br> (3) a file including IMGT genes and sequences used for parsing TCRB repertoire data (human_vj_allele_cdr3_nucseqs.tsv)<br> (4) a file including IMGT genes&nbsp;used for parsing TCRA&nbsp;repertoire data (human_vj_alleles_alpha.tsv)<br> (5)&nbsp;Parsed TCRA repertoire data (nicaragua_parsed_TCRA.tgz)<br> (6) Parsed TCRB repertoire data (nicaragua_parsed_TCRB.tgz)&nbsp;</p> <p><strong>Corresponding raw validation cohort TCR repertoire data is available here:</strong>&nbsp;https://www. ncbi.nlm.nih.gov/bioproject/PRJNA762269 (The BioProject database,&nbsp;accession number: PRJNA762269)</p> <p><strong>Software tools designed to work with these data are available here:</strong>&nbsp;https://github.com/phbradley/tcr-gwas</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Fig. 1, 1–10 in Song Repertoire And Comparative Analysis Of Song Structure Of Chaffinch, Fringilla Coelebs (Fringillidae), From The Northeast Of Balkan Region

Fig. 1, 1–10. Sonograms of most frequent song types of Balkan chaffinches.

opencc-by-4.0Jul 2014View details →
zenodo36/100

Fig. 1, 10–18 in Song Repertoire And Comparative Analysis Of Song Structure Of Chaffinch, Fringilla Coelebs (Fringillidae), From The Northeast Of Balkan Region

Fig. 1, 10–18. Sonograms of most frequent song types of Balkan chaffinches.

opencc-by-4.0Jul 2014View details →
dryad36/100

Data from: Olfaction written in bone: cribriform plate size parallels olfactory receptor gene repertoires in Mammalia

The evolution of mammalian olfaction is manifested in a remarkable diversity of gene repertoires, neuroanatomy, and skull morphology across living species. Olfactory receptor genes (ORG), which initiate the conversion of odorant molecules into odor perceptions and help an animal resolve the olfactory world, range in number from a mere handful to several thousand genes across species. Within the snout, each of these ORGs is exclusively expressed by a discrete population of olfactory sensory neurons (OSN), suggesting that newly evolved ORGs may be coupled with new OSN populations in the nasal epithelium. Because OSNs axon bundles leave high-fidelity perforations (foramina) in the bone as they traverse the cribriform plate (CP) to reach the brain, we predicted that taxa with larger ORG repertoires would have proportionately expanded footprints in the CP foramina. Previous work found a correlation between ORG number and absolute CP size that disappeared when body size effects were accounted for. Using updated, digital measurement data from high-resolution CT scans and reexamining the relationship between CP and body size, we report a striking linear correlation between relative CP area and number of functional ORGs across species from all mammalian superorders. This correlation suggests strong developmental links in the olfactory pathway between genes, neurons, and skull morphology. Furthermore, because ORG number is linked to olfactory discriminatory function, this correlation supports relative CP size as a viable metric for inferring olfactory capacity across modern and extinct species. By quantifying CP area from a fossil sabertooth cat (Smilodon fatalis) we predicted a likely ORG repertoire for this extinct felid.

opencc-zeroDec 2017View details →
zenodo36/100

Partis post-processed B cell receptor repertoires from BioProject PRJNA349143

<p>These files correspond to partis annotations of several datasets found in&nbsp;BioProject PRJNA349143. (DOI:&nbsp;10.5281/zenodo.821659).</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Figure 1 in The song structure and repertoire size of Daurian Redstarts (Phoenicurus auroreus) in South Korea

Figure 1. The map of recording sites of Daurian Redstart songs in South Korea.

opencc-by-4.0May 2019View details →
dryad36/100

A small vocal repertoire during the breeding season expresses complex behavioral motivations and individual signature in the Common Coot

<p><b>Backgroun</b><b>d:</b> Although acoustic communication plays an essential role in the social interactions of Rallidae, our knowledge of how Rallidae encode diverse types of information using simple vocalizations is limited. We recorded and examined the vocalizations of a Common Coot (<i>Fulica atra</i>) population during the breeding season to test the hypotheses that 1) different call types can be emitted under different behavioral contexts, and 2) variation in the vocal structure of a single call type may be influenced both by behavioral motivations and individual signature. We measured a total of 61 recordings of 30 adults while noting the behavioral activities in which individuals were engaged. We compared several acoustic parameters of the same call type emitted under different behavioral activities to determine how frequency and temporal parameters changed depending on behavioral motivations and individual differences.</p> <p><b>Results: </b>We found that adult Common Coots had a small vocal repertoire, including 4 types of call, composed of a single syllable that was used during 9 types of behaviors. The 4 calls significantly differed in both frequency and temporal parameters and can be clearly distinguished by discriminant function analysis. Minimum frequency of fundamental frequency (F<sub>0min</sub>) and duration of syllable (T) contributed the most to acoustic divergence between calls. Call <i>a</i> was the most commonly used (in 8 of the 9 behaviors detected), and maximum frequency of fundamental frequency (F<sub>0max</sub>) and interval of syllables (TI) contributed the most to variation in call <i>a</i>. Duration of syllable (T) in a single call <i>a</i> can vary with different behavioral motivations after individual vocal signature being controlled.</p> <p><b>Conclusions:</b> These results demonstrate that several call types of a small repertoire, and a single call with function-related changes in the temporal parameter in Common Coots could potentially indicate various behavioral motivations and individual signature. This study advances our knowledge of how Rallidae use "simple" vocal systems to express diverse motivations and provides new models for future studies on the role of vocalization in avian communication and behavior.</p>

opencc-zeroDec 2020View details →
zenodo36/100

Pandemic, epidemic, endemic: B cell repertoire analysis reveals unique anti-viral responses to SARS-CoV-2, Ebola and Respiratory Syncytial Virus

<p>VDJ gene usage, and associated amino acid sequences and properties, from healthy controls as well as patients with COVID-19, RSV or Ebola and Yellow fever vaccine&nbsp;recipients.</p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

Pre-processed IgH repertoire sequencing data from BioProject PRJNA748239

<p><strong>Data Processing</strong></p> <p>Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2).&nbsp;Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2)</p> <p>&nbsp;</p> <p>software_versions&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8</p> <p>quality_thresholds&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;FilterSeq.py pRESTO Q&gt;20</p> <p>paired_reads_assembly&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001</p> <p>primer_match_cutoffs&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;MaskPrimers.py pRESTO C primer &amp; V primer maxerror 0.2</p> <p>consensus_building&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5</p> <p>collapsing_method&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;CollapseSeq.py pRESTO</p> <p>germline_database&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;IMGT</p> <p>&nbsp;</p> <p>Format</p> <p>&nbsp;</p> <p>Processed sequences are provided in a tab delimited file format, including the following annotations:</p> <p>&nbsp;</p> <p><strong>C_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Isotype subclass</p> <p><strong>SEQUENCE_ID</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Sequence identifier</p> <p><strong>V_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;V segment gene and allele</p> <p><strong>D_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;D segment gene and allele</p> <p><strong>J_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;J segment gene and allele</p> <p><strong>JUNCTION_LENGTH</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Junction length</p> <p><strong>CONSCOUNT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence.</p> <p><strong>DUPCOUNT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;UMI count for the given unique sequence</p> <p><strong>ISOTYPE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Constant region primer (isotype)</p> <p><strong>MU_COUNT_CDR_R</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of replacement mutations in CDR region</p> <p><strong>MU_COUNT_CDR_S</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of silent mutations in CDR region</p> <p><strong>MU_COUNT_FWR_R</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of replacement mutations in FWR region</p> <p><strong>MU_COUNT_FWR_S</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of silent mutations in FWR region</p> <p><strong>MUT_TOTAL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Total number of mutations in V gene&nbsp;</p> <p><strong>NP_LENGTH</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Total number of N and P additions<strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</strong></p> <p><strong>SEQUENCE_INPUT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Full length sequence</p> <p><strong>SEQUENCE_IMGT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Gapped IMGT sequence</p> <p><strong>V_GERM_START_VDJ</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;position of the first nucleotide in ungapped V germline sequence alignment</p> <p><strong>JUNCTION</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Junction nucleotide sequence</p> <p><strong>GERMLINE_IMGT_D_MASK</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;IMGT-gapped germline nucleotide sequence with ns masking the NP1-D-NP2 regions</p> <p><strong>CDR3_AA_GRAVY</strong>&nbsp; &nbsp; &nbsp; &nbsp;CDR3 hydrophobicity</p> <p><strong>CDR3_AA_BULK</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; CDR3 bulkiness</p> <p><strong>CDR3_AA_ALIPHATIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Normalized aliphatic index</p> <p><strong>CDR3_AA_POLARITY</strong>&nbsp; &nbsp; &nbsp; &nbsp; CDR3 polarity</p> <p><strong>CDR3_AA_CHARGE</strong>&nbsp; &nbsp; &nbsp; &nbsp; normalised net&nbsp;charge</p> <p><strong>CDR3_AA_BASIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Basic side chain residue content</p> <p><strong>CDR3_AA_ACIDIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Acidic side chain residue content</p> <p><strong>CDR3_AA_AROMATIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;aromatic side chain conten</p> <p><strong>Subset</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Defined B cell subset&nbsp;</p> <p><strong>Repertoire</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG)</p> <p><strong>R_SCDR</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;R/S ratio in CDR region</p> <p><strong>R_SFWR</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;R/S ratio in FWR region</p> <p><strong>V_GENE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;V segment gene</p> <p><strong>D_GENE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;D segment gene</p> <p><strong>J_GENE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;J segment gene</p> <p><strong>V_FAM</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;V family gene</p> <p><strong>Run</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;ID of sequencing run</p> <p><strong>Sex</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Sex of the Subject</p> <p><strong>Age</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Age of the subject</p> <p><strong>UNIQUE_ID</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Subject identifier&nbsp;</p> <p><strong>SAMPLE</strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Sample identifier, linking back to raw data</p> <p><strong>Bcellno</strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Number of input B cells</p> <p><strong>Cells</strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Cell type&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>1.&nbsp;Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O&rsquo;Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein.&nbsp;2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires.&nbsp;<em>Bioinformatics</em>30: 1930&ndash;1932.</p> <p>2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein.&nbsp;2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data.&nbsp;<em>Bioinformatics</em>31: 3356&ndash;3358.</p> <p>3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool.&nbsp;<em>Nucleic Acids Res.</em>41.</p> <p>4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads.&nbsp;<em>Genome Res.</em>21: 936&ndash;939.</p>

opencc-by-4.0Aug 2021View details →
dryad36/100

Data from: Contextual inference underlies the learning of sensorimotor repertoires

<p>Humans spend a lifetime learning, storing and refining a repertoire of motor memories. For example, through experience, we become proficient at manipulating a large range of objects with distinct dynamical properties. However, it is unknown what principle underlies how our continuous stream of sensorimotor experience is segmented into separate memories and how we adapt and use this growing repertoire. Here we develop a theory of motor learning based on the key principle that memory creation, updating, and expression are all controlled by a single computation—contextual inference. Our theory reveals that adaptation can arise both by creating and updating memories (proper learning) and by changing how existing memories are differentially expressed (apparent learning). This insight allows us to account for key features of motor learning that had no unified explanation: spontaneous recovery, savings, anterograde interference, how environmental consistency affects learning rate and the distinction between explicit and implicit learning. Critically, our theory also predicts novel phenomena—evoked recovery and context-dependent single-trial learning—which we confirm experimentally. These results suggest that contextual inference, rather than classical single-context mechanisms, is the key principle underlying how a diverse set of experiences is reflected in our motor behaviour.</p>

opencc-zeroSep 2021View details →
zenodo36/100

Pre-processed IgH receptor repertoire data from MS patients after aHSCT from BioProject PRJNA763367

<p><strong>Data Processing</strong></p> <p>Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2).&nbsp;Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2)</p> <p>&nbsp;</p> <p><strong>software_versions</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8</p> <p><strong>quality_thresholds</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;FilterSeq.py pRESTO Q&gt;20</p> <p><strong>paired_reads_assembly</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001</p> <p><strong>primer_match_cutoffs</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;MaskPrimers.py pRESTO C primer &amp; V primer maxerror 0.2</p> <p><strong>consensus_building</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5</p> <p><strong>collapsing_method</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;CollapseSeq.py pRESTO</p> <p><strong>germline_database&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>IMGT</p> <p>&nbsp;</p> <p><strong>Format</strong></p> <p>Processed sequences are provided in a tab delimited file format, including the following annotations:</p> <p>&nbsp;</p> <p><strong>ISOTYPE_SUBCLASS &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Isotype subclass</p> <p><strong>SEQUENCE_ID&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sequence identifier</p> <p><strong>JUNCTION_LENGTH&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Junction length</p> <p><strong>CONSCOUNT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence.</p> <p><strong>DUPCOUNT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>UMI count for the given unique sequence</p> <p><strong>ISOTYPE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Constant region primer (isotype)</p> <p><strong>MUT_TOTAL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Total number of mutations in V gene&nbsp;</p> <p><strong>SAMPLE&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</strong>Sample identifier, linking back to raw data</p> <p><strong>JUNCTION&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Junction nucleotide sequence</p> <p><strong>Protein_seq &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Amino acid sequence</p> <p><strong>CDR3_AA_GRAVY&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>CDR3 hydrophobicity index</p> <p><strong>CDR3_AA_BULK &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 bulkiness</p> <p><strong>CDR3_AA_ALIPHATIC &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 aliphatic index</p> <p><strong>CDR3_AA_POLARITY &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 polarity</p> <p><strong>CDR3_AA_CHARGE &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 normalized net charge</p> <p><strong>CDR3_AA_BASIC &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 basic side chain residue content</p> <p><strong>CDR3_AA_ACIDIC &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 acidic side chain residue content</p> <p><strong>CDR3_AA_AROMATIC &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 aromatic side chain content</p> <p><strong>Subset&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Defined B cell subset&nbsp;</p> <p><strong>Repertoire&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG)</p> <p><strong>R_SCDR&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>R/S ratio in CDR region</p> <p><strong>R_SFWR&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>R/S ratio in FWR region</p> <p><strong>V_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>V segment gene</p> <p><strong>D_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>D segment gene</p> <p><strong>J_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>J segment gene</p> <p><strong>V_FAM&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>V family gene</p> <p><strong>Clust_REPRES&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster representative</p> <p><strong>Clust_SIZE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster size</p> <p><strong>Sex&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sex of the Subject</p> <p><strong>UNIQUE_ID&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sample identifier&nbsp;</p> <p><strong>Bcellno &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Input B cell number</p> <p><strong>Days_posttx &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Sampling time point relative to transplantation</p> <p><strong>Age_at_tx &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Age of the subject (at aHSCT)</p> <p><strong>Disease &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>MS subtype</p> <p><strong>Last_therapy &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Last therapy prior to aHSCT</p> <p><strong>Disease_duration &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Disease duration</p> <p><strong>CMV_reactivation &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Cytomegalovirus reactivation</p> <p><strong>Month_label &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Month post-aHSCT inverval bin</p> <p><strong>Patient_label &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Subject identifier</p> <pre> &nbsp;</pre> <p><strong>References</strong></p> <p>1.&nbsp;Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O&rsquo;Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein.&nbsp;2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires.&nbsp;<em>Bioinformatics</em>30: 1930&ndash;1932.</p> <p>2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein.&nbsp;2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data.&nbsp;<em>Bioinformatics</em>31: 3356&ndash;3358.</p> <p>3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool.&nbsp;<em>Nucleic Acids Res.</em>41.</p> <p>4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads.&nbsp;<em>Genome Res.</em>21: 936&ndash;939.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Figure 4 in Vocal repertoire and group-specific signature in the Smooth-billed Ani, Crotophaga ani Linnaeus, 1758 (Cuculiformes, Aves)

Figure 4. Boxplots (median and quartiles) of acoustic parameters of the similar vocalizations of Charqueada and Guararema groups of Smooth-billed Ani. Vocalizations:"Ahnee","Whine", "Pre-flight", "Flight" and "Vigil". Acoustic parameters: DUR = duration; MPF = maximum peak frequency; MFF = maximum fundamental frequency; MIF = minimum frequency; MAF = maximum frequency.

opencc-by-nc-4.0Jul 2021View details →
zenodo36/100

Figure 3 in Vocal repertoire and group-specific signature in the Smooth-billed Ani, Crotophaga ani Linnaeus, 1758 (Cuculiformes, Aves)

Figure 3. Spectrograms of the ten types of vocalizations of the Smooth-billed Ani: "Ahnee" (A), "Whine" (B, C, D, E, F and G), "Pre-flight" (H), "Shout" (I), "Flight" (J and K), "Hoot" (L), "Grunt" (M), "Ee-oo-ee" (N), "Vigil" (O),"INR" (P and Q).

opencc-by-nc-4.0Jul 2021View details →
zenodo36/100

Figure 1 in Vocal repertoire and group-specific signature in the Smooth-billed Ani, Crotophaga ani Linnaeus, 1758 (Cuculiformes, Aves)

Figure 1. Location of the studied groups of Smooth-billed Ani in the municipality of Alegre, ES, Brazil.

opencc-by-nc-4.0Jul 2021View details →
dryad36/100

A preliminary comparison of a songbird's song repertoire size and other song measures between an urban and a rural site

<p>Characteristics of birdsong, especially minimum frequency, have been shown to vary for some species between urban and rural populations and along urban-rural gradients. However, few urban-rural comparisons of song complexity—and none that we know of based on the number of distinct song types in repertoires—have occurred. Given the potential ability of song repertoire size to indicate bird condition, we primarily sought to determine if number of distinct song types displayed by Song Sparrows (<i>Melospiza melodia</i>) varied between an urban and a rural site. We determined song repertoire size of 24 individuals; 12 were at an urban ('human-dominated') site and 12 were at a rural ('agricultural') site. Then, we compared song repertoire size, note rate, and peak frequency between these sites. Song repertoire size and note rate did not vary between our human-dominated and agricultural sites. Peak frequency was greater at the agricultural site. Our finding that peak frequency was higher at the agricultural site compared to the human-dominated site, contrary to many previous findings pertaining to frequency shifts in songbirds, warrants further investigation. Results of our pilot study suggest that song complexity may be less affected by anthropogenic factors in Song Sparrows than are frequency characteristics. <a name="_Hlk80603521">Additional study, however, will be required to identify particular causal factors related to the trends that we report and to replicate, ideally via multiple urban-rural pairings, so that broader generalization is possible. </a></p>

opencc-zeroFeb 2023View details →
zenodo36/100

Data for "Analysis of Wilms' tumor protein 1 specific TCR repertoire in AML patients uncovers higher diversity in patients in remission than in relapsed"

<p>This folder holds the data for the paper &quot;Analysis of Wilms&#39; tumor protein 1 specific TCR repertoire in AML patients uncovers higher diversity in patients in remission than in relapsed&quot; (in submission) More information regarding this paper and the data is given in the GitHub repository (https://github.com/sgielis/WT1_TCR)</p> <p>The raw folder contains all MiXCR files for the two studied WT1 epitopes and two VZV epitopes. The VZV epitopes were not taken into account in this paper, but were used to build VZV-specific TCRex models for another paper [in submission]. Since all TCRs for the 4 epitopes were sequences together, this data was used for quality control purposes as explained in the paper. Following 4 folders are present:</p> <ul> <li>run1: TCR data from the first run for WT1-126, WT1-37 and IE62</li> <li>run1_orf18: TCR data from the first run for ORF18</li> <li>run2: WT1-37 data filtered on high and low threshold gating.</li> <li>run 3: extra TCR data for WT1-126, WT1-37 aligned with MiXCR</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Spatial database of the towns of the repertoire of all the roads of Spain in the year of grace of 1543 by Juan Villuga [shapefile]

<p>Este archivo proporciona una versi&oacute;n digital de los pueblos categorizados por Juan Villuga en 1543. El archivo consta de 1070 entidades geogr&aacute;ficas de tipo punto, cada entidad representa un pueblo documentado por Juan Villuga con su respectivo nombre. Se ha utilizado la proyecci&oacute;n ETRS89 30N. La vectorizaci&oacute;n ha sido realizada a partir de dos fuentes: 1) Villuga, Juan. 1543. Repertorio de todos los caminos de Espa&ntilde;a en el a&ntilde;o de gracia de 1543. Barcelona: Institut Cartografic i Geol&ograve;gic de Catalunya, R: RL 3419. (Consulta 14/05/2014). <a href="http://cartotecadigital.icgc.cat/cdm/singleitem/collection/espanya/id/2618/rec/1">http://cartotecadigital.icgc.cat/cdm/singleitem/collection/espanya/id/2618/rec/1</a>; y 2) Villuga, Juan. 1950 [1543]. Repertorio de todos los caminos de Espa&ntilde;a. Madrid: Reimpresiones Bibliogr&aacute;ficas.</p> <p>La tabla de atributos contiene informaci&oacute;n de la denominaci&oacute;n de cada pueblo dada por Juan Villuga (1543), de la categorizaci&oacute;n dada por Villuga, de la denominaci&oacute;n dada por&nbsp;Gonzalo Men&eacute;ndez Pidal (1951), de la denominaci&oacute;n actual y de la precisi&oacute;n de la digitalizaci&oacute;n.</p> <p>Cuanto a la metodolog&iacute;a, en primer lugar, se transcribieron los n&uacute;cleos y se relacionaron con la informaci&oacute;n de sus respectivas rutas y clasificaci&oacute;n definida por Villuga (capital, ciudad importante, ciudad peque&ntilde;a, venta) a trav&eacute;s de una tabla en formato .xml. Cada n&uacute;cleo, a su vez, se asoci&oacute; tambi&eacute;n con la denominaci&oacute;n actual y las establecidas en las cartograf&iacute;as de Villuga y Men&eacute;ndez Pidal. De esta forma, se registraron los cambios en la toponimia entre las cartograf&iacute;as de 1546 y 1941. Otra caracter&iacute;stica adicional que decidimos considerar fue el atributo de &quot;exactitud&quot; que se refiere a la calidad de la informaci&oacute;n espacial (las coordenadas). Cuando los datos de un n&uacute;cleo no eran exactos al georreferenciarlos, a&ntilde;adimos el atributo &quot;inexacto&quot;, de modo que la calidad de los datos tambi&eacute;n quedara registrada en la propia tabla de atributos.</p>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record