Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
zenodo36/100

Sewage - shallow sequencing data

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo36/100

p-IgGen Dataset: Cleaned paired and unpaired antibody sequence data for machine learning applications.

<p>This data is released alongside "p-IgGen: A Paired Antibody Generative Language Model", which contains full details on the data processing and cleaning.</p> <p>p-IgGen Paper: https://www.biorxiv.org/content/10.1101/2024.08.06.606780v1 .</p> <p>OAS: https://opig.stats.ox.ac.uk/webapps/oas/</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

IsoQuant graphs from PacBio and ONT Mouse sequencing data

<p>Graph files and simulated ONT BAM files.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Raw data for whole plasmid and whole genome sequencing

<p>Original data for plasmid and genomic DNA sequencing in the paper: Tailoring Microbial Fitness Through Computational Steering and CRISPRi-Driven Robustness Regulation</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Data from: Parallel mechanisms signal a hierarchy of sequence structure violations in the auditory cortex

<p>The brain predicts regularities in sensory inputs at multiple complexity levels, with neuronal mechanisms that remain elusive. Here, we monitored auditory cortex activity during the local-global paradigm, a protocol nesting different regularity levels in sound sequences. We observed that mice encode local predictions based on stimulus occurrence and stimulus transition probabilities, because auditory responses are boosted upon prediction violation. This boosting was due to both short-term adaptation and an adaptation-independent surprise mechanism resisting anesthesia. In parallel, and only in wakefulness, VIP interneurons responded to the omission of the locally expected sound repeat at sequence ending, thus providing a chunking signal potentially useful for establishing global sequence structure. When this global structure was violated, by either shortening the sequence or ending it with a locally expected but globally unexpected sound transition, activity slightly increased in VIP and PV neurons respectively. Hence, distinct cellular mechanisms predict different regularity levels in sound sequences.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Data for 'NovoRank: Refinement for De Novo Peptide Sequencing Based on Spectral Clustering and Deep Learning'

<p>The uploaded files include two archives for <a href="https://pubs.acs.org/doi/10.1021/acs.jproteome.4c00300" target="_blank" rel="noopener">NovoRank: Refinement for De Novo Peptide Sequencing Based on Spectral Clustering and Deep Learning</a>. The '<em>mgf_data</em>' archive contains all MGF files used in the study, while the '<em>sample_data</em>' archive includes sequencing data, clustering data generated using <code>MSCluster</code>, and XCorr calculation data computed with <code>CometX</code>, all of which were used in the research.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Geodetic model of the March 2021 Thessaly seismic sequence inferred from seismological and InSAR data

<p>A selection of Sentinel-1 (S1) wrapped and unwrapped measurements used in this study (from &quot;a&quot; to &quot;u&quot; files in tiff format as indicated in the word file attached). S1 data were processed by using our own internally developed InSAR&nbsp;processing chain.<br> <br> Earthquakes data locations.</p> <p><br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2021View details →
dryad36/100

Data from: RAD sequencing, genotyping error estimation and de novo assembly optimization for population genetic inference

Restriction site-associated DNA sequencing (RADseq) provides researchers with the ability to record genetic polymorphism across thousands of loci for non-model organisms, potentially revolutionising the field of molecular ecology. However, as with other genotyping methods, RADseq is prone to a number of sources of error that may have consequential effects for population genetic inferences, and these have received only limited attention in terms of the estimation and reporting of genotyping error rates. Here we use individual sample replicates, under the expectation of identical genotypes, to quantify genotyping error in the absence of a reference genome. We then use sample replicates to (1) optimize de novo assembly parameters within the program Stacks, by minimizing error and maximizing the retrieval of informative loci, and; (2) quantify error rates for loci, alleles and SNPs. As an empirical example we use a double digest RAD dataset of a non-model plant species, Berberis alpina, collected from high altitude mountains in Mexico.

opencc-zeroDec 2013View details →
dryad36/100

Data from: How "simple" methodological decisions affect interpretation of population structure based on reduced representation library DNA sequencing: a case study using the lake whitefish

Reduced representation (RRL) sequencing approaches (e.g., RADSeq, genotyping by sequencing) require decisions about how much to invest in genome coverage and sequencing depth (library quality), as well as choices of values for adjustable bioinformatics parameters. To empirically explore the importance of these "simple" decisions, we generated two independent sequencing libraries for the same 142 individual lake whitefish (Coregonus clupeaformis) using a nextRAD RRL approach: (1) A small number of loci and low sequencing depth (library A); and (2) more loci and higher sequencing depth (library B). The fish were selected from populations with different levels of expected genetic subdivision. Each library was analyzed using the STACKS pipeline followed by three types of population structure assessment (FST, DAPC and ADMIXTURE) with iterative increases in the stringency of sequencing depth and missing data requirements, as well as more specific a priori population maps. Library B was always able to resolve strong population differentiation in all three types of assessment regardless of the selected parameters. In contrast, library A produced more variable results; increasing the minimum sequencing depth threshold (-m) resulted in a reduced number of retained loci, and therefore lost resolution at high -m values for FST and ADMIXTURE, but not DAPC. FST and DAPC were robust to varying the population map and increasing the stringency of missing data requirements. In contrast, ADMIXTURE was unable to resolve strong population differentiation when increasing these same parameters in library A. Similarly, when examining fine scale population subdivision, library B was robust to changing parameters but library A lost resolution depending on the parameter set. We used library B to examine actual subdivision in our study populations. All three types of analysis found complete subdivision among populations in Lake Huron, ON and Dore Lake, SK, Canada using 10,640 SNP loci. Weak population subdivision was detected in Lake Huron with fish from sites in the north-west, Search Bay, North Point and Hammond Bay, showing slight differentiation. Overall, we show that apparently simple decisions about library quality and bioinformatics parameters can have potentially important impacts on the interpretation of population subdivision. Although costly, the early investment in a high-quality library and more conservative stringency settings on STACKS parameters lead to a final dataset that was more consistent and robust when examining both weak and strong population differentiation.

opencc-zeroMar 2020View details →
zenodo36/100

Training data for "Identification of allelic variants in SARS-CoV-2 from deep sequencing reads"

<p>Effectively monitoring global infectious disease crises, such as the COVID-19 pandemic, requires capacity to generate and analyze large volumes of sequencing data in near real time. These data have proven essential for monitoring the emergence and spread of new variants, and for understanding the evolutionary dynamics of the virus.</p> <p>Two sequencing platforms in combination with several established library preparation strategies are predominantly used to generate SARS-CoV-2 sequence data. However, data alone do not equal knowledge: they need to be analyzed. The Galaxy community developed analysis workflows to support the <strong>identification of allelic variants (AVs) in SARS-CoV-2 from deep sequencing reads</strong>.</p> <p>These workflows allow one to identify AVs and lineages in SARS-CoV-2 genomes with variant allele frequencies ranging from 5% to 100% (i.e., they detect variants with intermediate frequencies as well.</p> <p>In this tutorial we will see how to run these workflows for the different types of input data:</p> <ul> <li>Single end data derived from Illumina-based RNAseq experiments</li> <li>Paired end data derived from Illumina-based RNAseq experiments</li> <li>Paired-end data generated with Illumina-based Ampliconic (ARTIC) protocols</li> <li>ONT fastq files generated with Oxford nanopore (ONT)-based Ampliconic (ARTIC) protocols</li> </ul> <p>To illustrate the tutorial, we took some example datasets (paired-end data generated with Illumina-based Ampliconic (ARTIC) protocols) from COG-UK, the COVID-19 Genomics UK Consortium.</p>

opencc-by-4.0Jun 2021View details →
dryad36/100

Data from: A RAD-sequencing approach to genome-wide marker discovery, genotyping, and phylogenetic inference in a diverse radiation of primates

Until recently, most phylogenetic and population genetics studies of nonhuman primates have relied on mitochondrial DNA and/or a small number of nuclear DNA markers, which can limit our understanding of primate evolutionary and population history. Here, we describe a cost-effective reduced representation method (ddRAD-seq) for identifying and genotyping large numbers of SNP loci for taxa from across the New World monkeys, a diverse radiation of primates that shared a common ancestor ~20-26 mya. We also estimate, for the first time, the phylogenetic relationships among 15 of the 22 currently-recognized genera of New World monkeys using ddRAD-seq SNP data using both maximum likelihood and quartet-based coalescent methods. Our phylogenetic analyses robustly reconstructed three monophyletic clades corresponding to the three families of extant platyrrhines (Atelidae, Pitheciidae and Cebidae), with Pitheciidae as basal within the radiation. At the genus level, our results conformed well with previous phylogenetic studies and provide additional information relevant to the problematic position of the owl monkey (Aotus) within the family Cebidae, suggesting a need for further exploration of incomplete lineage sorting and other explanations for phylogenetic discordance, including introgression. Our study additionally provides one of the first applications of next-generation sequencing methods to the inference of phylogenetic history across an old, diverse radiation of mammals and highlights the broad promise and utility of ddRAD-seq data for molecular primatology.

opencc-zeroDec 2017View details →
zenodo36/100

Training data for the Sei framework sequence model

<p>Training data for the Sei framework deep learning model. The data contains chromatin profiles from the Cistrome Project: <strong>please agree to the terms of usage at the Cistrome Project (http://cistrome.org/db/#/bdown) before downloading.&nbsp;</strong></p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Supplementary Data for: Whole genome sequencing elucidates the species-wide diversity and evolution of fungicide resistance in the early blight pathogen Alternaria solani

<p>Supplementary Data for: Whole genome sequencing elucidates the species-wide diversity and evolution of fungicide resistance in the early blight pathogen Alternaria solani</p> <p>This repository contains:</p> <p>SNP call data / VCF file</p> <p>Scripts for all processing steps from mapping up to PCA and phylogenetic analyses (script.ts)<br> Scripts for population genomic analyses with LEA and PopGenome (scripts.SE)<br> All script names are self explanatory.</p>

opencc-by-4.0Jun 2021View details →
dryad36/100

Data from: Palynology of a short sequence of the Lower Devonian Beartooth Butte Formation at Cottonwood Canyon (Wyoming): Age, depositional environments and plant diversity

<p>The Beartooth Butte Formation hosts the most extensive Early Devonian macroflora of western North America.  The age of the flora at Cottonwood Canyon (Wyoming) has been constrained to the Lochkovian-Pragian interval, based on fish biostratigraphy and unpublished palynological data.  We present a detailed palynological analysis of the plant-bearing layers at Cottonwood Canyon.  The palynomorphs comprise 32 spore, five cryptospore, two prasinophycean algae and an acritarch species.  The stratigraphic ranges of these palynomorphs indicate a late Lochkovian - Pragian age, confirming previous age assignments.  Analyses on samples from three different depositional environments of the plant-bearing sequence – layers with in situ lycophyte populations, flood layers that buried those populations and an organic matter accumulation zone within a flood layer – demonstrate distinct palynofacies. Comparisons between palynomorph and plant macrofossil diversity reveal some discrepancies.  Whereas zosterophylls and lycophytes, most diverse and abundant among the macrofossils, have only one known corresponding spore type (assignable to zosterophylls) in the palynomorph assemblage, the trimerophytes, rare in the macrofossil assemblage, are represented by three spore types.  Some of these discrepancies reflect taphonomic differences between macrofossils and palynomorphs, others could be due to the fact that the parent plants of most palynomorph types in the Cottonwood Canyon assemblage are unknown.  These observations emphasize the need for concerted efforts to bring together the knowledge of macro- and microfloras within Early Devonian localities.  Nevertheless, given the palaeophytogeographic significance of the Beartooth Butte Formation flora, its palyno- and macrofossil assemblages, taken together, provide new data relevant to future discussions of Early Devonian biogeography.</p>

opencc-zeroJul 2021View details →
zenodo36/100

Pre-processed IgH repertoire sequencing data from BioProject PRJNA748239

<p><strong>Data Processing</strong></p> <p>Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2).&nbsp;Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2)</p> <p>&nbsp;</p> <p>software_versions&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8</p> <p>quality_thresholds&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;FilterSeq.py pRESTO Q&gt;20</p> <p>paired_reads_assembly&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001</p> <p>primer_match_cutoffs&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;MaskPrimers.py pRESTO C primer &amp; V primer maxerror 0.2</p> <p>consensus_building&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5</p> <p>collapsing_method&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;CollapseSeq.py pRESTO</p> <p>germline_database&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;IMGT</p> <p>&nbsp;</p> <p>Format</p> <p>&nbsp;</p> <p>Processed sequences are provided in a tab delimited file format, including the following annotations:</p> <p>&nbsp;</p> <p><strong>C_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Isotype subclass</p> <p><strong>SEQUENCE_ID</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Sequence identifier</p> <p><strong>V_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;V segment gene and allele</p> <p><strong>D_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;D segment gene and allele</p> <p><strong>J_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;J segment gene and allele</p> <p><strong>JUNCTION_LENGTH</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Junction length</p> <p><strong>CONSCOUNT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence.</p> <p><strong>DUPCOUNT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;UMI count for the given unique sequence</p> <p><strong>ISOTYPE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Constant region primer (isotype)</p> <p><strong>MU_COUNT_CDR_R</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of replacement mutations in CDR region</p> <p><strong>MU_COUNT_CDR_S</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of silent mutations in CDR region</p> <p><strong>MU_COUNT_FWR_R</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of replacement mutations in FWR region</p> <p><strong>MU_COUNT_FWR_S</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of silent mutations in FWR region</p> <p><strong>MUT_TOTAL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Total number of mutations in V gene&nbsp;</p> <p><strong>NP_LENGTH</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Total number of N and P additions<strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</strong></p> <p><strong>SEQUENCE_INPUT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Full length sequence</p> <p><strong>SEQUENCE_IMGT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Gapped IMGT sequence</p> <p><strong>V_GERM_START_VDJ</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;position of the first nucleotide in ungapped V germline sequence alignment</p> <p><strong>JUNCTION</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Junction nucleotide sequence</p> <p><strong>GERMLINE_IMGT_D_MASK</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;IMGT-gapped germline nucleotide sequence with ns masking the NP1-D-NP2 regions</p> <p><strong>CDR3_AA_GRAVY</strong>&nbsp; &nbsp; &nbsp; &nbsp;CDR3 hydrophobicity</p> <p><strong>CDR3_AA_BULK</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; CDR3 bulkiness</p> <p><strong>CDR3_AA_ALIPHATIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Normalized aliphatic index</p> <p><strong>CDR3_AA_POLARITY</strong>&nbsp; &nbsp; &nbsp; &nbsp; CDR3 polarity</p> <p><strong>CDR3_AA_CHARGE</strong>&nbsp; &nbsp; &nbsp; &nbsp; normalised net&nbsp;charge</p> <p><strong>CDR3_AA_BASIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Basic side chain residue content</p> <p><strong>CDR3_AA_ACIDIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Acidic side chain residue content</p> <p><strong>CDR3_AA_AROMATIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;aromatic side chain conten</p> <p><strong>Subset</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Defined B cell subset&nbsp;</p> <p><strong>Repertoire</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG)</p> <p><strong>R_SCDR</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;R/S ratio in CDR region</p> <p><strong>R_SFWR</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;R/S ratio in FWR region</p> <p><strong>V_GENE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;V segment gene</p> <p><strong>D_GENE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;D segment gene</p> <p><strong>J_GENE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;J segment gene</p> <p><strong>V_FAM</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;V family gene</p> <p><strong>Run</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;ID of sequencing run</p> <p><strong>Sex</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Sex of the Subject</p> <p><strong>Age</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Age of the subject</p> <p><strong>UNIQUE_ID</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Subject identifier&nbsp;</p> <p><strong>SAMPLE</strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Sample identifier, linking back to raw data</p> <p><strong>Bcellno</strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Number of input B cells</p> <p><strong>Cells</strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Cell type&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>1.&nbsp;Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O&rsquo;Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein.&nbsp;2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires.&nbsp;<em>Bioinformatics</em>30: 1930&ndash;1932.</p> <p>2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein.&nbsp;2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data.&nbsp;<em>Bioinformatics</em>31: 3356&ndash;3358.</p> <p>3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool.&nbsp;<em>Nucleic Acids Res.</em>41.</p> <p>4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads.&nbsp;<em>Genome Res.</em>21: 936&ndash;939.</p>

opencc-by-4.0Aug 2021View details →
dryad36/100

Alfalfa genotyping-by-sequencing (GBS) data

<p>Alfalfa (<i>Medicago</i> <i>sativa</i> L.) quantitative trait loci (QTL) mapping population (184 F<sub>1</sub>) derived from cultivars 3010 (cold-tolerant) as female parent and CW 100 (cold-sensitive) as male parent were genotyped using genotyping-by-sequencing (GBS). Polymorphic SNPs unique to either 3010 (AB x AA) or CW 1010 (AA x AB) were identified as single dose allele (SDA) markers and used to generate the genetic linkage maps. Two sets of linkage maps, a set for each parent, were used to map the traits and the QTL were identified. With the genotyping and phenotyping informations we were able to map various alfalfa traits such as fall dormancy, winter-hardiness, freezing tolerance, flowering time, yield and leaf-rust resistance. The raw sequence data were deposited at NCBI SRA with the accession number SRP150116. This study identified several genomic regions and associated markers that can be further utilized in marker-assisted breeding to improve the alfalfa. </p>

opencc-zeroAug 2021View details →
zenodo36/100

Data for "Unsupervised learning of sequence-specific aggregation behavior for a model copolymer"

<p>These are the data associated with the paper, &quot;Unsupervised learning of sequence-specific aggregation behavior for a model copolymer&quot; (DOI 10.1039/D1SM01012C). Each of the directories contains subdirectories with `GSD` files dumped from HOOMD. Each subdirectory roughly corresponds to one or two of the figures in the paper.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Simulated wastewater sequencing data for benchmarking SARS-CoV-2 variant abundance estimation

<p>To evaluate the accuracy of variant abundance&nbsp;predictions from wastewater sequencing, we built a collection of benchmarking datasets that resemble real wastewater samples. For each variant (B.1.1.7, B.1.351, B.1.427, B.1.429, P.1) we created a series of 33 benchmarks by simulating sequencing reads from a variant genome, as well as a collection of background (non-variant of concern/interest) sequences, such that the variant abundance ranges from 0.05% to 100%. Analogously, we created a second series of benchmarks, simulating reads only from the Spike gene of each SARS-CoV-2 genome. We refer to the first set of benchmarks as &quot;whole genome&quot; (WG)&nbsp;and to the second set of benchmarks as &quot;S-only&quot;. We repeated these simulations at different sequencing depths: 100x and 1000x coverage for the whole genome benchmarks, and 100x, 1000x, and 10,000x coverage for the S-only benchmarks.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Fig. 5 in Copelatus sibelaemontis sp. nov. (Coleoptera: Dytiscidae) from the Moluccas with generic assignment based on morphology and DNA sequence data

Fig. 5. Distribution of Copelatus sibelaemontis sp. nov.

opencc-by-4.0Dec 2010View details →
zenodo36/100

Fig. 23 in Phylogenetic Studies On Didelphid Marsupials Ii. Nonmolecular Data And New Irbp Sequences: Separate And Combined Analyses Of Didelphine Relationships With Denser Taxon Sampling

Fig. 23. Skull of Tlacuatzin canescens, a composite drawing based on USNM 125659 and 511261.

opencc-by-4.0Aug 2003View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record