Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Fig. 19 in Phylogenetic Studies On Didelphid Marsupials Ii. Nonmolecular Data And New Irbp Sequences: Separate And Combined Analyses Of Didelphine Relationships With Denser Taxon Sampling
Fig. 19. All equally mostparsimonious resolutions of the basal didelphine polytomy in figures 18 and 21. A, Resolution supported by 72 mostparsimonious trees (MPTs) from the IRBP1 analysis and 6 MPTs from the IRBP2 analysis; B, resolution supported by 72 MPTs from the IRBP1 analysis and 6 MPTs from the IRBP2 analysis; C, resolution supported by 36 MPTs from the IRBP1 analysis and 3 MPTs from the IRBP2 analysis; D, resolution supported by 36 MPTs from the IRBP1 analysis, 6 MPTs from the IRBP2 analysis, and 8 MPTs from the combined analysis; E, resolution supported by 36 MPTs from the IRBP1 analysis, 6 MPTs from the IRBP2 analysis, and 8 MPTs from the combined analysis; F, resolution supported by 18 MPTs from the combined analysis only.
Fig. 13 in Phylogenetic Studies On Didelphid Marsupials Ii. Nonmolecular Data And New Irbp Sequences: Separate And Combined Analyses Of Didelphine Relationships With Denser Taxon Sampling
Fig. 13. Anterolingual views of left M3 illustrating taxonomic differences in cingular morphology. Left, Marmosa murina (AMNH 272870) with preprotocrista and anterolabial cingulum joined to form a continuous shelf along the anterior margin of the tooth crown. Right, Monodelphis adusta (AMNH 272781) with separate crista and cingulum (no continuous shelf).
FIG. 1 in DNA Sequence Data from the Holotype of Marmosa elegans coquimbensis Tate, 1931 (Mammalia: Didelphidae) Resolve Its Disputed Relationships
FIG. 1. Bayesian phylogenetic tree of Thylamys cytochrome b sequences. Numbers at nodes indicate posterior probabilities (PP). Filled circles at nodes denote PPs equal to 1.0. Unmarked nodes received PPs less than 0.5. Within T. elegans, tips are labeled with country, region, specimen identifier, and, in parentheses, a GenBank accession number. The holotype of Marmosa elegans coquimbensis Tate, 1931, is in boldface type. For other species, tips of the phylogeny are collapsed and the outgroup is not shown. See appendix 1 for a full list of sequences included in the phylogeny. A full tree file corresponding to this topology is available on TreeBase (doi: http://purl.org/phylo/treebase/phylows/study/TB2:S25505).
CNN models and training, validation and test datasets for "PlotMI: interpretation of pairwise interactions and positional preferences learned by a deep learning model from sequence data"
<p>Convolutional neural network (CNN) models and their respective training, validation and test datasets used in manuscript:</p> <p>Tuomo Hartonen, Teemu Kivioja and Jussi Taipale, "PlotMI: interpretation of pairwise interactions and positional preferences learned by a deep learning model from sequence data"</p>
Data related to research article: Towards mouse genetic-specific RNA-sequencing read mapping
<p>This dataset contains data related to the research article: "Towards mouse genetic-specific RNA-sequencing read mapping".</p>
Results from the revision of MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data
<p>afr47041.zip, lat36378.zip, and eur115620.zip contain All of Us Summary Statistics used in the revised version of "MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data". Summary statistics for three cohorts are included: Afr47k, Lat36k, and Eur116k. These cohorts have not been downsampled to have equal levels of missingness.</p> <p>pips.tsv contains fine-mapped variants with PIP > 0.01 via MultiSuSiE from the revised version of "MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data". Subcohorts with the _unmatched suffix have not been downsampled to have equal levels of phenotyped missingness across ancestries. </p> <p>MultiSuSiE-main.zip contains the MultiSuSiE software packages (corresponds to the Github repo on 10/16/2025).</p> <p>Please cite:</p> <p>Rossen, Jordan, et al. "MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data." <em>medRxiv</em> (2024): 2024-05.</p> <p> </p>
Sequence data for 'Machine-driven parameter-space exploration of biochemical reactions'
<p>The development of complex, multi-step <em>omics</em> methods in molecular biology is a laborious, costly, iterative and often intuition-bound process where an optimum is sought in a parameter space through step-by-step optimisations. The the difficulty of miniaturising assays and the cost of the experiments limit the dynamic range and the number of parameters that can be explored. However, because of non-linearities of the response of biochemical systems to their reagent concentrations, a broad dynamic range is necessary. Here we demonstrate the use of a high-performance nanoliter handling platform (Labcyte Echo 525) and computer generation of liquid transfer programs to explore in quadruplicates more than 600 combination of 4 parameters of a biochemical reaction, which lead us to uncover non-linear responses, parameter interactions and novel mechanical insights. With the increased availability of « <em>cloud biology</em>» computer-driven laboratory platforms, our results participate in changing methods development for biotechnology towards reproducible, computer-aided exhaustive characterisation of biochemical systems.</p> <p>This dataset contains the raw sequencing data produced with an Illumina MiSeq instrument for this project. FASTQ files and sample sheets are found in the usual location (Data/Intensities/BaseCalls). The "Thumbnail_Images" and "L001" directories were deleted to save space.</p> <p>Run IDs: 171227_M00528_0321_000000000-B4GLP, 180403_M00528_0348_000000000-B4GP8, 180517_M00528_0364_000000000-BRGK6, 180123_M00528_0325_000000000-B4PCK, 180411_M00528_0351_000000000-BN3BL, 180606_M00528_0367_000000000-BN3FG, 180326_M00528_0346_000000000-B4GJR, 180501_M00528_0359_000000000-B4PJY, 180607_M00528_0368_000000000-BN9KM</p> <p> </p>
Additional files of scTensor paper "scTensor detects many-to-many cell-cell interactions from single cell RNA-sequencing data"
<p>Complex biological systems are described as a multitude of cell-cell interactions (CCIs). Recent single-cell RNA-sequencing studies focus on CCIs based on ligand-receptor (L-R) gene co-expression. However, the analytical methods are still not mature; such methods cannot detect CCIs and the related L-R pairs simultaneously or also are not appropriate to detect many-to-many CCIs.</p> <p>In this work, we propose scTensor, a novel method for extracting representative triadic relationships (or hypergraphs), which include ligand-expression, receptor-expression, and related L-R pairs. Through extensive studies with simulated and empirical datasets, we have shown that scTensor could detect some hypergraphs, which cannot be detected by conventional methods, especially when those CCIs are many-to-many relationships.</p>
Data from: Cost-saving population genomic investigation of Daphnia longispina complex resting eggs using whole genome amplification and pre-sequencing screening
<p>This dataset contains all paired MiSeq sequences that were generated for the study "Cost-saving population genomic investigation of<em> Daphnia longispina</em> complex resting eggs using whole genome amplification and pre-sequencing screening" by Nickel and Cordellier.</p> <p>The sample names used in the study and the associated file names are explained in the table<strong> </strong>"Study_sample_names.xlsx"</p> <p> </p>
Single-cell naïve IgM VH:VL sequence data from 22 Kymice
<p>Single-cell VH:VL sequencing data derived from naïve B-cells isolated from 22 Kymice. This dataset is published as part of the review process for the following preprint: https://www.biorxiv.org/content/10.1101/2022.06.27.497709v1. </p>
Data for: High-throughput profiling of sequence recognition by tyrosine kinases and SH2 domains using bacterial peptide display
<p>Tyrosine kinases and SH2 (phosphotyrosine recognition) domains have binding specificities that depend on the amino acid sequence surrounding the target (phospho)tyrosine residue. Although the preferred recognition motifs of many kinases and SH2 domains are known, we lack a quantitative description of sequence specificity that could guide predictions about signaling pathways or be used to design sequences for biomedical applications. Here, we present a platform that combines genetically-encoded peptide libraries and deep sequencing to profile sequence recognition by tyrosine kinases and SH2 domains. We screened several tyrosine kinases against a million-peptide random library and used the resulting profiles to design high-activity sequences. We also screened several kinases against a library containing thousands of human proteome-derived peptides and their naturally-occurring variants. These screens recapitulated independently measured phosphorylation rates and revealed hundreds of phosphosite-proximal mutations that impact phosphosite recognition by tyrosine kinases. We extended this platform to the analysis of SH2 domains and showed that screens could predict relative binding affinities. Finally, we expanded our method to assess the impact of non-canonical and post-translationally modified amino acids on sequence recognition. This specificity profiling platform will shed new light on phosphotyrosine signaling and could readily be adapted to other protein modification/recognition domains.</p>
Supplementary data of the paper 'Adaptive trends of sequence compositional complexity over pandemic time in the SARS CoV 2 coronavirus'
<p>Supplement of the paper<br> "Adaptive trends of sequence compositional complexity over pandemic time in the SARS-CoV-2 coronavirus”<br> During the spread of the COVID-19 pandemic, the SARS-CoV-2 coronavirus underwent mutation and recombination events that altered its genome compositional structure, thus providing an unprecedented opportunity to check an evolutionary process in real time. The mutation rate is known to be lower than expected for neutral evolution, suggesting natural selection and convergent evolution. We begin by summarizing the compositional heterogeneity of each viral genome by computing its Sequence Compositional Complexity (SCC). To analyze the full range of SCC diversity, we select random samples of high quality coronavirus genomes covering the full span of the pandemic. We then search for evolutionary trends that could inform us on the adaptive process of the virus to its human host by computing the phylogenetic ridge regression of SCC against time (i.e., the collection date of each viral isolate). In early samples, we find no statistical support for any trend in SCC values, although the viral genome appears to evolve faster than Brownian Motion (BM) expectation. However, in samples taken after the emergence of high fitness variants, and despite the brief time span elapsed, a driven decreasing trend for SCC and an increasing one for its absolute evolutionary rate are detected, pointing to a role for selection in the evolution of SCC in the coronavirus. We conclude that the higher fitness of variant genomes may have leads to adaptive trends of SCC over pandemic time in the coronavirus.</p> <p>Supplementary files</p> <table> <tbody> <tr> <td> <p><strong>File</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>SupplementaryTables S1-S19.zip</p> </td> <td> <p>Excel supplementary tables: The strain name, the collection date, and the SCC values for each analyzed genome.</p> </td> </tr> <tr> <td>nextstrain_ncov_open_global_timetree.nwk</td> <td>ML phylodynamic tree for the Nextstrain sample</td> </tr> <tr> <td> <p>SupplementaryTable S20.pdf</p> </td> <td> <p>A complete list acknowledging the authors, originating and submitting laboratories of the genetic sequences we used for the analysis of the Nextstrain sample.</p> </td> </tr> <tr> <td>Nextstrain_sample_fasta_3059.zip</td> <td>Nextstrain sample (sequences in Fasta format)</td> </tr> <tr> <td> <p>PhylogeneticTimetrees_NewickFormat.zip</p> </td> <td> <p>Phylogenetic timetrees (Newick format).</p> </td> </tr> </tbody> </table> <p> </p>
Data corresponding to: Evaluation of sequencing and PCR-based methods for the quantification of the viral genome formula
<p>Viruses show great diversity in their genome organisation. Multipartite viruses package their genome segments into separate particles, most or all of which are required to initiate infection in the host cell. The benefits of such seemingly inefficient genome organization are not well understood. One hypothesised benefit of multipartition is that it allows for flexible changes in gene expression by altering the frequency of each genome segment in different environments, such as encountering different host species. The ratio of the frequency of segments is termed the genome formula (GF). Thus far, formal studies quantifying the GF have been performed for well-characterised virus-host systems in experimental settings using RT-qPCR. However, to understand GF variation in natural populations or novel virus-host systems, a comparison of several methods for GF estimation including high-throughput sequencing (HTS) based methods is needed. Currently, it is unclear how HTS-methods compare a golden standard, such as RT-qPCR. Here we show a comparison of multiple GF quantification methods (RT-qPCR, RT-digital PCR, Illumina RNAseq and Nanopore direct RNA sequencing) using three host plants (<em>Nicotiana tabacum</em>, <em>Nicotiana benthamiana</em>, and <em>Chenopodium quinoa</em>) infected with cucumber mosaic virus (CMV), a tripartite RNA virus. Our results show that all methods give roughly similar results, though there is a significant method effect on genome formula estimates. While the RT-qPCR and RT-dPCR GF estimates are congruent, the GF estimates from HTS methods deviate from those found with PCR. Our findings emphasise the need to tailor the GF quantification method to the experimental aim, and highlight that it may not be possible to compare HTS and PCR-based methods directly. The difference in results between PCR-based methods and HTS highlights that the choice of quantification technique is not trivial.</p>
FIG. 7 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.
FIG. 7. — Median-joining networks based on trnSGG and rps4-trnS sequences: A, the Ptisana attenuata clade; B, the P. salicina/P. soluta comb. nov., stat. nov./P. smithii clade. The size of each circle is proportional to the haplotype frequency. Undetected intermediate haplotypes on nodes are shown as black circles and hatch marks represent mutational steps separating haplotypes.
FIG. 6 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.
FIG. 6. — Phylogram from the Bayesian phylogenetic analysis of the chloroplast DNA sequence data for Ptisana Murdock. Support values for branches are given in the order of Bayesian inference posterior probability; maximum parsimony bootstrap support; and maximum likelihood bootstrap support. Only values>0.80 PP and 60% BS are shown.
FIG. 4 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.
FIG. 4. — Distribution map for the New Caledonian endemic species of Ptisana attenuata (Labill.) Murdock (), P. rolandi-principis (Rosenst.) Christenh. (Δ), and P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov. (, with unvouchered field observations indicated by a broken outline). Shaded areas are ultramafic substrates. The collecting sites of the sequenced P. attenuata samples are indicated.
FIG. 5 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.
FIG. 5. — The holotype of Ptisana soluta (Compton) Murdock & Perrie, comb. nov., stat. nov. (Compton 1674, Ignambi, 1914, BM[BM000787128]) showing how the lamina transitions from 3-pinnate proximally to 2-pinnate distally. CC BY The Trustees of the Natural History Museum, London.
FIG. 3 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.
FIG. 3. — Ptisana soluta (Compton) Murdock & Perrie, comb. nov., stat. nov., field photos: A, frond with lamina 3-pinnate proximally and 2-pinnate distally. The frond at top-right is P. attenuata (Labill.) Murdock; B, stipes are greenish-brown at a distance; C, abaxial surface of costae and fertile lamina, showing transition from 3-pinnate to 2-pinnate; D, stipe greenish-brown and smooth; E, stipules around stipe bases. Photos: A-D, Leon Perrie from near Nouméa; E, Rémy Amice, from near Nouméa.
FIG. 1 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.
FIG. 1. — Ptisana attenuata (Labill.) Murdock, field photos: A, 3-pinnate frond; B, stipes are dark at a distance; C, abaxial surface of costae and lamina, with synangia; D, stipe dark and wrinkled; E, divided stipules around stipe bases. Photos: Leon Perrie, from near Nouméa.
Data from: Latent generative landscapes as maps of functional diversity in protein sequence space
<p>Variational autoencoders are unsupervised learning models with generative capabilities, when applied to protein data, they classify sequences by phylogeny and generate de novo sequences which preserve statistical properties of protein composition. While previous studies focus on clustering and generative features, here, we evaluate the underlying latent manifold in which sequence information is embedded. To investigate properties of the latent manifold, we utilize direct coupling analysis and a Potts Hamiltonian model to construct a latent generative landscape. We showcase how this landscape captures phylogenetic groupings, functional and fitness properties of several systems including Globins, β-lactamases, ion channels, and transcription factors. We provide support on how the landscape helps us understand the effects of sequence variability observed in experimental data and provides insights on directed and natural protein evolution. We propose that combining generative properties and functional predictive power of variational autoencoders and coevolutionary analysis could be beneficial in applications for protein engineering and design.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.