Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Figure 1 in Systematic revision of Sabellariidae (Polychaeta) and their relationships with other polychaetes using morphological and DNA sequence data
Figure 1. Trees resulting from parsimony analyses of Sabellariidae and previously related taxa (including members of Sabellida, Terebellida and Spionida). A, strict consensus after analyses based on 99 morphological features with jackknife support values. B, first of 25 most-parsimonious trees (TL 177, CI 0.58, RI 0.74) after analyses of morphological data with unambiguous changes marked on the topology. Numbers under nodes indicate jackknife values; black dots: synapomorphies, white dots: homoplastic character states. C, shortest tree (TL 9215, CI 0.54, RI 0.36) resulting from analysis of partial 18S, 28S and EF-1a sequences with jackknife support values. D, strict consensus of two most-parsimonious trees (TL 9434, CI 0.54 RI 0.37) of the combined dataset, with jackknife support values.
Figure 2 in Systematic revision of Sabellariidae (Polychaeta) and their relationships with other polychaetes using morphological and DNA sequence data
Figure 2. Strict consensus of 429 most-parsimonious trees after maximum-parsimony analysis of morphological data of members of Sabellariidae rooted with Spionidae (TL 73, CI 0.55, RI 0.86). Numbers under nodes indicate jackknife support values.
Figure 3 in Nucleotide sequence data confirm diagnosis and local endemism of variable morphospecies of Andean astroblepid catfishes (Siluriformes: Astroblepidae)
Figure 3. Results of the phylogenetic analysis of astroblepid morphospecies obtained from maximum likelihood analysis of the combined DNA sequence data set. Numerals at nodes represent bootstrap proportions (values less than 50% not shown); stars represent nodes supported by bootstrap values of 80% or greater. Sample numbers correspond with materials listed in Table 1. Letters designate morphospecies; shaded boxes denote monophyletic assemblages of population samples.
Figure 1 in Nucleotide sequence data confirm diagnosis and local endemism of variable morphospecies of Andean astroblepid catfishes (Siluriformes: Astroblepidae)
Figure 1. Variation in pigmentation in Astroblepus morphospecies A–I. A, morphospecies A, ANSP (Academy of Natural Sciences of Philadelphia) 180586 (4793), 51.6 mm standard length (SL), Araza River. B, morphospecies B, ANSP 180587 (4779), 75 mm SL, Araza River. C, morphospecies B, ANSP 180582 (4801), 80.4 mm SL, Araza drainage (Dr.) D, morphospecies B, ANSP 180582 (4800), 54.5 mm SL, Araza Dr. E, morphospecies C, ANSP 180581 (4805), 27.2 mm SL, Araza Dr. F, morphospecies C, ANSP 180586 (4794), 58 mm SL, Araza River. G, morphospecies D, ANSP 180599 (4822), 51.7 mm SL, Urubamba Dr. H, morphospecies D, ANSP 180602 (4499), 85 mm SL, Urubamba Dr. I, morphospecies H, ANSP 180618 (4423), 46.3 mm SL, Apurimac Dr. J, morphospecies H, ANSP 180616 (4436), 79.2 mm SL, Apurimac Dr. K, morphospecies E, ANSP 180595 (4785), 61.3 mm SL, Urubamba Dr. L, morphospecies E, ANSP 180605 (4490), 110.5 mm SL, Apurimac Dr. M, morphospecies F, ANSP 180606 (4487), 75.7 mm SL, Apurimac Dr. N, morphospecies F, ANSP 180601 (4759), 52.6 mm SL, Urubamba Dr. O, morphospecies G, ANSP 180588 (4787), 59.5 mm SL, Urubamba Dr. P, morphospecies I, ANSP 180607 (4477), 39.4 mm SL, Apurimac Dr. Photo in (A) by S. A. S.; photos in (B–P) by M. H. S. P.
Figure 2 in Nucleotide sequence data confirm diagnosis and local endemism of variable morphospecies of Andean astroblepid catfishes (Siluriformes: Astroblepidae)
Figure 2. Distribution of astroblepid morphospecies and study region. Circled letters correspond with the morphospecies designations (Table 1) and may represent more than one lot or collection locality.
KTU: K-mer Taxonomic Units improve the biological relevance of amplicon sequence variant microbiota data
<p>Testing datasets and files for the KTU algorithm</p>
SMRT sequencing data on four hypersuppressive Saccharomyces cereviciae mitochondrial DNAs
<p>Hypersuppressive mitochondrial DNAs are thought to be linear tandem repeats of the base unit of specific ORI regions on the mitochondrial genome. Here we confirm the linear tandem repeats using SMRT sequencing technology on four hypersuppressive clones. Mitochondrial DNA from four <em>Saccharomyces cervisiae</em> hypersuppressive mutants and a wild-type control was enriched and sequenced using SMRT sequencing technology.</p>
FIGURE. Median network analyses (MNA) of a subset of the C. trilobus aggregate (i.e. those in the clade A from Fig. 11) based on concatenated DNA sequence data from ITS, trnL-trnF and psbJ-petA. Stars and arrow indicate accessions discussed in the text. NI: North Island, SI: South Island. in Five new species of Corybas (Diurideae, Orchidaceae) endemic to New Zealand and phylogeny of the Nematoceras clade
FIGURE. Median network analyses (MNA) of a subset of the C. trilobus aggregate (i.e. those in the clade A from Fig. 11) based on concatenated DNA sequence data from ITS, trnL-trnF and psbJ-petA. Stars and arrow indicate accessions discussed in the text. NI: North Island, SI: South Island.
FIGURE. Bayesian tree of New Zealand spider orchids (Corybas) based on DNA sequence data from ITS, trnL-trnF and psbJ-petA. Major clades are indicated by open bars and capital letters, members of the C. trilobus aggregate are shaded, and posterior probabilities/ bootstrap percentages (≥50) indicated by numbers near each node. NI: North Island, SI: South Island, MCQI: Macquarie Island, CHI: Chatham Island in Five new species of Corybas (Diurideae, Orchidaceae) endemic to New Zealand and phylogeny of the Nematoceras clade
FIGURE. Bayesian tree of New Zealand spider orchids (Corybas) based on DNA sequence data from ITS, trnL-trnF and psbJ-petA. Major clades are indicated by open bars and capital letters, members of the C. trilobus aggregate are shaded, and posterior probabilities/ bootstrap percentages (≥50) indicated by numbers near each node. NI: North Island, SI: South Island, MCQI: Macquarie Island, CHI: Chatham Island
16S V4 raw read count data; 16S reads metadata; new MHC class II allele sequences
<p>Pathogen-mediated selection at the major histocompatibility complex (MHC) is thought to promote MHC-based mate choice in vertebrates. Mounting evidence implicates odour in conveying MHC genotype, but the underlying mechanisms remain uncertain. MHC effects on odour may be mediated by odour-producing symbiotic microbes whose community structure is shaped by MHC genotype. In birds, preen oil is the primary source of body odour and similarity at MHC predicts similarity in preen oil composition. Hypothesizing that this relationship is mediated by symbiotic microbes, we characterized MHC genotype, preen gland microbial communities, and preen oil chemistry of song sparrows (<i>Melospiza melodia</i>). Consistent with the microbial mediation hypothesis, pairwise similarity at MHC predicted similarity in preen gland microbiota. Overall microbial similarity did not predict chemical similarity of preen oil, counter to this hypothesis. However, permutation testing identified a maximally predictive set of microbial taxa that best reflect MHC genotype, and another set of taxa that best predict preen oil chemical composition. The relative strengths of relationships between MHC and microbes, microbes and preen oil, and MHC and preen oil suggest that MHC may affect host odour both directly and indirectly. Thus, birds may assess MHC genotypes based on both host-associated and microbially-mediated odours.</p>
Genotyping-by-Sequencing data of weedy and domesticated Brassica rapa L.
<p>The study of domestication contributes to our knowledge of evolution and crop genetic resources. Human selection has shaped wild <em>Brassica rapa</em> into diverse turnip, leafy, and oilseed crops. Despite its worldwide economic importance and potential as a model for understanding diversification under domestication, insights into the number of domestication events and initial crop(s) domesticated in <em>B. rapa</em> have been limited due to a lack of clarity about the wild or feral status of conspecific non-crop relatives. To address this gap and reconstruct the domestication history of <em>B. rapa</em>, we analyzed 68,468 genotyping-by-sequencing-derived SNPs for 416 samples in the largest diversity panel of domesticated and weedy <em>B. rapa</em> to date. To further understand the center of origin, we modeled the potential range of wild <em>B. rapa</em> during the mid-Holocene. Our analyses of genetic diversity across <em>B. rapa</em> morphotypes suggest that non-crop samples from the Caucasus, Siberia, and Italy may be truly wild, while those occurring in the Americas and much of Europe are feral. Clustering, tree-based analyses, and parameterized demographic inference further indicate that turnips were likely the first crop type domesticated, from which leafy types in East Asia and Europe were selected from distinct lineages. These findings clarify the domestication history and nature of wild crop genetic resources for <em>B. rapa</em>, which provides the first step toward investigating cases of possible parallel selection, the domestication and feralization syndrome, and novel germplasm for <em>Brassica</em> crop improvement.</p>
FIGURE. Phylogram of Panus generated from Maximum likelihood analysis of ITS sequence data. Lentinus crinitus (MK408650) was selected as the outgroup taxon. Maximum likelihood bootstrap values greater than 60% are indicated above the nodes. The new record Panus similis (HKAS 121668) is in black bold. in Yunnan-Guizhou Plateau: a mycological hotspot
FIGURE. Phylogram of Panus generated from Maximum likelihood analysis of ITS sequence data. Lentinus crinitus (MK408650) was selected as the outgroup taxon. Maximum likelihood bootstrap values greater than 60% are indicated above the nodes. The new record Panus similis (HKAS 121668) is in black bold.
Benchmarking the Autoencoder Design for Imputing Single-Cell RNA Sequencing Data
<p>This repository contains the real and synthetic datasets used in the paper "Benchmarking the Autoencoder Design for Imputing Single-Cell RNA Sequencing Data". The zip file includes three folders:</p> <p>1. overall imputation accuracy: the 12 real scRNA-seq datasets used in the evaluation of overall imputation accuracy.</p> <p>2. cell clustering: the 20 real scRNA-seq datasets with cell type labels used in the evaluation of cell clustering.</p> <p>3. DE gene: the 20 scRNA-seq syntehtic datasets with ground-truth DE genes used in the evaluation of DE gene analysis. These datasets are simulated by simulator scDesign and 20 real datasets. </p> <p> </p>
FIGURE 1. A in A note on the identity of the spikenard (Nardostachys jatamansi, Caprifoliaceae) based on DNA sequence data
FIGURE 1. A photograph of N. jatamansi with pink-colored flowers. Inset shows close-up of flowers. bar=10cm.
Simulated hepatitis B virus (HBV) sequencing data and HBV sequence variation graph materials
<p>Simulated HBV sequencing data (InSilicoSeq, HiSeq error model) and a sequence variation graph constructed using HBV genome sequences described in <a href="https://doi.org/10.1099/jgv.0.001387">https://doi.org/10.1099/jgv.0.001387</a></p>
Data for: Warming rates alter sequence of disassembly in experimental communities
<p class="MsoNormal"><a name="_Hlk71747934"></a><span>This study analyzed patterns of species loss in a community of four rotifers and six ciliates exposed to three rates of extreme warming.</span><span> Immediately prior to each +0.5°C increase in temperature, we sampled replicate communities and identified all surviving species. The sampling temperature at which no further organisms of a given species were observed alive represented the temperature of loss for that species and replicate. Recording the identity of all surviving species at each incremental change in temperature allowed us to both observe the sequence of species loss from the community (i.e., the sequence of disassembly) and also generate a corresponding table of the distinct communities surviving at each temperature. To test the effects of warming rate on the sequence of disassembly, we analyzed the sequence (order) of species loss of the ten species in each replicate. An order of "1" was assigned to the first species lost, "2" to the second, "3" to the third, and so on. To assess the contribution of warming rate to variability in community composition, we compared all of the distinct communities of surviving species that were observed in hourly-, daily- and weekly-rate treatments throughout the period of ramping temperature. The percentage of all distinct communities that were rate-specific (observed in at least one but not all rate treatments) provided a quantitative measure of the contribution of warming rate (per se) to variability in community composition. <span>Downloaded files include: a) temperature of species loss data for hourly, daily and weekly rate replicates, b) order of species loss data for hourly, daily and weekly rate replicates, and c) R code for Monte Carlo analyses. </span></span></p>
Supplemental data for: Classification of the Celastrales based on integration of genomic, morphological, and Sanger-sequence characters
<p>We present the best sampled phylogenetic analysis of Celastrales, with respect to both character and taxon sampling, and use it to present a natural classification of the order. Parnassiaceae are highly supported as sister to Celastraceae; we recognize both families as distinct. <em>Pottingeria</em> is highly supported as a member of an early derived lineage within Celastraceae. We recognize and circumscribe 13 subfamilies in Celastraceae, including the new subfamilies Crossopetaloideae, Maytenoideae, Microtropioideae, Monimopetaloideae, and Salaciopsioideae. We identified five genera that likely require generic recircumscriptions: <em>Cassine</em>, <em>Elachyptera</em>, <em>Gymnosporia</em>, <em>Salacia</em>, and <em>Semialarium</em>. Genera that had not been previously sampled in Sanger-sequence-based studies are resolved as follows: <em>Arnicratea</em> is sister to <em>Reissantia</em>, <em>Bequaertia</em> is in a clade with <em>Campylostemon</em> and <em>Tristemonanthus</em>, <em>Goniodiscus</em> is sister to <em>Wilczekra</em>, <em>Ptelidium</em> is nested within <em>Elaeodendron</em>, and <em>Tetrasiphon</em> is most closely related to <em>Gyminda</em>.</p>
Sequencing data for "Identifying and tracking mobile elements..." by van Dijk et al. (2023)
<p>The dataset contains the following directories:</p> <p><strong>All_raw_MAGs</strong>: all cross-assembled MAGs per community, with a text file for their BAT annotation results. Cross-assembly was done by combining all time points. <br> <strong>Candidatus_Saccharibacterium_MAGs</strong>: identified nanobacterium MAGs from various communities. <br> <strong>Cellvibrio_MAGs</strong>: identified cellvibrio spp. across communities. <br> <strong>Cellvibrio_Plasmids:</strong> Cp plasmids belonging to Cellvibrio (C1) which transferred from communities 8 to 4, 6, 9, and 10 (in horizontal communities)<br> <strong>Interactive_Dataset:</strong> HTML5 (Plotly) graphs to explore all xenotypic sequences and MAG abundance plots<br> <strong>Rscript_mock_data</strong>: R script used to design the mock data for the benchmark set. </p> <p>And a single spreadsheet (<strong>Xenoseq_mastersheet.xlsx, </strong>Supplementary Table 1) of data discussed in the main text.</p>
Data for: Combining target enrichment and Sanger sequencing data to clarify the systematics of the diverse Neotropical butterfly subtribe Euptychiina (Satyrinae, Nymphalidae)
<p>The diverse, largely Neotropical subtribe Euptychiina (Satyrinae, Nymphalidae) is widely regarded as one of the most taxonomically challenging groups among all butterflies. Over the last two decades, morphological and molecular studies have revealed widespread paraphyly and polyphyly among genera, and a comprehensive, robust phylogenetic hypothesis is needed to build a firm generic classification to support ongoing taxonomic revisions at the species level. Here, we generated a dataset which includes sequences for up to nine nuclear genes and the mitochondrial COI 'barcode' for a total of 1280 specimens representing 449 described and undescribed species of Euptychiina and 39 outgroups, resulting in the most complete phylogeny for the subtribe to date. In combination with a recently developed genomic backbone tree this dataset resulted in a topology with strong support for most branches. </p>
Synthetic data for Aligning Distant Sequences to Graphs using Long Seed Sketches
<p>Each directory inside the folder correponds to the number of levels used to generate the dataset (for more details, see the description written in the publication). Inside each directory, the files "reference_X" and "mutated_X" correpond to the sequences reference and mutated at rate X, respectively. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.