Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,109

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,109 results for “sequence analysis”

Learn how ShareScore rates datasets ↗
zenodo28/100

Figure 1 from: Fryssouli V, Zervakis GI, Polemis E, Typas MA (2020) A global meta-analysis of ITS rDNA sequences from material belonging to the genus Ganoderma (Basidiomycota, Polyporales) including new data from selected taxa. MycoKeys 75: 71-143. https://doi.org/10.3897/mycokeys.75.59872

Figure 1 a Initial labelling of 3908 Ganoderma sequences analysed in the present study: numbers in parentheses correspond to sequences deposited under the particular name in GenBank/ENA/DDBJ and UNITE, while species names appear underlined when ITS sequences derive from type material b final assigment of 3908 Ganoderma sequences to 80 species and six distinct groups as a result of the phylogenetic analyses performed in this study: numbers in parentheses correspond to the number of sequences grouped within each taxon (data deriving from Table 1 and Suppl. material 1: Tables S2, S4).

opencc-by-4.0Dec 2020View details →
zenodo28/100

Supplementary material 2 from: Fryssouli V, Zervakis GI, Polemis E, Typas MA (2020) A global meta-analysis of ITS rDNA sequences from material belonging to the genus Ganoderma (Basidiomycota, Polyporales) including new data from selected taxa. MycoKeys 75: 71-143. https://doi.org/10.3897/mycokeys.75.59872

Figure S1

opencc-zeroDec 2020View details →
zenodo28/100

Differential analysis of binarized single-cell RNA sequencing data captures biological variation

<p>Processed datasets used for binary differential analysis experiments.</p>

opencc-by-4.0Jan 2021View details →
zenodo28/100

Figure 4 from: Shan W, Tursun M, Zhou S, Zhang Y, Dai H (2021) Complete mitochondrial genome sequence of Lepus yarkandensis Günther, 1875 (Lagomorpha, Leporidae): characterization and phylogenetic analysis. ZooKeys 1012: 135-150. https://doi.org/10.3897/zookeys.1012.59035

Figure 4 A schematic of the structural organization of the mitochondrial control region in Lepus yarkandensis. Control region flanking genes tRNA-Phe and tRNA-Pro presented in red. Conserved elements in the control region denoted by gray boxes: TAS, termination associated sequence; CD, central conserved domain; CSB, conserved sequence block. SR, short repeat; LR, long repeat.

opencc-by-4.0Feb 2021View details →
zenodo28/100

Figure 5 from: Shan W, Tursun M, Zhou S, Zhang Y, Dai H (2021) Complete mitochondrial genome sequence of Lepus yarkandensis Günther, 1875 (Lagomorpha, Leporidae): characterization and phylogenetic analysis. ZooKeys 1012: 135-150. https://doi.org/10.3897/zookeys.1012.59035

Figure 5 Neighbor-joining and Bayes trees based on the complete mtDNA sequences of 25 lagomorphs. Values separated by slash (/) represent bootstrap support values for the NJ and Bayes trees.

opencc-by-4.0Feb 2021View details →
zenodo28/100

Supplementary material 1 from: Shan W, Tursun M, Zhou S, Zhang Y, Dai H (2021) Complete mitochondrial genome sequence of Lepus yarkandensis Günther, 1875 (Lagomorpha, Leporidae): characterization and phylogenetic analysis. ZooKeys 1012: 135-150. https://doi.org/10.3897/zookeys.1012.59035

Figure S1a, S1b

opencc-zeroFeb 2021View details →
zenodo28/100

Figure 1 from: Shan W, Tursun M, Zhou S, Zhang Y, Dai H (2021) Complete mitochondrial genome sequence of Lepus yarkandensis Günther, 1875 (Lagomorpha, Leporidae): characterization and phylogenetic analysis. ZooKeys 1012: 135-150. https://doi.org/10.3897/zookeys.1012.59035

Figure 1 Complete mitochondrial genome map of Lepus yarkandensis. Genes encoded on the heavy and light strands are shown outside and inside the circle, respectively.

opencc-by-4.0Feb 2021View details →
dryad28/100

Data from: A comprehensive analysis of teleost MHC class I sequences

Background: MHC class I (MHCI) molecules are the key presenters of peptides generated through the intracellular pathway to CD8-positive T-cells. In fish, MHCI genes were first identified in the early 1990′s, but we still know little about their functional relevance. The expansion and presumed sub-functionalization of cod MHCI and access to many published fish genome sequences provide us with the incentive to undertake a comprehensive study of deduced teleost fish MHCI molecules. Results: We expand the known MHCI lineages in teleosts to five with identification of a new lineage defined as P. The two lineages U and Z, which both include presumed peptide binding classical/typical molecules besides more derived molecules, are present in all teleosts analyzed. The U lineage displays two modes of evolution, most pronouncedly observed in classical-type alpha 1 domains; cod and stickleback have expanded on one of at least eight ancient alpha 1 domain lineages as opposed to many other teleosts that preserved a number of these ancient lineages. The Z lineage comes in a typical format present in all analyzed ray-finned fish species as well as lungfish. The typical Z format displays an unprecedented conservation of almost all 37 residues predicted to make up the peptide binding groove. However, also co-existing atypical Z sub-lineage molecules, which lost the presumed peptide binding motif, are found in some fish like carps and cavefish. The remaining three lineages, L, S and P, are not predicted to bind peptides and are lost in some species. Conclusions: Much like tetrapods, teleosts have polymorphic classical peptide binding MHCI molecules, a number of classical-similar non-classical MHCI molecules, and some members of more diverged MHCI lineages. Different from tetrapods, however, is that in some teleosts the classical MHCI polymorphism incorporates multiple ancient MHCI domain lineages. Also different from tetrapods is that teleosts have typical Z molecules, in which the residues that presumably form the peptide binding groove have been almost completely conserved for over 400 million years. The reasons for the uniquely teleost evolution modes of peptide binding MHCI molecules remain an enigma.

opencc-zeroDec 2014View details →
dryad28/100

Data from: ASSET: analysis of sequences of synchronous events in massively parallel spike trains

With the ability to observe the activity from large numbers of neurons simultaneously using modern recording technologies, the chance to identify sub-networks involved in coordinated processing increases. Sequences of synchronous spike events (SSEs) constitute one type of such coordinated spiking that propagates activity in a temporally precise manner. The synfire chain was proposed as one potential model for such network processing. Previous work introduced a method for visualization of SSEs in massively parallel spike trains, based on an intersection matrix that contains in each entry the degree of overlap of active neurons in two corresponding time bins. Repeated SSEs are reflected in the matrix as diagonal structures of high overlap values. The method as such, however, leaves the task of identifying these diagonal structures to visual inspection rather than to a quantitative analysis. Here we present ASSET (Analysis of Sequences of Synchronous EvenTs), an improved, fully automated method which determines diagonal structures in the intersection matrix by a robust mathematical procedure. The method consists of a sequence of steps that i) assess which entries in the matrix potentially belong to a diagonal structure, ii) cluster these entries into individual diagonal structures and iii) determine the neurons composing the associated SSEs. We employ parallel point processes generated by stochastic simulations as test data to demonstrate the performance of the method under a wide range of realistic scenarios, including different types of non-stationarity of the spiking activity and different correlation structures. Finally, the ability of the method to discover SSEs is demonstrated on complex data from large network simulations with embedded synfire chains. Thus, ASSET represents an effective and efficient tool to analyze massively parallel spike data for temporal sequences of synchronous activity.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Integrating sequence evolution into probabilistic orthology analysis

Orthology analysis, that is, finding out whether a pair of homologous genes are orthologs — stemming from a speciation — or paralogs — stemming from a gene duplication - is of central importance in computational biology, genome annotation, and phylogenetic inference. In particular, an orthologous relationship makes functional equivalence of the two genes highly likely. A major approach to orthology analysis is to reconcile a gene tree to the corresponding species tree, (most commonly performed using the most parsimonious reconciliation, MPR). However, most such phylogenetic orthology methods infer the gene tree without considering the constraints implied by the species tree and, perhaps even more importantly, only allow the gene sequences to influence the orthology analysis through the a priori reconstructed gene tree. We propose a sound, comprehensive Bayesian Markov chain Monte Carlo-based method, DLRSOrthology, to compute orthology probabilities. It efficiently sums over the possible gene trees and jointly takes into account the current gene tree, all possible reconciliations to the species tree, and the, typically strong, signal conveyed by the sequences. We compare our method with PrIME-GEM, a probabilistic orthology approach built on a probabilistic duplication-loss model, and MRBAYESMPR, a probabilistic orthology approach that is based on conventional Bayesian inference coupled with MPR. We find that DLRSOrthology outperforms these competing approaches on synthetic data as well as on biological data sets and is robust to incomplete taxon sampling artifacts.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Sequence analysis of European maize inbred line F2 provides new insights into molecular and chromosomal characteristics of presence/absence variants

Maize is well known for its exceptional structural diversity, including copy number variants (CNVs) and presence/absence variants (PAVs), and there is growing evidence for the role of structural variation in maize adaptation. While PAVs have been described in this important crop species, the extent of presence/absence variation and the relative position of inbred-specific regions remain to be elucidated. De novo genome sequencing of the F2 maize inbred line which played a key role in European breeding programs over the past 50 years revealed thousands of novel genomic regions, making up 88Mb of DNA, that are present in the F2 but not in B73. Comparison of B73 and F2 PAV localization revealed contrasted chromosomal distributions between the two inbreds and specific evolutionary dynamics of PAVs as compared to SNPs. Detailed sequence and functional annotation of F2 PAV sequences revealed hundreds of new genes with transcriptional support, but also a large fraction of repetitive sequences. Detailed analysis of sequence breakpoint highlights the role of double strand break repair, but also transposon insertion in PAV generation. Typing of the B73 and F2 PAVs in maize temperate inbreds revealed that some PAVs are found only in European Flint material, thus pinpointing structural features that may be at the origin of adaptive traits involved in the success of this material. Linkage disequilibrium (LD) analysis revealed that LD is strong within PAVs, as expected by the absence of recombination in crosses where PAV is missing in one of the parents.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Assembly and comparative analysis of transposable elements from low coverage genomic sequence data in Asparagales

The research field of comparative genomics is moving from a focus on genes to a more holistic view including the repetitive complement. This study aimed to characterize relative proportions of the repetitive fraction of large, complex genomes in a non-model system. The monocotyledonous plant order Asparagales (onion, asparagus, agave) comprises some of the largest angiosperm genomes and represents variation in both genome size and structure (karyotype). Anonymous, low coverage, single-end Illumina data from eleven exemplar Asparagales taxa were assembled using a de novo method. Resulting contigs were annotated using a reference library of available monocot repetitive sequences. Mapping reads to contigs provided rough estimates of relative proportions of each type of transposon in the nuclear genome. The results were parsed into general repeat types and synthesized with genome size estimates and a phylogenetic context to describe the pattern of transposable element evolution among these lineages. The major finding is that while some lineages in Asparagales exhibit conservation in repeat proportions, there is generally wide variation in types and frequency of repeats. This approach is an appropriate first step in characterizing repeats in evolutionary lineages with a paucity of genomic resources.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Analysis of expressed sequence tags from the placenta of the live-bearing fish Poeciliopsis (Poeciliidae)

Matrotrophic fish in the genus Poeciliopsis (Poeciliidae) have a placenta-like structure used in post-fertilization maternal provisioning of the developing embryo. To understand better the structure and function of the Poeciliopsis placenta we derived cDNA libraries from the maternal follicular placenta of two matrotrophic Poeciliopsis sister species, P. turneri and P. presidionis. These species inherited their placenta from a common ancestor and represent one of three independent origins of placentas in Poeciliopsis. Expressed sequence tags were generated and putative function was determined using BLASTX homology searches and Gene Ontology annotation. Reverse transcriptase-PCR was used to verify placenta tissue expression of a putative candidate gene, alpha-2 macroglobulin. 1956 (71.5% of the total submitted ESTs) and 924 (71.0% of the total submitted ESTs) unique transcripts were identified for the P. turneri and P. presidionis placenta, respectively. Homology search and Gene Ontology annotation revealed putative genes whose products may be involved in specific transport functions of the maternal follicle. These putative genes are excellent candidates for future research on the evolution of the placenta. We discuss our results in light of the parent-offspring conflict theory of placental evolution and in terms of the Poeciliid placenta structure and function.

opencc-zeroDec 2010View details →
dryad28/100

Data from: Genome sequencing and comparative analysis of three Chlamydia pecorum strains associated with different pathogenic outcomes

Background: Chlamydia pecorum is the causative agent of a number of acute diseases, but most often causes persistent, subclinical infection in ruminants, swine and birds. In this study, the genome sequences of three C. pecorum strains isolated from the faeces of a sheep with inapparent enteric infection (strain W73), from the synovial fluid of a sheep with polyarthritis (strain P787) and from a cervical swab taken from a cow with metritis (strain PV3056/3) were determined using Illumina/Solexa and Roche 454 genome sequencing. Results: Gene order and synteny was almost identical between C. pecorum strains and C. psittaci. Differences between C. pecorum and other chlamydiae occurred at a number of loci, including the plasticity zone, which contained a MAC/perforin domain protein, two copies of a &gt;3400 amino acid putative cytotoxin gene and four (PV3056/3) or five (P787 and W73) genes encoding phospholipase D. Chlamydia pecorum contains an almost intact tryptophan biosynthesis operon encoding trpABCDFR and has the ability to sequester kynurenine from its host, however it lacks the genes folA, folKP and folB required for folate metabolism found in other chlamydiae. A total of 15 polymorphic membrane proteins were identified, belonging to six pmp families. Strains possess an intact type III secretion system composed of 18 structural genes and accessory proteins, however a number of putative inc effector proteins widely distributed in chlamydiae are absent from C. pecorum. Two genes encoding the hypothetical protein ORF663 and IncA contain variable numbers of repeat sequences that could be associated with persistence of infection. Conclusions: Genome sequencing of three C. pecorum strains, originating from animals with different disease manifestations, has identified differences in ORF663 and pseudogene content between strains and has identified genes and metabolic traits that may influence intracellular survival, pathogenicity and evasion of the host immune system.

opencc-zeroDec 2013View details →
zenodo28/100

FIGURE 2 in Systematic position of Dinidoridae within the superfamily Pentatomoidea (Hemiptera: Heteroptera) revealed by the Bayesian phylogenetic analysis of the mitochondrial 12S and 16S rDNA sequences

FIGURE 2. Phylogenetic tree obtained from the Bayesian inference analysis of the 16S rDNA dataset.

opennotspecifiedAug 2012View details →
zenodo28/100

FIGURE 1 in Systematic position of Dinidoridae within the superfamily Pentatomoidea (Hemiptera: Heteroptera) revealed by the Bayesian phylogenetic analysis of the mitochondrial 12S and 16S rDNA sequences

FIGURE 1. Phylogenetic tree obtained from the Bayesian inference analysis of the 12S rDNA dataset.

opennotspecifiedAug 2012View details →
zenodo28/100

FIG. 3 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations

FIG. 3.—Populationstructure inference based on STRUCTURE analysis of 199,921 sites for individual bats from four Hawaiian Islands. (A) Ad hoc statistic delta Kanalysis indicates a peak at the Κ = 5; (B) STRUCTURE population inference with Κ = 3, 4, 5. Sample information included in supplementary table S4, Supplementary Material online.

opencc-by-4.0Aug 2020View details →
zenodo28/100

Healthy woodchuck genome with viral sequences appended used for single-cell RNA-seq analysis

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo28/100

Curated RNA Sequencing Data for Zebrafish (Danio rerio) snoRNA Expression Analysis

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
dryad28/100

UCE and Sanger sequenced data for phylogenetic analysis of jumping spiders (Baviini and Nungia, Salticidae)

<p>The systematics and taxonomy of the tropical Asian jumping spiders of the tribe Baviini is reviewed, with a molecular phylogenetic study (UCE sequence capture, traditional Sanger sequencing) guiding a reclassification of the group's genera. The well-studied members of the group are placed into six genera: <i>Bavia</i> Simon, 1877, <i>Indopadilla</i> Caleb &amp; Sankaran, 2019, <i>Padillothorax</i> Simon, 1901, <i>Piranthus</i> Thorell, 1895, <i>Stagetillus</i> Simon, 1885, and one new genus, <i>Maripanthus</i> Maddison. The identity of <i>Padillothorax</i> is clarified, and <i>Bavirecta</i> Kanesharatnam &amp; Benjamin, 2018 synonymized with it. <i>Hyctiota</i> Strand, 1911 is synonymized with <i>Stagetillus</i>. The molecular phylogeny divides the baviines into three clades, the <i>Piranthus</i> clade with a long embolus (<i>Piranthus</i>, <i>Maripanthus</i>), the genus <i>Padillothorax</i> with a flat body and short embolus, and the <i>Bavia</i> clade with a higher body and (usually) short embolus (remaining genera). In general, morphological synapomorphies support or extend the molecularly-delimited groups. Eighteen new species are described (all with taxonomic authority W. Maddison): <i>Bavia nessagyna</i>, <i>Indopadilla bamilin</i>, <i>I. kodagura</i>, <i>I. nesinor</i>, <i>I. redunca</i>, <i>I. redynis</i>, <i>I. sabivia</i>, <i>I. vimedaba</i>, <i>Maripanthus draconis</i> (type species of <i>Maripanthus</i>), <i>M. jubatus</i>, <i>M. reinholdae</i>, <i>Padillothorax badut</i>, <i>P. mulu</i>, <i>Piranthus api</i>, <i>P. bakau</i>, <i>P. kohi</i>, <i>P. mandai</i>, and <i>Stagetillus irri</i>. The distinctions between baviines and the astioid <i>Nungia</i> Żabka, 1985 are reviewed, leading to four species being moved into <i>Nungia</i> from <i>Bavia</i> and other genera<i>. </i>Fifteen new combinations are established, and one combination is restored. Five of these new or restored combinations correct previous errors of placing species in genera that have superficially similar palps but extremely different body forms, in fact belonging in distantly related tribes — emphasizing that the general shape of male palps should be used with caution in determining relationships. A little-studied genus, <i>Padillothorus</i> Prószyński, 2018, is tentatively assigned to the Baviini. <i>Ligdus</i> Thorell, 1895 is assigned to the Ballini.</p>

opencc-zeroOct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record