Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,574

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,574 results for “genome sequencing”

Learn how ShareScore rates datasets ↗
zenodo28/100

Figure 5 from: Sun G, Zhao C, Xia T, Wei Q, Yang X, Feng S, Sha W, Zhang H (2020) Sequence and organisation of the mitochondrial genome of Japanese Grosbeak (Eophona personata), and the phylogenetic relationships of Fringillidae. ZooKeys 995: 67-80. https://doi.org/10.3897/zookeys.995.34432

Figure 5 The phylogenetic tree generated for 17 species of Fringillidae. The values indicated at the nodes are Bayesian posterior probabilities (left) and ML bootstrap proportions (right).

opencc-by-4.0Nov 2020View details →
zenodo28/100

Figure 1 from: Sun G, Zhao C, Xia T, Wei Q, Yang X, Feng S, Sha W, Zhang H (2020) Sequence and organisation of the mitochondrial genome of Japanese Grosbeak (Eophona personata), and the phylogenetic relationships of Fringillidae. ZooKeys 995: 67-80. https://doi.org/10.3897/zookeys.995.34432

Figure 1 Circular map of the mitochondrial genome of Eophona personata. tRNAs are denoted as one-letter symbols according to IUPAC-IUB single-letter amino acid codes; L1 = UUR, L2 = CUN, S1 = UCN, S2 = AGY.

opencc-by-4.0Nov 2020View details →
zenodo28/100

Figure 4 from: Shan W, Tursun M, Zhou S, Zhang Y, Dai H (2021) Complete mitochondrial genome sequence of Lepus yarkandensis Günther, 1875 (Lagomorpha, Leporidae): characterization and phylogenetic analysis. ZooKeys 1012: 135-150. https://doi.org/10.3897/zookeys.1012.59035

Figure 4 A schematic of the structural organization of the mitochondrial control region in Lepus yarkandensis. Control region flanking genes tRNA-Phe and tRNA-Pro presented in red. Conserved elements in the control region denoted by gray boxes: TAS, termination associated sequence; CD, central conserved domain; CSB, conserved sequence block. SR, short repeat; LR, long repeat.

opencc-by-4.0Feb 2021View details →
zenodo28/100

Figure 5 from: Shan W, Tursun M, Zhou S, Zhang Y, Dai H (2021) Complete mitochondrial genome sequence of Lepus yarkandensis Günther, 1875 (Lagomorpha, Leporidae): characterization and phylogenetic analysis. ZooKeys 1012: 135-150. https://doi.org/10.3897/zookeys.1012.59035

Figure 5 Neighbor-joining and Bayes trees based on the complete mtDNA sequences of 25 lagomorphs. Values separated by slash (/) represent bootstrap support values for the NJ and Bayes trees.

opencc-by-4.0Feb 2021View details →
zenodo28/100

Supplementary material 1 from: Shan W, Tursun M, Zhou S, Zhang Y, Dai H (2021) Complete mitochondrial genome sequence of Lepus yarkandensis Günther, 1875 (Lagomorpha, Leporidae): characterization and phylogenetic analysis. ZooKeys 1012: 135-150. https://doi.org/10.3897/zookeys.1012.59035

Figure S1a, S1b

opencc-zeroFeb 2021View details →
zenodo28/100

Figure 1 from: Shan W, Tursun M, Zhou S, Zhang Y, Dai H (2021) Complete mitochondrial genome sequence of Lepus yarkandensis Günther, 1875 (Lagomorpha, Leporidae): characterization and phylogenetic analysis. ZooKeys 1012: 135-150. https://doi.org/10.3897/zookeys.1012.59035

Figure 1 Complete mitochondrial genome map of Lepus yarkandensis. Genes encoded on the heavy and light strands are shown outside and inside the circle, respectively.

opencc-by-4.0Feb 2021View details →
dryad28/100

Data from: Low-coverage, whole-genome sequencing of Artocarpus camansi (Moraceae) for phylogenetic marker development and gene discovery

Premise of the study: We used moderately low-coverage (17×) whole-genome sequencing of Artocarpus camansi (Moraceae) to develop genomic resources for Artocarpus and Moraceae. Methods and Results: A de novo assembly of Illumina short reads (251,378,536 pairs, 2 × 100 bp) accounted for 93% of the predicted genome size. Predicted coding regions were used in a three-way orthology search with published genomes of Morus notabilis and Cannabis sativa. Phylogenetic markers for Moraceae were developed from 333 inferred single-copy exons. Ninety-eight putative MADS-box genes were identified. Analysis of all predicted coding regions resulted in preliminary annotation of 49,089 genes. An analysis of synonymous substitutions for pairs of orthologs (Ks analysis) in M. notabilis and A. camansi strongly suggested a lineage-specific whole-genome duplication in Artocarpus. Conclusions: This study substantially increases the genomic resources available for Artocarpus and Moraceae and demonstrates the value of low-coverage de novo assemblies for nonmodel organisms with moderately large genomes.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Plastid genome sequences of legumes reveal parallel inversions and multiple losses of rps16 in papilionoids

To date, publicly available plastid genomes of legumes have for the most part been limited to the subfamily Papilionoideae. Here we report 13 new plastid genomes of legumes spanning all three subfamilies. The genomes representing Caesalpinioideae and Mimosoideae are highly conserved in gene content and gene order, similar to the ancestral angiosperm genome organization. Genomes within the Papilionoideae, however, have reduced sizes due to deletions in nine intergenic spacers primarily in the large single copy region. Our study also indicates that rps16 has been independently lost at least five times in legumes, with additional gene and intron losses scattered among the papilionoids. Additionally, genera from two distinct lineages within the papilionoids, Lupinus and Robinia, have a parallel inversion of 36 kb and 39 kb, respectively. This parallel inversion is novel as it appears to be caused by a 29 bp repeat within two trnS genes. This repeat is present in all available legume plastid genomes indicating that there is the potential for this inversion to be present in more species. This case of a homoplasious inversion is also evidence that some inversion events may not be reliable phylogenetic markers.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Reconstructing phylogeny from reduced-representation genome sequencing data without assembly or alignment

Reduced-representation genome sequencing such as RADseq aids the analysis of genomes by reducing the quantity of data, thereby lowering both sequencing costs and computational burdens. RADseq was initially designed for studying genetic variation across genomes at the population level, but has also proved to be suitable for interspecific phylogeny reconstruction. RADseq data pose challenges for standard phylogenomic methods, however, due to incomplete coverage of the genome and large amounts of missing data. Alignment-free methods are both efficient and accurate for phylogenetic reconstructions with whole genomes and are especially practical for non-model organisms; nonetheless, alignment-free methods have not been applied with reduced genome sequencing data. Here, we test a full-genome assembly and alignment-free method, AAF, in application to RADseq data and propose two procedures for reads selection to remove reads from restriction sites that were not found in taxa being compared. We validate these methods using both simulations and real datasets. Reads selection improved the accuracy of phylogenetic construction in every simulated scenario and the two real datasets, making AAF as good or better than a comparable alignment-based method, even though AAF had much lower computational burdens. We also investigated the sources of missing data in RADseq and their effects on phylogeny reconstruction using AAF. The AAF pipeline modified for RADseq or other reduced-representation sequencing data, phyloRAD, is available on github (https://github.com/fanhuan/phyloRAD).

opencc-zeroDec 2017View details →
dryad28/100

Data from: Raw whole Drosophila genome sequence traces have contaminant sequences from bacterial symbionts

Many Drosophila genomes have been sequenced and assembled recently, and many more genome sequencing projects are in progress. However, Drosophila have bacterial, fungal, and protozoan symbionts, and the DNA of these symbionts may be isolated in the process of sequencing Drosophila genomes. Here, we assess how much sequence is isolated from these symbionts and if the sequence contamination affected how these Drosophila genomes were assembled. We do find raw sequence from bacterial symbionts and humans in Drosophila genome sequence traces analyzed. Surprisingly, the four most-common contaminant species were shared among the Drosophila genomes. However, we do not find evidence of bacterial sequences in two published Drosophila genome assemblies.

opencc-zeroDec 2009View details →
dryad28/100

Data from: Genome sequence of Striga asiatica provides insight into the evolution of plant parasitism

Parasitic plants in the genus Striga, commonly known as witchweeds, cause major crop losses in sub-Saharan Africa and pose a threat to agriculture worldwide. An understanding of Striga parasite biology, which could lead to agricultural solutions, has been hampered by the lack of genome information. Here we report the draft genome sequence of Striga asiatica with 34,577 predicted protein-coding genes, which reflects gene family contractions and expansions that are consistent with a three-phase model of parasitic plant genome evolution. Striga seeds germinate in response to host-derived strigolactones (SLs) and then develop a specialised penetration structure, the haustorium, to invade the host root. A family of SL receptors has undergone a striking expansion, suggesting a molecular basis for the evolution of broad host range among Striga spp. We found that genes involved in lateral root development in non-parasitic model species are coordinately induced during haustorium development in Striga, suggesting a pathway that was partly co-opted during the evolution of the haustorium. In addition, we found evidence for horizontal transfer of host genes as well as retrotransposons, indicating gene flow to S. asiatica from hosts. Our results provide valuable insights into the evolution of parasitism and a key resource for the future development of Striga control strategies.

opencc-zeroSep 2019View details →
dryad28/100

Data from: Testing models of speciation from genome sequences: divergence and asymmetric admixture in Island Southeast Asian Sus species during the Plio-Pleistocene climatic fluctuations

In many temperate regions, ice ages promoted range contractions into refugia resulting in divergence (and potentially speciation), while warmer periods led to range expansions and hybridization. However, the impact these climatic oscillations had in many parts of the tropics remains elusive. Here, we investigate this issue using genome sequences of three pig (Sus) species, two of which are found on islands of the Sunda-shelf shallow seas in Island Southeast Asia (ISEA). A previous study revealed signatures of inter-specific admixture between these Sus species (Frantz et al. (2013) Genome sequencing reveals fine scale diversification and reticulation history during speciation in Sus. Genome biology, 14, R107). However, the timing, directionality and extent of this admixture remain unknown. Here we use a likelihood based model comparison to more finely resolve this admixture history and test whether it was mediated by humans or occurred naturally. Our analyses suggest that inter-specific admixture between Sunda-shelf species was most likely asymmetric and occurred long before the arrival of humans in the region. More precisely, we show that these species diverged during the late Pliocene but around 23% of their genomes have been affected by admixture during the later Pleistocene climatic transition. In addition, we show that our method provides a significant improvement over D-statistics which are uninformative about the direction of admixture.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Chlamydomonas genome resource for laboratory strains reveals a mosaic of sequence variation, identifies true strain histories, and enables strain-specific studies

Chlamydomonas reinhardtii is a widely used reference organism in studies of photosynthesis, cilia, and biofuels. Most research in this field uses a few dozen standard laboratory strains that are reported to share a common ancestry, but exhibit substantial phenotypic differences. In order to facilitate ongoing Chlamydomonas research and explain the phenotypic variation, we mapped the genetic diversity within these strains using whole-genome resequencing. We identified 524,640 single nucleotide variants and 4812 structural variants among 39 commonly used laboratory strains. Nearly all (98.2%) of the total observed genetic diversity was attributable to the presence of two, previously unrecognized, alternate haplotypes that are distributed in a mosaic pattern among the extant laboratory strains. We propose that these two haplotypes are the remnants of an ancestral cross between two strains with ∼2% relative divergence. These haplotype patterns create a fingerprint for each strain that facilitates the positive identification of that strain and reveals its relatedness to other strains. The presence of these alternate haplotype regions affects phenotype scoring and gene expression measurements. Here, we present a rich set of genetic differences as a community resource to allow researchers to more accurately conduct and interpret their experiments with Chlamydomonas.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Congruent deep relationships in the grape family (Vitaceae) based on sequences of chloroplast genomes and mitochondrial genes via genome skimming

Vitaceae is well-known for having one of the most economically important fruits, i.e., the grape (Vitis vinifera). The deep phylogeny of the grape family was not resolved until a recent phylogenomic analysis of 417 nuclear genes from transcriptome data. However, it has been reported extensively that topologies based on nuclear and organellar genes may be incongruent due to differences in their evolutionary histories. Therefore, it is important to reconstruct a backbone phylogeny of the grape family using plastomes and mitochondrial genes. In this study, next-generation sequencing data sets of 27 species were obtained using genome skimming with total DNAs from silica-gel preserved tissue samples on an Illumina HiSeq 2500 instrument. Plastomes were assembled using the combination of de novo and reference genome (of V. vinifera) methods. Sixteen mitochondrial genes were also obtained via genome skimming using the reference genome of V. vinifera. Extensive phylogenetic analyses were performed using maximum likelihood and Bayesian methods. The topology based on either plastome data or mitochondrial genes is congruent with the one using hundreds of nuclear genes, indicating that the grape family did not exhibit significant reticulation at the deep level. The results showcase the power of genome skimming in capturing extensive phylogenetic data: especially from chloroplast and mitochondrial DNAs.

opencc-zeroDec 2015View details →
dryad28/100

Data from: De novo transcriptome characterization and development of genomic tools for Scabiosa columbaria L. using next-generation sequencing techniques.

Next-generation sequencing (NGS) technologies are increasingly applied in many organisms, including non-model organisms that are important for ecological and conservation purposes. Illumina and 454 sequencing are among the most used NGS technologies and have been shown to produce optimal results at reasonable costs when used together. Here, we describe the combined application of these two NGS technologies to characterize the transcriptome of a plant species of ecological and conservation relevance for which no genomic resource is available, Scabiosa columbaria. We obtained 528,557 reads from a 454 GS-FLX run and a total of 28,993,627 reads from two lanes of an Illumina GAII single run. After reads trimming, the de novo assembly of both types of reads produced 109,630 contigs. Both the contigs and the >75 bp remaining singletons were blasted against Uniprot/Swissprot database, resulting in 29,676 and 10,515 significant hits, respectively. Based on sequence similarity with known gene products, these sequences represent at least 12,516 unique genes, most of which are well covered by contig sequences. In addition, we identified 4,320 microsatellite loci, of which 856 had flanking sequences suitable for PCR primer design. We also identified 75,054 putative SNPs. This annotated sequence collection and the relative molecular markers represent a main genomic resource for S. columbaria which should contribute to future research in conservation and population biology studies. Our results demonstrate the utility of NGS technologies as starting point for the development of genomic tools in nonmodel but ecologically important species.

opencc-zeroDec 2009View details →
dryad28/100

Data from: Genome-level homology and phylogeny of Vibrionaceae (Gammaproteobacteria: Vibrionales) with three new complete genome sequences

Background: Phylogenetic hypotheses based on complete genome data are presented for the Gammaproteobacteria family Vibrionaceae. Two taxon samplings are presented: one including all those taxa for which the genome sequences are complete in terms of arrangement (chromosomal location of fragments; 19 taxa) and one for which the genome sequences contain multiple contigs (44 taxa). Analyses are presented under the Maximum Parsimony and Maximum Likelihood optimality criteria for total evidence datasets, the two chromosomes separately, and individual analyses of locally collinear blocks. Three of the genomes included in the 44 taxon dataset, those of Vibrio gazogenes, Salinivibrio costicola, and Aliivibrio logei have been newly sequenced and their genome sequences are documented here. Results: Phylogenetic results for the 19-taxon datasets show similar levels of collinear subset of dataset incongruence as a previous study of 22 taxa from the sister family Shewanellaceae, while also echoing the strong phylogenetic performance of random subsets of data also shown in this study. Phylogenetic results for both the 19-taxon and 44-taxon datasets corroborate previous hypotheses about the placement of Photobacterium and Aliivibrio within Vibrionaceae and also highlight problems with how Photobacterium is delimited and indicate that it likely should be dissolved into Vibrio to produce a phylogenetic taxonomy. The 19-taxon and 44-taxon trees based on the large chromosome are congruent for the majority of taxa that are present in both datasets. Analyses of the 44-taxon sampling based on the second, small chromosome are quite different from those based on the large chromosome, which is not surprising given the dramatically divergent nature of the small chromosome and the difficulty in postulating primary homologies. Conclusions: The phylogenetic analyses presented here represent the most comprehensive genome-level phylogenetic analyses in terms of taxa and data. Based on the availability of genome data for many bacterial species on GenBank, many other bacterial groups would also be amenable to similar genome-scale phylogenetic analyses even when present in multiple contigs. The result that collinear subsets of data are incongruent with the concatenated dataset and with each other while random data subsets show very little incongruence echoes the result of previous work on Shewanellaceae. The 44-taxon phylogenetic analysis presented here thus represents the future of phylogenomic analyses in scope and complexity.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Targeted sequencing of venom genes from cone snail genomes improves understanding of conotoxin molecular evolution

To expand our capacity to discover venom sequences from the genomes of venomous organisms, we applied targeted sequencing techniques to selectively recover venom gene superfamilies and non-toxin loci from the genomes of 32 cone snail species (family, Conidae), a diverse group of marine gastropods that capture their prey using a cocktail of neurotoxic peptides (conotoxins). We were able to successfully recover conotoxin gene superfamilies across all species with high confidence (> 100X coverage) and used these data to provide new insights into conotoxin evolution. First, we found that conotoxin gene superfamilies are composed of 1-6 exons and are typically short in length (mean = ~85bp). Second, we expanded our understanding of the following genetic features of conotoxin evolution: (a) positive selection, where exons coding the mature toxin region were often three times more divergent than their adjacent noncoding regions, (b) expression regulation, with comparisons to transcriptome data showing that cone snails only express a fraction of the genes available in their genome (24%-63%), and (c) extensive gene turnover, where Conidae species varied from 120-859 conotoxin gene copies. Finally, using comparative phylogenetic methods, we found that while diet specificity did not predict patterns of conotoxin evolution, dietary breadth was positively correlated with total conotoxin gene diversity. Overall, the targeted sequencing technique demonstrated here has the potential to radically increase the pace at which venom gene families are sequenced and studied, reshaping our ability to understand the impact of genetic changes on ecologically relevant phenotypes and subsequent diversification.

opencc-zeroDec 2017View details →
dryad28/100

Data from: "You are not what you eat: massive parallel sequencing reveals that gut microbiome is not diet-related in larval Dilophus febrilis (Diptera: Bibionidae)" in Genomic Resources Notes Accepted 1 June 2015 to 31 July 2015

This article documents the public availability of metagenome sequence data from 454 amplicon sequencing of larval dipteran gut (Dilophus febrilis) and their potential food sources dwarf shrub litter (Vaccinium gaultheroides), grass litter (Dactylis glomerata), and cow dung (Bos primigenius taurus).

opencc-zeroDec 2014View details →
dryad28/100

Data from: Chromosomal inversions and ecotypic differentiation in Anopheles gambiae: the perspective from whole-genome sequencing

The molecular mechanisms and genetic architecture that facilitate adaptive radiation of lineages remain elusive. Polymorphic chromosomal inversions, due to their recombination-reducing effect, are proposed instruments of ecotypic differentiation. Here we study an ecologically diversifying lineage of An. gambiae, known as the Bamako chromosomal form based on its unique complement of three chromosomal inversions, to explore the impact of these inversions on ecotypic differentiation. We used pooled and individual genome sequencing of Bamako, typical (non-Bamako) An. gambiae, and the sister species An. coluzzii to investigate evolutionary relationships and genome-wide patterns of nucleotide diversity and differentiation among lineages. Despite extensive shared polymorphism and limited differentiation from the other taxa, Bamako clusters apart from the other taxa, and forms a maximally supported clade in neighbor-joining trees based on whole genome data (including inversions) or solely on collinear regions. Nevertheless, FST outlier analysis reveals that the majority of differentiated regions between Bamako and typical An. gambiae are located inside chromosomal inversions, consistent with their role in the ecological isolation of Bamako. Exceptionally differentiated genomic regions were enriched for genes implicated in nervous system development and signaling. Candidate genes associated with a selective sweep unique to Bamako contain substitutions not observed in sympatric samples of the other taxa, and several insecticide resistance gene alleles shared between Bamako and other taxa segregate at sharply different frequencies in these samples. Bamako represents a useful window into the initial stages of ecological and genomic differentiation from sympatric populations in this important group of malaria vectors.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Whole-genome sequencing approaches for conservation biology: advantages, limitations, and practical recommendations

Whole-genome resequencing (WGR) is a powerful method for addressing fundamental evolutionary biology questions that have not been fully resolved using traditional methods. WGR includes four approaches: the sequencing of individuals to a high depth of coverage with either unresolved (huWGR) or resolved haplotypes (hrWGR), the sequencing of population genomes to a high depth by mixing equimolar amounts of unlabelled-individual DNA (Pool-seq), and the sequencing of multiple individuals from a population to a low depth (lcWGR). These techniques require the availability of a reference genome. This, along with the still high cost of shotgun sequencing and the large demand for computing resources and storage, has limited their implementation in non-model species with scarce genomic resources and in fields such as conservation biology. Our goal here is to describe the various WGR methods, their pros and cons, and potential applications in conservation biology. WGR offers an unprecedented marker density and surveys a wide diversity of genetic variations not limited to single nucleotide polymorphisms (e.g. structural variants and mutations in regulatory elements), increasing their power for the detection of signatures of selection and local adaptation as well as for the identification of the genetic basis of phenotypic traits and diseases. Currently though, no single WGR approach fulfills all requirements of conservation genetics, and each method has its own limitations and sources of potential bias. We discuss proposed ways to minimize such biases. We envision a not distant future where the analysis of whole genomes becomes a routine task in many non-model species and fields including conservation biology.

opencc-zeroDec 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record