Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,293

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,293 results for “gene sequencing”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Population structure, relatedness and ploidy levels in an apple gene bank revealed through genotyping-by-sequencing

In recent years, new genome-wide marker systems have provided highly informative alternatives to low density marker systems for evaluating plant populations. To date, most apple germplasm collections have been genotyped using low-density markers such as simple sequence repeats (SSRs), whereas only a few have been explored using high-density genome-wide marker information. We explored the genetic diversity of the Pometum gene bank collection (University of Copenhagen, Denmark) of 349 apple accessions using over 15,000 genome-wide single nucleotide polymorphisms (SNPs) and 15 SSR markers, in order to compare the strength of the two approaches for describing population structure. We found that 119 accessions shared a clonal relationship with at least one other accession in the collection, resulting in the identification of 272 (78%) unique accessions. Of these unique accessions, over half (52%) share a first-degree relationship with at least one other accession. There is therefore a high degree of clonal and family relatedness in the Danish apple gene bank. We find significant genetic differentiation between Malus domestica and its supposed primary wild ancestor, M. sieversii, as well as between accessions of Danish origin and all others. Overall, we found strong concordance between analyses based on the genome-wide SNPs and the 15 SSR loci. However, we argue that GBS is superior to traditional SSR approaches because it allowed the estimation of ploidy levels that were in accordance with flow cytometry results, and can be further exploited in genome-wide association studies (GWAS). Finally, we compare GBS with SSR for the purposes of characterizing a diverse apple gene bank and discuss the advantages and constraints of the two approaches.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Recombination-dependent replication and gene conversion homogenize repeat sequences and diversify plastid genome structure

PREMISE OF THE STUDY: There is a misinterpretation in the literature regarding the variable orientation of the small single copy region of plastid genomes (plastomes). The common phenomenon of small and large single copy inversion, hypothesized to occur through intramolecular recombination between inverted repeats (IR) in a circular, single unit-genome, in fact more likely occurs through recombination-dependent replication (RDR) of linear plastome templates. If RDR can be primed through both intra- and intermolecular recombination, then this mechanism could not only create inversion isomers of so-called single copy regions, but also an array of alternative sequence arrangements. METHODS: We used Illumina paired-end and PacBio single-molecule real-time (SMRT) sequences to characterize repeat structure in the plastome of Monsonia emarginata L'Hér. (Geraniaceae). We used OrgConv and inspected nucleotide alignments to infer ancestral nucleotides and identify gene conversion among repeats and mapped long (>1 kb) SMRT reads against the unit-genome assembly to identify alternative sequence arrangements. RESULTS: Although M. emarginata lacks the canonical IR, we found that large repeats (>1 kilobase; kb) represent ~22% of the plastome nucleotide content. Among the largest repeats (>2 kb) we identified GC-biased gene conversion and mapping filtered, long SMRT reads to the M. emarginata unit-genome assembly revealed alternative, substoichiometric sequence arrangements. CONCLUSION: We offer a model based on RDR and gene conversion between long repeated sequences in the M. emarginata plastome, and provide support that both intra-and intermolecular recombination between large repeats, particularly in repeat-rich plastomes, varies unit-genome structure while homogenizing the nucleotide sequence of repeats.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Targeted gene enrichment and high-throughput sequencing for environmental biomonitoring: a case study using freshwater macroinvertebrates

Recent studies have advocated biomonitoring using DNA techniques. In this study, two high-throughput sequencing (HTS)-based methods were evaluated: amplicon metabarcoding of the cytochrome C oxidase subunit I (COI) mitochondrial gene and gene enrichment using MYbaits (targeting nine different genes including COI). The gene-enrichment method does not require PCR amplification and thus avoids biases associated with universal primers. Macroinvertebrate samples were collected from 12 New Zealand rivers. Macroinvertebrates were morphologically identified and enumerated, and their biomass determined. DNA was extracted from all macroinvertebrate samples and HTS undertaken using the illumina miseq platform. Macroinvertebrate communities were characterized from sequence data using either six genes (three of the original nine were not used) or just the COI gene in isolation. The gene-enrichment method (all genes) detected the highest number of taxa and obtained the strongest Spearman rank correlations between the number of sequence reads, abundance and biomass in 67% of the samples. Median detection rates across rare (<1% of the total abundance or biomass), moderately abundant (1–5%) and highly abundant (>5%) taxa were highest using the gene-enrichment method (all genes). Our data indicated primer biases occurred during amplicon metabarcoding with greater than 80% of sequence reads originating from one taxon in several samples. The accuracy and sensitivity of both HTS methods would be improved with more comprehensive reference sequence databases. The data from this study illustrate the challenges of using PCR amplification-based methods for biomonitoring and highlight the potential benefits of using approaches, such as gene enrichment, which circumvent the need for an initial PCR step.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Comparative population genetic analysis of bocaccio rockfish Sebastes paucispinis using anonymous and gene-associated simple sequence repeat loci

Comparative population genetic analyses of traditional and emergent molecular markers aid in determining appropriate use of new technologies. The bocaccio rockfish Sebastes paucispinis is a high-gene-flow marine species off the west coast of North America that experienced strong population decline over the past three decades. We used 18 anonymous and 13 gene associated simple sequence repeat loci (EST-SSRs) to characterize range-wide population structure with temporal replicates. No FST-outliers were detected using the LOSITAN program, suggesting that neither balancing nor divergent selection affected the loci surveyed. Consistent hierarchical structuring of populations by geography or year class was not detected regardless of marker class. The EST-SSRs were less variable than the anonymous SSRs, but no correlation between FST and variation or marker class was observed. General Linear Model analysis showed that low EST-SSR variation was attributable to low mean repeat number. Comparative genomic analysis with Gasterosteus aculeatus, Takifugu rubripes, and Oryzias latipes showed consistently lower repeat number in EST-SSRs than SSR loci that were not in ESTs. Purifying selection likely imposed functional constraints on EST-SSRs resulting in low repeat numbers that affected diversity estimates, but did not affect the observed pattern of population structure.

opencc-zeroDec 2011View details →
dryad32/100

Data from: A universal probe set for targeted sequencing of 353 nuclear genes from any flowering plant designed using k-medoids clustering

Sequencing of target-enriched libraries is an efficient and cost-effective method for obtaining DNA sequence data from hundreds of nuclear loci for phylogeny reconstruction. Much of the cost of developing targeted sequencing approaches is associated with the generation of preliminary data needed for the identification of orthologous loci for probe design. In plants, identifying orthologous loci has proven difficult due to a large number of whole-genome duplication events, especially in the angiosperms (flowering plants). We used multiple sequence alignments from over 600 angiosperms for 353 putatively single-copy protein-coding genes identified by the One Thousand Plant Transcriptomes Initiative to design a set of targeted sequencing probes for phylogenetic studies of any angiosperm group. To maximize the phylogenetic potential of the probes while minimizing the cost of production, we introduce a k-medoids clustering approach to identify the minimum number of sequences necessary to represent each coding sequence in the final probe set. Using this method, five to 15 representative sequences were selected per orthologous locus, representing the sequence diversity of angiosperms more efficiently than if probes were designed using available sequenced genomes alone. To test our approximately 80,000 probes, we hybridized libraries from 42 species spanning all higher-order groups of angiosperms, with a focus on taxa not present in the sequence alignments used to design the probes. Out of a possible 353 coding sequences, we recovered an average of 283 per species and at least 100 in all species. Differences among taxa in sequence recovery could not be explained by relatedness to the representative taxa selected for probe design, suggesting that there is no phylogenetic bias in the probe set. Our probe set, which targeted 260 kbp of coding sequence, achieved a median recovery of 137 kbp per taxon in coding regions, a maximum recovery of 250 kbp, and an additional median of 212 kbp per taxon in flanking non-coding regions across all species. These results suggest that the Angiosperms353 probe set described here is effective for any group of flowering plants and would be useful for phylogenetic studies from the species level to higher-order groups, including the entire angiosperm clade itself.

opencc-zeroDec 2017View details →
zenodo32/100

FIGURE 1. Bayesian tree inferred from SSU gene DNA sequences. Posterior probabilities exceeding 50 in A review of the genus Tripylina Brzeski, 1963 (Nematoda: Triplonchida), with descriptions of five new species from New Zealand

FIGURE 1. Bayesian tree inferred from SSU gene DNA sequences. Posterior probabilities exceeding 50% are given on appropriate clades. Nematode species, GenBank numbers, locations are listed for each taxon if known.

opennotspecifiedDec 2009View details →
zenodo32/100

FIGURE 2. Bayesian tree inferred from LSU gene DNA sequences. Posterior probabilities exceeding 50 in A review of the genus Tripylina Brzeski, 1963 (Nematoda: Triplonchida), with descriptions of five new species from New Zealand

FIGURE 2. Bayesian tree inferred from LSU gene DNA sequences. Posterior probabilities exceeding 50% are given on appropriate clades. Nematode species, GenBank numbers, locations are listed for each taxon if known.

opennotspecifiedDec 2009View details →
zenodo32/100

FIGURE 5 Bayesian phylogenetic tree inferred from SSU gene DNA sequences. Posterior probabilities great than 50 in New Zealand species of the genus Tripyla Bastian, 1865 (Nematoda: Triplonchida: Tripylidae). I: A new species, a new record and key to long-tailed species

FIGURE 5 Bayesian phylogenetic tree inferred from SSU gene DNA sequences. Posterior probabilities great than 50% are given on appropriate clades. Nematode species, GenBank numbers, locations are listed for each taxon if known.

opennotspecifiedDec 2009View details →
zenodo32/100

FIGURE 6 Bayesian phylogenetic tree inferred from LSU gene DNA sequences. Posterior probabilities greater than 50 in New Zealand species of the genus Tripyla Bastian, 1865 (Nematoda: Triplonchida: Tripylidae). I: A new species, a new record and key to long-tailed species

FIGURE 6 Bayesian phylogenetic tree inferred from LSU gene DNA sequences. Posterior probabilities greater than 50% are given on appropriate clades. Nematode species, GenBank numbers, locations are listed for each taxon if known.

opennotspecifiedDec 2009View details →
zenodo32/100

Development and validation of a 36-gene sequencing assay for hereditary cancer risk assessment

<p>Sanger and MLPA datasets for the article "Development and validation of a 36-gene sequencing assay for hereditary cancer risk assessment" (http://dx.doi.org/10.1101/088252).</p>

opencc-by-nc-4.0Dec 2016View details →
zenodo32/100

FIGURE 3 in Morphology and SSU rRNA gene sequence of the new brackish water ciliate, Anteholosticha pseudomonilata n. sp. (Ciliophora, Hypotrichida, Holostichidae) from Korea

FIGURE 3. The alignment of variable sites for SSU rRNA gene sequences of Anteholosticha pseudomonilata n. sp. and five congeners. Nucleotide position number is marked at the top of each aligned line. Missing sites are indicate by gaps (–) and the sites matched with A. pseudomonilata are represented with dots.

opennotspecifiedDec 2011View details →
zenodo32/100

FIGURE 2 in Morphology and SSU rRNA gene sequence of the new brackish water ciliate, Anteholosticha pseudomonilata n. sp. (Ciliophora, Hypotrichida, Holostichidae) from Korea

FIGURE 2. Photomicrographs of Anteholosticha pseudomonilata n. sp. from live cells (A–E), and after protargol staining (F– H). (A, B) Ventral view of typical individuals. (C) Partial of dorsal view, to show the distribution of cortical granules (arrowheads) and dorsal bristles (arrows). (D) Ventral view showing macronuclear nodules (arrows) and micronuclei (arrowheads). (E) Posterior body, arrows mark the inclusions within cytoplasm. (F) Ventral view of the holotype specimen (the same individual as illustrated in 1D, E and 2G, H), arrow denotes anterior end of left marginal row. (G) Ventral view of mid-posterior portion, arrows and arrowheads show micronuclei and macronuclear nodules, respectively. (H) Anterior portion of ventral side, showing the buccal cirrus (arrow), frontoterminal cirri (arrowheads), and the right frontal cirrus (double-arrowheads). Scale bars = 50 µm.

opennotspecifiedDec 2011View details →
zenodo32/100

FIGURE 1. Anteholosticha pseudomonilata n in Morphology and SSU rRNA gene sequence of the new brackish water ciliate, Anteholosticha pseudomonilata n. sp. (Ciliophora, Hypotrichida, Holostichidae) from Korea

FIGURE 1. Anteholosticha pseudomonilata n. sp. from live cells (A–C), and after protargol impregnation (D, E). (A) Ventral view of a typical individual. (B) Ventral view, arrows indicate the inclusions within cytoplasm at both cell ends. (C) Noting arrangement of cortical granules (arrows). (D, E) Ventral and dorsal views of the holotype specimen, showing the general infraciliature. Arrow in D marks the posterior end of the midventral complex. Arrowheads in E depict the "extra" dikinetids ahead of the right marginal row. AZM = adoral zone of membranelles; BC = buccal cirrus; EM = endoral membrane; FC = frontal cirri; FTC = frontoterminal cirri; LMR = left marginal row; Ma = macronuclei; Mi = micronuclei; MP = midventral pairs; PM = paroral membrane; PTC = pretransverse ventral cirri; RMR = right marginal row; TC = transverse cirri; 1-4 = dorsal kineties. Scale bars in (A) = 40 µm; in (D, E) = 30 µm.

opennotspecifiedDec 2011View details →
zenodo32/100

FIGURE 3 Bayesian tree inferred from SSU gene rDNA sequences. Posterior probabilities exceeding 50 in A review of the genus Trischistoma Cobb, 1913 (Nematoda: Enoplida), with descriptions of four new species from New Zealand

FIGURE 3 Bayesian tree inferred from SSU gene rDNA sequences. Posterior probabilities exceeding 50% are given on appropriate clades. Nematode species, GenBank numbers are listed for each taxon.

opennotspecifiedDec 2011View details →
zenodo32/100

FIGURE 4 Bayesian tree inferred from LSU gene rDNA sequences. Posterior probabilities exceeding 50 in A review of the genus Trischistoma Cobb, 1913 (Nematoda: Enoplida), with descriptions of four new species from New Zealand

FIGURE 4 Bayesian tree inferred from LSU gene rDNA sequences. Posterior probabilities exceeding 50% are given on appropriate clades. Nematode species and GenBank numbers are listed for each taxon if known.

opennotspecifiedDec 2011View details →
zenodo32/100

FIGURE 4. Bayesian tree inferred from LSU gene DNA sequences. Posterior probabilities exceeding 50 in Laimaphelenchus persicus n. sp. (Nematoda: Aphelenchoididae) from Iran

FIGURE 4. Bayesian tree inferred from LSU gene DNA sequences. Posterior probabilities exceeding 50% are given on appropriate clades. Nematode species and GenBank numbers are listed for each taxon.

opennotspecifiedDec 2012View details →
zenodo32/100

FIGURE 1. Bayesian phylogenetic tree inferred from SSU gene DNA sequences. Posterior probabilities great than 50 in New Zealand species of the genus Tripyla Bastian, 1865 (Nematoda: Triplonchida: Tripylidae). II: Two new, a known species and key to species

FIGURE 1. Bayesian phylogenetic tree inferred from SSU gene DNA sequences. Posterior probabilities great than 50% are given on appropriate clades. Nematode species, GenBank numbers, locations are listed for each taxon if known.

opennotspecifiedDec 2013View details →
zenodo32/100

FIGURE 2 in The morphology and SSU rRNA gene sequence analysis of a poorly-known brackish water ciliate, Pinacocoleps tesselatus (Kahl, 1930) (Ciliophora, Colepidae) from Hangzhou Bay, China

FIGURE 2. Photomicrographs of Pinacocoleps tesselatus (Kahl, 1930) from live cells (A–G), after silver carbonate impregnation (H, J), and after protargol impregnation (I). (A) Lateral view of a typical individual. (B) Anterior secondary plate. (C) Anterior main plate. (D) Posterior main plate. (E) Posterior secondary plate. (F) A squashed specimen showing the arrangement of plates. (G) Posterior view, arrows mark the posterior spines. (H) Lateral view, arrows denote the oral basket, arrowheads mark the extrusomes. (I) Anterior view, showing the oral structure, arrows indicate the adoral organelles. (J) Lateral view, showing the ciliary pattern. Scale bars (in A, J) = 30 μm; in (B–H) = 10 μm.

opennotspecifiedDec 2013View details →
zenodo32/100

FIGURE 1 in The morphology and SSU rRNA gene sequence analysis of a poorly-known brackish water ciliate, Pinacocoleps tesselatus (Kahl, 1930) (Ciliophora, Colepidae) from Hangzhou Bay, China

FIGURE 1. Morphology and infraciliature of Pinacocoleps tesselatus (Kahl, 1930) Foissner et al. 2008 (A–E), P. similis (Kahl, 1933) Chen et al. 2010 (F, G), P. heteracanthus (Noland, 1937) Chen et al. 2010 (H), P. arenarius (Bock, 1952) Chen et al. 2010 (I), P. spiralis (Noland, 1937) Chen et al. 2010 (J), P. i n c u r v u s (Ehrenberg, 1833) Foissner et al. 2008 (K), and P. pulcher (Spiegel, 1926) Foissner et al. 2008 (L). (A) Lateral view of typical individual. (B) One row of plates. The circumoral plate is omitted. (C) Ciliary pattern at apical end of body. (D) Ciliary pattern of P. t e s s e l a t u s. (E) P. tesselatus (Kahl, 1930) (from Kahl 1930). (F) P. similis (Kahl, 1933) (from Chen et al. 2010). (G) P. similis (Kahl, 1933) (from Borror 1972). (H) P. heteracanthus (Noland, 1937) (from Noland 1937). (I) P. arenarius (Bock, 1952) (from Bock 1952). (J) P. spiralis (Noland, 1937) (from Noland 1937). (K) P. i n c u r v u s (Ehrenberg, 1933) (from Kahl 1930). (L) P. pulcher (Spiegel, 1926) (from Kahl 1930). AO = adoral organelle; AS = anterior spine; CC = caudal cilium; CK = circumoral kinety; Ma = macronucleus; Mi = micronucleus; PC = perioral ciliature; PS = posterior spine; SK = somatic kinety. Scale bars = 30 μm.

opennotspecifiedDec 2013View details →
zenodo32/100

FIGURE 3 in The morphology and SSU rRNA gene sequence analysis of a poorly-known brackish water ciliate, Pinacocoleps tesselatus (Kahl, 1930) (Ciliophora, Colepidae) from Hangzhou Bay, China

FIGURE 3. Maximum likelihood (ML) phylogenetic tree based on the small subunit (SSU) rDNA of Pinacocoleps tesselatus and other colepids. Numbers at branching points show bootstrap values of 1,000 replicates for ML tree and posterior probability for Bayesian (BI) tree, respectively. Fully supported (100%/1.00) branches are marked with solid circles. The scale bar corresponds to 10 substitutions per 100 nucleotide positions. Taxonomic classification mainly follows Lynn (2008). GenBank numbers follow species names.

opennotspecifiedDec 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record