Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

219

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

219 results for “genotyping‐by‐sequencing”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Identification and characterization of sex-associated loci in sockeye salmon using genotyping-by-sequencing and comparison with a sex-determining assay based on the sdY gene

Loci that can be used to screen for sex in salmon can provide important information for study of both wild and cultured populations. Here, we tested for associations between sex and genotypes at thousands of loci available from a genotyping-by-sequencing (GBS) dataset to discover sex-associated loci in sockeye salmon (Oncorhynchus nerka). We discovered seven sex-associated loci, developed high-throughput assays for two loci, and tested the utility of these two assays in eight collections of sockeye salmon sampled throughout North America. We also screened an existing assay based on the master sex-determining gene in salmon (sdY) in these collections. The ability of GBS-derived loci to assign fish to their phenotypic sex varied substantially among collections suggesting that recombination between the loci that we discovered and the sex-determining gene has occurred. Assignment accuracy to phenotypic sex was much higher with the sdY assay but was still less than 100%. Alignment of sequences from GBS-derived loci to draft genomes for two salmonids provided strong evidence that many of these loci are found on the chromosome orthologous to the known sex chromosome in sockeye salmon. Our study is the first to describe the approximate location of the sex-determining region in sockeye salmon and indicates that sdY is also the master sex-determining gene in this species. However, discordances between sdY genotypes and phenotypic sex and the variable performance of GBS-derived loci warrant more research.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Genotyping by sequencing reveals contrasting patterns of population structure, ecologically mediated divergence and long-distance dispersal in North American palms

Comparative studies can provide powerful insights into processes that affect population divergence and thereby help to elucidate the mechanisms by which contemporary populations may respond to environmental change. Furthermore, approaches such as genotyping by sequencing (GBS) provide unprecedented power for resolving genetic differences among species and populations. We therefore used GBS to provide a genome-wide perspective on the comparative population structure of two palm genera, Washingtonia and Brahea, on the Baja California peninsula, a region of high landscape and ecological complexity. First, we used phylogenetic analysis to address taxonomic uncertainties among five currently recognised species. We resolved three main clades, the first corresponding to W. robusta and W. filifera, the second to B. brandegeei and B. armata, and the third to B. edulis from Guadalupe Island. Focusing on the first two clades, we then delved deeper by investigating the underlying population structure. Striking differences were found, with GBS uncovering four distinct Washingtonia populations and identifying a suite of loci associated with temperature, consistent with ecologically mediated divergence. By contrast, individual mountain ranges could be resolved in Brahea and few loci were associated with environmental variables, implying a more prominent role of neutral divergence. Finally, evidence was found for long-distance dispersal events in Washingtonia but not Brahea, in line with knowledge of the dispersal mechanisms of these palms including the possibility of human-mediated dispersal. Overall, our study demonstrates the power of GBS together with a comparative approach to elucidate markedly different patterns of genome-wide divergence mediated by multiple effectors.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Genotyping-by-sequencing of genome-wide microsatellite loci reveals fine-scale harvest composition in a coastal Atlantic salmon fishery

Individual assignment and genetic mixture analysis are commonly utilized in contemporary wildlife and fisheries management. Although microsatellite loci provide unparalleled numbers of alleles per locus, their use in assignment applications is increasingly limited. However, next-generation sequencing, in conjunction with novel bioinformatic tools allows large numbers of microsatellite loci to be simultaneously genotyped, presenting new opportunities for individual assignment and genetic mixture analysis. Here we scanned the published Atlantic salmon genome to identify 706 microsatellite loci, from which we developed a final panel of 101 microsatellites distributed across the genome (average 3.4 loci per chromosome). Using samples from 35 Atlantic salmon populations (n=1485 individuals) from coastal Labrador, Canada, a region characterized by low levels of differentiation in this species, this panel identified 844 alleles (average of 8.4 alleles per locus). Simulation-based evaluations of assignment and mixture identification accuracy revealed unprecedented resolution, clearly identifying 26 rivers or groups of rivers spanning 500 km of coastline. This baseline was used to examine the stock composition of 696 individuals harvested in the Labrador Atlantic salmon fishery and revealed that coastal fisheries largely targeted regional groups (<300km). This work suggests that the development and application of large sequenced microsatellite panels presents great potential for stock resolution in Atlantic salmon and more broadly in other exploited anadromous and marine species.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Genotyping HapSTR loci: phase determination from direct sequencing of PCR products

HapSTRs combine information from a microsatellite (or simple tandem repeat, STR) with one or more single nucleotide polymorphisms (SNPs) in the DNA sequence immediately flanking the STR. These loci may offer increased power for the estimation of demographic parameters, but also present some challenges for data collection and analysis. We describe a process for inferring HapSTR alleles, including the flanking haplotypes, STR alleles, and their phase relative to each other, directly from DNA sequence electropherograms of PCR products from heterozygous individuals. Our approach eliminates the need for more costly and time-consuming processes such as cloning or acrylamide gel electrophoresis to separate alleles prior to sequencing.

opencc-zeroDec 2010View details →
dryad32/100

Data from: Development of a genotype-by-sequencing immunogenetic assay as exemplified by screening for variation in red fox with and without endemic rabies exposure

Pathogens are recognized as major drivers of local adaptation in wildlife systems. By determining which gene variants are favored in local interactions among populations with and without disease, spatially explicit adaptive responses to pathogens can be elucidated. Much of our current understanding of host responses to disease comes from a small number of genes associated with an immune response. High-throughput sequencing (HTS) technologies, such as genotype-by-sequencing (GBS), facilitate expanded explorations of genomic variation among populations. Hybridization-based GBS techniques can be leveraged in systems not well characterized for specific variants associated with disease outcome to "capture" specific genes and regulatory regions known to influence expression and disease outcome. We developed a multiplexed, sequence capture assay for red foxes to simultaneously assess ~300-kbp of genomic sequence from 116 adaptive, intrinsic, and innate immunity genes of predicted adaptive significance and their putative upstream regulatory regions along with 23 neutral microsatellite regions to control for demographic effects. The assay was applied to 45 fox DNA samples from Alaska, where three arctic rabies strains are geographically restricted and endemic to coastal tundra regions, yet absent from the boreal interior. The assay provided 61.5% on-target enrichment with relatively even sequence coverage across all targeted loci and samples (mean = 50×), which allowed us to elucidate genetic variation across introns, exons, and potential regulatory regions (4,819 SNPs). Challenges remained in accurately describing microsatellite variation using this technique; however, longer-read HTS technologies should overcome these issues. We used these data to conduct preliminary analyses and detected genetic structure in a subset of red fox immune-related genes between regions with and without endemic arctic rabies. This assay provides a template to assess immunogenetic variation in wildlife disease systems.

opencc-zeroDec 2016View details →
dryad32/100

Construction of genetic linkage map based on SNP markers, QTL mapping and detection of candidate genes of growth-related traits in Pacific abalone using genotyping-by-sequencing

<p><a name="_Hlk72585736"><span>Pacific abalone (<i>Haliotis discus hannai</i>) is a commercially important high valued molluscan species. Its wild population has decreased in recent years. Pacific abalone is widely cultured in Korea. Traditional breeding programs have been implemented for hatchery production of abalone seeds. To obtain more genetic information for the molecular breeding program, a high-density linkage map and quantitative trait locus (QTL) for three growth-related traits was constructed for Pacific abalone. F1 cross population with two parents were sampled to construct the linkage map using genotyping by sequencing (GBS). A total of 664,630,534 clean reads and 56,686 SNPs were generated. In sum, 3,345 segregating SNPs were used to construct a consensus linkage map. The map spanned 1,747.023 cM with 18 linkage groups and an average interval of 0.55 cM. QTL analysis revealed two significant QTL in LG10 on the consensus linkage map in each growth-related trait. Both the QTLs are located in the telomere region of the chromosome. Moreover, four potential candidate genes for growth-related traits were identified in the QTL region. Expression analysis revealed that identified genes are involved in growth regulation of abalone. The newly constructed genetic linkage map, growth-related QTLs and potential candidate genes identified in the present study can be used as valuable genetic resources and will be useful for marker-assisted selection (MAS) of Pacific abalone in molecular breeding program.</span></a></p>

opencc-zeroJun 2021View details →
dryad32/100

Data from: Population structure, relatedness and ploidy levels in an apple gene bank revealed through genotyping-by-sequencing

In recent years, new genome-wide marker systems have provided highly informative alternatives to low density marker systems for evaluating plant populations. To date, most apple germplasm collections have been genotyped using low-density markers such as simple sequence repeats (SSRs), whereas only a few have been explored using high-density genome-wide marker information. We explored the genetic diversity of the Pometum gene bank collection (University of Copenhagen, Denmark) of 349 apple accessions using over 15,000 genome-wide single nucleotide polymorphisms (SNPs) and 15 SSR markers, in order to compare the strength of the two approaches for describing population structure. We found that 119 accessions shared a clonal relationship with at least one other accession in the collection, resulting in the identification of 272 (78%) unique accessions. Of these unique accessions, over half (52%) share a first-degree relationship with at least one other accession. There is therefore a high degree of clonal and family relatedness in the Danish apple gene bank. We find significant genetic differentiation between Malus domestica and its supposed primary wild ancestor, M. sieversii, as well as between accessions of Danish origin and all others. Overall, we found strong concordance between analyses based on the genome-wide SNPs and the 15 SSR loci. However, we argue that GBS is superior to traditional SSR approaches because it allowed the estimation of ploidy levels that were in accordance with flow cytometry results, and can be further exploited in genome-wide association studies (GWAS). Finally, we compare GBS with SSR for the purposes of characterizing a diverse apple gene bank and discuss the advantages and constraints of the two approaches.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Genotyping-by-sequencing for estimating relatedness in non-model organisms: avoiding the trap of precise bias

There has been remarkably little attention to using the high resolution provided by genotyping-by-sequencing (i.e. RADseq and similar methods) datasets for assessing relatedness in wildlife populations. A major hurdle is the genotyping error, especially allelic dropout, often found in this type of dataset that could lead to downward-biased, yet precise, estimates of relatedness. Here we assess the applicability of genotyping-by-sequencing datasets for relatedness inferences given their relatively high genotyping error rates. Individuals of known relatedness were simulated under genotyping error, allelic dropout, and missing data scenarios based on an empirical ddRAD dataset, and their true relatedness was compared to that estimated by seven relatedness estimators. We found that an estimator chosen through such analyses can circumvent the influence of genotyping error, with the estimator of Ritland (1996) shown to be unaffected by allelic dropout and to be the most accurate when there is genotyping error. We also found that the choice of estimator should not rely solely on the strength of correlation between estimated and true relatedness as a strong correlation does not necessarily mean estimates are close to true relatedness. We also demonstrated how even a large SNP dataset with genotyping error (allelic dropout or otherwise) or missing data still performs better than a perfectly genotyped microsatellite dataset of tens of markers. The simulation-based approach used here can be easily implemented by others on their own genotyping-by-sequencing datasets to confirm the most appropriate and powerful estimator for their dataset.

opencc-zeroDec 2016View details →
dryad32/100

Data from: A high-density linkage map for Astyanax mexicanus using genotyping-by-sequencing technology

The Mexican tetra, Astyanax mexicanus, is a unique model system consisting of cave-adapted and surface-dwelling morphotypes which diverged &gt;1My ago. This remarkable natural experiment has enabled powerful genetic analyses of cave adaptation. Here, we describe the application of next-generation sequencing technology to the creation of a high-density linkage map. Our map comprises over 2200 markers populating 25 linkage groups constructed from genotypic data generated from a single genotyping-by-sequencing project. We leveraged emergent genomic and transcriptomic resources to anchor hundreds of anonymous Astyanax markers to the genome of the zebrafish (Danio rerio), the most closely related model organism to our study species. This facilitated the identification of 784 distinct connections between our linkage map and the Danio rerio genome, highlighting several regions of conserved genomic architecture between the two species despite ~150My of divergence. Using a Mendelian cave-associated trait as a proof-of-principle, we successfully recovered the genomic position of the albinism locus near the gene Oca2. Further, our map successfully informed the positions of unplaced Astyanax genomic scaffolds within particular linkage groups. This ability to identify the relative location, orientation and linear order of unaligned genomic scaffolds will facilitate ongoing efforts to improve upon the current early draft and assemble future versions of the Astyanax physical genome. Moreover, this improved linkage map will enable higher resolution genetic analyses and catalyze the discovery of the genetic basis for cave-associated phenotypes.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Genotype by sequencing identifies natural selection as a driver of intraspecific divergence in Atlantic populations of the high dispersal marine invertebrate, Macoma petalum

Mitochondrial DNA analyses indicate that the Bay of Fundy population of the intertidal tellinid bivalve Macoma petalum is genetically divergent from coastal populations in the Gulf of Maine and Nova Scotia. To further examine the evolutionary forces driving this genetic break, we performed double digest genotype by sequencing (GBS) to survey the nuclear genome for evidence of both neutral and selective processes shaping this pattern. The resulting reads were mapped to a partial transcriptome of its sister species, M. balthica, to identify single nucleotide polymorphisms (SNPs) in protein-coding genes. Population assignment tests, principle components analyses, analysis of molecular variance, and outlier tests all support differentiation between the Bay of Fundy genotype and the genotypes of the Gulf of Maine, Gulf of St. Lawrence, and Nova Scotia. Although both neutral and non-neutral patterns of genetic subdivision were significant, genetic structure among the regions was nearly 20 times higher for loci putatively under selection, suggesting a strong role for natural selection as a driver of genetic diversity in this species. Genetic differences were the greatest between the Bay of Fundy and all other population samples, and some outlier proteins were involved in immunity-related processes. Our results suggest that in combination with limited gene flow across the mouth of the Bay of Fundy, local adaptation is an important driver of intraspecific genetic variation in this marine species with high dispersal potential.

opencc-zeroDec 2016View details →
zenodo32/100

FIGURE 40 in A revision of the genus Lepidobrya Womersley (Collembola: Entomobryidae) based on morphology and sequence data of the genotype

FIGURE 40. Neighbor-Joining phylogram (below) and pairwise distance matrix (above) using patristic (P) distance for haplotypes of Lepidobrya mawsoni from Macquarie I. (MI) and Auckland Islands (Adams I.; ADI), and Lepidobrya spp. from Campbell I. (CI) and Antipodes I. (AI). Inset images: Lepidobrya sp. from Adam I. (left) and L. mawsoni from Macquarie I. (right). Codes, haplotypes (number of sequences with same haplotype in parentheses), locations and GenBank accessions refer to Table 1.

opennotspecifiedDec 2017View details →
zenodo32/100

FIGURES 34‒39. Lepidobrya mawsoni. 34 in A revision of the genus Lepidobrya Womersley (Collembola: Entomobryidae) based on morphology and sequence data of the genotype

FIGURES 34‒39. Lepidobrya mawsoni. 34, inner differentiated tibiotarsal chaetae; 35, distal part of manubrium ventrally; 36, S-chaetae on Abd. I; 37‒39, bothriotricha and accessory scales; 37, Abd. II laterally; 38, Abd. III laterally; 39, Abd. IV.

opennotspecifiedDec 2017View details →
zenodo32/100

FIGURES 30‒33 in A revision of the genus Lepidobrya Womersley (Collembola: Entomobryidae) based on morphology and sequence data of the genotype

FIGURES 30‒33. Tergal chaetotaxy in Lepidobrya mawsoni, left side. 30, thorax; 31, Abd. I‒III; 32, Abd. IV, only partial sens illustrated; 33, Abd. V.

opennotspecifiedDec 2017View details →
zenodo32/100

FIGURES 16‒24. Lepidobrya mawsoni. 16 in A revision of the genus Lepidobrya Womersley (Collembola: Entomobryidae) based on morphology and sequence data of the genotype

FIGURES 16‒24. Lepidobrya mawsoni. 16, Ant. III apical organ; 17, labrum; 18, clypeal chaetae; 19; dorsal cephalic chaetotaxy; 20, right labial papillae E, dorsal view; 21, labial and postlabial chaetae; 22, trochanteral organ; 23, hind claw, posterior view; 24, anterior face of ventral tube. Symbols representing chaetal elements used in this paper are as follows: large circle, macrochaeta; small circle, microchaeta; cross, bothriotrichum; circle with a slash, pseudopore.

opennotspecifiedDec 2017View details →
zenodo32/100

FIGURES 4‒9 in A revision of the genus Lepidobrya Womersley (Collembola: Entomobryidae) based on morphology and sequence data of the genotype

FIGURES 4‒9. Scales/chaetae in the posterior row along tergal margin in Lepidobrya mawsoni, left side except Fig. 9. 4, Th. II; 5, Th. III; 6, Abd. I; 7, Abd. II; 8, Abd. III; 9, Abd. IV. Scale bars: 4‒8, 50 µm; 9, 120 µm.

opennotspecifiedDec 2017View details →
zenodo32/100

FIGURES 10‒15. Lepidobrya mawsoni. 10 in A revision of the genus Lepidobrya Womersley (Collembola: Entomobryidae) based on morphology and sequence data of the genotype

FIGURES 10‒15. Lepidobrya mawsoni. 10, chaetae on Abd. V (right side); 11, ventral scales on manubrium; 12, dorsal side of manubrium; 13, dental scales; 14, base of left Ant. I, dorsal view; 15, base of left Ant. I, ventral view. Scale bars: 50 µm.

opennotspecifiedDec 2017View details →
zenodo32/100

FIGURES 25‒29. Lepidobrya mawsoni. 25 in A revision of the genus Lepidobrya Womersley (Collembola: Entomobryidae) based on morphology and sequence data of the genotype

FIGURES 25‒29. Lepidobrya mawsoni. 25, posterior face and lateral flap of ventral tube; 26, male genital plate; 27, manubrial plaque; 28, distal part of manubrium ventrally; 29, mucro.

opennotspecifiedDec 2017View details →
zenodo32/100

FIGURE 41 in A revision of the genus Lepidobrya Womersley (Collembola: Entomobryidae) based on morphology and sequence data of the genotype

FIGURE 41. Abundance in numbers per trap day of Lepidobrya mawsoni on Macquarie Island from December 1992 until December 1993.

opennotspecifiedDec 2017View details →
dryad32/100

Genotype likelihoods for low-coverage whole-genome sequencing data of yellow warblers

<p>The following datasets include the required input files used to empirically test population assignment in WGSassign on Yellow Warbler data. The file "yewa.known.ind105.ds_2x.beagle.gz" includes the filtered variants of 105 Yellow Warbler individuals output as genotype likelihoods and stored in a Beagle-formatted file. The ID file, "yewa.known.ind105.reference.IDs.txt", is a tab-delimited file with 2 columns, the first being the sample ID, and the second being the known reference population. The sample order in the ID file should match that of the input beagle file. To measure the assignment accuracy of WGSassign, we used leave-one-out cross validation using the input beagle file and our ID file.</p>

opencc-zeroJan 2024View details →
dryad32/100

Data from: Genome-wide sequence-based genotyping supports a nonhybrid origin of Castanea alabamensis

<p>The genus Castanea in North America contains multiple tree and shrub taxa of conservation concern. The two species within the group, American chestnut (Castanea dentata) and chinquapin (C. pumila sensu lato), display remarkable morphological diversity across their distributions in the eastern United States and southern Ontario. Previous investigators have hypothesized that hybridization between C. dentata and C. pumila has played an important role in generating morphological variation in wild populations. A putative hybrid taxon, Castanea alabamensis, was identified in northern Alabama in the early 20th century; however, the question of its hybridity has been unresolved. We tested the hypothesized hybrid origin of C. alabamensis using genome-wide sequence-based genotyping of C. alabamensis, all currently recognized North American Castanea taxa, and two Asian Castanea species at &gt;100,000 single-nucleotide polymorphism (SNP) loci. With these data, we generated a high-resolution phylogeny, tested for admixture among taxa, and analyzed population genetic structure of the study taxa. Bayesian clustering and principal components analysis provided no evidence of admixture between C. dentata and C. pumila in C. alabamensis genomes. Phylogenetic analysis of genome-wide SNP data indicated that C. alabamensis forms a distinct group within C. pumila sensu lato. Our results are consistent with the model of a nonhybrid origin for C. alabamensis. Our finding of C. alabamensis as a genetically and morphologically distinct group within the North American chinquapin complex provides further impetus for the study and conservation of the North American Castanea species.</p>

opencc-zeroFeb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record