Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

219

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

219 results for “genotyping‐by‐sequencing”

Learn how ShareScore rates datasets ↗
dryad36/100

Data from: Targeted genotyping-by-sequencing of potato and data analysis with R/polyBreedR

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad36/100

A genotype-by-sequencing dataset and identity-by-state matrix of genetic variation in Pinus radiata from 16 counties

Open the record for dataset details and reuse information.

publicAug 2025View details →
dryad36/100

Data from: RapidRat: development, validation and application of a genotyping-by-sequencing panel for rapid biosecurity and invasive species management

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad36/100

Data from: RAD sequencing, genotyping error estimation and de novo assembly optimization for population genetic inference

Open the record for dataset details and reuse information.

publicJun 2014View details →
dryad36/100

Data from: Genotyping by sequencing and genome–environment associations in wild common bean predict widespread divergent adaptation to drought

Open the record for dataset details and reuse information.

publicJan 2019View details →
dryad36/100

Data from: A comparison of non-destructive visceral swab and tissue biopsy sampling methods for genotyping-by-sequencing in the freshwater mussel Fusconaia askewi

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad36/100

Genotypes of Aedes aegypti mosquitoes derived from SNP chip and low-coverage whole genome sequencing for platform cross-validation

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad36/100

Data from: Evaluating genotyping-in-thousands by sequencing as a genetic monitoring tool for a climate sentinel mammal using non-invasive and archival samples

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad36/100

Genotyping-by-sequencing Single-nucleotide Polymorphism Dataset for Corynorhinus rafinesquii (CORA) and Myotis austroriparius (MYAU)

Open the record for dataset details and reuse information.

publicAug 2024View details →
dryad36/100

Data from: A genotyping-in-thousands by sequencing panel to inform invasive deer management using non-invasive fecal and hair samples

Open the record for dataset details and reuse information.

publicMay 2022View details →
dryad36/100

Genotyping-by-sequencing of Canada’s Apple Biodiversity Collection

Open the record for dataset details and reuse information.

publicMay 2022View details →
dryad32/100

Data from: Genotyping-in-Thousands by sequencing reveals marked population structure in Western Rattlesnakes to inform conservation status

<p>Delineation of units below the species level is critical for prioritizing conservation actions for species at-risk. Genetic studies play an important role in characterizing patterns of population connectivity and diversity to inform the designation of conservation units, especially for populations that are geographically isolated. The northernmost range margin of Western Rattlesnakes (<em>Crotalus oreganus</em>) occurs in British Columbia, Canada, where it is federally classified as threatened and restricted to five geographic regions. In these areas, Western Rattlesnakes hibernate (den) communally, raising questions about connectivity within and between den complexes. At present, Western Rattlesnake conservation efforts are hindered by a complete lack of information on genetic structure and degree of isolation at multiple scales, from the den to the regional level. To fill this knowledge gap, we used Genotyping-in-Thousands by sequencing (GT-seq) to genotype an optimized panel of 362 single nucleotide polymorphisms (SNPs) from individual samples (n = 461) collected across the snake's distribution in western Canada and neighboring Washington (USA). Hierarchical STRUCTURE analyses found evidence for population structure within and among the five geographic regions in BC, as well as in Washington. Within these regions, 11 genetically distinct complexes of dens were identified, with some regions having multiple complexes. No significant pattern of isolation-by-distance and generally low levels of migration were detected among den complexes across regions. Additionally, snakes within dens generally were more related than those among den complexes within a region, indicating limited movement. Overall, our results suggest that the single, recognized designatable unit for Western Rattlesnakes in Canada should be re-assessed to proactively focus conservation efforts on preserving total genetic variation detected range wide. More broadly, our study demonstrates a novel application of GT-seq for investigating patterns of diversity in wild populations at multiple scales to better inform conservation management.</p>

opencc-zeroDec 2019View details →
dryad32/100

Data from: Imputation of canine genotype array data using 365 whole-genome sequences improves power of genome-wide association studies

Genomic resources for the domestic dog have improved with the widespread adoption of a 173k SNP array platform and updated reference genome. SNP arrays of this density are sufficient for detecting genetic associations within breeds but are underpowered for finding associations across multiple breeds or in mixed-breed dogs, where linkage disequilibrium rapidly decays between markers, even though such studies would hold particular promise for mapping complex diseases and traits. Here we introduce an imputation reference panel, consisting of 365 diverse, whole-genome sequenced dogs and wolves, which increases the number of markers that can be queried in genome-wide association studies approximately 130-fold. Using previously genotyped dogs, we show the utility of this reference panel in identifying potentially novel associations, including a locus on CFA20 significantly associated with cranial cruciate ligament disease, and fine-mapping for canine body size and blood phenotypes, even when causal loci are not in strong linkage disequilibrium with any single array marker. This reference panel resource will improve future genome-wide association studies for canine complex diseases and other phenotypes.

opencc-zeroAug 2020View details →
zenodo32/100

Whole genome sequence analysis of porcine astroviruses reveals novel genetically diverse genotypes circulating in East African smallholder pig farms

<p>Supplementary materials for the porcine astrovirus study in East Africa.</p> <p><strong>Table S1</strong>: Pairwise comparison of nucleotide sequence identities of the complete (near complete, U460) genomes of the seven (7) astrovirus field strains (bold) and with sequences of other astroviruses available in GenBank&nbsp;</p> <p><strong>Table S2</strong>. Summary of nucleotide sequence identity matrix of the capsid protein (ORF2) among the seven (7) astroviruses field strains (bold) and the known reference strains in the GenBank using Clustal Omega</p> <p><strong>Table S3</strong>. Summary of amino acid sequence identity matrix of the capsid protein (ORF2) among the 7 astroviruses field strains (bold) and the known reference strains in the GenBank using Clustal Omega</p> <p><strong>Table S4</strong>: Estimates of evolutionary divergence between the East African PoAstVs and selected known AstV in the GenBank based on the amino acid sequences of complete ORF2 protein. The number of amino acid differences per site from between sequences is shown. Standard error estimate(s) are shown above the diagonal for our strains.</p> <p><strong>Table S5</strong>. Recommended potential linear antigenic epitopes predicted inside capsid protein (ORF2) of our field strains by SVMTriP web-based tool and corresponding antigenicity predicted by VaxiJen software</p>

opencc-by-4.0Sep 2020View details →
dryad32/100

Assigning the sex-specific markers via genotyping-by-sequencing onto the Y chromosome for a torrent frog Amolops mantzorum

<p><span>We use a genotyping-by-sequencing (GBS) approach to identify sex-linked markers in a torrent frog (<i>Amolops mantzorum</i>) using wild-caught individuals of 21 males and 19 females from the same population. A total of 141 putatively sex-linked markers were screened from 1,015,964 GBS tags through three approaches, respectively based on sex differences in allele frequencies, sex difference in heterozygosity, and sex-limited occurrence. With validations, 69 sex-linked markers were confirmed, all of which point to male heterogamety. The male specificity of eight sex markers was further verified by PCR amplifications with a large number of additional individuals covering the whole geographic distribution of the species. Y chromosome (No. 5) was microdissected under a light microscope, amplified by whole-genome amplification, and assembled a draft Y genome. 55 out of 69 sex-linked markers could be mapped to the Y chromosome assembly (i.e 79.7 %). Thus the chromosome 5 could be added as candidate chromosomes that particularly favored to recruit for sex determination than others among frogs. Three sex-linked markers that mapped on Y chromosome were aligned to three different promoter regions of <i>Rana rugosa</i> CYP19A1 gene, which might be considered as a candidate gene to trigger sex determination in <i>A</i>.<i> mantzorum</i>.</span></p>

opencc-zeroOct 2020View details →
dryad32/100

The 49,890 SNPgenotype derived from genotyping-by-sequencing strategy for the NIP/9311 backcross inbred lines population

<p>Transmission ratio distortion (TRD) refers to a widespread phenomenon in which one allele is transmitted by heterozygotes more frequently to the progeny than the opposite allele. TRD is considered as a mark suggesting the presence of reproductive barrier. However, the genetic and molecular mechanisms underlying TRD in rice remain largely unknown. In the present study, a population of backcross inbred lines (BILs) derived from the cross of a japonica cultivar Nipponbare and an indica variety 9311 was utilized to study the genetic base of TRD. A total of eighteen genomic regions were identified for TRD in the BILs. Among them, twelve and six regions showed indica (9311) and japonica (Nipponbare) alleles with preferential transmission, respectively. A series of F2 populations were used to confirm the TRD effects, including six genomic regions that were confirmed by chromosome segment substitution line (CSSL)-derived F2 populations from inter-subspecific allelic combinations. However, none of the regions was confirmed by the CSSL-derived populations from intra-subspecific allelic combination. Furthermore, significant epistatic interaction was found between TRD1.3 and TRD8.1 suggesting that TRD could positively contribute to breaking inter-subspecific reproductive barriers. Our results have laid the foundation for identifying the TRD genes and provide an effective strategy to breakdown TRD for breeding wide-compatible lines, which will be further utilized in the inter-subspecific hybrid breeding programs.</p>

opencc-zeroOct 2020View details →
dryad32/100

Microsatellite genotypes and ITS2 DNA sequence data for Seriatopora hystrix

<p>Coral reefs provide essential goods and services but are degrading at an alarming rate due to local and global anthropogenic stressors. The main limitation that prevents the implementation of adequate conservation measures is that connectivity and genetic structure of populations are poorly known. Here, the genetic diversity and connectivity of the brooding scleractinian coral, <i>Seriatopora hystrix</i> were assessed at two scales by genotyping ten microsatellite markers for 356 individual colonies. Seriatopora hystrix showed high differentiation, both at large scale between the Red Sea and the Western Indian Ocean (WIO), and at smaller scale along the coast of East Africa.As such high levels of differentiation might indicate the presence of more than one species, a haploweb analysis was conducted with the nuclear marker ITS2, confirming that the Red Sea populations are genetically distinct from the WIO ones.Based on microsatellite analyses three groups could be distinguished within the WIO: (I) north Madagascar, (II) south-west Madagascar together with one site in northern Mozambique (Nacala), and (III) all other sites in northern Mozambique, Tanzania and Kenya. These patterns of restricted connectivity could be explained by the short pelagic larval duration of <i>S. hystrix,</i> and/or by oceanographic factors, such as eddies in the Mozambique Channel (causing larval retention in northern Madagascar but facilitating dispersal from northern Mozambique towards south-west Madagascar). This study provides an additional line of evidence supporting the conservation priority status of the Northern Mozambique Channel and should inform coral reef management decisions in the region.</p> <p> </p>

opencc-zeroDec 2019View details →
dryad32/100

Data from: Genotyping-by-sequencing for Populus population genomics: an assessment of genome sampling patterns and filtering approaches

Continuing advances in nucleotide sequencing technology are inspiring a suite of genomic approaches in studies of natural populations. Researchers are faced with data management and analytical scales that are increasing by orders of magnitude. With such dramatic advances comes a need to understand biases and error rates, which can be propagated and magnified in large-scale data acquisition and processing. Here we assess genomic sampling biases and the effects of various population-level data filtering strategies in a genotyping-by-sequencing (GBS) protocol. We focus on data from two species of Populus, because this genus has a relatively small genome and is emerging as a target for population genomic studies. We estimate the proportions and patterns of genomic sampling by examining the Populus trichocarpa genome (Nisqually-1), and demonstrate a pronounced bias towards coding regions when using the methylation-sensitive ApeKI restriction enzyme in this species. Using population-level data from a closely related species (P. tremuloides), we also investigate various approaches for filtering GBS data to retain high-depth, informative SNPs that can be used for population genetic analyses. We find a data filter that includes the designation of ambiguous alleles resulted in metrics of population structure and Hardy-Weinberg equilibrium that were most consistent with previous studies of the same populations based on other genetic markers. Analyses of the filtered data (27,910 SNPs) also resulted in patterns of heterozygosity and population structure similar to a previous study using microsatellites. Our application demonstrates that technically and analytically simple approaches can readily be developed for population genomics of natural populations.

opencc-zeroDec 2013View details →
dryad32/100

Data from: A high-density exome capture genotype-by-sequencing panel for forestry breeding in Pinus radiata

Development of genome-wide resources for application in genomic selection or genome-wide association studies, in the absences of full reference genomes, present a challenge to the forestry industry, where longer breeding cycles could benefit from the accelerated selection possible through marker-based breeding value predictions. In particular, large conifer megagenomes require a strategy to reduce complexity, whilst ensuring genome-wide coverage is achieved. Using a transcriptome-based reference template, we have successfully developed a high density exome capture genotype-by-sequencing panel for radiata pine (Pinus radiata D.Don), capable of capturing in excess of 80,000 single nucleotide polymorphism (SNP) markers with a minor allele frequency above 0.03 in the population tested. This represents approximately 29,000 gene models from a core set of 48,914 probes. A set of 704 SMP markers capable of pedigree reconstruction and differentiating individual genotypes were tested within two full-sib mapping populations. While as few as 70 markers could reconstruct parentage in almost all cases, the impact of missing genotypes was noticeable in several offspring. Therefore, sets of 60 sets of 110 randomly selected SNP markers were compared for both parentage reconstruction and clone differentiation. The performance in parentage reconstruction showed little variation over 60 iterations. However, there was notable variation in discriminatory power between closely related individuals, indicating a higher density SNP marker panel may be required to elucidate hidden relationships in complex pedigrees.

opencc-zeroOct 2019View details →
dryad32/100

Data from: Finding the right coverage: The impact of coverage and sequence quality on SNP genotyping error rates

Restriction-enzyme-based sequencing methods enable the genotyping of thousands of single nucleotide polymorphism (SNP) loci in non-model organisms. However, in contrast to traditional genetic markers, genotyping error rates in SNPs derived from restriction-enzyme-based methods remain largely unknown. Here, we estimated genotyping error rates in SNPs genotyped with double digest RAD sequencing from Mendelian incompatibilities in known mother-offspring dyads of Hoffman's two-toed sloth (Choloepus hoffmanni) across a range of coverage and sequence quality criteria, for both reference-aligned and de novo-assembled datasets. Genotyping error rates were more sensitive to coverage than sequence quality and low coverage yielded high error rates, particularly in de novo-assembled datasets. For example, coverage ≥5 yielded median genotyping error rates of ≥0.03 and ≥0.11 in reference-aligned- and de novo-assembled datasets, respectively. Genotyping error rates declined to ≤0.01 in reference-aligned datasets with a coverage &gt;30, but remained &gt;0.04 in the de novo-assembled datasets. We observed approximately 10- and 13-fold declines in the number of loci sampled in the reference-aligned and de novo-assembled datasets when coverage was increased from &gt;5 to &gt;30 at quality score ≥30, respectively. Finally, we assessed the effects of genotyping coverage on a common population genetic application, parentage assignments, and showed that the proportion of incorrectly assigned maternities was relatively high at low coverage. Overall, our results suggest that the tradeoff between sample size and genotyping error rates be considered prior to building sequencing libraries, reporting genotyping error rates become standard practice, and that effects of genotyping errors on inference be evaluated in restriction-enzyme-based SNP studies.

opencc-zeroDec 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record