Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

42

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

42 results for “array genotyping”

Learn how ShareScore rates datasets ↗
zenodo44/100

Genotyping of the Chinese Spring x Renan mapping population with the TaBW280K SNP array

<p>The TaBW280K SNP array (Rimbert et al., PLoS ONE 2018) was used to genotype 430 Single Seed Descent (SSD) individuals<br> derived from a cross between Chinese Spring and Renan (CsRe; Choulet et al., Science 2014). Out of the 280,226 probesets, 85,276 were found to be polymorphic between the two parental lines and PHR on the population. Eventually, 83,721 (98.2%) SNPs were genetically mapped in 21 linkage groups corresponding to the 21 chromosomes of bread wheat, with no unlinked markers. This file contains the genotyping data of the 430 SSD lines.<br> &nbsp;</p>

opencc-by-4.0Jan 2018View details →
zenodo40/100

Development of a high-density 665 K SNP array for rainbow trout genome-wide genotyping. Supplemental VCF file

<p>Single nucleotide polymorphism (SNP) arrays, also named &laquo; SNP chips &raquo;, enable very large numbers of individuals to be genotyped at a targeted set of thousands of genome-wide identified markers. We used preexisting variant datasets from USDA, a French commercial line and 30X-coverage whole genome sequencing of INRAE isogenic lines to develop an Affymetrix 665 K SNP array (HD chip) for rainbow trout. In total, we identified 32,372,492 SNPs that were polymorphic in the USDA or INRAE databases. A subset of identified SNPs were selected for inclusion on the chip, prioritizing SNPs whose flanking sequence uniquely aligned to the Swanson reference genome, with homogenous repartition over the genome and the highest Minimum Allele Frequency in both USDA and French databases. Of the 664,531 SNPs which passed the Affymetrix quality filters and were manufactured on the HD chip, 65.3% and 60.9% passed filtering metrics and were polymorphic in two other distinct French commercial populations in which, respectively, 288 and 175 sampled fish were genotyped. Only 576,118 SNPs mapped uniquely on both Swanson and Arlee reference genomes, and 12,071 SNPs did not map at all on the Arlee reference genome. Among those 576,118 SNPs, 38,948 SNPs were kept from the&nbsp; commercially available medium-density 57K SNP chip. We demonstrate the utility of the HD chip by describing the high rates of&nbsp; linkage disequilibrium at 2 kb to 10 kb in the rainbow trout genome in comparison to the linkage disequilibrium observed at 50 kb to&nbsp; 100 kb which are usual distances between markers of the medium-density chip.</p> <p>&nbsp;</p> <p>File submitted correspond to the supplementary data 1 of the publication (under submission) : INRAE_USDA_MAF1.vcf.gz</p>

opencc-by-4.0Jun 2022View details →
dryad36/100

Complex feline disease mapping using a dense genotyping array

<p>The current feline genotyping array of 63k single nucleotide polymorphisms has proven its utility within breeds, and its use has led to the identification of variants associated with Mendelian traits in purebred cats. However, compared to single gene disorders, association studies of complex diseases, especially with the inclusion of random bred cats with relatively low linkage disequilibrium, require a denser genotyping array and an increased sample size to provide statistically significant associations. Here, we undertook a multi-breed study of 1,122 cats, most of which were admitted and phenotyped for nine common complex feline diseases at the Cornell University Hospital for Animals. Using a proprietary 340k single nucleotide polymorphism mapping array, we identified significant genome-wide associations with hyperthyroidism, diabetes mellitus, and eosinophilic keratoconjunctivitis. These results provide genomic locations for variant discovery and candidate gene screening for these important complex feline diseases, which are relevant not only to feline health, but also to the development of disease models for comparative studies.</p>

opencc-zeroApr 2022View details →
dryad36/100

SolCAP 8K array genotyping of population 15143

<p><span></span></p> <p>The dataset is a tab-deliminted text file with genotyping information for a diploid potato progeny population derived from a cross of diploid potato clones 12120-03 X 07506-01. The SolCap 8303 Infinium Chip was used for SNP genotyping.  It was originally generated for study on genetic mapping of Verticillium wilt resistance.  The data was used again for analysis of segregation distortion in the most recent manuscript.  </p>

opencc-zeroJul 2022View details →
zenodo36/100

Illumina HD genotypes for 3,092 cattle from Burkina Faso, Ghana, Nigeria and Tanzania for: "Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle"

<p>Raw HD data for Riggio&nbsp;et al. 2022:&nbsp;Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle</p> <p>This repository contains the raw Illumina HD genotypes (i.e., 777,962 SNPs) mapped to the bovine UMD3.1 genome assembly for 3,092 animals from four African countries (namely Burkina Faso, Ghana, Nigeria and Tanzania).&nbsp;</p>

opencc-by-4.0Jul 2022View details →
dryad36/100

SolCAP 8K array genotyping of population 15143

Open the record for dataset details and reuse information.

publicJul 2022View details →
dryad36/100

Complex feline disease mapping using a dense genotyping array

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad32/100

Data from: Imputation of canine genotype array data using 365 whole-genome sequences improves power of genome-wide association studies

Genomic resources for the domestic dog have improved with the widespread adoption of a 173k SNP array platform and updated reference genome. SNP arrays of this density are sufficient for detecting genetic associations within breeds but are underpowered for finding associations across multiple breeds or in mixed-breed dogs, where linkage disequilibrium rapidly decays between markers, even though such studies would hold particular promise for mapping complex diseases and traits. Here we introduce an imputation reference panel, consisting of 365 diverse, whole-genome sequenced dogs and wolves, which increases the number of markers that can be queried in genome-wide association studies approximately 130-fold. Using previously genotyped dogs, we show the utility of this reference panel in identifying potentially novel associations, including a locus on CFA20 significantly associated with cranial cruciate ligament disease, and fine-mapping for canine body size and blood phenotypes, even when causal loci are not in strong linkage disequilibrium with any single array marker. This reference panel resource will improve future genome-wide association studies for canine complex diseases and other phenotypes.

opencc-zeroAug 2020View details →
dryad32/100

Data from: Imputation of canine genotype array data using 365 whole-genome sequences improves power of genome-wide association studies

Open the record for dataset details and reuse information.

publicOct 2019View details →
dryad28/100

Data from: Characterization of the gray whale Eschrichtius robustus genome and a genotyping array based on single-nucleotide polymorphisms in candidate genes

Genetic and genomic approaches have much to offer in terms of ecology, evolution, and conservation. To better understand the biology of the gray whale Eschrichtius robustus (Lilljeborg, 1861), we sequenced the genome and produced an assembly that contains ∼95% of the genes known to be highly conserved among eukaryotes. From this assembly, we annotated 22,711 genes and identified 2,057,254 single-nucleotide polymorphisms (SNPs). Using this assembly, we generated a curated list of candidate genes potentially subject to strong natural selection, including genes associated with osmoregulation, oxygen binding and delivery, and other aspects of marine life. From these candidate genes, we queried 92 autosomal protein-coding markers with a panel of 96 SNPs that also included 2 sexing and 2 mitochondrial markers. Genotyping error rates, calculated across loci and across 69 intentional replicate samples, were low (0.021%), and observed heterozygosity was 0.33 averaged over all autosomal markers. This level of variability provides substantial discriminatory power across loci (mean probability of identity of 1.6 × 10−25 and mean probability of exclusion &gt;0.999 with neither parent known), indicating that these markers provide a powerful means to assess parentage and relatedness in gray whales. We found 29 unique multilocus genotypes represented among our 36 biopsies (indicating that we inadvertently sampled 7 whales twice). In total, we compiled an individual data set of 28 western gray whales (WGSs) and 1 presumptive eastern gray whale (EGW). The lone EGW we sampled was no more or less related to the WGWs than expected by chance alone. The gray whale genomes reported here will enable comparative studies of natural selection in cetaceans, and the SNP markers should be highly informative for future studies of gray whale evolution, population structure, demography, and relatedness.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Vitis phylogenomics: hybridization intensities from a SNP array outperform genotype calls

Understanding relationships among species is a fundamental goal of evolutionary biology. Single nucleotide polymorphisms (SNPs) identified through next generation sequencing and related technologies enable phylogeny reconstruction by providing unprecedented numbers of characters for analysis. One approach to SNP-based phylogeny reconstruction is to identify SNPs in a subset of individuals, and then to compile SNPs on an array that can be used to genotype additional samples at hundreds or thousands of sites simultaneously. Although powerful and efficient, this method is subject to ascertainment bias because applying variation discovered in a representative subset to a larger sample favors identification of SNPs with high minor allele frequencies and introduces bias against rare alleles. Here, we demonstrate that the use of hybridization intensity data, rather than genotype calls, reduces the effects of ascertainment bias. Whereas traditional SNP calls assess known variants based on diversity housed in the discovery panel, hybridization intensity data survey variation in the broader sample pool, regardless of whether those variants are present in the initial SNP discovery process. We apply SNP genotype and hybridization intensity data derived from the Vitis9kSNP array developed for grape to show the effects of ascertainment bias and to reconstruct evolutionary relationships among Vitis species. We demonstrate that phylogenies constructed using hybridization intensities suffer less from the distorting effects of ascertainment bias, and are thus more accurate than phylogenies based on genotype calls. Moreover, we reconstruct the phylogeny of the genus Vitis using hybridization data, show that North American subgenus Vitis species are monophyletic, and resolve several previously poorly known relationships among North American species. This study builds on earlier work that applied the Vitis9kSNP array to evolutionary questions within Vitis vinifera and has general implications for addressing ascertainment bias in array-enabled phylogeny reconstruction.

opencc-zeroDec 2012View details →
dryad28/100

Data from: The mouse universal genotyping array: from substrains to subspecies

Genotyping microarrays are an important resource for genetic mapping, population genetics, and monitoring of the genetic integrity of laboratory stocks. We have developed the third generation of the Mouse Universal Genotyping Array (MUGA) series, GigaMUGA, a 143,259-probe Illumina Infinium II array for the house mouse (Mus musculus). The bulk of the content of GigaMUGA is optimized for genetic mapping in the Collaborative Cross and Diversity Outbred populations, and for substrain-level identification of laboratory mice. In addition to 141,090 single nucleotide polymorphism probes, GigaMUGA contains 2006 probes for copy number concentrated in structurally polymorphic regions of the mouse genome. The performance of the array is characterized in a set of 500 high-quality reference samples spanning laboratory inbred strains, recombinant inbred lines, outbred stocks, and wild-caught mice. GigaMUGA is highly informative across a wide range of genetically diverse samples, from laboratory substrains to other Mus species. In addition to describing the content and performance of the array, we provide detailed probe-level annotation and recommendations for quality control.

opencc-zeroDec 2014View details →
dryad28/100

Data from: A 34K SNP genotyping array for Populus trichocarpa: Design, application to the study of natural populations and transferability to other Populus species

Open the record for dataset details and reuse information.

publicJan 2013View details →
dryad28/100

Data from: The mouse universal genotyping array: from substrains to subspecies

Open the record for dataset details and reuse information.

publicNov 2016View details →
dryad28/100

Data from: Vitis phylogenomics: hybridization intensities from a SNP array outperform genotype calls

Open the record for dataset details and reuse information.

publicOct 2014View details →
dryad28/100

Data from: The Mouse Universal Genotyping Array: from substrains to subspecies

Open the record for dataset details and reuse information.

publicSep 2016View details →
dryad28/100

Data from: Characterization of the gray whale Eschrichtius robustus genome and a genotyping array based on single-nucleotide polymorphisms in candidate genes

Open the record for dataset details and reuse information.

publicJun 2018View details →
dryad28/100

Data from: Development of SNP genotyping arrays in two shellfish species

Open the record for dataset details and reuse information.

publicJan 2014View details →
geo24/100

Low-pass sequencing increases the power of GWAS and decreases measurement error of polygenic risk scores compared to genotyping arrays

GEO Series GSE165845. Homo sapiens. 360 samples. Type: Genome variation profiling by array.

openGEO-OpenJan 2021View details →
geo24/100

Gene expression in xylem tissue on an Eucalyptus pseudo-testcross population: genotyping subset of discovery array probes

GEO Series GSE24195. Eucalyptus grandis x Eucalyptus urophylla; Eucalyptus. 136 samples. Type: Expression profiling by array; Other.

openGEO-OpenSep 2010View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record