Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,019

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,019 results for “SNP”

Learn how ShareScore rates datasets ↗
dryad36/100

Estimating the inbreeding level and genetic relatedness in an isolated population of critically endangered Sichuan taimen (Hucho bleekeri) using genome wide SNP markers

<p>Sichuan taimen (Hucho bleekeri) is critically endangered fish listed in The Red List of Threatened Species compiled by the International Union for Conservation of Nature (IUCN). Specific locus amplified fragment sequencing (SLAF-seq)-based genotyping was performed for Sichuan taimen with 43 yearling individuals from 3 locations in Taibai River (a tributary of Yangtze River) that has been sequestered from its access to the ocean for more than 30 years since late 1980s. Applying the inbreeding level and genetic relatedness estimation using 15,396 genome wide SNP markers, we found that the inbreeding level of this whole isolated population was at a low level (average F=2.6×10-3±0.079), and the means of coancestry coefficients within and between the three sampling locations were all very low (close to 0), too. Genomic differentiation was negatively correlated with the geographical distances between the sampling locations (p &lt; 0.001) and the 43 individuals could be considered as genetically independent two groups. The low levels of genomic inbreeding and relatedness indicated a relatively large number of sexually mature individuals were involved in reproduction in Taibai River. This study suggested a genomic-relatedness-guided breeding and conservation strategy for wild fish species without pedigree information records.</p>

opencc-zeroJan 2021View details →
dryad36/100

Population structure of five native sheep breeds of Sweden estimated with high density SNP genotypes

Background <p>Native Swedish sheep breeds are part of the North European short-tailed sheep group; characterized in part by their genetic uniqueness. Our objective was to study the population structure of native Swedish sheep. Five breeds were genotyped using the 600 K SNP array. Dalapäls and Klövsjö sheep are from the middle of Sweden; Gotland and Gute sheep from Gotland, an island in the Baltic Sea; and Fjällnäs sheep from northern Sweden. We studied population structure by: principal component analysis (PCA), cluster-based analysis of admixture, and an estimated population tree.</p> Results <p>The analyses of the five Swedish breeds revealed that these breeds are five distinct breeds, while Gute and Gotland are more closely related to each other as seen in all analyses. All breeds had long branch lengths in the population tree indicating they've been subjected to drift. We repeated our analyses using 39 K SNP and including 50 K SNP genotypes from other European and southwestern Asian breeds from the Sheep HapMap project and 600 K SNP genotypes from a dataset of French sheep. Results arranged breeds into five groups: south-west Asia, south-west Europe, central Europe, north Europe and north European short-tailed sheep. Within this last group, Norwegian and Icelandic breeds, Finn and Romanov sheep, Scottish breeds, and Gute and Gotland sheep were more closely related while the remaining Swedish breeds and Ouessant sheep were distinct from all breeds and had longer branches in the population tree.</p> Conclusions <p>We showed population structure of five Swedish breeds and their structure within European and southwestern Asian breeds. Swedish breeds are unique, distinct breeds that have been subjected to drift but group with other north European short-tailed sheep.</p>

opencc-zeroDec 2019View details →
dryad36/100

Data from: Phylogenetic relationships, breeding implications, and cultivation history of Hawaiian taro (Colocasia esculenta) through genome-wide SNP genotyping

Taro, Colocasia esculenta, is one of the world's oldest root crops and of particular economic and cultural significance in Hawai'i, where historically more than 150 different landraces were grown. We developed a genome-wide set of more than 2400 high-quality single nucleotide polymorphism (SNP) markers from 70 taro accessions of Hawaiian, South Pacific, Palauan, and mainland Asian origins, with several objectives: (a) uncover the phylogenetic relationships between Hawaiian and other Pacific landraces, (b) shed light on the history of taro cultivation in Hawai'i, and (c) develop a tool to discriminate among Hawaiian and other taros. We found that almost all existing Hawaiian landraces fall into five monophyletic groups that are largely consistent with the traditional Hawaiian classification based on morphological characters, e.g., leaf shape and petiole color. Genetic diversity was low within these clades but considerably higher between them. Population structure analyses further indicated that the diversification of taro in Hawai'i most likely occurred by a combination of frequent somatic mutation and occasional hybridization. Unexpectedly, the South Pacific accessions were found nested within the clades mainly composed of Hawaiian accessions, rather than paraphyletic to them. This suggests that the origin of clades identified here preceded the colonization of Hawai'i, and that early Polynesian settlers brought taro landraces from different clades with them. In the absence of a sequenced genome, this marker set provides a valuable resource towards obtaining a genetic linkage map, and to study the genetic basis of phenotypic traits of interest to taro breeding such as disease resistance.

opencc-zeroDec 2016View details →
dryad36/100

Data from: Structure and phylogeography of two tropical predators, spinner (Stenella longirostris) and pantropical spotted (S attenuata) dolphins, from SNP data

Little is known about global patterns of genetic connectivity in pelagic dolphins, including how circumtropical pelagic dolphins spread globally following the rapid and recent radiation of the subfamily delphininae. In this study, we tested phylogeographic hypotheses for two circumtropical species, the spinner dolphin (Stenella longirostris) and the pantropical spotted dolphin (Stenella attenuata), using &gt;3,000 nuclear DNA single nucleotide polymorphisms (SNPs) in each species. Analyses for population structure indicated significant genetic differentiation between almost all subspecies and populations in both species. Bayesian phylogeographic analyses of spinner dolphins showed deep divergence between Indo-Pacific, Atlantic, and eastern tropical Pacific Ocean (ETP) lineages. Despite high morphological variation, our results show very close relationships between endemic ETP spinner subspecies in relation to global diversity. The dwarf spinner dolphin is a monophyletic subspecies nested within a major clade of pantropical spinner dolphins from the Indian and western Pacific Ocean populations. Population-level division among the dwarf spinner dolphins was detected–with the northern Australia population being very different from that in Indonesia. In contrast to spinner dolphins, the major boundary for spotted dolphins is between offshore and coastal habitats in the ETP, supporting the current subspecies-level taxonomy. Comparing these species underscores the different scale at which population structure can arise, even in species that are similar in habitat (i.e., pelagic) and distribution.

opencc-zeroDec 2017View details →
dryad36/100

Implementing large genomic SNP datasets in phylogenetic network reconstructions: a case study of particularly rapid radiations of cichlid fish

<p><span><span><span><span><span><span><span><span><span><span><span>The Midas cichlids of the <i>Amphilophus</i> <i>citrinellus </i>spp<i>.</i> species complex from Nicaragua, are an extraordinary adaptive rapid radiation (&lt;24,000 years old; 13 described species). These cichlids are a very challenging group to infer its evolutionary history in phylogenetic analyses, due to the apparent prevalence of ILS, as well as past and current gene flow. Assuming solely a vertical transfer of genetic material from an ancestral lineage to new lineages is not appropriate in many cases of genes transferred horizontally in nature. Recently developed methods to infer phylogenetic networks under such circumstances might be able to circumvent these problems. These models accommodate not just incomplete lineage sorting, but also gene flow, under the multispecies network coalescent model (MSNC), processes that are at work in young, hybridizing, and/or rapidly diversifying lineages. There are currently only a few programs available that implement MSNC for estimating phylogenetic networks. Here, we present a novel way to incorporate single nucleotide polymorphism (SNP) data into the currently available PhyloNetworks program. Based on simulations, we demonstrate that SNPs can provide enough power to recover the true phylogenetic network. Moreover, our approach results in a faster algorithm compared to the original pipeline in PhyloNetworks, without losing power. We also applied our new approach to infer the phylogenetic network of Midas cichlid radiation. We implemented the most comprehensive genomic dataset to date (RADseq dataset of 679 individuals and &gt;37K SNPs from 19 ingroup lineages) <span><span>and present estimated phylogenetic networks for this extremely young and fast-evolving radiation of cichlid fish. </span></span>We demonstrate that the MSNC is more appropriate than the multispecies coalescent alone for the analysis of this rapid radiation. </span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroFeb 2020View details →
dryad36/100

SNP data for Northern Alligator Lizards

<p>Understanding the processes that shape genetic diversity by either promoting or preventing population divergence can help identify geographic areas that either facilitate or limit gene flow. Furthermore, broadly distributed species allow us to understand how biogeographic and ecogeographic transitions affect gene flow. We investigated these processes using genomic data in the Northern Alligator Lizard (<em>Elgaria coerulea</em>), which is widely distributed in Western North America across diverse ecoregions (California Floristic Province and Pacific Northwest) and mountain ranges (Sierra Nevada, Coastal Ranges, and Cascades). We collected single nucleotide polymorphism (SNP) data from 120 samples of <em>E. coerulea</em>. Biogeographic analyses of squamate reptiles with similar distributions have identified several shared diversification patterns that provide testable predictions for <em>E. coerulea</em>, including deep genetic divisions in the Sierra Nevada, demographic stability of southern populations, and recent post-Pleistocene expansion into the Pacific Northwest. We use genomic data to test these predictions by estimating the structure, connectivity, and phylogenetic history of populations. At least ten distinct populations are supported, with mixed-ancestry individuals situated at most population boundaries. A species tree analysis provides strong support for the early divergence of populations in the Sierra Nevada Mountains and recent diversification into the Pacific Northwest. Admixture and migration analyses detect gene flow among populations in the Lower Cascades and Northern California, and a spatial analysis of gene flow identified significant barriers to gene flow across both the Sierra Nevada and Coast Ranges. The distribution of genetic diversity in <em>E. coerulea</em> is uneven, patchy, and interconnected at population boundaries. The biogeographic patterns seen in <em>E. coerulea</em> are consistent with predictions from co-distributed species.</p>

opencc-zeroNov 2023View details →
zenodo36/100

Mingrelian SNP Genotype Data

<p>This dataset contains data from 645,337 single nucleotide polymorphisms (SNPs) that were genotyped on GenoChip 2+ microarrays. The SNP data were ascertained from the mtDNA, Y-chromosome and autosomes for each individual, depending on their biological sex. In total, 5,205 mtDNA and 10,272 Y-chromosome SNPs were extracted from the array data. These data files have been uploaded as .csv files and also be uploaded as plink-formatted files. Details about the analysis of the SNP data can be found in the associated manuscript:</p><p>Theodore G Schurr, Ramaz Shengelia, Michel Shamoon-Pour, David Chitanava, Shorena Laliashvili, Irma Laliashvili, Redate Kibret, Yanu Kume-Kangkolo, Irakli Akhvlediani, Lia Bitadze, Iain Mathieson, Aram Yardumian, Genetic Analysis of Mingrelians Reveals Long-Term Continuity of Populations in Western Georgia (Caucasus),&nbsp;<i>Genome Biology and Evolution</i>, 2023; evad198,&nbsp;<a href="https://doi.org/10.1093/gbe/evad198">https://doi.org/10.1093/gbe/evad198</a></p>

opencc-by-4.0Oct 2023View details →
dryad36/100

SNP genotypes from Magallanes

<p>Hybrid zones among mussel species have been extensively studied in the northern hemisphere. In South America, it has only recently become possible to study the natural hybrid zones, due to the clarification of the taxonomy of native mussels of the <em>Mytilus</em> genus. Analyzing 54 SNP markers, we show the genetic species composition and admixture in the hybrid zone between <em>M. chilensis </em>and <em>M. platensis</em> in the southern end of South America. Bayesian, non-Bayesian clustering and re-assignment algorithms showed that the natural hybrid zone between <em>M. chilensis </em>and <em>M. platensis </em>in the Strait of Magellan, Isla Grande de Tierra del Fuego, and the Falkland Islands shows complex architecture. It can be divided into three different areas: the first one is on the Atlantic coast where only pure <em>M. platensis</em> and hybrid were found. In the second one, inside the Strait of Magellan, pure individuals of both species and mussels with variable degrees of hybridization coexist. In the last area at the Strait in front of Punta Arenas City, fjords on the Isla Grande de Tierra del Fuego, and at the Beagle Channel, only <em>M. chilensis</em> and a low number of hybrids were found.  According to the proportion of hybrids, bays with protected conditions away from strong currents would give better conditions for hybridization. We do not find evidence of any other mussel species such as <em>M. edulis, M. galloprovincialis, M. planulatus, </em>or <em>M. trossulus </em>in the zone</p>

opencc-zeroDec 2023View details →
dryad36/100

SNP data for F2 population derived from Oryza rufipogon and O. nivara

<p>To elucidate the genetic architecture underlying phenotypic divergence is essential to the understanding of ecological adaptation and speciation. Two wild rice species (<em>O. rufipogon</em> and <em>O. nivara</em>) are a progenitor-daughter species pair with ecological divergence and provide a unique system for studying ecological adaptation/speciation. Here, we constructed a high-resolved linkage map and conducted a quantitative trait locus (QTL) analysis of 19 phenotypic traits using an F<sub>2</sub> population generated from a cross between the perennial <em>O. rufipogon</em> and annual <em>O. nivara</em>. We identified 113 QTLs associated with interspecific divergence of 16 quantitative traits, with effect sizes ranging from 1.61% to 34.1% in terms of the percentage of variation explained (PVE). The distribution of effect sizes of QTLs followed a negative exponential, suggesting that a few genes of large effect and many genes of small effect were responsible for the phenotypic divergence. We observed 18 clusters of QTLs (QTL hotspots), with each involving multiple adaptive traits, demonstrating the importance of coinheritance of loci/genes in ecological adaptation/speciation. Analysis of effect direction and <em>v</em>-test statistics revealed that interspecific differentiation of most traits was driven by divergent natural selection, supporting the argument that ecological adaptation/speciation would proceed rapidly under coordinated selection on multiple traits.</p>

opencc-zeroJan 2024View details →
dryad36/100

Mapping of the QTLs governing grain micronutrients and thousand kernel weight in wheat (Triticum aestivum L.) using high density SNP markers

<p>The mapping population consists of 166 recombinant inbred lines (RILs) derived from a cross between HD3086 and HI1500.</p> <p><strong>Phenotypic data</strong><br>The RILs population along with parents were evaluated under four conditions namely timely sown irrigation (TSIR) taken as control, timely sown restricted irrigation (TSRI), late sown irrigation (LSIR), and late sown restricted irrigation (LSRI) conditions at Delhi, and under restricted irrigation condition at Indore. From each plot, 20 random spikes were harvested and spikes from each plot were threshed separately. While cleaning, care was taken to prevent metal and dust contamination. The grain iron concentration (GFeC) and grain zinc concentration (GZnC) were measured using Energy Dispersive X-ray Fluorescence (ED-XRF) machine (model X-Supreme 8000 M/s Oxford Inc, USA).  The thousand kernel weight (TKW) was recorded by counting 1000 grains manually and weighted with an electronic balance.</p> <p><strong>Genotypic data</strong><br>DNA was extracted from 21 days old seedlings using CTAB method (Murray and Thompson, 1980). Genomic DNA quality was determined using 0.8% agarose gel electrophoresis with λ DNA as the standard and quantified using nanodrop. The 35K SNP Axiom breeders' array was used for genotyping of parents and the RILs population.</p>

opencc-zeroJan 2024View details →
dryad36/100

SNP markers used for QTL mapping in the inbred lines

<p><span>Young leaves of the 175 inbred lines and their seven parents were collected from seedlings grown in a greenhouse. </span><span>About 200 mg bulk leaf sample from three plants of a line was placed in 2 ml safe-lock </span><span>Eppendorf tube and stored at ‒80 </span><span>˚C for one night prior to crushing using a Mixer Mill (TissueLyser II, Qiagen, Germany). Genomic DNA was extracted using SIGMA DNA extraction kit (Sigma-Aldrich, St. Louis, MO, USA) following the manufacturer's instruction. DNA concentration and purity of the samples were assessed using a NanoDrop 2000c spectrophotometer (Thermo Scientific, Wilmington, DE, USA). The samples were processed and sequenced using tunable genotyping-by-sequencing (tGBS®) method by Data2Bio (Ames, IW, USA). Genomic DNA was digested using two restriction enzymes NSpI (5′-RCATG^Y-3′) and BfuCI/Sau3AI (5′-^GATC-3′) which created 3´and 5´overhangs, respectively. Two single-stranded oligos, one containing a sample-specific internal barcode and the other a universal oligo, were ligated to the complementary 3´ and 5´ overhangs, respectively. </span>All 175 inbred lines' and seven parents' treated DNA was pooled for construction of the tGBS library and sequencing. The raw sequence data were demultiplexed by barcode, which was subsequently removed bioinformatically from each sequence. The barcode-trimmed sequence reads of genotype were further trimmed using the trimming software, Lucy (Chou &amp; Holmes, 2001; Li &amp; Chou 2004) to remove low-quality reads based on Phred quality scores of Q15.</p>

opencc-zeroFeb 2024View details →
zenodo36/100

Panmixia and active colonisation of the invasive palm Trachycarpus fortunei (Arecaceae) in Southern Switzerland and Northern Italy as inferred by microsatellites and SNP markers

<p>Dataset for the paper named &quot;Panmixia and active colonisation of the invasive palm Trachycarpus fortunei (Arecaceae) in Southern Switzerland and Northern Italy as inferred by microsatellites and SNP markers&quot;</p> <p>GBS analysis:</p> <p>- variants.vcf.gz : compressed non filtered VCF file with 208 samples and 73685 markers on 36195 loci</p> <p>- variants.filt.vcf.gz:&nbsp; Filtered Variant call file (compressed) - Samples with &gt; 50% missing genotypes, and variants with genotype calls in less than 80% of samples are removed; variants with maf &lt; 1% are removed -207 samples and 31312 markers on 19301 loci - 1 samples removed 6CL</p> <p>Microsatellites:</p> <p>TFT.fortunei_Microsatellites_FSTATFINAL_Pop.dat</p> <p>Samples file</p> <p>-Trachycarpus_Samples_sheet.xlsx : list of samples used (lab extractions)&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
dryad36/100

Genomics of humic adaptation in Eurasian perch (Perca fluviatilis): SNP genotypes of 32 perch individuals, supplementary figures and tables

<p>Extreme <span>environments are inhospitable to the majority of species, but some organisms are able to survive in such hostile conditions due to evolutionary adaptations. For example, </span><span>m</span><span>odern bony fishes have colonized various aquatic environments, including perpetually dark,</span><span> hypoxic, hypersaline and toxic habitats</span><span>. </span><span>Eurasian perch (</span><em>Perca fluviatilis</em><span>) is among the few fish species of northern latitudes that is able to live in very acidic humic lakes. Such lakes represent almost "nocturnal" environments; they contain high levels of dissolved organic matter, which in addition to creating a challenging visual environment, also affects a large number of other habitat parameters and biotic interactions. To reveal the genomic targets of humic-associated selection, we performed whole-genome sequencing of perch originating from 16 humic and 16 clear-water lakes in northern Europe. We identified over 800,000 SNPs, of which &gt;10,000 were identified as potential candidates under selection (associated with &gt;3,000 genes) using multiple outlier approaches. Our findings suggest that adaptation to the humic environment may involve hundreds of regions scattered across the genome. Putative signals of adaptation were detected in genes and gene families with diverse functions, including organism development and ion transportation. The observed excess of variants under selection in regulatory regions highlights the importance of adaptive evolution via regulatory elements, rather than via protein sequence modification. Our study demonstrates the power of whole-genome analysis to illuminate multifaceted nature of humic adaptation and provides the foundation for further investigation of causal mutations underlying phenotypic traits of ecological and evolutionary importance.</span></p>

opencc-zeroMar 2022View details →
dryad36/100

Data from: Targeted genome-wide SNP genotyping in feral horses using non-invasive fecal swabs

<p>The development of high-throughput sequencing has prompted a transition in wildlife genetics from using microsatellites toward sets of Single Nucleotide Polymorphisms (SNPs). However, genotyping large numbers of targeted SNPs using non-invasive samples remains challenging due to relatively large DNA input requirements. Recently, target enrichment has emerged as a promising approach requiring little template DNA. We assessed the efficacy of Tecan Genomics' Allegro Targeted Genotyping (ATG) for generating genome-wide SNP data in feral horses using DNA isolated from fecal swabs. Total and host-specific DNA were quantified for 989 samples collected as part of a long-term individual-based study of feral horses on Sable Island, Nova Scotia, Canada, using dsDNA fluorescence and a host-specific qPCR assay, respectively. Forty-eight samples representing 44 individuals containing at least 10ng of host DNA (ATG's recommended minimum input) were genotyped using a custom multiplex panel targeting 279 SNPs. Genotyping accuracy and consistency were assessed by contrasting ATG genotypes with those obtained from the same individuals with SNP microarrays, and from multiple samples from the same horse, respectively. 62% of swabs yielded the minimum recommended amount of host DNA for ATG. Ignoring samples that failed to amplify, ATG recovered an average of 86.7% targeted sites per sample, while genotype concordance between ATG and SNP microarrays was 98.5%. The repeatability of genotypes from the same individual approached unity with an average of 99.9%. This study demonstrates the suitability of ATG for genome-wide, non-invasive targeted SNP genotyping, and will facilitate further ecological and conservation genetics research in equids and related species.</p>

opencc-zeroApr 2022View details →
dryad36/100

Aedes aegypti in North America (Microsatellite and SNP array)

<p>The <em>Aedes aegypti</em> mosquito first invaded the Americas about 500 years ago and today is a widely distributed invasive species and the primary vector for viruses causing dengue, chikungunya, Zika, and yellow fever. Here we test the hypothesis that the North American colonization by <em>Ae. aegypti</em> occurred via a series of founder events. We present findings on genetic diversity, structure, and demographic history using data from 70 <em>Ae. aegypti</em> populations in North America genotyped at 12 microsatellite loci and/or ~20,000 single nucleotide polymorphisms (SNPs), the largest genetic study of the region to date. We find evidence consistent with a colonization driven by serial founder effect (SFE), with Florida as the putative source for a series of westward invasions. This scenario was supported by 1) a decrease in the genetic diversity of <em>Ae. aegypti </em>populations moving west, 2) a correlation between pairwise genetic and geographic distances, and 3) demographic analysis based on allele frequencies. A few <em>Ae. aegypti</em> populations on the west coast do not follow the general trend, likely due to a recent and distinct invasion history. We argue that SFE provides a helpful albeit simplified model for the movement of <em>Ae. aegypti </em>across North America, with outlier populations warranting further investigation.</p>

opencc-zeroMay 2022View details →
zenodo36/100

Dartseq SNP data of Uganda sorghum germplasm

<p>SNP data generated&nbsp;DArTseq for&nbsp;Ugandan S. bicolor germplasm accessions (UG set) from the Plant Genetic Resources Centre at the Uganda National Genebank.</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Documentation and code for reproducing analyses presented in: vcferr: Development, Validation, and Application of a SNP Genotyping Error Simulation Framework

<p>Documentation and code for reproducing analyses presented in: vcferr: Development, Validation, and Application of a SNP Genotyping Error Simulation Framework. Please see the <strong>README.pdf</strong> for step-by-step instructions for reproducing the entire analysis described in the paper.</p>

opencc-by-4.0Jun 2022View details →
dryad36/100

PAM-altering SNP-based allele-specific CRISPR-Cas9 therapeutic strategies for Huntington's disease

<p>Huntington's disease (HD) is caused by an expanded CAG repeat in huntingtin (<em>HTT</em>). Since HD is dominant, and loss of <em>HTT </em>leads to neurological abnormalities, safe therapeutic strategies require selective inactivation of mutant <em>HTT</em>. Previously, we proposed a concept of CRISPR-Cas9 using mutant-specific PAM sites generated by SNPs to selectively inactivate mutant <em>HTT</em>. Aiming at revealing suitable targets for clinical development, we analyzed the largest HD genotype dataset to reveal target <strong>P</strong>AM-<strong>a</strong>ltering <strong>S</strong>NPs (PAS) and subsequently evaluated their allele specificities. The gRNAs based on the PAM sites generated by rs2857935, rs16843804, and rs16843836 showed high levels of allele specificity in patient-derived cells. Simultaneous use of two gRNAs based on rs2857935-rs16843804 or rs2857935-rs16843836 produced selective genomic deletions in mutant <em>HTT </em>and prevented the transcription of mutant <em>HTT </em>mRNA without impacting the expression of normal counterpart or re-integration of the excised fragment elsewhere in the genome. RNAseq and off-target analysis confirmed high levels of allele specificity and the lack of recurrent off-targeting. Approximately 60% of HD subjects are eligible for mutant-specific CRISPR-Cas9 strategies of targeting one of these 3 PAS in conjunction with one non-allele-specific site, supporting high applicability of PAS-based allele-specific CRISPR approaches in the HD patient population.</p>

opencc-zeroAug 2022View details →
dryad36/100

Comparing mixed models and Random Forest association tests using naturalGWAS and a Striped Bass SNP dataset

<p>In this study, we used the phenotype simulation package naturalGWAS to test the performance of Zhao's Random Forest method in comparison to an uncorrected Random Forest test, latent factor mixed models (LFMM), genome-wide efficient mixed models (GEMMA), and confounder adjusted linear regression (CATE). We created 400 sets of phenotypes, corresponding to five effect sizes and 2, 5, 15, or 30 causal loci, simulated from two empirical datasets containing SNPs from Striped Bass representing three and 13 populations. All association methods were evaluated for their ability to detect genotype-phenotype associations based on power, false discovery rates, and number of false positives. Genomic inflation was highest for uncorrected Random Forest and LFMM tests and lowest for Gemma and Zhao's Random Forest. All association tests had similar power to detect causal loci, and Zhao's Random Forest had the lowest false discovery rate in all scenarios. To measure the performance of association tests in small datasets with few loci surrounding a causal gene we also ran analyses again after removing causal loci from each dataset. All association tests were only able to find true positives, defined as loci located within 30k bp of a causal locus, in 3%–18% of simulations. In contrast, at least one false positive was found in 17%–44% of simulations. Zhao's Random Forest again identified the fewest false positives of all association tests studied. The ability to test the power of association tests for individual empirical datasets can be an extremely useful first step when designing a GWAS study.</p>

opencc-zeroAug 2022View details →
zenodo36/100

Formatted Public GWAS Summary Statistics for 16 Traits and LD pruned SNP sets

<p>This data set includes 16 files with formatted GWAS summary statistic and a csv file gwas_info.csv. The csv provides the original download link and publication for each study. The data in this repository were created by downloading raw summary statistics for each study and processing them using Joe Marcus&#39; GWAS pipeline (https://github.com/jhmarcus/gwass). The resulting data set have consistent allele orientation and column headers making them convenient for analysis. We use them in an MR analysis of pairs of GWAS traits described in Section 2.3 of&nbsp; Morrison et al (2019) (https://www.biorxiv.org/content/10.1101/682237v3) and here&nbsp;https://jean997.github.io/cause/gwas_pairs.html.</p> <p>New Sep 2022: I have added LD pruned SNP sets for each pair of traits.For each exposure/outcome pair, the set of snps in</p> <p>snps_&lt;exposure&gt;__&lt;outcome&gt;.txt</p> <p>were generated by LD pruning using LD estimated using LD Shrink (available https://zenodo.org/record/1464357/) at a threshold of r^2 &lt; 0.1. LD pruning was performed using the ld_prune function in the cause R package (github.com/jean997/cause). These are the SNP sets used in the analysis in the paper</p>

opencc-by-4.0Jul 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record