Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

123

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

123 results for “SNP genotyping”

Learn how ShareScore rates datasets ↗
zenodo44/100

Genotyping of the Chinese Spring x Renan mapping population with the TaBW280K SNP array

<p>The TaBW280K SNP array (Rimbert et al., PLoS ONE 2018) was used to genotype 430 Single Seed Descent (SSD) individuals<br> derived from a cross between Chinese Spring and Renan (CsRe; Choulet et al., Science 2014). Out of the 280,226 probesets, 85,276 were found to be polymorphic between the two parental lines and PHR on the population. Eventually, 83,721 (98.2%) SNPs were genetically mapped in 21 linkage groups corresponding to the 21 chromosomes of bread wheat, with no unlinked markers. This file contains the genotyping data of the 430 SSD lines.<br> &nbsp;</p>

opencc-by-4.0Jan 2018View details →
zenodo44/100

SNP and indel discovery and genotyping in next-generation sequencing data

<p>Code, logs and data for discovery and genotyping of SNPs and indels, in the the D.melanogaster genome, using GATK HaplotypeCaller. Code is in the zipped folder named code.zip. Run logs for this code as in the zipped folder named logs.zip. The unfiltered vcf genotypes file is named lhm_rg_HC_2015-09-15.vcf.gz. The filtered vcf genotypes file is named f1.lhm_rg_HC_raw.vcf.gz. The vcf submitted to NCBI dbSNP (filtered, and with indels &gt;50bp and variants with null alternate alleles both removed) is named dbSNP.lhm_rg_HC_raw.vcf.gz. The folder local_reference.zip contains the reference assembly files against which genotypes were called against, and includes the code used to format the data prior to use. Also included is genotypes data from the two in-house reference line samples sequenced (BDGP6+ISO1 mito/dm6, Bloomington <em>Drosophila</em> Stock Center no. 2057)</p> <p>Samples are 220 Sussex-LH<sub>M</sub> hemiclones, and 2 RG. The first run did not include chromosome 4 and the mitochondrial genome, so these were genotyped separately, and then added to the rest of the results.</p> <p>The link for the NCBI dbSNP record is currently https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461and the submitter handle is MORROW_EBE_SUSSEX.</p> <p>At the time of writting, the NCBI D.melanogaster build is still being updated, and therefore ss identifiers, but not rs identifers are available.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>

opencc-by-4.0Oct 2016View details →
zenodo40/100

High density SNP genotypes (Infinium Human CytoSNP-850K v1.2 BeadChip) of ERAP2-WT and ERAP2-KO Birdshot LCL

<p>SNP genotype data was&nbsp;performed on DNA isolated from WT and CRISPR-Cas9 edited LCLs (ERAP2-KO) according to standard procedures using the Infinium Human CytoSNP-850K v1.2 BeadChip (Illumina, San Diego, CA, USA). SNP-array results and data analysis were carried out using NxClinical software v5.1 (BioDiscovery, Los Angeles, CA, USA). Human genome build Feb. 2009 GRCh37/hg19 was used.&nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Development of a high-density 665 K SNP array for rainbow trout genome-wide genotyping. Supplemental VCF file

<p>Single nucleotide polymorphism (SNP) arrays, also named &laquo; SNP chips &raquo;, enable very large numbers of individuals to be genotyped at a targeted set of thousands of genome-wide identified markers. We used preexisting variant datasets from USDA, a French commercial line and 30X-coverage whole genome sequencing of INRAE isogenic lines to develop an Affymetrix 665 K SNP array (HD chip) for rainbow trout. In total, we identified 32,372,492 SNPs that were polymorphic in the USDA or INRAE databases. A subset of identified SNPs were selected for inclusion on the chip, prioritizing SNPs whose flanking sequence uniquely aligned to the Swanson reference genome, with homogenous repartition over the genome and the highest Minimum Allele Frequency in both USDA and French databases. Of the 664,531 SNPs which passed the Affymetrix quality filters and were manufactured on the HD chip, 65.3% and 60.9% passed filtering metrics and were polymorphic in two other distinct French commercial populations in which, respectively, 288 and 175 sampled fish were genotyped. Only 576,118 SNPs mapped uniquely on both Swanson and Arlee reference genomes, and 12,071 SNPs did not map at all on the Arlee reference genome. Among those 576,118 SNPs, 38,948 SNPs were kept from the&nbsp; commercially available medium-density 57K SNP chip. We demonstrate the utility of the HD chip by describing the high rates of&nbsp; linkage disequilibrium at 2 kb to 10 kb in the rainbow trout genome in comparison to the linkage disequilibrium observed at 50 kb to&nbsp; 100 kb which are usual distances between markers of the medium-density chip.</p> <p>&nbsp;</p> <p>File submitted correspond to the supplementary data 1 of the publication (under submission) : INRAE_USDA_MAF1.vcf.gz</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

SNP genotypes for 524 wild mice and selected laboratory strains

<p>SNP genotypes from the Mouse Universal Genotyping Array for 524 wild mice and 12 selected laboratory strains. &nbsp;Data are provided in PLINK binary format (*.bed/*.bim/*.fam files) with an accompanying sample manifest (comma-separated text.)</p>

opencc-by-4.0May 2017View details →
zenodo40/100

Genome-wide SNP discovery in native American and Hungarian Robinia pseudoacacia genotypes using next-generation double-digest restriction-site-associated DNA sequencing (ddRAD-Seq)

<p>Initial filtered ddRADseq dataset with highly variable SNP markers from native American and Hungarian <em>Robinia pseudoacacia</em> L. individuals</p>

opencc-by-4.0Nov 2022View details →
dryad36/100

Population structure of five native sheep breeds of Sweden estimated with high density SNP genotypes

Background <p>Native Swedish sheep breeds are part of the North European short-tailed sheep group; characterized in part by their genetic uniqueness. Our objective was to study the population structure of native Swedish sheep. Five breeds were genotyped using the 600 K SNP array. Dalapäls and Klövsjö sheep are from the middle of Sweden; Gotland and Gute sheep from Gotland, an island in the Baltic Sea; and Fjällnäs sheep from northern Sweden. We studied population structure by: principal component analysis (PCA), cluster-based analysis of admixture, and an estimated population tree.</p> Results <p>The analyses of the five Swedish breeds revealed that these breeds are five distinct breeds, while Gute and Gotland are more closely related to each other as seen in all analyses. All breeds had long branch lengths in the population tree indicating they've been subjected to drift. We repeated our analyses using 39 K SNP and including 50 K SNP genotypes from other European and southwestern Asian breeds from the Sheep HapMap project and 600 K SNP genotypes from a dataset of French sheep. Results arranged breeds into five groups: south-west Asia, south-west Europe, central Europe, north Europe and north European short-tailed sheep. Within this last group, Norwegian and Icelandic breeds, Finn and Romanov sheep, Scottish breeds, and Gute and Gotland sheep were more closely related while the remaining Swedish breeds and Ouessant sheep were distinct from all breeds and had longer branches in the population tree.</p> Conclusions <p>We showed population structure of five Swedish breeds and their structure within European and southwestern Asian breeds. Swedish breeds are unique, distinct breeds that have been subjected to drift but group with other north European short-tailed sheep.</p>

opencc-zeroDec 2019View details →
dryad36/100

Data from: Phylogenetic relationships, breeding implications, and cultivation history of Hawaiian taro (Colocasia esculenta) through genome-wide SNP genotyping

Taro, Colocasia esculenta, is one of the world's oldest root crops and of particular economic and cultural significance in Hawai'i, where historically more than 150 different landraces were grown. We developed a genome-wide set of more than 2400 high-quality single nucleotide polymorphism (SNP) markers from 70 taro accessions of Hawaiian, South Pacific, Palauan, and mainland Asian origins, with several objectives: (a) uncover the phylogenetic relationships between Hawaiian and other Pacific landraces, (b) shed light on the history of taro cultivation in Hawai'i, and (c) develop a tool to discriminate among Hawaiian and other taros. We found that almost all existing Hawaiian landraces fall into five monophyletic groups that are largely consistent with the traditional Hawaiian classification based on morphological characters, e.g., leaf shape and petiole color. Genetic diversity was low within these clades but considerably higher between them. Population structure analyses further indicated that the diversification of taro in Hawai'i most likely occurred by a combination of frequent somatic mutation and occasional hybridization. Unexpectedly, the South Pacific accessions were found nested within the clades mainly composed of Hawaiian accessions, rather than paraphyletic to them. This suggests that the origin of clades identified here preceded the colonization of Hawai'i, and that early Polynesian settlers brought taro landraces from different clades with them. In the absence of a sequenced genome, this marker set provides a valuable resource towards obtaining a genetic linkage map, and to study the genetic basis of phenotypic traits of interest to taro breeding such as disease resistance.

opencc-zeroDec 2016View details →
zenodo36/100

Mingrelian SNP Genotype Data

<p>This dataset contains data from 645,337 single nucleotide polymorphisms (SNPs) that were genotyped on GenoChip 2+ microarrays. The SNP data were ascertained from the mtDNA, Y-chromosome and autosomes for each individual, depending on their biological sex. In total, 5,205 mtDNA and 10,272 Y-chromosome SNPs were extracted from the array data. These data files have been uploaded as .csv files and also be uploaded as plink-formatted files. Details about the analysis of the SNP data can be found in the associated manuscript:</p><p>Theodore G Schurr, Ramaz Shengelia, Michel Shamoon-Pour, David Chitanava, Shorena Laliashvili, Irma Laliashvili, Redate Kibret, Yanu Kume-Kangkolo, Irakli Akhvlediani, Lia Bitadze, Iain Mathieson, Aram Yardumian, Genetic Analysis of Mingrelians Reveals Long-Term Continuity of Populations in Western Georgia (Caucasus),&nbsp;<i>Genome Biology and Evolution</i>, 2023; evad198,&nbsp;<a href="https://doi.org/10.1093/gbe/evad198">https://doi.org/10.1093/gbe/evad198</a></p>

opencc-by-4.0Oct 2023View details →
dryad36/100

SNP genotypes from Magallanes

<p>Hybrid zones among mussel species have been extensively studied in the northern hemisphere. In South America, it has only recently become possible to study the natural hybrid zones, due to the clarification of the taxonomy of native mussels of the <em>Mytilus</em> genus. Analyzing 54 SNP markers, we show the genetic species composition and admixture in the hybrid zone between <em>M. chilensis </em>and <em>M. platensis</em> in the southern end of South America. Bayesian, non-Bayesian clustering and re-assignment algorithms showed that the natural hybrid zone between <em>M. chilensis </em>and <em>M. platensis </em>in the Strait of Magellan, Isla Grande de Tierra del Fuego, and the Falkland Islands shows complex architecture. It can be divided into three different areas: the first one is on the Atlantic coast where only pure <em>M. platensis</em> and hybrid were found. In the second one, inside the Strait of Magellan, pure individuals of both species and mussels with variable degrees of hybridization coexist. In the last area at the Strait in front of Punta Arenas City, fjords on the Isla Grande de Tierra del Fuego, and at the Beagle Channel, only <em>M. chilensis</em> and a low number of hybrids were found.  According to the proportion of hybrids, bays with protected conditions away from strong currents would give better conditions for hybridization. We do not find evidence of any other mussel species such as <em>M. edulis, M. galloprovincialis, M. planulatus, </em>or <em>M. trossulus </em>in the zone</p>

opencc-zeroDec 2023View details →
dryad36/100

Genomics of humic adaptation in Eurasian perch (Perca fluviatilis): SNP genotypes of 32 perch individuals, supplementary figures and tables

<p>Extreme <span>environments are inhospitable to the majority of species, but some organisms are able to survive in such hostile conditions due to evolutionary adaptations. For example, </span><span>m</span><span>odern bony fishes have colonized various aquatic environments, including perpetually dark,</span><span> hypoxic, hypersaline and toxic habitats</span><span>. </span><span>Eurasian perch (</span><em>Perca fluviatilis</em><span>) is among the few fish species of northern latitudes that is able to live in very acidic humic lakes. Such lakes represent almost "nocturnal" environments; they contain high levels of dissolved organic matter, which in addition to creating a challenging visual environment, also affects a large number of other habitat parameters and biotic interactions. To reveal the genomic targets of humic-associated selection, we performed whole-genome sequencing of perch originating from 16 humic and 16 clear-water lakes in northern Europe. We identified over 800,000 SNPs, of which &gt;10,000 were identified as potential candidates under selection (associated with &gt;3,000 genes) using multiple outlier approaches. Our findings suggest that adaptation to the humic environment may involve hundreds of regions scattered across the genome. Putative signals of adaptation were detected in genes and gene families with diverse functions, including organism development and ion transportation. The observed excess of variants under selection in regulatory regions highlights the importance of adaptive evolution via regulatory elements, rather than via protein sequence modification. Our study demonstrates the power of whole-genome analysis to illuminate multifaceted nature of humic adaptation and provides the foundation for further investigation of causal mutations underlying phenotypic traits of ecological and evolutionary importance.</span></p>

opencc-zeroMar 2022View details →
dryad36/100

Data from: Targeted genome-wide SNP genotyping in feral horses using non-invasive fecal swabs

<p>The development of high-throughput sequencing has prompted a transition in wildlife genetics from using microsatellites toward sets of Single Nucleotide Polymorphisms (SNPs). However, genotyping large numbers of targeted SNPs using non-invasive samples remains challenging due to relatively large DNA input requirements. Recently, target enrichment has emerged as a promising approach requiring little template DNA. We assessed the efficacy of Tecan Genomics' Allegro Targeted Genotyping (ATG) for generating genome-wide SNP data in feral horses using DNA isolated from fecal swabs. Total and host-specific DNA were quantified for 989 samples collected as part of a long-term individual-based study of feral horses on Sable Island, Nova Scotia, Canada, using dsDNA fluorescence and a host-specific qPCR assay, respectively. Forty-eight samples representing 44 individuals containing at least 10ng of host DNA (ATG's recommended minimum input) were genotyped using a custom multiplex panel targeting 279 SNPs. Genotyping accuracy and consistency were assessed by contrasting ATG genotypes with those obtained from the same individuals with SNP microarrays, and from multiple samples from the same horse, respectively. 62% of swabs yielded the minimum recommended amount of host DNA for ATG. Ignoring samples that failed to amplify, ATG recovered an average of 86.7% targeted sites per sample, while genotype concordance between ATG and SNP microarrays was 98.5%. The repeatability of genotypes from the same individual approached unity with an average of 99.9%. This study demonstrates the suitability of ATG for genome-wide, non-invasive targeted SNP genotyping, and will facilitate further ecological and conservation genetics research in equids and related species.</p>

opencc-zeroApr 2022View details →
zenodo36/100

Documentation and code for reproducing analyses presented in: vcferr: Development, Validation, and Application of a SNP Genotyping Error Simulation Framework

<p>Documentation and code for reproducing analyses presented in: vcferr: Development, Validation, and Application of a SNP Genotyping Error Simulation Framework. Please see the <strong>README.pdf</strong> for step-by-step instructions for reproducing the entire analysis described in the paper.</p>

opencc-by-4.0Jun 2022View details →
dryad36/100

Genotypes of Aedes aegypti mosquitoes derived from SNP chip and low-coverage whole genome sequencing for platform cross-validation

<p>The mosquito <em>Aedes aegypti </em>is the primary vector of many human arboviruses such as dengue, yellow fever, chikungunya, and Zika, which affect millions of people world-wide. Population genetics studies on this mosquito have been important in understanding its invasion pathways and success as a vector of human disease. The Axiom aegypti1 SNP chip was developed from a sample of geographically diverse <em>Ae. aegypti </em>populations to facilitate genomic studies on this species. Here we evaluate the utility of the Axiom aegypti1 SNP chip for population genetics and compare it with a low-depth shot-gun sequencing approach using mosquitoes from the species' native (Africa) and invasive range (outside Africa). These analyses indicate that the results from the SNP chip are highly reproducible and have a higher sensitivity to capture alternative alleles than a low-coverage whole-genome sequencing approach. Although the SNP chip suffers from ascertainment bias, results from population structure, ancestry, demographic, and phylogenetic analyses using the SNP chip were congruent with those derived from low coverage whole genome sequencing, and consistent with previous reports on Africa and outside Africa populations using microsatellites. More importantly, we identified a subset of SNPs that can be reliably used to generate merged databases, opening the door to combined analyses. We conclude that the Axiom aegypti1 SNP chip is a convenient, more accurate, low-cost alternative to low-depth whole genome sequencing for population genetic studies of <em>Ae. aegypti</em> that do not rely on full allelic frequency spectra. Whole genome sequencing and SNP chip data can be easily merged, extending the usefulness of both approaches. </p>

opencc-zeroApr 2024View details →
dryad36/100

Autosomal SNP-genotype data of brown bears (Ursus arctos) in Finland

<p>Harmonising methodology between countries is crucial in transborder population monitoring. However, immediate application of alleged, established DNA-based methods across the extended area can entail drawbacks and may lead to biases. Therefore, genetic methods need to be tested across the whole area before being deployed. Around 4,500 brown bears (<em>Ursus arctos</em>) live in Norway, Sweden, and Finland and they are divided into the western (Scandinavian) and eastern (Karelian) population. Both populations have recovered and are connected via asymmetric migration. DNA-based population monitoring in Norway and Sweden uses the same set of genetic markers. With Finland aiming to implement monitoring, we tested the available SNP-panel developed to assess brown bears in Norway and Sweden, on tissue samples from a representative set of 93 legally harvested individuals from Finland. The aim was to test for ascertainment bias and evaluate its suitability for DNA-based transnational-monitoring covering all three countries. We compared results to the performance of microsatellite genotypes of the same individuals in Finland and against SNP-genotypes from individuals sampled in Sweden (<em>N</em>=95) and Norway (<em>N</em>=27). In Finland, a higher resolution for individual identification was obtained for SNPs (PI=1.18E-27) compared to microsatellites (PI=4.2E-11). Compared to Norway and Sweden, probability of identity of the SNP-panel was slightly higher and expected heterozygosity lower in Finland indicating ascertainment bias. Yet, our evaluation show that the available SNP-panel outperforms the microsatellite panel currently applied in Norway and Sweden. The SNP-panel represents a powerful tool that could aid improving transnational DNA-based monitoring of brown bears across these three countries.</p>

opencc-zeroMay 2024View details →
dryad36/100

SNP genotyping of indigenous goats of Uganda based on the Goat_IGGC_65K_v2 illumina chip

<p>Uganda's indigenous goats are characterised based on ethnic communities that raise them, average mature weight, and hair coat characteristics. Uganda's indigenous goats have  been genotyped based on the Goat_IGGC_65K_v2 illumina chip to study their population structure and genetic characteristics. Information generated from this data is vital for the sustainable utilisation, development, and conservation of Uganda's goat genetic resources.</p>

opencc-zeroMay 2024View details →
dryad36/100

SNP genotype dataset from brown and anadromous trout

<p>Populations of anadromous brown trout, also known as sea trout, have suffered recent marked declines in abundance due to multiple factors, including climate change and human activities. While much is known about their freshwater phase, less is known about the species' marine feeding migrations. This situation is hindering the effective management and conservation of anadromous trout in the marine environment. Using a panel of 95 single nucleotide polymorphism markers we developed a genetic baseline, which demonstrated strong regional structuring of genetic diversity in trout populations around the English Channel and adjacent waters. Extensive baseline testing showed this structuring allowed the high-confidence assignment of known-origin individuals to the region of origin. This study presents new data on the movements of anadromous trout in the English Channel and southern North Sea. Assignment of anadromous trout sampled from 12 marine and estuarine localities highlighted contrasting results for these areas. The majority of these fisheries are composed predominately of stocks local to the sampling location. However, there were multiple cases of long-distance movements of anadromous trout, with several individuals originating from rivers in northeast England being caught in the English Channel and southern North Sea, in some cases more than 1000 km from their natal region. These results have implications for the management of sea trout in inshore waters around the English Channel and southern North Sea.</p>

opencc-zeroJul 2024View details →
zenodo36/100

SNP genotype matrix for GWAS and Machine Learning analyses

<p><strong>SNP datasets used for GWAS and Machine Learning analyses</strong></p> <p>All datasets come from the easyGWAS website: <a href="https://easygwas.ethz.ch/down/1/">https://easygwas.ethz.ch/down/1/</a></p> <p>&nbsp;</p> <p><strong>=== Horton et al. 2012 ===</strong></p> <p><strong>1307 Arabidopsis genotypes x 214,057</strong> <strong>SNPs</strong></p> <p><strong>1) In the form of a genotype matrix </strong></p> <p>The file is called <a href="https://zenodo.org/api/files/d862e79f-02f2-4176-9b8e-04e48a2cf72c/horton2012.raw?versionId=2b05fb8d-b486-4f00-b7bf-bb024b892dc9">Horton2012.raw</a></p> <p><a href="https://www.nature.com/articles/ng.1042">https://www.nature.com/articles/ng.1042</a></p> <p>Preview of the first lines and columns:</p> <p>FID&nbsp;&nbsp; &nbsp;Chr1_657_T&nbsp;&nbsp; &nbsp;Chr1_3102_G&nbsp;&nbsp; &nbsp;Chr1_4648_A&nbsp;&nbsp; &nbsp;Chr1_4880_T&nbsp;&nbsp; &nbsp;Chr1_5975_G&nbsp;&nbsp; &nbsp;Chr1_6063_T&nbsp;&nbsp; &nbsp;Chr1_6449_C<br> 9381&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0<br> 9380&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;2<br> 9378&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0<br> 9371&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0<br> 9367&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0<br> 9363&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0<br> 9356&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0<br> 9355&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0<br> 9354&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;0</p> <p>...etc...</p> <p>PLINK 1.9 was used to convert the .ped and .map file to a .raw format with:&nbsp;</p> <pre><code class="language-bash">plink --file original_data/genotype --recodeA --tab</code></pre> <p>Genotypes are encoded as 0, 1 or 2 with:</p> <pre> SNP SNP_A --- ----- A A -&gt; 0 A C -&gt; 1 C C -&gt; 2 0 0 -&gt; NA </pre> <p>Then only the Family ID was kept (same as individual ID) and other columns (Paternal ID, Maternal ID, Sex, Phenotype) were removed.</p> <p>The corresponding PLINK manual page used is here: <a href="https://zzz.bwh.harvard.edu/plink/dataman.shtml#recode">https://zzz.bwh.harvard.edu/plink/dataman.shtml#recode</a></p> <p><strong>1) In the form of set of files compatible with PLINK out of the box</strong></p> <p>The archive file is called <a href="https://zenodo.org/api/files/b34fd40e-2db1-47b0-92c8-0ad51ad92d46/AtPolyDB_call_method_75_Horton2012.tar.gz">AtPolyDB_call_method_75_Horton2012.tar.gz</a> and contains three files:</p> <ul> <li>genotype.ped: pedigree information from the 1307 ecotypes</li> <li>genotype.map: the SNP positions on the genome</li> <li>phenotypes.pheno: the phenotype value of the 1307 ecotypes</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →
dryad36/100

Data from: Targeted genome-wide SNP genotyping in feral horses using non-invasive fecal swabs

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad36/100

SNP genotypes from Magallanes

Open the record for dataset details and reuse information.

publicDec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record