Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,785
datasets available to search
ShareScore release 0.9.0
Dataset results
2,785 results for “Genotype”
Soybean Root Phenotype and Genotype Data from the Piney Purdue Agricultural Center (PPAC), Indiana
<p>This data repository contains records of root phenotypes collected in the Pinney Purdue Agricultural Center (PPAC) (Wanatah, Indiana, USA) on 24 soybean genotypes in 2022 along with their genotype information from the intersection of the BARCSoySNP6K and SoySNP50K assays.</p> <p>The repository contains the following files:<br>Bogati_soybean_root_phenotype_data1.xlsx</p> <p>6k_and_50k_geno.map</p> <p>6k_and_50k_geno.ped</p> <p>The map file contains chromosome number, SNP ID, Genetic Distance, and Base pair position.</p> <p>The ped file contains sample name (first two columns) with genotype data corresponding to the .map file beginning in column 7.</p> <p> </p> <p>Acknowledgements: To-Chia Ting, Luis Vargas, and Sajad Jamshidi assisted in the collection of root phenotype data. Chance Clark helped extract DNA for genotyping.</p> <p> </p>
HLA and KIR allele genotyping for HPRC-frz2 haplotype assemblies
<p>HLA/KIR annotation for available long reads assemblies, constructed by the pipeline : https://github.com/YingZhou001/Immuannot</p> <p>IPD-KIR version: V2.13.0, IPD-IMGT/HLA version: V3.59.0</p> <p>472 haploid assemblies included</p>
Clostridioides difficile in Honduras: a genomic and phenotypic characterization of the persistent RT027 and emergent RT002 genotypes
<p>Supplementary dataset to the manuscript: Clostridioides difficile in Honduras: a genomic and phenotypic characterization of the persistent RT027 and emergent RT002 genotypes, by Mauricio Andino-Molina, Mostafa Abdel-Glil, Fanny Hidalgo-Villeda, Edgardo Tzoc, Gernot Schmoock, Mathias W. Pletz, Heinrich Neubauer & Christian Seyboldt. </p> <p>Dataset analysed with clostyper (https://gitlab.com/FLI_Bioinfo/clostyper)</p>
Genetic architecture of alcohol consumption identified by a genotype-stratified GWAS, and impact on esophageal cancer risk in Japanese
<p><span>An East Asian-specific variant on <em>aldehyde</em> <em>dehydrogenase</em> <em>2</em> (<em>ALDH2</em> rs671, G>A) is the major genetic determinant of alcohol consumption. We performed an rs671 genotype-stratified genome-wide association study (GWAS) meta-analysis in up to 175,672 Japanese individuals to uncover additional loci associated with alcohol consumption in an rs671-dependent manner. Three loci (<em>GCKR</em>, <em>KLB</em>, and <em>ADH1B</em>)</span> <span>satisfied the genome-wide significance threshold in wild-type homozygotes (GG), whereas six loci (<em>GCKR</em>, <em>ADH1B</em>, <em>ALDH1B1</em>, <em>ALDH1A1</em>, <em>ALDH2</em>, and <em>GOT2</em>) did so in heterozygotes (GA). Of these, five loci showed genome-wide significant interaction with rs671. Genetic correlation analyses revealed ancestry-specific genetic architecture in heterozygotes. Subsequent polygenic risk scoring depicted interactions highlighted by stratified GWAS. Further, most discovered loci showed significant effects on risk of esophageal cancer, a representative alcohol-related disease, and multiple other phenotypes. Our results identify the genotype-specific genetic architecture of alcohol consumption and reveal its potential impact on alcohol-related disease risk.</span></p>
Topographic barriers drive the pronounced genetic subdivision of a range-limited fossorial rodent - nuclear genomic genotypes
<p>Genotype likelihood (GMR_minind40_mmaf0.05_glf2.beagle.gz) and pseudohaploid calls (GMR_minind40_mmaf0.05_glf2.haplo.gz) used for the nuclear genomic analyses in the manuscript "Topographic barriers drive the pronounced genetic subdivision of a range-limited fossorial rodent".</p><p>The order of the individuals in each file is listed in Names.txt</p><p> </p>
Genotypes and geographic positions of 5797 European white oaks from 636 locations genotyped at 355 nuclear SNPs and 28 maternally inherited SNPs of the chloroplast and mitochondria
<p class="MsoNormal"><span>The data set is the result of genetic inventory on 5797 white oaks collected at 636 locations all over Europe. The oaks trees were assigned in forest inventories as <em>Quercus robur</em> </span><em>L.</em> <span>(3342), <em>Quercus petraea </em></span><em>Matt</em>. <span>(2090), <em>Quercus pubescens </em></span><em>Willd</em>. <span>(170) or as unspecified <em>Quercus</em>. spp. (195). The sampling had a focus on central and east Europe as well as the Black Sea and Caucasus region. All individuals were genotyped at 355 nuclear SNPs and 28 maternally inherited SNPs of the chloroplast and mitochondria. The combination of the maternally inherited SNPs resulted in 26 different haplotypes. </span></p> <p class="MsoNormal"><span>The genotype of each individual is one row in the csv-file "genotypes". The genotypes at the nuclear markers are diploid and represented by two columns per gene marker. The genetic information at the organelle genome is haploid. For each of these gene markers one column is used. Genotypes are coded by Arabic numbers. The meaning of the numbers is explained in the table "coding genotypes" in a second csv-file. Each Individual has a unique "Genotype_ID" and a "Thuenen_Sample_ID". The "Thuenen_Sample_ID" is a unique ID that serves to identify the sample in our depository at the Thuenen Institute of Forest Genetics. Each individual has data on the geographic origin given as "Longitude" and "Latitude" in decimal degrees. For each individual the putative oak species ("Putative species") as it has been assigned in the forest inventories is given. The numbers of the "Haplotype" represent the multilocus combination of the mitochondrial and chloroplast SNPs of that individual.</span></p>
Mingrelian SNP Genotype Data
<p>This dataset contains data from 645,337 single nucleotide polymorphisms (SNPs) that were genotyped on GenoChip 2+ microarrays. The SNP data were ascertained from the mtDNA, Y-chromosome and autosomes for each individual, depending on their biological sex. In total, 5,205 mtDNA and 10,272 Y-chromosome SNPs were extracted from the array data. These data files have been uploaded as .csv files and also be uploaded as plink-formatted files. Details about the analysis of the SNP data can be found in the associated manuscript:</p><p>Theodore G Schurr, Ramaz Shengelia, Michel Shamoon-Pour, David Chitanava, Shorena Laliashvili, Irma Laliashvili, Redate Kibret, Yanu Kume-Kangkolo, Irakli Akhvlediani, Lia Bitadze, Iain Mathieson, Aram Yardumian, Genetic Analysis of Mingrelians Reveals Long-Term Continuity of Populations in Western Georgia (Caucasus), <i>Genome Biology and Evolution</i>, 2023; evad198, <a href="https://doi.org/10.1093/gbe/evad198">https://doi.org/10.1093/gbe/evad198</a></p>
Data for: A diverse parasite pool can improve effectiveness of biological control constrained by genotype-by-genotype interactions
<p>The outcomes of biological control programs can be highly variable, with natural enemies often failing to establish or spread in pest populations. This variability has posed a major obstacle in use of the bacterial parasite <em>Pasteuria</em> <em>penetrans</em> for biological control of <em>Meloidogyne</em> species, economically devastating plant-parasitic nematodes for which there are limited management options. A leading hypothesis for this variability in control is that infection is successful only for specific combinations of bacterial and nematode genotypes. Under this hypothesis, failure of biological control results from the use of <em>P</em>. <em>penetrans</em> genotypes that cannot infect local <em>Meloidogyne</em> genotypes. We tested this hypothesis using isofemale lines of <em>M</em>. <em>arenaria</em> derived from a single field population and multiple sources of <em>P</em>. <em>penetrans</em> from the same and nearby fields. In strong support of the hypothesis, susceptibility to infection depended on the specific combination of host line and parasite source, with lines of <em>M</em>. <em>arenaria</em> varying substantially in which <em>P</em>. <em>penetrans</em> source could infect them. In light of this result, we tested whether using a diverse pool of <em>P</em>. <em>penetrans</em> could increase infection and thereby control. We found that increasing the diversity of the <em>P</em>. <em>penetrans</em> inoculum from one to eight sources more than doubled the fraction of <em>M</em>. <em>arenaria</em> individuals susceptible to infection and reduced variation in susceptibility across host lines. Together, our results highlight genotype-by-genotype specificity as an important cause of variation in biological control and call for the maintenance of genetic diversity in natural enemy populations.</p>
SNP genotypes from Magallanes
<p>Hybrid zones among mussel species have been extensively studied in the northern hemisphere. In South America, it has only recently become possible to study the natural hybrid zones, due to the clarification of the taxonomy of native mussels of the <em>Mytilus</em> genus. Analyzing 54 SNP markers, we show the genetic species composition and admixture in the hybrid zone between <em>M. chilensis </em>and <em>M. platensis</em> in the southern end of South America. Bayesian, non-Bayesian clustering and re-assignment algorithms showed that the natural hybrid zone between <em>M. chilensis </em>and <em>M. platensis </em>in the Strait of Magellan, Isla Grande de Tierra del Fuego, and the Falkland Islands shows complex architecture. It can be divided into three different areas: the first one is on the Atlantic coast where only pure <em>M. platensis</em> and hybrid were found. In the second one, inside the Strait of Magellan, pure individuals of both species and mussels with variable degrees of hybridization coexist. In the last area at the Strait in front of Punta Arenas City, fjords on the Isla Grande de Tierra del Fuego, and at the Beagle Channel, only <em>M. chilensis</em> and a low number of hybrids were found. According to the proportion of hybrids, bays with protected conditions away from strong currents would give better conditions for hybridization. We do not find evidence of any other mussel species such as <em>M. edulis, M. galloprovincialis, M. planulatus, </em>or <em>M. trossulus </em>in the zone</p>
Template-specific optimization of NGS genotyping pipelines reveals allele-specific variation in MHC gene expression
<p>Using high-throughput sequencing for precise genotyping of multi-locus gene families, such as the Major Histocompatibility Complex (MHC), remains challenging, due to the complexity of the data and difficulties in distinguishing genuine from erroneous variants. Several dedicated genotyping pipelines for data from high-throughput sequencing, such as next-generation sequencing (NGS), have been developed to tackle the ensuing risk of artificially inflated diversity. Here, we thoroughly assess three such multi-locus genotyping pipelines for NGS data, the DOC method, AmpliSAS and ACACIA, using MHC class IIβ datasets of three-spined stickleback gDNA, cDNA, and "artificial" plasmid samples with known allelic diversity. We show that genotyping of gDNA and plasmid samples at optimal pipeline parameters was highly accurate and reproducible across methods. However, for cDNA data, gDNA-optimal parameter configuration yielded decreased overall genotyping precision and consistency between pipelines. Further adjustments of key clustering parameters were required tο account for higher error rates and larger variation in sequencing depth per allele, highlighting the importance of template-specific pipeline optimization for reliable genotyping of multi-locus gene families. Through accurate paired gDNA-cDNA typing and MHC-II haplotype inference, we show that MHC-II allele-specific expression levels correlate negatively with allele number across haplotypes. Lastly, sibship-assisted cDNA-typing of MHC-I revealed novel variants linked in haplotype blocks and a higher-than-previously-reported individual MHC-I allelic diversity. In conclusion, we provide novel genotyping protocols for the three-spined stickleback MHC-I and -II genes and evaluate the performance of popular NGS-genotyping pipelines. We also show that fine-tuned genotyping of paired gDNA-cDNA samples facilitates amplification bias-corrected MHC allele expression analysis.</p>
A public mid-density genotyping platform for alfalfa (Medicago sativa L.)
<p>Small public breeding programs have many barriers to adopting technology, particularly creating, and using genetic marker panels for genomic-based decisions in selection. Here we report the creation of a DArTag panel of 3,000 loci distributed across the alfalfa genome for use in molecular breeding and genomic prediction. The creation of this marker panel brings cost-effective and rapid genotyping capabilities to public breeding programs. The open access provided by this platform will allow genetic data sets generated on the marker panel to be compared and joined across projects, institutions, and countries. This genotyping resource has the power to bring genotyping equity to breeders in alfalfa. This is the first installment of a series of papers on creating affordable public genotyping resources for underserved agricultural plant and animal species.</p>
Data from: Evaluating genotyping-in-thousands by sequencing as a genetic monitoring tool for a climate sentinel mammal using non-invasive and archival samples
<p>Genetic tools for wildlife monitoring can provide valuable information on spatiotemporal population trends and connectivity, particularly in systems experiencing rapid environmental change. Though many DNA sequencing approaches still require high quality and quantity of DNA obtained from traditional sources (e.g. blood and tissue), rapid genotyping tools such as Genotyping-in-Thousands by sequencing (GT-seq) have improved our ability to make use of degraded and less concentrated DNA commonly obtained from non-invasive and archival samples. Here, we developed a multi-purpose GT-seq panel (307 single nucleotide polymorphisms) for a climate sentinel mammal (the American pika, <em>Ochotona princeps</em>) for use as a genetic tool for monitoring populations in the Canadian Rocky Mountains. We optimized the panel using contemporary tissue samples (n = 77) and subsequently applied it to archival tissue (n = 17) and contemporary fecal pellet samples (n = 129) to evaluate its effectiveness at identifying individuals and sex, estimating relatedness, and inferring population structure. The panel demonstrated high efficacy with contemporary and archival tissue samples (94.7% and 90.5% genotyping success, respectively) and negligible genotyping error (0.001% and 0.0%, respectively). Despite relatively high genotyping success for fecal pellet samples (79.7%), high genotyping error (28.4%) limited its power as a monitoring tool to assess genetic variation using non-invasive samples and highlighted the need for further optimization around sample and data collection.</p>
Original genotype data of 159 wheat samples
<p>This dataset contains all 55K SNP original genotype data from 159 wheat samples used in the study, including SNP site IDs, chromosomes and positions, allele information, and genotype information for each material at each SNP site.</p>
First insights into population structure and genetic diversity versus host specificity in trypanorhynch tapeworms using multiplexed shotgun genotyping
<p>Theory predicts relaxed host specificity and high host vagility should contribute to reduced genetic structure in parasites while strict host specificity and low host vagility should increase genetic structure. Though these predictions are intuitive, they have never been explicitly tested in a population genomic framework. Trypanorhynch tapeworms, which parasitize sharks and rays (elasmobranchs) as definitive hosts, are the only order of elasmobranch tapeworms that exhibit considerable variability in their definitive host specificity. This allows for unique combinations of host use and geographic range, making trypanorhynchs ideal candidates for studying how these traits influence population-level structure and genetic diversity. Multiplexed shotgun genotyping (MSG) datasets were generated to characterize component population structure and infrapopulation diversity for a representative of each trypanorhynch suborder: the ray-hosted <em>Rhinoptericola megacantha</em> (Trypanobatoida) and the shark-hosted Callitetrarhynchus gracilis (Trypanoselachoida). Adults of <em>R. megacantha</em> are more host-specific and less broadly distributed than adults of <em>C. gracilis</em>, allowing correlation between these factors and genetic structure. Replicate tapeworm specimens were sequenced from the same host individual, from multiple conspecific hosts within and across geographic regions, and from multiple definitive host species. For <em>R. megacantha</em>, population structure coincided with geography rather than host species. For <em>C. gracilis</em>, limited population structure was found, suggesting a potential link between degree of host specificity and structure. Conspecific trypanorhynchs from the same host individual were found to be as, or more, genetically divergent from one another as from conspecifics from different host individuals. For both species, high levels of homozygosity and positive FIS values were documented.</p>
Supplementary data: Effect of genotype by environment interaction (GEI) analysis for potato tuber yield and their quality traits in organic multi-environment domains of Poland
<p>Climate and raw data supplementary to the related publication in the journal Agriculture (ISSN 2077-0472).</p>
Genotype data of 1970 Pedunculate oak trees (Quercus robur L.) in 13 European countries at 381 gene loci covering the nuclear and organelle genome
<p>The data set is the result of genetic inventory on 1970 Pedunculate oak trees from 197 locations in Europe. The samples are from the countries: Belarus, Bosnia, Bulgaria, Croatia, Finland, France, Germany, Hungary, Italy, Latvia, Poland, Russia and Ukraine. At each location ten individual trees were collected. The data set includes the location ID and geographic coordinates of each sampled tree (longitude and latitude in decimal degrees) and the genotype data. All samples were screened with a targeted sequencing approach on a set of 381 polymorphic loci (356 nuclear SNPs, 3 nuclear InDels, 17 chloroplast SNPs and five mitochondrial SNPs).</p> <p>The genotype of each individual is one row in the csv-file "genotypes". The genotypes at the nuclear markers are diploid and represented by two columns per gene marker. The genetic information at the organelle genome is haploid. For each of these gene markers one column is used. Genotypes are coded by Arabic numbers. The meaning of the numbers is explained in the table "coding genotypes" in a second csv-file.</p>
Title of Dataset: NMR metabolomic analysis of Drosophila head extracts: 2 genotypes (control, paraKO)
<p>Characterization of <em>para<sup>ko</sup></em> model in <em>Drosophila melanogaster</em> showed homeostasis disturbances as heat-induced phenotype, neuromuscular and cognitive alterations. Moreover, preliminary results during starvation assay revealed possibly differences in metabolism. To assess that NMR spectroscopy was made, comparing heads from young control and mutant flies.</p>
Data from: Host-parasite dynamics shaped by temperature and genotype: quantifying the role of underlying vital rates
<p>1. Global warming challenges the persistence of local populations, not only through heat-induced stress, but also through indirect biotic changes. We study the interactive effects of temperature, competition and parasitism in the water flea <i>Daphnia magna</i>.</p> <p>2. We carried out a common garden experiment monitoring the dynamics of <i>Daphnia</i> populations along a temperature gradient. Halfway through the experiment, all populations became infected with the ectoparasite <i>Amoebidium parasiticum</i>, enabling us to study interactive effects of temperature and parasite dynamics. We combined Integral Projection Models with epidemiological models, parameterized using the experimental data on the performance of individuals within dynamic populations. This enabled us to quantify the contribution of different vital rates and epidemiological parameters to population fitness across temperatures and <i>Daphnia</i> clones originating from two latitudes.</p> <p>3. Interactions between temperature and parasitism shaped competition, where Belgian clones performed better under infection than Norwegian clones, mainly due to higher survival. Infected <i>Daphnia</i> populations performed better at higher than at lower temperatures, mainly due to an increased host capability of reducing parasite loads. Temperature strongly affected individual vital rates, but effects largely cancelled out on a population-level. In contrast, parasitism strongly reduced fitness through consistent negative effects on all vital rates. As a result, temperature-mediated parasitism was more important than the direct effects of temperature in shaping population dynamics. Both the outcome of the competition treatments and the observed extinction patterns support our modeling results.</p> <p>4. Our study highlights that shifts in biotic interactions can be equally or more important for responses to warming than direct physiological effects of warming, emphasizing that we need to include such interactions in our studies to predict the competitive ability of natural populations experiencing global warming.</p>
Supplementary Table 1. Raw data of egg quality parameters for 990 egg samples with ATOL (Animal Trait Ontology for Livestock) descriptors, as function of hen age, pen no. and genotype in 15 replicates.
<p>Data table of egg quality parameters</p>
Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities: validation cohort meta data and parsed TCR repertoire data
<p>Meta data corresponding the the validation cohort for the paper, "Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities" by Magdalena L Russell, Aisha Souquette, David M Levine, Stefan A Schattgen, E Kaitlynn Allen, Guillermina Kuan, Noah Simon, Angel Balmaseda, Aubree Gordon, Paul G Thomas, Frederick A Matsen IV, and Philip Bradley. These meta data include: </p> <p>(1) SNP genotypes for the two SNPs which overlap with the discovery cohort<br> - (nicaragua_snp_genotypes_ints.tsv) -- SNP genotypes as integers<br> - (nicaragua_snp_genotypes_strings.tsv) -- SNP genotypes as allele strings <br> (2) the ancestry PCs for each individual in the validation cohort (nicaragua_snp_ancestry_PCA.tsv)<br> (3) a file including IMGT genes and sequences used for parsing TCRB repertoire data (human_vj_allele_cdr3_nucseqs.tsv)<br> (4) a file including IMGT genes used for parsing TCRA repertoire data (human_vj_alleles_alpha.tsv)<br> (5) Parsed TCRA repertoire data (nicaragua_parsed_TCRA.tgz)<br> (6) Parsed TCRB repertoire data (nicaragua_parsed_TCRB.tgz) </p> <p><strong>Corresponding raw validation cohort TCR repertoire data is available here:</strong> https://www. ncbi.nlm.nih.gov/bioproject/PRJNA762269 (The BioProject database, accession number: PRJNA762269)</p> <p><strong>Software tools designed to work with these data are available here:</strong> https://github.com/phbradley/tcr-gwas</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.