Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
131
datasets available to search
ShareScore release 0.7.1
Dataset results
131 results for “genetic loci”
Pathogen lifestyle determines host genetic signature of quantitative disease resistance loci in oilseed rape (Brassica napus)
<p>Supplemental datasets associated with publication: Pathogen lifestyle determines host genetic signature of quantitative disease resistance loci in oilseed rape (<em>Brassica napus</em>)</p> <p><strong>Abstract</strong></p> <ul> <li>Crops are affected by several pathogens, but these are rarely studied in parallel to identify common and unique genetic factors controlling diseases. Broad-spectrum quantitative disease resistance (QDR) is desirable for crop breeding as it confers resistance to several pathogen species.</li> <li>Here, we use associative transcriptomics (AT) to identify candidate gene loci associated with <em>Brassica napus</em> constitutive QDR to four contrasting fungal pathogens: <em>Alternaria brassicicola</em>, <em>Botrytis cinerea</em>, <em>Pyrenopeziza</em><em> brassicae</em> and <em>Verticillium longisporum. </em>We did not identify any loci associated with broad-spectrum QDR to fungal pathogens with contrasting lifestyles. Instead, we observed QDR dependent on the lifestyle of the pathogen—hemibiotrophic and necrotrophic pathogens had distinct QDR responses and associated loci, including some loci associated with early immunity. Furthermore, we identify a genomic deletion associated with resistance to <em>V. longisporum </em>and potentially broad-spectrum QDR.</li> <li>This is the first time AT has been used for several pathosystems simultaneously to identify host genetic loci involved in broad-spectrum QDR. We highlight constitutively expressed candidate loci for broad-spectrum QDR with no antagonistic effects on susceptibility to the other pathogens studies as candidates for crop breeding. In conclusion, this study represents and advancement in our understanding if broad-spectrum QDR in <em>B. napus </em>and is a significant resource for the scientific community. </li> </ul> <p><strong>Description of data files</strong></p> <p><strong>Full dataset for input into AT analysis </strong>Full datasets (infection phenotypes for <em>A. brassicicola, B. cinerea, </em>or <em>V.longisporum, </em>ROS measurements for chitin, flg22, or elf18) and link to original <em>P. brassicae </em>dataset. These datasets were used for input into the Associative Transcriptomics pipeline (Nichols, 2022, <a href="https://github.com/bsnichols/GAGA. https://zenodo.org/badge/latestdoi/512807075">https://github.com/bsnichols/GAGA. https://zenodo.org/badge/latestdoi/512807075</a>). </p> <p><strong>Table S1 </strong>Mean, normalized phenotype data for resistance to pathogens (<em>Alternaria brassicicola, Botrytis cinerea, Pyrenopeziza brassicae </em>and <em>Verticillium longisporum</em>) and ROS response induced by PAMPS (chitin, flg22, and elf18). These data were used for association transcriptomic analysis.<strong> </strong></p> <p><strong>Table S2 </strong>Full list of single nucleotide polymorphism (SNP) markers and significance levels from genome-wide association (GWA) analyses for resistance to pathogens (<em>Alternaria brassicicola, Botrytis cinerea, Pyrenopeziza brassicae </em>and <em>Verticillium longisporum</em>) and ROS response induced by PAMPS (chitin, flg22, and elf18). Each excel tab contains the analyses for a single trait. The best fit model for GWA analysis is indicated in the tab title. Manhattan plots showing marker-trait association are included for data visualization; x-axis indicates SNP location along the chromosome; the y-axis indicates the -log10(p) (P value). Qqplots are included to demonstrate model fit.</p> <p><strong>Table S3</strong> Full list of gene expression markers (GEMs) and significance levels from GEM analyses for resistance to pathogens (<em>Alternaria brassicicola, Botrytis cinerea, Pyrenopeziza brassicae and Verticillium longisporum</em>) and ROS response induced by PAMPS (chitin, flg22, and elf18). Each excel tab contains the analyses for a single trait. Manhattan plots showing marker-trait association are included for data visualization; x-axis indicates GEM location along the chromosome; the y-axis indicates the -log10(p) (P value). </p> <p><strong>Table S4 </strong>184 gene expression markers (GEMs) associated with chitin-induced ROS compared with GEMs associated with resistance to pathogens (<em>Alternaria brassicicola, Botrytis cinerea, Pyrenopeziza brassicae </em>and<em> Verticillium longisporum</em>) and ROS response induced by flg22, and elf18. Lists correspond to Venn diagrams in Fig. 2. The first tab includes all 184 GEMs associated with chitin-induced ROS. The subsequent tabs include lists of shared GEMs associated with chitin-induced ROS response and each additional trait (quantitative disease resistance (QDR) to each fungal pathogen or additional PAMP-induced ROS responses). The title of each tab indicates the data included in each comparison and the number of shared GEMs. Predicted <em>Arabidopsis thaliana</em> orthologs and corresponding descriptions are shown where possible. </p> <p><strong>Table S5</strong> Enrichment analyses to determine if the number of gene expression markers (GEMs) shared between different lists is greater than the number of GEMs that would be expected by chance (e.g., lists of quantitative disease resistance (QDR) GEMs for two fungal pathogens). The representation factor is the number of overlapping GEMs divided by the expected number of overlapping GEMs drawn from two independent groups (traits), considering the total number of GEMs sequenced (53884). A representation factor > 1 indicates more overlap than expected of two groups, a representation factor < 1 indicates less overlap than expected, and a representation factor of 1 indicates that the two groups by the number of genes expected for independent groups of genes. </p> <p><strong>Table S6 R</strong>esults from Weighted Co-expression Gene Network Analysis (WGCNA). The first tab indicates significant modules from WGCNA analysis. Black and magenta modules are associated with antagonistic effects on resistance/susceptibility to all four pathogens. The second tab includes a full list of the GEM markers (Table S3), which are in significant WGCNA modules. The third, fourth and, fifth tabs indicate all significant GEMs in the black module, GO terms associated with GEMs in the black module, and all GO terms associated with the black module, respectively. The sixth, seventh and, eighth tabs indicate all significant GEMs in the magenta module, GO terms associated with GEMs in the magenta module, and all GO terms associated with the magenta module, respectively.</p> <p><strong>Table S7 </strong>Shared gene expression markers (GEMs) associated with resistance to different pathogens (<em>Alternaria brassicicola, Botrytis cinerea, Pyrenopeziza brassicae </em>and <em>Verticillium longisporum</em>). Lists correspond to matrices and Venn diagrams in Fig. 3. The first tab includes all GEMs associated quantitative disease resistance (QDR) to the fungal pathogens. The subsequent tabs include lists of shared GEMs associated with QDR to two or more fungal pathogens. The title of each tab indicates the data included in each comparison and the number of shared GEMs. Predicted <em>Arabidopsis thaliana</em> orthologs and corresponding descriptions are shown where possible. </p> <p><strong>Table S8 </strong>List of genes in linkage disequilibrium with the top marker for <em>Verticillium longisporum</em> resistance from genome-wide association (GWA) analysis on chromosome A09 (107 genes)(Tab 1) and the homoeologous region on C08 (Tab 2). Their percentage identity and query coverage in <em>Brassica napus</em> reference genotypes Quinta, Tapidor, Westar and Zhongshuang 11 compared to the <em>B. napus</em> pantranscriptome is indicated. Predicted <em>Arabidopsis thaliana</em> orthologs and corresponding descriptions are shown where possible. </p> <p> </p>
Data from: Do genetic loci that cause reproductive isolation in the lab inhibit gene flow in nature?
<p>The genetic dissection of reproductive barriers between diverging lineages provides enticing clues into the origin of species. One strategy uses linkage analysis in experimental crosses to identify genomic locations involved in phenotypes that mediate reproductive isolation. A second framework searches for genomic regions that show reduced rates of exchange across natural hybrid zones. It is often assumed that these approaches will point to the same loci, but this assumption is rarely tested. In this perspective, we discuss the factors that determine whether loci connected to postzygotic reproductive barriers in the laboratory are inferred to reduce gene flow in nature. We synthesize data on the genetics of postzygotic isolation in house mice, one of the most intensively studied systems in speciation genetics. In a rare empirical comparison, we measure the correspondence of loci tied to postzygotic barriers via genetic mapping in the laboratory and loci at which gene flow is inhibited across a natural hybrid zone. We find no evidence that the two sets of loci overlap beyond what is expected by chance. In light of these results, we recommend avenues for empirical and theoretical research to resolve the potential incongruence between the two predominant strategies for understanding the genetics of speciation.</p>
Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities: discovery cohort meta data and parsed TCR repertoire data
<p>Meta data corresponding the the discovery cohort for the paper, "Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities" by Magdalena L Russell, Aisha Souquette, David M Levine, Stefan A Schattgen, E Kaitlynn Allen, Guillermina Kuan, Noah Simon, Angel Balmaseda, Aubree Gordon, Paul G Thomas, Frederick A Matsen IV, and Philip Bradley. These meta data include: </p> <p>(1) a file mapping the SNP data subject IDs to the TCR repertoire data subject IDs (gwas_id_mapping.tsv)<br> (2) a file including the PCAir PCs, self-reported ancestry, and genomic ancestry for each subject (all_pc_air.txt)<br> (3) a file including the PCAir variance explained by each PC (all_pc_air_variance.txt)<br> (3) a file including the SNP ID, chromosome, hg19 position, allele, rsid, and quality control metrics for each SNP in the SNP array (emerson_snp_rs_data.tsv)<br> (4) a file including IMGT genes and sequences used for parsing TCRB repertoire data (human_vj_allele_cdr3_nucseqs.tsv)<br> (5) a file including predicted TRBD2 allele genotypes for each subject (emerson_trbd2_alleles.tsv)<br> (6) Parsed TCRB repertoire data. These raw data were first published in Emerson et. al, <em>Nature Genetics </em>2017. (emerson_parsed_tcrb.tgz)</p> <p><strong>Corresponding discovery cohort raw TCR repertoire data is available here: </strong>https: //doi.org/10.21417/B7001Z (ImmuneACCESS database)<br> <strong>Corresponding discovery cohort SNP data is available here:</strong> https: //www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs001918.v1.p1 (The database of Genotypes and Phenotypes, accession number: phs001918)<br> <br> <strong>Software tools designed to work with these data are available here:</strong> https://github.com/phbradley/tcr-gwas</p>
Fig. 1 in The Dynamics Of Genetic Structure Of Round G O B Y N E O G O B I U S M E L A N O S T O M U S (Pa L L A S) Groupings In The Odessa Bay Of The Black Sea Utilizing Biochemical Marker Loci
Fig. 1. Frequencies of S-alleles by polymorphic locus Es2 in round goby groupings from different parts of the Odessa Bay. * – significant deviation of allele frequencies in round goby groupings from the south and the north part of the Odessa Bay (Р = 0,05); # – significant deviation of allele frequencies in round goby groupings in the south part of the Odessa Bay in 2015-2016 in comparison to 2013-2014 (Р = 0,05).
Fig. 2 in The Dynamics Of Genetic Structure Of Round G O B Y N E O G O B I U S M E L A N O S T O M U S (Pa L L A S) Groupings In The Odessa Bay Of The Black Sea Utilizing Biochemical Marker Loci
Fig. 2. Frequencies of S-alleles by polymorphic locus of myogene 3 in round goby groupings from different parts of the Odessa Bay * – significant deviation of allele frequencies in round goby groupings from the south and the north part of the Odessa Bay in 2013 and 2014 (Р = 0,05); # – significant deviation of allele frequencies in round goby groupings from the south part of the Odessa Bay in 2013-2014 in comparison to 2015 (Р = 0,05).
FIGURE 1 in Development of microsatellite loci and population genetics in the bumblebee catfish species Pseudopimelodus atricaudus and Pseudopimelodus magnus (Siluriformes: Pseudopimelodidae)
FIGURE 1 | Sampling sites of Pseudopimelodus magnus and P. atricaudus in the middle and lower sectors of the Cauca River.
Fig. 2 in Testing microsatellite loci and preliminary genetic study for Eurasian otter in South Korea
Fig. 2. Locations of sampling for tissue (1. Hoengseong-gun, Gangwon-do, 2. Uljin-gun, Gyeongsangbuk-do, 3. Jeongeup-si, Jeollabukdo, 4. Muju-gun, Jeollabuk-do, 5. Hampyeong-gun, Jeollanam-do).
Microsat Data for 'Simulated Disperser Analysis: determining the number of loci required to genetically identify dispersers'
<p>Microsattelite data from 94 samples (<em>Stunus vulgaris</em>) from 3 populations and including 29 loci. Used in the paper 'Simulated Disperser Analysis: determining the number of loci required to genetically identify dispersers'. </p>
FIGURE 2 in Development of microsatellite loci and population genetics of the catfish Pimelodus yuma (Siluriformes: Pimelodidae)
FIGURE 2 | Discriminant analysis of principal components for nine microsatellite loci and 138 individuals of Pimelodus yuma in three sections (S4/5, S6 and S7/8) of the Cauca River.
FIGURE 1 in Development of microsatellite loci and population genetics of the catfish Pimelodus yuma (Siluriformes: Pimelodidae)
FIGURE 1 | Studied sampling sites of Pimelodus yuma along the lower sections (S4–S8) of the Cauca River. The pentagons indicate sampling sites in floodplain lakes and the stars indicate sites along the main channel of the river.
FIGURE 2 in Population genetics of the endangered catfish Pseudoplatystoma magdaleniatum (Siluriformes: Pimelodidae) based on species-specific microsatellite loci
FIGURE 2 | Results of Structure (A, B) and Discriminant analysis of principal components (C) for Pseudoplatystoma magdaleniatum. A: K = 1; B: K = 2; M: Margento, PC: Punta Cartagena, PB: Puerto Berrío, SN: Samaná Norte.
FIGURE 3 in Development of microsatellite loci and population genetics of the catfish Pimelodus yuma (Siluriformes: Pimelodidae)
FIGURE 3 | STRUCTURE results for Pimelodus yuma showing K= 2 genetic stocks in three sections (S4/5, S6 and S7/8) of the Cauca River.
Key triggers of adaptive genetic variability of sessile oak [Q. petraea (Matt.) Liebl.] from the Balkan refugia: outlier detection and association of SNP loci from ddRAD-seq data
<p>Knowledge on the genetic composition of <em>Quercus petraea</em> in south-eastern Europe is limited despite the species' significant role in the re-colonisation of Europe during the Holocene, and the diverse climate and physical geography of the region. Therefore, it is imperative to conduct research on adaptation in sessile oak to better understand its ecological significance in the region. While large sets of SNPs have been developed for the species, there is a continued need for smaller sets of SNPs that are highly informative about the possible adaptation to this varied landscape. By using double digest restriction site associated DNA sequencing data from our previous study, we mapped RAD-tag sequences to the <em>Quercus robur</em> reference genome and identified a set of SNPs putatively related to drought stress-response. A total of 179 individuals from eighteen natural populations at sites covering heterogeneous climatic conditions in the southeastern natural distribution range of <em>Q. petraea</em> were genotyped. The detected highly polymorphic variant sites revealed three genetic clusters with a generally low level of genetic differentiation and balanced diversity among them but showed a north–southeast gradient. Selection tests showed nine outlier SNPs positioned in different functional regions. Genotype-environment association analysis of these markers yielded a total of 53 significant associations, explaining 2.4–16.6% of the total genetic variation. Our work exemplifies that adaptation to drought may be under natural selection in the examined <em>Q. petraea</em> populations.</p>
Data from: A few essential genetic loci distinguish Penstemon species with flowers adapted to pollination by bees or hummingbirds
<p>In the formation of species, adaptation by natural selection generates distinct combinations of traits that function well together. The maintenance of adaptive trait combinations in the face of gene flow depends on the strength and nature of selection acting on the underlying genetic loci. Floral pollination syndromes exemplify the evolution of trait combinations adaptive for particular pollinators. The North American wildflower genus <em>Penstemon</em> displays remarkable floral syndrome convergence, with at least 20 separate lineages that have evolved from ancestral bee pollination syndrome (wide blue-purple flowers that present a landing platform for bees and small amounts of nectar) to hummingbird pollination syndrome (bright red narrowly tubular flowers offering copious nectar). Related taxa that differ in floral syndrome offer an attractive opportunity to examine the genomic basis of complex trait divergence. In this study, we characterized genomic divergence among 229 individuals from a <em>Penstemon </em>species complex that includes both bee and hummingbird floral syndromes. Field plants are easily classified into species based on phenotypic differences and hybrids displaying intermediate floral syndromes are rare. Despite unambiguous phenotypic differences, genomewide differentiation between species is minimal. Hummingbird-adapted populations are more genetically similar to nearby bee-adapted populations than to geographically distant hummingbird-adapted populations, in terms of genomewide <em>d<sub>XY</sub>.</em> However, a small number of genetic loci are strongly differentiated between species. These ~ 20 "species-diagnostic loci", which appear to have nearly fixed differences between pollination syndromes, are sprinkled throughout the genome in high recombination regions. Several map closely to previously established floral trait QTLs. The striking difference between the diagnostic loci and the genome as whole suggests strong selection to maintain distinct combinations of traits, but with sufficient gene flow to homogenize the genomic background. A surprisingly small number of alleles confer phenotypic differences that form the basis of species identity in this species complex.</p>
Data from: Do genetic loci that cause reproductive isolation in the lab inhibit gene flow in nature?
Open the record for dataset details and reuse information.
Data from: A few essential genetic loci distinguish Penstemon species with flowers adapted to pollination by bees or hummingbirds
Open the record for dataset details and reuse information.
Genome-wide association mapping to identify genetic loci for cold tolerance and cold recovery during germination in rice
<p>To investigate the genetic architecture underlying cold tolerance during germination in rice (<i>Oryza sativa</i>), we conducted a genome-wide association study (GWAS) using a novel diversity panel of 257 rice accessions from around the world and 5,185 SNP markers from a 7K SNP marker array. Genotyping was performed using a 7K Illumina iSelect custom-designed array by following the Infinium HD Array Ultra Protocol. The 7K array, called the C7AIR, was designed by Dr. Susan McCouch's Lab at Cornell University and consists of 7,098 SNPs (Morales et al. 2020, under review). After genotyping 257 rice accessions with the 7K array (C7AIR), poor-performing SNP markers (SNPs of call rate <90%; minor allele frequency <5%; or heterozygosity >20%) were removed from the dataset. For our study, a subset of 5,185 high-quality SNP markers obtained after filtering was used to perform the genome-wide association analysis. The dataset representing the genotype data of 5,185 SNP markers by 257 rice accessions is presented here.</p>
Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities: validation cohort meta data and parsed TCR repertoire data
<p>Meta data corresponding the the validation cohort for the paper, "Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities" by Magdalena L Russell, Aisha Souquette, David M Levine, Stefan A Schattgen, E Kaitlynn Allen, Guillermina Kuan, Noah Simon, Angel Balmaseda, Aubree Gordon, Paul G Thomas, Frederick A Matsen IV, and Philip Bradley. These meta data include: </p> <p>(1) SNP genotypes for the two SNPs which overlap with the discovery cohort<br> - (nicaragua_snp_genotypes_ints.tsv) -- SNP genotypes as integers<br> - (nicaragua_snp_genotypes_strings.tsv) -- SNP genotypes as allele strings <br> (2) the ancestry PCs for each individual in the validation cohort (nicaragua_snp_ancestry_PCA.tsv)<br> (3) a file including IMGT genes and sequences used for parsing TCRB repertoire data (human_vj_allele_cdr3_nucseqs.tsv)<br> (4) a file including IMGT genes used for parsing TCRA repertoire data (human_vj_alleles_alpha.tsv)<br> (5) Parsed TCRA repertoire data (nicaragua_parsed_TCRA.tgz)<br> (6) Parsed TCRB repertoire data (nicaragua_parsed_TCRB.tgz) </p> <p><strong>Corresponding raw validation cohort TCR repertoire data is available here:</strong> https://www. ncbi.nlm.nih.gov/bioproject/PRJNA762269 (The BioProject database, accession number: PRJNA762269)</p> <p><strong>Software tools designed to work with these data are available here:</strong> https://github.com/phbradley/tcr-gwas</p>
Modelling the genetic aetiology of complex disease: human-mouse conservation of noncoding features and disease-associated loci
<p>Understanding the genetic aetiology of loci associated with disease is crucial for developing preventative measures and effective treatments. Mouse models are used extensively to understand human pathobiology and mechanistic functions of disease-associated loci. However, the utility of mouse models is limited by evolutionary divergence in transcription regulation for pathways of interest. Here, we summarise the conservation of genomic (exonic and multi-cell regulatory) features and complex disease associated variant sites between humans and mice. Our results highlight the importance of understanding evolutionary divergence in transcription regulation when interpreting functional studies using mice as models for human disease variants.</p>
Fig. 1 in Testing microsatellite loci and preliminary genetic study for Eurasian otter in South Korea
Fig. 1. Spraints collection sites along Ungokcheon Stream, Bonghwa-gun, Gyeongsangbuk-do.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.