Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
450
datasets available to search
ShareScore release 0.7.1
Dataset results
450 results for “Candidate Genes”
Targeted Re-sequencing Identifies Candidate Fusiform Rust Resistance Genes in Loblolly Pine
<p>A fasta file containing the subset of the v2.01 Pita genome in addition to the novel NLR genes that were targeted by hybridization probes. </p> <p>A bed file describing the intervals targeted by the hybridization probes.</p> <p>Trinity assemblies of the 30 RNAseq libraries along with predictions by transdecoder of CDS and peptide sequences from those trinity assemblies. </p>
Ontology based text mining of gene-phenotype associations: application to candidate gene prediction
<p>Gene-phenotype associations play an important role in understanding<br> the disease mechanisms which is a requirement for treatment<br> development. A portion of gene-phenotype associations are observed<br> mainly experimentally and made publicly available through several<br> standard resources such as MGI. However, there is still a vast<br> amount of gene--phenotype associations buried in the biomedical<br> literature. Given the large amount of literature data, we need<br> automated text mining tools to alleviate the burden in manual<br> curation of gene-phenotype associations and to develop<br> comprehensive resources. We developed an ontology based<br> approach in combination with statistical methods to text mine<br> gene-phenotype associations from literature. Our method achieved<br> AUC values of 0.90 and 0.75 in recovering known gene-phenotype<br> associations from HPO and MGI respectively. We posit that candidate<br> genes and their relevant diseases should be expressed with similar<br> phenotypes in publications. Thus, we demonstrate the utility of our<br> approach by predicting disease candidate genes based on the semantic<br> similarities of phenotypes associated with genes and diseases. We evaluated our disease candidate prediction model on<br> the gene-disease associations from MGI. Our model achieved AUC<br> values of 0.90 and 0.87 on OMIM (human) and MGI (mouse) datasets of<br> gene-disease associations respectively. Our manual analysis on the<br> text mined data revealed that, our method can accurately extract<br> gene-phenotype associations which are not currently covered by the<br> existing public gene-phenotype resources. Overall, results indicate<br> that our method can precisely extract known as well as new<br> gene-phenotype associations from literature. This released dataset at Zenodo covers our gene-phenotype extracts from the literature. All the methods used to extract the data are available at https://github.com/bio-ontology-research-group/genepheno.</p>
GWAS to single cell: Intersecting single-cell transcriptomics and genome wide association studies identifies crucial cell-populations and candidate genes for atherosclerosis.
<p><strong>Background</strong></p> <p>Genome-wide association studies (GWAS) have discovered hundreds of common genetic variants for atherosclerotic disease and cardiovascular risk factors. The translation of susceptibility loci into biological mechanisms and targets for drug discovery remains challenging. Intersecting genetic and gene expression data has led to identification of candidate genes. However, the assayed tissues are often non-diseased and heterogeneous in cell composition confounding the candidate prioritization. We collected single-cell transcriptomics (scRNA-seq) from atherosclerotic plaques and aimed to identify cell-type-specific expression of disease-associated genes. </p> <p> </p> <p><strong>Methods and Results</strong></p> <p>To identify disease-associated candidate genes, we applied gene-based analyses using GWAS summary statistics from 46 atherosclerotic, cardiometabolic, and other traits. Next we intersected these candidates with single-cell transcriptomics (scRNA-seq) to identify those genes that are specifically expressed in individual cell (sub)populations of atherosclerotic plaques. We derive an enrichment score and show that loci that associated with coronary artery disease demonstrated a prominent substrate in plaque smooth muscle cells (<em>SKI</em>, <em>KANK2</em>, <em>SORT1</em>), endothelial cells (<em>SLC44A1</em>, <em>ATP2B1</em>), and macrophages (<em>APOE</em>, <em>HNRNPUL1</em>). Further sub clustering of SMC-subtypes revealed genes in risk loci for coronary calcification specifically enriched in a synthetic cluster of SMCs. To verify the robustness of our approach, we used liver-derived scRNAseq-data and showed enrichment of circulating lipids-associated loci in hepatocytes.</p> <p><br> <strong>Conclusion</strong></p> <p>We confirm known gene-cell pairs relevant for atherosclerotic disease, and discovered novel pairs pointing to new biological mechanisms amenable for therapy. We present an intuitive single-cell transcriptomics driven workflow rooted in human large-scale genetic studies to identify putative candidate genes and affected cells associated with cardiovascular traits.</p> <p> </p>
Differential Gene Expression Datasets for "Identification of candidate repurposable drugs to combat COVID‑19 using a signature‑based approach"
<p>This dataset has the unfiltered transcriptome differential expression results used in the paper "Identification of candidate repurposable drugs to combat COVID‑19 using a signature‑based approach". </p>
TPE-OLD Candidate Genes Suggested by Geposan
<p>This dataset contains a ranking of human genes suggesting candidates for TPE-OLD. This ranking was produced using <a href="https://github.com/johrpan/geposan" target="_blank" rel="noopener">geposan</a>, an R package for analyzing gene position data in comparison to a set of reference genes. The same dataset is also available in an interactive form at <a href="https://tpe-old.uni-rostock.de" target="_blank" rel="noopener">tpe-old.uni-rostock.de</a>.</p>
Genome-wide analysis identified candidate variants and genes associated with heat stress adaptation in Egyptian sheep breeds
<p>The current study was conducted from 2009 to 2019 in three hot and dry agroecological zones in Egypt: Western Desert coastal zone, New Valley desert oasis, and hot-dry Upper Egypt. Within these zones, three local sheep breeds were studied: Barki (83 ewes), Wahati (55 ewes) and Saidi (68 ewes). During the study period, the animals exercised under natural heat stress (simulating summer grazing on poor pasture). Meteorological and physiological parameters were measured and recorded. The heat tolerance index of the animals was calculated to identify animals with high and low heat tolerance based on the animals' response to the five main physiological parameters (scale from 0 to 5). DNA samples were extracted for genomic analysis. The genetic diversity measurements showed a significant influence of breed and location on the populations. The influence of breed is more significant than that of location. The inbreeding analysis shows that the desert breeds (Wahati and Barki) have lower values than the urban breed (Saidi). The high rate of sub-clustering indicates the process of sub-population through inbreeding pressure. Wahati and Barki are very distinct breeds with strong identification, while Saidi breed has crosses with other breeds. The most significant SNPs associated with heat tolerance were found in MYO5A, PRKG1, GSTCD, and RTN1 genes (P < 0.0001). MYO5A had an effect of 0.74 on the trait heat tolerance in the studied population. It produces a protein that is widely distributed in the melanin-producing neural crest of the skin. Genetic association between genetic and phenotypic variations showed that OAR1 18300122.1, located in ST3GAL3, had the greatest positive effect on heat tolerance. GWAS analysis identified SNPs associated with heat tolerance in the PLCB1, STEAP3, KSR2, UNC13C , PEBP4, and GPAT2 genes.</p>
From common gardens to candidate genes: Exploring local adaptation to climate in red spruce
<p><span>Local adaptation to climate is common in plant species and has been studied in a range of contexts, from improving crop yields to predicting population maladaptation to future conditions. The genomic era has brought new tools to study this process, which was historically explored through common garden experiments. </span></p> <p><span>In this study, we combine genomic methods and common gardens to investigate local adaptation in red spruce and identify environmental gradients and loci involved in climate adaptation. We first use climate transfer functions to estimate the impact of climate change on seedling performance in three common gardens. We then explore the use of multivariate gene-environment association (GEA) methods to identify genes underlying climate adaptation, with particular attention to the implications of conducting genome scans with and without correction for neutral population structure.</span></p> <p><span>This integrative approach uncovered phenotypic evidence of local adaptation to climate and identified a set of putatively adaptive genes, some of which are involved in three main adaptive pathways found in other temperate and boreal coniferous species: drought tolerance, cold hardiness, and phenology. These putatively adaptive genes segregated into two "modules" associated with different environmental gradients.</span></p> <p><span>This study nicely exemplifies the multivariate dimension of adaptation to climate in trees. </span></p>
Figure 7 in Identification of candidate genes involved with dicamba resistance in waterhemp (Amoronthus tuberculotus) via transcriptomics analyses
Figure 7. Temporal quantitative PCR results for (A) GST-nt, (B) ABC10, (C) PEROX12, and (D) GST-ct before and at multiple time points after dicamba treatment. Asterisks (*) indicate comparisons that were significant (t-test P-value <0.05), with error bars indicating variability across replicates.
Figure 4 in Identification of candidate genes involved with dicamba resistance in waterhemp (Amoronthus tuberculotus) via transcriptomics analyses
Figure 4. Genomic distribution in sliding 50-kb window plots of differentially expressed genes (DEs). The y-axis units refer to windows in mega base pairs (Mbp) Each plot represents one of the 16 pseudo-chromosomes of Amoronthus tuberculotus. Peaks represent clusters of DEs. Blue dashed lines represent previously identified hot-spot locations for 2,4-D resistance (Giacomini et al. 2020).
Figure 3 in Identification of candidate genes involved with dicamba resistance in waterhemp (Amoronthus tuberculotus) via transcriptomics analyses
Figure 3. Biological process GO-term enrichment analysis. Circle size represents the significance of overrepresented enrichment, and color gradient represents the significance of conditional enrichment. The x axis represents the number of genes annotated with each GO-term in the y axis.
Figure 2 in Identification of candidate genes involved with dicamba resistance in waterhemp (Amoronthus tuberculotus) via transcriptomics analyses
Figure 2. Volcano plot of all genes with key differentially expressed genes highlighted. Major genes with potential involvement in dicamba resistance are labeled according to their homologous UniprotKB ID. Genes in red and blue were significantly up- and downregulated, respectively, in dicamba-resistant relative to sensitive plants. The y axis refers to −log10 false discovery rate (FDR), and the x axis refers to the log2 expression fold change (FC).
Figure 8 in Identification of candidate genes involved with dicamba resistance in waterhemp (Amoronthus tuberculotus) via transcriptomics analyses
Figure 8. Proposed dicamba resistance mechanisms in the CHR population. Currently, knowledge about the synthetic auxin effect on plants indicates an overproduction of abscisic acid (ABA), leading to a large production of reactive oxygen species (ROS) and plant death (Christoffoleti et al. 2015; Gaines 2020). The proposed resistance mechanism is that enhanced response to oxidative stress via peroxidases and glutathione S-transferases alleviates dicamba toxicity. Other putative resistance mechanisms, such as glycosylation of dicamba and ABA, are also proposed with transport via ATP-binding cassette (ABC) transporters for further degradation. Overproduction of salicylic acid is also proposed as a potential tool for alleviating oxidative stress. Created with BioRender.com.
Figure 5 in Identification of candidate genes involved with dicamba resistance in waterhemp (Amoronthus tuberculotus) via transcriptomics analyses
Figure 5. Weighted co-expression network analysis results. (A) Gene expression dendrogram for module assignment where a total of 33 modules were identified. (B) Traitmodule correlation plot with values outside parentheses representing Pearson correlation and values inside parentheses representing the significance correlation P-values. Correlation values range from −1 to 1, with red values indicating a positive association and blue values indicating a negative association with dicamba resistance. ME refers to modules followed by their color code.
Figure 1 in Identification of candidate genes involved with dicamba resistance in waterhemp (Amoronthus tuberculotus) via transcriptomics analyses
Figure 1. Plant selection and phenotype classification for RNA-seq. Photos show the differences in phenotypes of some of the individuals selected for sequencing: (A) resistant plants and (B) sensitive plants. Selection was done based on visual damage estimation, biomass, and plant area measured via image analysis (Bobadilla et al. 2022). Photos were taken 14 d after treatment with dicamba at 560 g ai ha−1. The graph shows the relationship between biomass and plant area across resistant and sensitive individuals.
Rare Genomic Copy Number Variants Implicate New Candidate Genes for Bicuspid Aortic Valve
<p>Whole genome genotyping data in dbGAP format and copy number variant calls in dbVar format.</p> <p>dbVar data includes CNV calls from cases with early onset bicuspid aortic valve disease (EBAV), cases from the International BAV Consortium (BAVCon), and controls from the dbGAP Wisconsin Longitudinal Study on Aging dataset (WLS).</p> <p>dbGAP files are divided into 12 batches of genotypes from EBAV subjects (EBAV1-12):</p> <p>1) PLINK output files (.map and .ped)</p> <p>2) GenomeStudio Final Report files</p> <p>3) One master pedigree file</p> <p>4) dbGAP subject mapping files</p> <p> </p>
Selective sweeps identification in distinct groups of cultivated rye (Secale cereale L.) germplasm provides potential candidate genes for crop improvement
<p><strong>Background</strong></p> <p>During domestication and subsequent improvement, plants were subjected to intensive positive selection for desirable traits. Identification of selection targets is important with respect to the future targeted broadening of diversity in breeding programmes. Rye (<em>Secale</em> <em>cereale</em> L.) is a cereal that is closely related to wheat, and it is an important crop in Central, Eastern and Northern Europe. The aim of the study was (i) to identify diverse groups of rye accessions based on high-density, genome-wide analysis of genetic diversity within a set of 478 rye accessions, covering a full spectrum of diversity within the genus, from wild accessions to inbred lines used in hybrid breeding, and (ii) to identify selective sweeps in the established groups of cultivated rye germplasm and putative candidate genes targeted by selection. <strong> </strong></p> <p><strong>Results</strong></p> <p>Population structure and genetic diversity analyses based on high-quality SNP (DArTseq) markers revealed the presence of three complexes in the <em>Secale</em> genus: <em>S. sylvestre, S. strictum </em>and<em> S. cereale/vavilovii</em>, a relatively narrow diversity of <em>S. sylvestre</em>, very high diversity of <em>S. strictum</em>, and signatures of strong positive selection in <em>S. vavilovii</em>. Within cultivated ryes, we detected the presence of genetic clusters and the influence of improvement status on the clustering. Rye landraces represent a reservoir of variation for breeding, and especially a distinct group of landraces from Turkey should be of special interest as a source of untapped variation. Selective sweep detection in cultivated accessions identified 133 outlier positions within 13 sweep regions and 170 putative candidate genes related, among others, to response to various environmental stimuli (such as pathogens, drought, cold), plant fertility and reproduction (pollen sperm cell differentiation, pollen maturation, pollen tube growth), and plant growth and biomass production.</p> <p><strong>Conclusions</strong></p> <p>Our study provides valuable information for efficient management of rye germplasm collections, which can help to ensure proper safeguarding of their genetic potential and provides numerous novel candidate genes targeted by selection in cultivated rye for further functional characterisation and allelic diversity studies.</p>
Data from: Signatures of local adaptation in candidate genes of oaks (Quercus spp.) in respect to present and future climatic conditions
Open the record for dataset details and reuse information.
Selective sweeps identification in distinct groups of cultivated rye (Secale cereale L.) germplasm provides potential candidate genes for crop improvement
Open the record for dataset details and reuse information.
From common gardens to candidate genes: Exploring local adaptation to climate in red spruce
Open the record for dataset details and reuse information.
Data from: Candidate gene SNP variation in floodplain populations of pedunculate oak (Quercus robur L.) near the species' southern range margin: weak differentiation yet distinct associations with water availability
<p>Populations residing near species' low-latitude range margins (LLM) often occur in warmer and drier environments than those in the core range. Thus, their genetic composition could be shaped by climatic drivers that differ from those occurring at higher latitudes, resulting in potentially adaptive variants of conservation value. Such variants could facilitate the adaptation of populations from other portions of the geographic range to similar future conditions anticipated under ongoing climate change. However, very few studies have assessed standing genetic variation at potentially adaptive loci in natural LLM populations. We investigated standing genetic variation at SNPs located within 117 candidate genes and its links to putative climatic selection pressures across 19 pedunculate oak (Quercus robur L.) populations distributed along a regional climatic gradient near the species' southern range margin in southeastern Europe. These populations are restricted to floodplain forests along large lowland rivers, whose hydric regime is undergoing significant shifts under modern rapid climate change. The populations showed very weak geographic structure, suggesting extensive genetic connectivity and gene flow or shared ancestry. We identified eight (6.2%) positive FST-outlier loci, and genotype-environment association analyses revealed consistent associations between SNP allele frequencies and several climatic variables linked to water availability. A total of 61 associations involving 37 SNPs (28.5%) from 35 annotated genes provided important insights into putative functional mechanisms in our system. Our findings provide empirical support for the role of LLM populations as sources of potentially adaptive variation that could enhance species' resilience to climate change-related pressures.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.