Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
51
datasets available to search
ShareScore release 0.9.0
Dataset results
51 results for “Genome wide association mapping”
Combining genome-wide studies of breast, prostate, ovarian and endometrial cancers maps cross-cancer susceptibility loci and identifies new genetic associations
<p>Data set linked to the paper, "Combining genome-wide studies of breast, prostate, ovarian and endometrial cancers maps cross-cancer susceptibility loci and identifies new genetic associations". Pre-print of the paper is here: <a href="https://doi.org/10.1101/2020.06.16.146803">https://doi.org/10.1101/2020.06.16.146803</a>.</p> <p> </p> <p>cross_cancer_sum_stats.txt.gz contains summary genome-wide association statistics for susceptibility to single cancers (breast (BR), prostate (PR), ovarian (OV), endometrial (EN), estrogen receptor (ER)-positive breast (POS), ER-negative breast (NEG), and high-grade serous ovarian (HGS) cancers) and from the cross-cancer meta-analysis (main [main] and subtype-focused [sub]). EA in the header refers to the effect allele, OA is the other allele, EAF is the effect allele frequency in the largest of the single cancer data sets (BR), IMPR2 is the imputation quality in the largest of the single cancer data sets (BR), SE is the standard error, PVAL is the P-value, RE2Cs1 is the RE2C statistic mean effect part, RE2Cs2 is the RE2C statistic heterogeneity part, RE2Cp* is the RE2C* P-value. More on RE2Cp* can be found here: <a href="http://software.buhmhan.com/RE2C/index.php?mid=contact&act=dispBoardWrite">http://software.buhmhan.com/RE2C/index.php?mid=contact&act=dispBoardWrite</a> and in <a href="https://academic.oup.com/bioinformatics/article/33/14/i379/3953957">https://academic.oup.com/bioinformatics/article/33/14/i379/3953957</a> SNP names in cross_cancer_sum_stats.txt.gz include the chromosome and build 37 position.</p> <p> </p> <p>main_tetrachoric_corr_matrix.txt and subtype_tetrachoric_corr_matrix.txt provide the tetrachoric correlation matrices used in the main and subtype-focused meta-analyses. These were also used to specify the cryptic.cor argument of the exh.abf function of MetABF. More on MetABF can be found here: <a href="https://github.com/trochet/metabf">https://github.com/trochet/metabf</a> and in <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/gepi.22202">https://onlinelibrary.wiley.com/doi/abs/10.1002/gepi.22202</a></p> <p> </p> <p>prior_sigmas_for_metabf.txt contains the values used to specify the prior.sigma argument of the exh.abf function in MetABF.</p> <p> </p> <p>The breast cancer data used are described in <a href="https://pubmed.ncbi.nlm.nih.gov/29059683/"><strong>PMID 29059683</strong></a> and can be downloaded from <a href="http://bcac.ccge.medschl.cam.ac.uk/bcacdata/oncoarray/oncoarray-and-combined-summary-result/gwas- summary-results-breast-cancer-risk-2017/">http://bcac.ccge.medschl.cam.ac.uk/bcacdata/oncoarray/oncoarray-and-combined-summary-result/gwas- summary-results-breast-cancer-risk-2017/</a> (this link also includes acknowledgements). The prostate cancer data are described in <a href="https://pubmed.ncbi.nlm.nih.gov/29892016/"><strong>PMID 29892016</strong></a> and can be downloaded from: <a href="http://practical.icr.ac.uk/blog/?page_id=8164">http://practical.icr.ac.uk/blog/?page_id=8164</a> (this link also includes acknowledgements). The ovarian cancer data used are described in <a href="https://pubmed.ncbi.nlm.nih.gov/28346442/"><strong>PMID 28346442</strong></a> and can be downloaded from <a href="https://www.ebi.ac.uk/gwas/studies/GCST004415">https://www.ebi.ac.uk/gwas/studies/GCST004415</a>. The endometrial cancer data are described in <a href="https://pubmed.ncbi.nlm.nih.gov/30093612/"><strong>PMID 30093612</strong></a> and can be downloaded from <a href="https://www.ebi.ac.uk/gwas/studies/GCST006464">https://www.ebi.ac.uk/gwas/studies/GCST006464</a>. These links point to the same data that form the basis of the cross_cancer_sum_stats.txt.gz file.</p> <p> </p> <p><strong>The sample size and precision of the data presented should preclude identification of any individual study participant. However, in downloading these data, you undertake not to attempt to identify individual study participant and not to re-post these data to a third-party website. Please cite the PMIDs highlighted above along with the appropriate acknowledements if you use the cross_cancer_sum_stats.txt.gz file.</strong></p> <p> </p> <p>If you have any questions about this repository, please email Siddhartha Kar at siddhartha dot kar at bristol dot ac dot uk</p>
Data from: Genome-wide association mapping of quantitative traits in a breeding population of sugarcane
Background: Molecular markers associated with relevant agronomic traits could significantly reduce the time and cost involved in developing new sugarcane varieties. Previous sugarcane genome-wide association analyses (GWAS) have found few molecular markers associated with relevant traits at plant-cane stage. The aim of this study was to establish an appropriate GWAS to find molecular markers associated with yield related traits consistent across harvesting seasons in a breeding population. Sugarcane clones were genotyped with DArT (Diversity Array Technology) and TRAP (Target Region Amplified Polymorphism) markers, and evaluated for cane yield (CY) and sugar content (SC) at two locations during three successive crop cycles. GWAS mapping was applied within a novel mixed-model framework accounting for population structure with Principal Component Analysis scores as random component. Results: A total of 43 markers significantly associated with CY in plant-cane, 42 in first ratoon, and 41 in second ratoon were detected. Out of these markers, 20 were associated with CY in 2 years. Additionally, 38 significant associations for SC were detected in plant-cane, 34 in first ratoon, and 47 in second ratoon. For SC, one marker-trait association was found significant for the 3 years of the study, while twelve markers presented association for 2 years. In the multi-QTL model several markers with large allelic substitution effect were found. Sequences of four DArT markers showed high similitude and e-value with coding sequences of Sorghum bicolor, confirming the high gene microlinearity between sorghum and sugarcane. Conclusions: In contrast with other sugarcane GWAS studies reported earlier, the novel methodology to analyze multi-QTLs through successive crop cycles used in the present study allowed us to find several markers associated with relevant traits. Combining existing phenotypic trial data and genotypic DArT and TRAP marker characterizations within a GWAS approach including population structure as random covariates may prove to be highly successful. Moreover, sequences of DArT marker associated with the traits of interest were aligned in chromosomal regions where sorghum QTLs has previously been reported. This approach could be a valuable tool to assist the improvement of sugarcane and better supply sugarcane demand that has been projected for the upcoming decades.
Data from: Linkage disequilibrium clustering-based approach for association mapping with tightly linked genome-wide data
Open the record for dataset details and reuse information.
Data from: Genome-wide association mapping of quantitative traits in a breeding population of sugarcane
Open the record for dataset details and reuse information.
Data from: Genome-wide association mapping in a wild avian population identifies a link between genetic and phenotypic variation in a life-history trait
Open the record for dataset details and reuse information.
Genome-Wide Mapping of RNA-Protein Associations via Sequencing
GEO Series GSE270010. Homo sapiens. 5 samples. Type: Other.
Genome-wide map of long non-coding RNA steroid receptor RNA activator (SRA) and its associated RNA heliase p68 in human pluripotent stem cells NTERA2
GEO Series GSE58641. Homo sapiens. 3 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Exome sequencing and genome-wide copy number variant mapping reveal novel associations with sensorineural hereditary hearing loss
GEO Series GSE64088. Homo sapiens. 308 samples. Type: Genome variation profiling by genome tiling array.
Genome-wide targeted methyl-seq: Allele-specific DNA methylation is increased in cancers and its dense mapping in normal plus neoplastic cells increases the yield of disease-associated regulatory SNPs
GEO Series GSE137287. Homo sapiens. 14 samples. Type: Methylation profiling by high throughput sequencing.
Genome-wide maps of Yan chromatin association
GEO Series GSE34040. Drosophila melanogaster. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Chromatin-associated RNA sequencing (ChAR-seq) maps genome-wide RNA-to-DNA contacts
GEO Series GSE97131. Drosophila melanogaster. 5 samples. Type: Other.
Genome-wide mapping of i-Motifs reveals their association with transcription regulation in live human cells [RNA-seq]
GEO Series GSE220881. Homo sapiens. 3 samples. Type: Expression profiling by high throughput sequencing.
Unbiased, Genome-wide in vivo Mapping of Transcriptional Regulatory Elements Reveals Sex Differences in Chromatin Structure Associated with Sex-specific Liver Gene Expression
GEO Series GSE21777. Mus musculus. 10 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Genome-wide mapping of cytosine methylation revealed dynamic DNA methylation patterns associated with rice centromeres
GEO Series GSE21414. Oryza sativa. 1 samples. Type: Methylation profiling by high throughput sequencing.
Genome-wide mapping of i-Motifs reveals their association with transcription regulation in live human cells
GEO Series GSE220882. Homo sapiens. 17 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Expression profiling by high throughput sequencing.
Genome-wide mapping of m6A modified chromatin associated RNA
GEO Series GSE173517. Mus musculus. 8 samples. Type: Other.
Genome wide mapping of AR and its associated factors in prostate cancer
GEO Series GSE58428. Homo sapiens. 30 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Genome-wide Association Mapping for Heat Shock Tolerance in Mercenaria mercenaria through SNP Microarray Analysis
GEO Series GSE290453. Mercenaria mercenaria. 633 samples. Type: SNP genotyping by SNP array.
Genome-wide mapping to profile histone H3-associated epigenetic marks and chromatin conformation in senescent human stromal cells
GEO Series GSE163105. Homo sapiens. 24 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Genome-wide mapping of Ccp1 association sites in fission yeast
GEO Series GSE95046. Schizosaccharomyces pombe. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.