Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

720

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

720 results for “genomics study”

Learn how ShareScore rates datasets ↗
zenodo52/100

A genome-wide study of ruminants reveals two endogenous retrovirus families still active in goats

<p>Additional file - A genome-wide study of ruminants reveals two endogenous retrovirus families still active in goats&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Genome-Wide DNA Methylation in Peripheral Blood and Long-Term Exposure to Source-Specific Transportation Noise and Air Pollution: The SAPALDIA Study (Supplementary Data)

<p>The zip file contains supplementary data for the publication - Genome-Wide DNA Methylation in Peripheral Blood and Long-Term Exposure to Source-Specific Transportation Noise and Air Pollution: The SAPALDIA Study, accepted for publication in Environmental Health Perspectives (DOI: 10.1289/EHP6174).</p> <p>The description of the files are noted below:</p> <p><strong>1. Readme File for SAPALDIA Noise and Air Pollution EWAS Single Exposure.zip </strong></p> <p>This zip file contains all the results of the association between source-specific transportation noise (aircraft, railway and road traffic), air pollution (NO<sub>2</sub> and PM<sub>2.5</sub>), and genome-wide DNA methylation, derived from multi-exposure models.</p> <p><strong>SAPALDIA_EWAS_SingleExposure_AircraftLden.txt</strong> contains the results for aircraft noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_RailwayLden.txt</strong> contains the results for railway noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_RoadtrafficLden.txt</strong> contains the results for road traffic noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_NO2.txt</strong> contains the results for nitrogen dioxide</p> <p><strong>SAPALDIA_EWAS_SingleExposure_PM25.txt</strong> contains the results for fine particulate matter</p> <p>&nbsp;</p> <p><strong>General footnote for all files:</strong>SAPALDIA: Swiss cohort study on air pollution and lung and heart diseases in adults. CpG: Cytosine-phosphate-Guanine. CHR: chromosome. SE: standard error. Lden: day-evening-night noise level. NO<sub>2</sub>: nitrogen dioxide. PM<sub>2.5</sub>: particulate matter with aerodynamic diameter &lt;2.5 &micro;m. Beta coefficients represent increase or decrease in DNA methylation per 10 dB increase in aircraft, railway or road traffic Lden or 10 &micro;g/m<sup>3</sup> increase in NO<sub>2</sub> or PM<sub>2.5</sub>. All estimates were from single exposure epigenome-wide linear mixed models, with random intercept at the level of participant. Each model was adjusted for age, sex, educational level, area, and neighborhood socio-economic status, greenness index, smoking status and pack years, exposure to passive smoke, consumption of fruits, vegetables and alcohol, nested study, asthma status, survey, source-specific noise truncation indicator (for Lden models) and leukocyte composition. In a preliminary step, DNA methylation &beta;-values were regressed on the Illumina control probe-derived first 30 principal components to correct for correlation structures and technical bias, and residuals of these regressions covering 430,477 CpGs were used as the technical bias-corrected methylation level at the CpG sites.</p> <p>Extreme values of the residuals (lying beyond three times the interquartile range below the first quartile and above the third quartile at each CpG site) were replaced with their corresponding detection threshold value (&ldquo;modified winsorization&rdquo;). The &ldquo;winsorized&rdquo; data were then used as the dependent variables in the epigenome-wide association study.</p> <p>&nbsp;</p> <p><strong>2. Readme File for SAPALDIA Noise and Air Pollution EWAS Multi Exposure.zip </strong></p> <p>This zip file contains all the results of the association between source-specific transportation noise (aircraft, railway and road traffic), air pollution (NO<sub>2</sub> and PM<sub>2.5</sub>), and genome-wide DNA methylation, derived from multi-exposure models.</p> <p><strong>SAPALDIA_EWAS_MultiExposure_AircraftLden.txt</strong> contains the results for aircraft noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_RailwayLden.txt</strong> contains the results for railway noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_RoadtrafficLden.txt</strong> contains the results for road traffic noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_NO2.txt</strong> contains the results for nitrogen dioxide</p> <p><strong>SAPALDIA_EWAS_MultiExposure_PM25.txt</strong> contains the results for fine particulate matter</p> <p><strong>General table footnotes: </strong>SAPALDIA: Swiss cohort study on air pollution and lung and heart diseases in adults. CpG: Cytosine-phosphate-Guanine. CHR: chromosome. SE: standard error. Lden: day-evening-night noise level. NO<sub>2</sub>: nitrogen dioxide. PM<sub>2.5</sub>: particulate matter with aerodynamic diameter &lt;2.5 &micro;m. Beta coefficients represent increase or decrease in DNA methylation per 10 dB increase in aircraft, railway or road traffic Lden or 10 &micro;g/m<sup>3</sup> increase in NO<sub>2</sub> or PM<sub>2.5</sub>. All estimates were from multi-exposure epigenome-wide linear mixed models, with random intercept at the level of participant, and were adjusted for age, sex, educational level, area, and neighborhood socio-economic status, greenness index, smoking status and pack years, exposure to passive smoke, consumption of fruits, vegetables and alcohol, nested study, asthma status, survey, source-specific noise truncation indicator and leukocyte composition. Multi-exposure models included all five exposures (Aircraft, railway, road traffic Lden and respective truncation indicators, NO<sub>2</sub> and PM<sub>2.5</sub>) at the same time. In a preliminary step, DNA methylation &beta;-values were regressed on the Illumina control probe-derived first 30 principal components to correct for correlation structures and technical bias, and residuals of these regressions covering 430,477 CpGs were used as the technical bias-corrected methylation level at the CpG sites. Extreme values of the residuals (lying beyond three times the interquartile range below the first quartile and above the third quartile at each CpG site) were replaced with their corresponding detection threshold value (&ldquo;modified winsorization&rdquo;). The &ldquo;winsorized&rdquo; data were then used as the dependent variables in the epigenome-wide association study.</p>

opencc-by-4.0May 2020View details →
zenodo48/100

Genome-wide association study suggests that variation at the RCOR1 locus is associated with tinnitus in UK Biobank

<p>The dataset contains results of a genome-wide association studies for age-related hearing impairment (ARHI)-related traits as described in the following publication:<br> Wells, H.R.R., Abidin, F.N.Z., Freidin, M.B. et al. Genome-wide association study suggests that variation at the RCOR1 locus is associated with tinnitus in UK Biobank. Sci Rep 11, 6470 (2021). https://doi.org/10.1038/s41598-021-85871-6</p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Summary statistics accompanying the article "Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency" in Scientific Reports (2022)

<p>Summary statistics for genome-wide association studies reported in:</p> <p>Bell, S., Tozer, D.J., &amp; Markus H.S. (2022). Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency. <em>Scientific Reports</em>, DOI: <a href="https://dx.doi.org/10.1038/s41598-022-19106-7">10.1038/s41598-022-19106-7</a>.&nbsp;</p> <p><strong>Abstract</strong></p> <p>Complex brain networks play a central role in integrating activity across the human brain, and such networks can be identified in the absence of any external stimulus. We performed 10 genome-wide association studies of resting state network measures of intrinsic brain activity in up to 36,150 participants of European ancestry in the UK Biobank. We found that the heritability of global network efficiency was largely explained by blood oxygen level-dependent (BOLD) resting state fluctuation amplitudes (RSFA), which are thought to reflect the vascular component of the BOLD signal. RSFA itself had a significant genetic component and we identified 24 genomic loci associated with RSFA, 157 genes whose predicted expression correlated with it, and 3 proteins in the dorsolateral prefrontal cortex and 4 in plasma. We observed correlations with cardiovascular traits, and single-cell RNA specificity analyses revealed enrichment of vascular related cells. Our analyses also revealed a potential role of lipid transport, store-operated calcium channel activity, and inositol 1,4,5-trisphosphate binding in resting-state BOLD fluctuations. We conclude that that the heritability of global network efficiency is largely explained by the vascular component of the BOLD response as ascertained by RSFA, which itself has a significant genetic component.</p> <p>&nbsp;</p> <p>Further information on the files uploaded here can be found in the README. Users interested in bulk downloading these summary statistics may find <a href="https://github.com/dvolgyes/zenodo_get">zenodo_get</a> helpful.</p>

opencc-by-4.0Aug 2022View details →
zenodo48/100

Bin-assembled Escherichia coli genomes from a study in Punjab, Pakistan

<h2>Bin-assembled <em>Escherichia coli</em> genomes from Punjab, Pakistan</h2> <p>These assemblies are a part of a cross-sectional study conducted in Punjab, Pakistan aimed at investigating <em>E. coli</em> colonisation diversity in healthy carriage with the use of CLED enrichment plates.</p> <h3><strong>About</strong></h3> <h4><strong>Version history</strong></h4> <p><strong>v0.1.1 (current version)</strong></p> <ul> <li>Added reference to the study.</li> </ul> <p><strong>v0.1.0</strong></p> <ul> <li>Added brief description with a few missing parts.</li> </ul> <h4><strong>Distribution</strong></h4> <p>If you use these assemblies in your study please cite the source as appropriate.&nbsp;These assemblies are made available under a CC-BY 4.0 license.</p> <h4><strong>Citation</strong></h4> <p>Khawaja, T., M&auml;klin, T., Kallonen, T. et al. Deep sequencing of <em>Escherichia coli</em> exposes colonisation diversity and impact of antibiotics in Punjab, Pakistan. Nature Communications 15, 5196 (2024).&nbsp;<a href="https://doi.org/10.1038/s41467-024-49591-5">https://doi.org/10.1038/s41467-024-49591-5</a></p> <h3><strong>Methods briefly</strong></h3> <h4><strong>Species identification</strong></h4> <p>Sequencing data from the ENA project <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB36642">PRJEB36642</a> was error-corrected with <a href="https://github.com/opengene/fastp">fastp</a> and pseudoaligned with <a href="https://github.com/algbio/themisto">Themisto</a> against a species-level index (available from <a href="https://doi.org/10.5281/zenodo.6656881">https://doi.org/10.5281/zenodo.6656881</a>). Reads were assigned to species using the <a href="https://doi.org/10.1099%2Fmgen.0.000691">mSWEEP/mGEMS pipeline</a> as described in <a href="https://www.nature.com/articles/s41467-022-35178-5">https://www.nature.com/articles/s41467-022-35178-5</a>.</p> <h4><strong>Lineage identification</strong></h4> <p>Read from the species-level bins were again pseudoaligned with Themisto against an <em>E. coli</em> index (will be made available in a later version). Lineage-level assignment was performed using mSWEEP and mGEMS at the level of <a href="https://genome.cshlp.org/content/29/2/304">PopPUNK</a> sequence clusters. The created bins were screened with <a href="https://github.com/tmaklin/coreutils_demix_check">demix_check</a> and bins that received a score of 1 or 2 were kept. Data in the kept bins were assembled with <a href="https://github.com/tseemann/shovill">shovill</a> and the bin-assembled genomes (BAGs) were quality controlled with <a href="https://genome.cshlp.org/content/25/7/1043">checkm</a> for &gt;= 90% completeness and &lt;= 10% contamination. Finally, BAGs shorter than 4 Mb or longer than 6 Mb were removed.</p> <h3><strong>Contact</strong></h3> <p>Tommi M&auml;klin &lt;tommi'at'maklin.fi&gt;.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Genome-wide association study of full body nevus count in the Brisbane Twin Nevus Study (BTNS)

<p>The project uses the Brisbane Twin Nevus Study (BTNS) (N=3863)) to compare nevus counts on different anatomical sites to assess which anatomical site serves as best proxy for counting nevi on the whole body.In the project, a GWAS of nevus count on the whole body and GWAS of nevus count on the outer arm are performed.Here is the GWAS of total nevus count.</p> <p>Sample: GWAS analysis only includes samples of European ancestry. total nevus count were counted by trained research nurse.</p> <p>Genotype: All genotypes were imputed to a human haplotype map (HapMap) reference panel. Genome-wide association analyses were performed using Genome-wide Efficient Mixed model Association (GEMMA), which can account for genetically related individuals such as twins and siblings. Sex, age, age2, sex*age, sex*age2, sunburn, BSA, sun exposed hours weighted by UV index and 5 PCs, additionally two batch effect variables; were included as covariates. SNP imputation quality filter retained SNP with an INFO &gt; 0.3. minor allele frequency frequency filter was applied to retain SNP MAF &gt; 0.1</p> <p>Columns include:</p> <p>CHR: Chromosome</p> <p>BP: Base pair</p> <p>SNP: rsID</p> <p>A1: Effect allele</p> <p>A2: Non-effect allele</p> <p>A1FQ: Effect allele frequency</p> <p>HWE: Hardy-Weinberg Equilibrium</p> <p>BETA: Effect estimate (of effect allele_</p> <p>SEB: Standard error of beta</p> <p>PRB: P value</p> <p>N: Per SNP sample size</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Dataset for publication: An inter-laboratory study characterizes the impact of bioinformatic approaches on genome-based cluster detection for foodborne bacterial pathogens

<p>This dataset is part of a dry-lab interlaboratory study conducted across Germany, regarding bacterial outbreak detection based on NGS data, with a focus on bioinformatic analysis of four species to identify potential variability caused by different data analysis approaches and human interpretation. Participants were asked to follow their usual in-house protocols while adhering to the general guidelines. A quality assessment (with sample exclusion) was followed by 7-gene Multilocus-Sequence Typing (MLST), core genome Multilocus Sequencing Typing (cgMLST), and SNP calling. The participants were then asked to identify clusters. The study was not intended to resemble a standard proficiency test with a passing/failing grade, but rather to investigate and quantify obvious variability in the results and, where possible, the reasons for it. For this purpose, the datasets included borderline cases in terms of quality.</p>

opencc-by-4.0Oct 2025View details →
zenodo44/100

A chromosome-level genome resource for studying virulence mechanisms and evolution of the coffee rust pathogen Hemileia vastatrix

<p>Recurrent epidemics of coffee leaf rust, caused by the fungal pathogen <em>Hemileia vastatrix,</em> have constrained the sustainable production of Arabica coffee for over 150 years. The ability of <em>H. vastatrix </em>to overcome resistance in coffee cultivars and evolve new races is inexplicable for a pathogen that supposedly only utilizes clonal reproduction. Understanding the evolutionary complexity between <em>H. vastatrix</em> and its only known host, including determining how the pathogen evolves virulence so rapidly is crucial for disease management. Achieving such goals relies on the availability of a comprehensive and high-quality genome reference assembly. To date, two reference genomes have been assembled and published for <em>H. vastatrix</em> that, while useful, remain fragmented and do not represent chromosomal scaffolds. Here, we present a complete scaffolded pseudochromosome-level genome resource for <em>H. vastatrix </em>strain 178a (Hv178a). Our initial assembly revealed an unusually high degree of gene duplication (over 50% BUSCO basidiomycota_odb10 genes). Upon inspection, this was predominantly due to a single scaffold that itself showed 91.9% BUSCO Completeness. Taxonomic analysis of predicted BUSCO genes placed this scaffold in Exobasidiomycetes and suggests it is a distinct genome, which we have named Hv178a associated fungal genome (Hv178a AFG). The high depth of coverage and close association with Hv178a raises the prospect of symbiosis, although we cannot completely rule out contamination at this time. The main Ca. 546 Mbp Hv178a genome was primarily (97.7%) localised to 11 pseudochromosomes (51.5 Mb N50), building the foundation for future advanced studies of genome structure and organization. Citation:&nbsp;https://doi.org/10.1101/2022.07.29.502101</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction

<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the&nbsp;<em>P. teres&nbsp;</em>f.<em>&nbsp;maculata&nbsp;</em>isolate FGOB10Ptm-1.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction

<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the&nbsp;<em>P. teres&nbsp;</em>f.<em>&nbsp;maculata&nbsp;</em>isolate P-A14.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

False discovery rate calculations for genome-wide association study of reproductive fitness in Drosophila melanogaster (Sussex LHM sample)

<p>R code and results of applying false discovery (FDR) rate calculations to establish statistical signficance in a genome-wide association study of reproductive fitness in Drosophila melanogaster. Phenotype values were generated on hemiclone female and male lines from an outbred, laboratory adapted population. Thus, GWAS were previously performed seperately on the phenotype values for each sex, and also using a bivariate GWAS implemented in the R package multiPhen.</p> <p>FDR calculations were performed using the R package 'fdrtool' on all SNPs, and on LD-independent SNPs, the latter of which was used to determine p-value thresholds for genome-wide significance when all SNPs were considered.</p> <p>This version differs from the original in that: i) Some gene positions/names have been reassigned for accuaracy in the input data. ii) A file containing the p-value thresholds corresponding to an FDR of 0.1 has been added. The 95% credible intervals for each SNP association have been added to the results data files.</p>

opencc-by-4.0Sep 2017View details →
zenodo44/100

Genomic incongruence accompanies the evolution of flower symmetry in Eudicots: a case study in the poppy family (Papaveraceae, Ranunculales)

<p>Nuclear and plastid datasets and phylogenomic workflow associated to "Genomic Incongruence Accompanies the Evolution of Flower Symmetry in Eudicots: a case study in the poppy family (Papaveraceae, Ranunculales)", published in&nbsp;<em>Frontiers in Plant Science </em>15:1340056.<br>This compressed file (poppy_repo.zip) contains a markdown readme file (poppy_readme.md) describing the phylogenomic workflow followed, as well as two dataset folders (poppy_nuc and poppy_pl) divided into four (aln_nuc, gtr_nuc, sptr_nuc, and chrono_nuc) and three (aln_pl, sptr_pl, and chrono_pl) subfolders, respectively.<br>The nuclear folder (poppy_nuc) comprises shrunk and trimmed alignments (aln_nuc), ML gene trees (gtr_nuc), coalescent species trees (sptr_nuc), and a time tree (chrono_nuc).<br>The plastid folder (poppy_pl) comprises shrunk and trimmed alignments (aln_pl), a concatenated ML species tree (sptr_pl), and a time tree (chrono_pl).<br>The research article is available at https://www.frontiersin.org/journals/plant-science/articles/10.3389/fpls.2024.1340056 (doi: 10.3389/fpls.2024.1340056).</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Applications and raw data for SciPipe genomics and transcriptomics case studies

<p>Accompanying applications and raw data for the genomics and transcriptomics (RNA-Seq) case studies for SciPipe [1] available at&nbsp;https://github.com/pharmbio/scipipe-demo&nbsp;</p> <p>[1]&nbsp;http://scipipe.org</p>

opencc-by-4.0Jul 2018View details →
zenodo44/100

GWAS to single cell: Intersecting single-cell transcriptomics and genome wide association studies identifies crucial cell-populations and candidate genes for atherosclerosis.

<p><strong>Background</strong></p> <p>Genome-wide association studies (GWAS) have discovered hundreds of common genetic variants for atherosclerotic disease and cardiovascular risk factors. The translation of susceptibility loci into biological mechanisms and targets for drug discovery remains challenging. Intersecting genetic and gene expression data has led to identification of candidate genes. However, the assayed tissues are often non-diseased and heterogeneous in cell composition confounding the candidate prioritization. We collected single-cell transcriptomics (scRNA-seq) from atherosclerotic plaques and aimed to identify cell-type-specific expression of disease-associated genes.&nbsp;</p> <p>&nbsp;</p> <p><strong>Methods and Results</strong></p> <p>To identify disease-associated candidate genes, we applied gene-based analyses using GWAS summary statistics from 46 atherosclerotic, cardiometabolic, and other traits. Next we intersected these candidates with single-cell transcriptomics (scRNA-seq) to identify those genes that are specifically expressed in individual cell (sub)populations of atherosclerotic plaques. We derive an enrichment score and show that loci that associated with coronary artery disease demonstrated a prominent substrate in plaque smooth muscle cells (<em>SKI</em>, <em>KANK2</em>, <em>SORT1</em>), endothelial cells (<em>SLC44A1</em>, <em>ATP2B1</em>), and macrophages (<em>APOE</em>, <em>HNRNPUL1</em>). Further sub clustering of SMC-subtypes revealed genes in risk loci for coronary calcification specifically enriched in a synthetic cluster of SMCs. To verify the robustness of our approach, we used liver-derived scRNAseq-data and showed enrichment of circulating lipids-associated loci in hepatocytes.</p> <p><br> <strong>Conclusion</strong></p> <p>We confirm known gene-cell pairs relevant for atherosclerotic disease, and discovered novel pairs pointing to new biological mechanisms amenable for therapy. We present an intuitive single-cell transcriptomics driven workflow rooted in human large-scale genetic studies to identify putative candidate genes and affected cells associated with cardiovascular traits.</p> <p>&nbsp;</p>

opencc-by-4.0May 2021View details →
zenodo44/100

From genomics to integrative species delimitation? The case study of the Indo-Pacific Pocillopora corals

<p>With the advent of genomics, sequencing thousands of loci from hundreds of individuals now appears feasible at reasonable costs, allowing complex phylogenies to be resolved. This is particularly relevant for cnidarians, for which insufficient data is available due to the small number of currently available markers and obscures species boundaries. Difficulties in inferring gene trees and morphological incongruences further blur the study and conservation of these organisms. Yet, can genomics alone be used to delimit species? Here, focusing on the coral genus <em>Pocillopora</em>, whose colonies play key roles in Indo-Pacific reef ecosystems but have challenged taxonomists for decades, we explored and discussed the usefulness of multiple criteria (genetics, morphology, biogeography and symbiosis ecology) to delimit species of this genus. Phylogenetic inferences, clustering approaches and species delimitation methods based on genome-wide single-nucleotide polymorphisms (SNP) were first used to resolve <em>Pocillopora</em> phylogeny and propose genomic species hypotheses from 356 colonies sampled across the Indo-Pacific (western Indian Ocean, tropical southwestern Pacific and south-east Polynesia). These species hypotheses were then compared to other lines of evidence based on genetic, morphology, biogeography and symbiont associations. Out of 21 species hypotheses delimited by genomics, 13 were strongly supported by all approaches, while six could represent either undescribed species or nominal species that have been synonymised incorrectly. Altogether, our results support (1) the obsolescence of macromorphology (i.e., overall colony and branches shape) but the relevance of micromorphology (i.e., corallite structures) to refine <em>Pocillopora</em> species boundaries, (2) the relevance of the mtORF (coupled with other markers in some cases) as a diagnostic marker of most species, (3) the requirement of molecular identification when species identity of colonies is absolutely necessary to interpret results, as morphology can blur species identification in the field, and (4) the need for a taxonomic revision of the genus <em>Pocillopora</em>. These results give new insights into the usefulness of multiple criteria for resolving <em>Pocillopora</em>, and more widely, scleractinian species boundaries, and will ultimately contribute to the taxonomic revision of this genus and the conservation of its species.</p> <p>&nbsp;</p> <p>This deposit contains the data related to&nbsp;Oury N, No&euml;l C, Mona S, Aurelle D, Magalon H (2023) From genomics to integrative species delimitation? The case study of the Indo-Pacific <em>Pocillopora </em>corals. Mol Phylogenet Evol 107803. doi:10.1016/j.ympev.2023.107803</p> <p>See 0_README.txt for more content details.</p>

opencc-by-4.0May 2023View details →
dryad40/100

Data from: Can the genomics of ecological speciation be predicted across the divergence continuum from host races to species? A case study in Rhagoletis

<p>Studies assessing the predictability of evolution typically focus on short-term adaptation within populations or the repeatability of change among lineages. A missing consideration in speciation research is to determine whether natural selection predictably transforms standing genetic variation within populations into differences between species. Here, we test whether host-related selection on diapause timing anticipates genome-wide differentiation during ecological speciation by comparing ancestral hawthorn and newly formed apple-infesting host races of <i>Rhagoletis pomonella </i>to their sibling species <i>R. mendax</i> that attacks blueberries. The responses of 57,857 single nucleotide polymorphisms in a diapause study on the hawthorn race strongly predicted the direction and magnitude of genomic divergence among the three flies at a field site in Fennville, Michigan, USA. As anticipated, the apple race and <i>R. mendax</i> show parallel changes in the frequencies of putative inversions on three chromosomes associated with the earlier fruiting times of apples and blueberries compared to hawthorns. A diapause experiment on <i>R. mendax</i> revealed compensatory mutations throughout the genome accounting for the earlier eclosion of blueberry, but not apple flies. Thus, a degree of predictability, although not complete, exists in the genomics of diapause across the ecological speciation continuum in <i>Rhagoletis</i>. The generality of this result is placed in the context of other similar systems.</p>

opencc-zeroAug 2020View details →
zenodo40/100

Genome-wide association study identifies RNF123 locus as associated with chronic widespread musculoskeletal pain

<p>The dataset (CWP_GWAS_EU_ANCESTRY_UKB.txt) contains summary statistics for discovery GWAS of chronic widespread pain based on northern Europeans from UK Biobank comprising 6,914 cases of chronic widespread musculoskeletal pain and 242,929 controls. The&nbsp;sensitivity GWAS (CWP_sensitivity_GWAS_EU_ANCESTRY_UKB.txt) dataset contains summary statistics derived from 6,914 cases of chronic widespread musculoskeletal pain and 223,606 controls. Methodological details available here,&nbsp;https://doi.org/10.1101/2020.11.30.20241000&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Full Summary Statistics - TBDAR Genome-to-genome Study

<p><strong>TBDAR_G2G_Full_Summary_Stats.tar.gz:&nbsp;</strong>Full summary statistics (See README for details)</p> <p><strong>Mtb_Human_IDs.txt:&nbsp;</strong>Mapping between M.tb and human sequencing IDs, to faciliate joint analyses.&nbsp;</p> <p><strong>Supple_Data1.csv</strong>: <span lang="EN">G2G associations that meet the significance threshold of 5 &times; 10⁻⁸ </span></p>

opencc-by-4.0May 2023View details →
dryad40/100

Growth traits of a tropical timber species at Southeast Asia, Shorea macrophylla, and scripts for genome wide association study and genomic prediction

<p><em><span>Shorea macrophylla</span></em><span> is a commercially important tropical tree species grown for timber and oil. It is amenable to plantation forestry due to its fast initial growth. Genomic selection (GS) has been used in tree breeding studies to shorten long breeding cycles but has not previously been applied to <em>S. macrophylla</em>. To build genomic prediction models for GS, leaves and growth trait data were collected from a half-sib progeny population of <em>S. macrophylla</em> in Sari Bumi Kusuma forest concession, central Kalimantan, Indonesia. 18037 SNP markers were identified in two ddRAD-seq libraries. Genomic prediction models based on these SNPs were then generated for breast height and total height in the 7th year from planting (D7 and H7). These traits were chosen because of their relatively high narrow-sense genomic heritability and because seven years was considered long enough to assess initial growth. Genomic prediction models were built using 12 methods with the full set of identified SNPs and subsets of 48, 96, and 192 SNPs selected based on the results of a genome-wide association study (GWAS). The GBLUP and RKHS methods gave the highest predictive ability (PA) for D7 and H7 and showed that D7 has an additive genetic architecture while H7 has an epistatic genetic architecture. LightGBM and CNN1D also achieved high PA for D7 with 48 and 96 selected SNPs, and for H7 with 96 and 192 selected SNPs, showing that gradient boosting decision trees and deep learning can be useful in genomic prediction. For almost all methods and both traits, PA was higher when SNPs were selected based on their GWAS P-values than when using the full set of SNPs. These results suggest that GS with GWAS-based SNP selection could be used in <em>S. macrophylla </em>breeding to improve initial growth and reduce genotyping costs for next generation seedlings.</span></p>

opencc-zeroOct 2023View details →
zenodo40/100

Data from: Unravelling cucumber resistance to several viruses via genome-wide association studies highlighted resistance hotspots and new QTLs

<p>The mapping and introduction of sustainable resistance to viruses in crops is a major challenge in modern breeding, especially regarding vegetables. We hence assembled a panel of cucumber elite lines and landraces from different horticultural groups for testing with six virus species. We mapped 18 quantitative trait loci (QTL) with a multiloci genome wide association studies (GWAS), some of which have already been described in the literature. We detected two resistance hotspots, one on chromosome 5 for resistance to the cucumber mosaic virus (CMV), cucumber vein yellowing virus (CVYV), cucumber green mottle mosaic virus (CGMMV) and watermelon mosaic virus (WMV), colocalizing with the RDR1 gene, and another on chromosome 6 for resistance to the zucchini yellowing mosaic virus (ZYMV) and papaya ringspot virus (PRSV) close to the putative VPS4 gene location. We observed clear structuring of resistance among horticultural groups due to plant virus coevolution and modern breeding which have impacted linkage disequilibrium (LD) in resistance QTLs. The inclusion of genetic structure in GWAS models enhanced the GWAS accuracy in this study. The dissection of resistance hotspots by local LD and haplotype construction helped gain insight into the panel&rsquo;s resistance introduction history. ZYMV and CMV resistance were both introduced from different donors in the panel, resulting in multiple resistant haplotypes at same locus for ZYMV, and in multiple resistant QTLs for CMV.</p>

opencc-by-4.0Aug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record