Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,604

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,604 results for “omics”

Learn how ShareScore rates datasets ↗
zenodo40/100

Integrative spatial omics reveals distinct tumor-promoting multicellular niches and immunosuppressive mechanisms in African American and European American patients with TNBC (Spatial Transcriptomic 10X Visium portion)

<p>Racial disparities in triple-negative breast cancer (TNBC) outcomes have been reported. However, the biological mechanisms underlying these disparities remain unclear. We integrated imaging mass cytometry and spatial transcriptomics, to characterize the tumor microenvironment (TME) of African American (AA) and European American (EA) patients with TNBC. The TME in AA patients was characterized by interactions between endothelial cells, macrophages, and mesenchymal-like cells, which were associated with poor patient survival. In contrast, the EA TNBC-associated niche is enriched in T-cells and neutrophils suggestive of an exhaustion and suppression of otherwise active T cell responses. Ligand-receptor and pathway analyses of race-associated niches found AA TNBC to be &ldquo;immune cold&rdquo; and hence immunotherapy resistant tumors, and EA TNBC as &lsquo;inflamed&rsquo; tumors that evolved a distinctive immunosuppressive mechanism. Our study revealed the presence of racially distinct tumor-promoting and immunosuppressive microenvironments in AA and EA patients with TNBC, which may explain the poor clinical outcomes.</p> <p>&nbsp;</p> <p>This dataset contains the 10X Visium Spatial Transcriptomic data of TNBC patients. There are two cohorts.</p> <p>&nbsp;</p> <p><strong>Baylor Scott and White (BSW) cohort</strong>: <strong>10x.visium.tar.gz</strong>, containing 10 patients with TNBC from Baylor Scott and White affiliated Hospital.&nbsp;</p> <p>Each sample is made of Space Ranger processed spot-separated gene expression data (processed to HDF5 AnnData file). There are also H&amp;E images, and spot coordinate files available.&nbsp;</p> <p>&nbsp;</p> <p>For&nbsp;<strong>Georgia validation cohort</strong>, 400 genes used for validation of ESG signatures (associated with BA-Community 1 and WA-Community-1) were obtained and provided by Ritu Aneja's lab. These 400 genes' spot-based expression data across Black and White TNBC patients are provided. See file&nbsp;<strong>georgia.validation.visium.tar.gz</strong>. Expression was normalized by total counts per spot, followed by log-normalization by Giotto.</p> <p>&nbsp;</p> <p>As well in our paper, we integrated a published racial TNBC cohort for deriving some of initial results in the paper. This refers to the Bassiouni et al (Cancer Research) paper in Carpten's group. <strong>GSM_giotto_processed.tar.gz</strong> refers to this dataset, which we deposit here. The data were normalized by Giotto using standard procedure.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Population-scale skeletal muscle single-nucleus multi-omic profiling reveals extensive context specific genetic regulation

<p>Data accompanying the manuscript "Population-scale skeletal muscle single-nucleus multi-omic profiling reveals extensive context specific genetic regulation".</p> <p>Note: For ATAC fragment files, e,caQTL full cis scan summary files, clustering objects, please see the CMDGA portal (https://cmdga.org/search/?searchTerm=stephen-parker%3AVarshney2024)<br>For raw data including fastq files, please see dbGaP repo phs001048.v3.p1</p> <p>Data in this repository includes:</p> <p>Filename: Description</p> <p>1. list of 8,666 genes for which exon-only counts were considered. See methods section "Adjusting RNA counts for overlapping gene annotations" in the manuscript.</p> <p>2. nucleus_sample_cluster_map.tsv: nucleus-sample-cluster map with other QC info.&nbsp;<br># index: nucleus identified syntax &lt;modality&gt;.&lt;batch&gt;.NM.&lt;10X channel&gt;.&lt;barcode&gt;&nbsp;<br># UMAP_1, UMAP_2: UMAP coordinates for visualization<br># modality: rna or atac<br># batch: processing batch identifier<br># hqaa_umi: high quality autosomal alignments (HQAA) for atac nuclei, unique molecular identifier (UMI) for tna&nbsp;<br># fraction_mitochondrial: fraction of reads mapping to the mitochondrial genome<br># cohort: sample cohort<br># tss_enrichment: TSS enrichment for atac nuclei<br># coarse_cluster_name: cluster name</p> <p>3. peaks.tar.gz: snATAC peak features including:<br># consensus-summits.bed: consensus summits along with the cell type that the summits was highest in.<br># narrow peaks in clusters<br># consensus summit feature (summit +- 150bp) identified in each cluster - these were used in GWAS enrichments.</p> <p>4. snrna-cell-type-specific-genes.tsv: Normalized expression scores for genes in each cell-type cluster</p> <p>5. eqtl_permute.tar.gz: Permutation scan eQTL in each cell-type cluster. Columns:&nbsp;<br># variant: syntax &lt;chrom&gt;:&lt;hg38 pos&gt;:&lt;ref&gt;:&lt;alt&gt;<br># effect_allele: effect allele (was the alt allele)<br># other_allele: non-effect allele<br># feature: gene name<br># featureCoordinates_tss: gene TSS<br># p-value: nominal p value<br># beta: slope/beta of the linear regression. Keyed on the alt allele<br># se: standard error of the slope<br># snp: SNP ID<br># strand: gene strand<br># n_variants_tested: number of variants tested for the gene<br># distance_var_pheno: distance of the variant with the gene TSS<br># n_effective_tests: number of effective tests<br># p_beta: beta distribution adjusted p value<br># qvalue: qvalue (Storey)</p> <p>6. caqtl_permute.tar.gz: # Permutation scan caQTL in each cell-type cluster. Columns:&nbsp;<br># variant: syntax &lt;chrom&gt;:&lt;hg38 pos&gt;:&lt;ref&gt;:&lt;alt&gt;<br># effect_allele: effect allele (was the alt allele)<br># other_allele: non-effect allele<br># feature: peak feature coordinates<br># p-value: nominal p value<br># beta: slope/beta of the linear regression. Keyed on the alt allele<br># se: standard error of the slope<br># snp: SNP ID<br># n_variants_tested: number of variants tested for the gene<br># distance_var_pheno: distance of the variant with the gene TSS<br># n_effective_tests: number of effective tests<br># p_beta: beta distribution adjusted p value<br># qvalue: qvalue (Storey)</p> <p>7. eqtl_credible_sets.tar.gz: # eQTL credible set. The file name denotes the egene and the signal hit id. Bed file columns:&nbsp;<br># 1: snp chromosome<br># 2: snp start<br># 3: snp end<br># 4: snp chrom_pos_ref_alt<br># 5: Bayes Factor&nbsp;<br># 6: PIP<br># 7: SNP rsid</p> <p>8. caqtl_credible_sets.tar.gz: # caqtl credible set. The file name denotes the capeak and the signal hit id. Bed file columns:&nbsp;<br># 1: snp chromosome<br># 2: snp start<br># 3: snp end<br># 4: snp chrom_pos_ref_alt<br># 5: Bayes Factor&nbsp;<br># 6: PIP<br># 7: SNP rsid</p> <p>9. cicero_all.tar.gz # Cicero coaccessibility results. Columns<br># Peak 1: Macs2 narrowpeak coordinate for peak 1<br># Peak 2: Macs2 narrowpeak coordinate for peak 2<br># coaccess: Cicero coaccessibility score</p> <p>10. cicero_gene_tss.tar.gz: Cicero coaccessibility results between peak and genes. Macs2 narrow peaks in the TSS+1kb upstream region are assigned that gene name. Columns<br># Cicero coaccessibility results between peak and genes. Macs2 narrow peaks in the TSS+1kb upstream region are assigned that gene name.Columns<br># Peak 1: Macs2 narrowpeak coordinate for peak 1<br># gene_name: Assigned gene<br># Peak 2: Macs2 narrowpeak coordinate for peak 2<br># coaccess: Cicero coaccessibility score<br>## &nbsp;Peak1 is the narrowpeak in the TSS region, peak2 is the distal peak</p> <p>11. mash.tar.gz Mashr results for e/caQTL - lfsr, posterior means and posterior SD for each tested eSNP-eGene, caSNP-caPeak pair.&nbsp;</p> <p>12. cellregmap.tar.gz: Cellregmap results for endothelial nucleus-level eQTL scans.<br>## Persistent genetic effect beta_g was calculated in a simple association model.&nbsp;<br>## An interaction model was fit to test for GxC effect. columns:<br># rho1, g2, e1, and eps2 are variance component measures outputs from CellRegMap corresponding to interaction, genetic, environment and residual variance components.&nbsp;<br># p_nominal: nominal p from cellRegMap<br># kind: model kind in CellRegMap - simple association or interaction<br># beta_g: &nbsp;Persistent genetic effect<br># gene_name: gene name for eQTL or peak feature name for caQTL<br># context: context used either factors (continuous) or subclusters (discrete)<br># snp: index snp for which model is fit. This is the most significant identified snp from our standard e,caQTL scans. chrom-hg38pos-rsid</p> <p><br>13. coloc-eqtl-caqtl.tsv: # Summary of eQTL-caQTL coloc in each cluster. Columns:<br># nsnps: Number of SNPs in the region<br># eqtl_hit: SNP with the highest Bayes factor in the SuSiE eQTL credible set<br># caqtl_hit: SNP with the highest Bayes factor in the SuSiE caQTL credible set<br># PP.H0.abf: Coloc posterior probability for no signal<br># PP.H1.abf: Coloc posterior probability for signal in dataset 1<br># PP.H2.abf: Coloc posterior probability for signal in dataset 2<br># PP.H3.abf: Coloc posterior probability for different signals in datasets 1 and 2<br># PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2<br># idx1: Index of the SuSiE credible set for dataset 1<br># idx2: Index of the SuSiE credible set for dataset 2<br># cluster: cluster name<br># egene: eGene name<br># capeak: caPeak coordinates</p> <p>14. cit-mrs-summary.tsv: &nbsp;Summary from CIT and MR Steiger directionality tests. Columns:<br># cluster: cluster name<br># egene: eGene name<br># capeak: caPeak coordinates<br># eqhit: SNP with the highest Bayes factor in the SuSiE eQTL credible set<br># cahit: SNP with the highest Bayes factor in the SuSiE caQTL credible set<br># p.cit_c_c-e: P value for CIT causal cahit-ca-to-e model<br># q.cit_c_c-e: q value for CIT causal cahit-ca-to-e model<br># p.cit_rc_c-e: P value for CIT reverse-causal eqhit-ca-to-e model&nbsp;<br># q.cit_rc_c-e: value for CIT reverse-causal eqhit-ca-to-e model&nbsp;<br># p.cit_c_e-c: P value for CIT causal eqhit-e-to-ca model<br># q.cit_c_e-c: q value for CIT causal eqhit-e-to-ca model<br># p.cit_rc_e-c: P value for CIT reverse-causal cahit-e-to-ca model&nbsp;<br># q.cit_rc_e-c: q value for CIT reverse-causal cahit-e-to-ca model&nbsp;<br># cit_direction: Direction inferred from CIT &nbsp;<br># correct_causal_direction--ca-to-e: MR Steiger directionality test - is ca-to-e direction correct?<br># correct_causal_direction--e-to-ca: MR Steiger directionality test - is e-to-ca direction correct?<br># sensitivity_ratio--ca-to-e: MR Steiger Sensitivity ratio for ca-to-e model&nbsp;<br># sensitivity_ratio--e-to-ca: &nbsp;MR Steiger Sensitivity ratio for e-to-ca model<br># steiger_test--ca-to-e: MR Steiger directionality test P value for ca-to-e model<br># steiger_test--e-to-ca: MR Steiger directionality test P value for e-to-ca model<br># steiger_q--ca-to-e: MR Steiger directionality test q value for ca-to-e model<br># steiger_q--e-to-ca: MR Steiger directionality test q value for e-to-ca model<br># mrs_direction: Direction inferred from MR Steiger<br># direction: Direction inferred requiring consistent results between CIT and MR Steiger directionality test</p> <p>15. coloc-gwas-eqtl.tsv and<br>16. coloc-gwas-caqtl.tsv # Summary of e/caQTL coloc with GWAS in each cluster. Columns:<br># nsnps: Number of SNPs in the region<br># gwas_hit: SNP with the highest bayes factor in the SuSiE GWAS credible set<br># eqtl_hit: SNP with the highest bayes factor in the SuSiE eQTL credible set<br># caqtl_hit: SNP with the highest bayes factor in the SuSiE caQTL credible set<br># PP.H0.abf: Coloc posterior probability for no signal<br># PP.H1.abf: Coloc posterior probability for signal in dataset 1<br># PP.H2.abf: Coloc posterior probability for signal in dataset 2<br># PP.H3.abf: Coloc posterior probability for different signal in datasets 1 and 2<br># PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2<br># idx1: Index of the SuSiE credible set for dataset 1<br># idx2: Index of the SuSiE credible set for dataset 2<br># cluster: cluster name<br># egene: eGene name<br># capeak: caPeak coordinates<br># p12min: Min prior p12 where the PP H4 &gt; 0.5. Lower this value, more robust is the colocalization<br># trait: GWAS trait name<br># gwas_locus: GWAS locus name for the coloc test - a 250kb left and right flanking genomic window on this SNP was considered for testing coloc between all pairs of GWAS/QTL signals identified in this region &nbsp;<br># traitname: Expanded GWAS trait name<br># variable_type: GWAS type&nbsp;<br># source: Source of GWAS - either UKBB or other study</p> <p>17. supplementary_tables.xlsx: Supplementary tables from the manuscript.<br>Information included in sheets:<br>1. "marker_genes": Marker genes known from literature used to annotate clusters<br>2. "n_nuclei": n pass-QC nuclei per modality-sample-cluster</p> <p>2. "snrna_GO_enrichment": GO term enrichment: matrix of cluster vs top 2 GO terms</p> <p>3. "qtl_scan_info": &nbsp;e/caQTL scan info<br>cluster: cluster<br>ntested_eqtl: N genes tested for eQTL<br>nsig_eqtl: N significant (5% FDR) eGenes<br>n_pheno_pcs_eqtl: N phenotype PCs considered for eQTL<br>ratio_eqtl: Ratio of N eGenes/N genes tested<br>nsig_caqtl: &nbsp;N peaks tested for caQTL<br>ntested_caqtl: N significant (5% FDR) caPeaks<br>n_pheno_pcs_caqtl: N phenotype PCs considered for caQTL<br>ratio_caqtl: Ratio of N caPeaks/N peaks tested<br>nsamples_eqtl: N samples for eQTL<br>nsamples_caqtl: N samples for caQTL</p> <p>4. "gwas_trait_list": GWAS trait info<br>trait: GWAS trait ID<br>traitname: GWAS trait description<br>variable_type: GWAS type. case/control (cc), continuous_irnt=continuous inverse-normal transformed<br>source: GWAS source<br>doi: GWAS study DOI</p> <p>5. "traits_in_ldsc_baseline" - list of annotations included in the baseline model for LDSC</p> <p>6. "gwas_enrichment_in_peaks" GWAS enrichment in cluster peaks (S-LDSC)</p> <p>7. "gwas_enrichment_in_qtl_peaks" GWAS enrichment in QTL peaks (fGWAS) # fGWAS results comparing GWAS enrichment in type 1 annotations<br>CI_lower_ln, estimate_ln, CI_upper_ln: natural log of lower confidence interval, estimate, and upper confidence interval<br>trait: trait id<br>traitname: trait name<br>annotation: annotation<br>sig: 1 if CIs don't overlap 0, otherwise 0</p> <p>8. t2d_gwas_caqtl_coloc and<br>9. t2d_gwas_eqtl_coloc:<br>Summary of e,caQTL coloc with T2D GWAS in each cluster, along with target gene nominations. Columns:<br>nsnps: Number of SNPs in the region<br>gwas_hit: SNP with the highest bayes factor in the SuSiE GWAS credible set<br>eqtl_hit: SNP with the highest bayes factor in the SuSiE eQTL credible set<br>caqtl_hit: SNP with the highest bayes factor in the SuSiE caQTL credible set<br>PP.H0.abf: Coloc posterior probability for no signal<br>PP.H1.abf: Coloc posterior probability for signal in dataset 1<br>PP.H2.abf: Coloc posterior probability for signal in dataset 2<br>PP.H3.abf: Coloc posterior probability for different signal in datasets 1 and 2<br>PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2<br>idx1: Index of the SuSiE credible set for dataset 1<br>idx2: Index of the SuSiE credible set for dataset 2<br>cluster: cluster name<br>egene: eGene name<br>capeak: caPeak coordinates<br>p12min: Min prior p12 where the PP H4 &gt; 0.5. Lower this value, more robust is the colocalization<br>trait: GWAS trait id<br>diamante_gwas_locus: GWAS signal from the DIAMANTE 2018 study. Some signals that our SuSiE runs identified were not present in the original study in which case this column is NA<br>traitname: Expanded GWAS trait name<br>capeak_in_tss: caPeak in TSS + 1kb upstream region of a gene<br>gene_target_standard_cicero: caPeak coaccessible with TSS peak of a gene considering nuclei from all samples for co-accessibility<br>gene_target_allelic_cicero: &nbsp;caPeak coaccessible with TSS peak of a gene considering nuclei from samples homozygous for the caSNP allele associated with increased accessibility<br>gwashit_nominal_egene: gwas_hit nominally associated with these genes nominated in the columns capeak_in_tss, gene_target_standard_cicero, and &nbsp;gene_target_allelic_cicero</p> <p>10. MPRA results for the C2CD4A locus</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

txtools use cases omic references and processed data

<p><strong>Dataset</strong></p> <p>This dataset entry is meant to be downloaded programmatically while rendering the txtools_useCases.Rmd notebooks, to facilitate their replication using the provided genomic references. Processed data is also provided to show ready-to-use examples of data processed by txtools.</p> <p><strong>Abstract</strong></p> <p>We present txtools, an R package that enables the processing, analysis, and visualization of RNA-seq data at the nucleotide-level resolution, seamlessly integrating alignments to the genome with transcriptomic representation. txtools&rsquo; main inputs are BAM files and a transcriptome annotation, and the main output is a table, capturing mismatches,&nbsp; deletions, and the number of reads beginning and ending at each nucleotide in the transcriptomic space. txtools further facilitates downstream visualization and analyses. We showcase, using examples from the epitranscriptomic field, how a few calls to txtools functions can yield insightful and ready-to-publish results. txtools is of broad utility also in the context of structural mapping and RNA:protein interaction mapping. By providing a simple and intuitive framework, we believe that txtools will be a useful and convenient tool and pave the path for future discovery.&nbsp; txtools is available for installation from its GitHub repository at <a href="https://github.com/AngelCampos/txtools">https://github.com/AngelCampos/txtools</a>.&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

quickSparseM: a library for memory- and time-efficient computation on large, sparse matrices with application to omics data

<p>This page contains the code and datasets used in "quickSparseM: a library for memory- and time-efficient computation on large, sparse matrices with application to omics data".</p> <p>File <strong>test_datasets.zip</strong> containes three datasets:</p> <ul> <li><em>D1.RData</em>: scRNA-seq omics data derived from Salcher et. al (2022)</li> <li><em>D2.RData</em>: scRNA-seq omics data derived from Pineda et al. (2024)</li> <li><em>D3.RData</em>: in silico WGS SNP data.</li> </ul> <p>File <strong>test_scripts.zip</strong> containes the code to reproduce the results.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Meta-omics-aided isolation of elusive anaerobic arsenic-methylating soil bacteria

<p>Data pertaining to the manuscript &quot;<strong>Meta-omics-aided isolation of an elusive anaerobic arsenic-methylating soil bacterium&quot;</strong>&nbsp;by Karen Viacava, Jiangtao, Qiao, Andrew Janowczyk, Suresh Poudel, Nicolas Jacquemin, Karin Lederballe Meibom, Him K. Shrestha, Matthew C. Reid, Robert L. Hettich and&nbsp;Rizlan Bernier-Latmani published in ISME journal.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Deciphering colorectal cancer genetics through multi-omic analysis of 100,204 cases and 154,587 controls of European and East Asian ancestries

<p><strong>Colorectal cancer (CRC) is a leading cause of mortality worldwide. We conducted a genome-wide association study meta-analysis of 100,204 CRC cases and 154,587 controls of European and Asian ancestry, identifying 205 independent risk associations, of which 50 were unreported. We performed integrative genomic, transcriptomic and methylomic analyses across large bowel mucosa and other tissues. Transcriptome- and methylome-wide association studies revealed an additional 53 risk associations. We identified 155 high confidence effector genes functionally linked to CRC risk, many of which had no previously established role in CRC. These have multiple different functions, and specifically indicate that variation in normal colorectal homeostasis, proliferation, cell adhesion, migration, immunity and microbial interactions determines CRC risk. Cross-tissue analyses indicated that over a third of effector genes most likely act outside the colonic mucosa. Our findings provide insights into colorectal oncogenesis, and highlight potential targets across tissues for new CRC treatment and chemoprevention strategies.</strong></p> <p><strong>The data submitted here are expression and methylation models with LD reference data for the&nbsp;transcriptome-wide (TWAS), methylome-wide (MWAS) and&nbsp;transcript isoform-wide association study (TIsWAS)&nbsp;as described in the&nbsp;manuscript &quot;Deciphering colorectal cancer genetics through multi-omic analysis of 100,204 cases and 154,587 controls of European and East Asian ancestries&quot;. Details of the methods are presented in the method section and supplementary information file.&nbsp;</strong></p> <p><strong>TWAS analysis&nbsp;</strong></p> <p>Gene expression models for the six in-house expression datasets were generated using the PredictDB v7 pipeline for a total of 1,077 participants. Elastic net model building with 10-fold cross-validation was performed independently for each dataset. The elastic net models for GTEx v8 Colon Transverse were obtained from the PredictDB data repository (<a href="http://predictdb.org/">http://predictdb.org/</a>) and had been generated using the same pipeline. Models were computed using HapMap2 SNPs &plusmn;1Mb from each gene, together with covariate factors estimated using PEER32, clinical covariates when appropriate (age, sex and, where appropriate, case-control status, type of polyp and anatomic location in the colorectum), and three PCs from the individual dataset&rsquo;s SNP genotype data.</p> <p>Transcript-based TWAS analyses (TIsWAS) were likewise performed by using transcript-level data from the SOCCS, BarcUVa-Seq and GTEx Colon Transverse datasets.</p> <p><strong>MWAS analysis&nbsp;</strong></p> <p>Methylation beta values were calculated based on the manufacturer&rsquo;s standard, ranging from 0 to 1. Quality control and data normalization were performed in R using the ChAMP software pipeline for the EPIC and 450K arrays. Briefly, we filtered out failed probes with detection P &gt; 0.02 in &gt;5% of samples, probes with &lt;3 reads in &gt;5% of samples per probe and all non-CpG probes. Samples with failed probes &gt;0.1 were also excluded from downstream analyses. We discarded all probes with SNPs within 10bp of the interrogated CpG (from 1,000 Genomes Project, CEU population)34, and probes that ambiguously mapped to multiple locations in the human genome with up to two mismatches33. We only considered probes mapping to autosomes and those overlapping between the EPIC and the 450K arrays. Normalization was achieved using the Beta MIxture Quantile (BMIQ) method. Per probe methylation models were created using the PredictDB pipeline on the normalized methylation matrix and the genotypes as per TWAS eQTL analysis. To optimize power, we restricted our analysis to 263,341-238,443 (for the 450K array) and 377,678 (for the EPIC array) probes annotated to Islands, Shores and Shelves, and discarded &ldquo;Open Sea&rdquo; regions.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 3. Designing a in Omics in Weed Science: A Perspective from Genomics, Transcriptomics, and Metabolomics Approaches

Figure 3. Designing a metabolomics study. (A) The various approaches for performing a metabolomics experimental study. GC-MS, gas chromatography–mass spectrometry; HILIC-LC-MS/MS, hydrophilic interaction chromatography for liquid chromatography–tandem mass spectrometry; LC-MS/MS, liquid chromatography–tandem mass spectrometry. (B) The general metabolomics workflow. It involves formulating a biological question, setting up an experimental design to test the hypothesis, sample treatment and harvest, metabolite extraction, clean-up, chromatographic separation, identification, statistical validation, and functional interpretation.

opencc-by-4.0Aug 2018View details →
zenodo40/100

Figure 1 in Omics in Weed Science: A Perspective from Genomics, Transcriptomics, and Metabolomics Approaches

Figure 1. Classical systems biology concept and omics organization. The central dogma of molecular biology covers the progressive functionalization of the genotype to the phenotype. The omics techniques track and capture various molecular entities across the biological system.

opencc-by-4.0Aug 2018View details →
zenodo40/100

Dataset: Singular Genomics Systems, Inc. (OMIC) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Single-cell -omics datasets containing a trajectory

<p>Contains datasets to compare trajectory inference methods. See&nbsp;<a href="https://benchmark.dynverse.org">https://benchmark.dynverse.org</a></p> <p>These files can be opened in R using `readRDS` .</p> <p>Each file is a list containing at least:</p> <ul> <li>expression&nbsp;and counts: filtered log<sub>2</sub>&nbsp;normalised expression files and raw counts</li> <li>milestone_network, milestone_percentages and divergence_regions: the trajectory model</li> </ul>

opencc-by-sa-4.0Oct 2018View details →
zenodo40/100

Integrated Omics-Based Discovery of Novel Genes, Secondary Metabolites Clusters, and Small Molecules in Penicillium spp. with Disparate Fungal Isolates

<p><em><span>Penicillium expansum</span></em><span> is a ubiquitous postharvest pathogen of pome fruit that causes blue mold decay of apple fruit while another member of the genus, <em>P. chrysogenum</em><span>,</span><em> </em>is a well-studied saprophyte used for antibiotic and small molecule production. While these two fungi have been investigated individually, the recent discovery of <em>P. chrysogenum </em>hindering <em>P. expansum</em> apple fruit infection has not been well studied. To shed light on this interaction between the two species, we conducted a comparative transcriptomic, metabolomic, and genomic study. Global transcriptional and metabolomic outputs were disparate between the species, nearly identical for the <em>P. chrysogenum </em>isolates, and different between <em>P. expansum </em>isolates. Further, the two <em>P. chrysogenum</em> genomes revealed secondary metabolite gene clusters that differed from <em>P. expansum</em>. This included the absence of an intact patulin gene cluster in <em>P. chrysogenum</em>, which corroborates the metabolomic data regarding the species&rsquo; inability to produce patulin. Additionally, <em>P. expansum </em>virulence gene homologues were identified in <em>P. chrysogenum </em>and were similarly transcriptionally regulated <em>in vitro</em>. Molecules with potential antimicrobial activity, and phytohormones like indole-3-acetic acid (IAA), were detected for the first time in <em>P. expansum</em> while pharmacological compounds like the well-studied antibiotic penicillin G were identified in <em>P. chrysogenum</em> culture supernatants. Our findings provide new omics-based resources that enable the study of small molecule production of interest, the potential of <em>Penicillium</em>-derived antimicrobials for postharvest decay control, and <em>P.</em> <em>expansum&rsquo;s</em> metabolites roles in host-pathogen interactions. </span></p>

opencc-by-4.0Oct 2024View details →
dryad40/100

Data for: Multi-omics analysis identifies symbionts and pathogens of blacklegged ticks (Ixodes scapularis) from a Lyme disease hotspot in southeastern Ontario, Canada

<p>Ticks in the family Ixodidae are recognized as important vectors of zoonoses including Lyme disease (LD), which is caused by spirochete bacteria from the <em>Borreliella</em> (<em>Borrelia</em>) <em>burgdorferi</em> sensu lato (<em>Bbsl</em>) complex. The blacklegged tick (<em>Ixodes scapulars</em>) continues to expand across Canada, creating hotspots of elevated LD risk at the leading edge of its expansion range. Current efforts to understand the risk of pathogen transmission associated with <em>I. scapularis</em> in Canada focus primarily on targeted screens, while variation in the tick microbiome remains poorly understood. Using multi-omics consisting of 16S metabarcoding and ribosome-depleted, whole-shotgun RNA transcriptome sequencing, we examined the microbial communities associated with adult <em>I. scapularis</em> (N = 32), sampled from four tissue types (whole tick, salivary glands, midgut, and viscera) and three geographical locations within an LD hotspot near Kingston, Ontario. The communities consisted of both endosymbiotic and known or potentially pathogenic microbes, including RNA viruses, bacteria, and a <em>Babesia</em> sp. intracellular parasite. We show that β-diversity is significantly higher between individual tick salivary gland and midgut bacterial communities, compared to whole ticks; while linear discriminant analysis (LDA) effect size (LEfSe) determined that the three potentially pathogenic bacteria detected by V4 16S rDNA sequencing were also discriminatory for dissected tissues only, including a <em>Borrelia</em> from the <em>Bbsl</em> complex, <em>Borrelia miyamotoi</em>, and <em>Anaplasma phagocytophilum. </em>Importantly, we find co-infection of <em>I. scapularis</em> by multiple microbes, in contrast to diagnostic protocols for LD, which typically focus on infection from a single pathogen of interest (<em>B. burgdorferi</em> sensu stricto).</p>

opencc-zeroNov 2022View details →
zenodo40/100

Extended data: Tissue-specific multi-omics analysis of atrial fibrillation

<p>Summary statistics and result repository for the publication Tissue-specific multi-omics analysis of atrial fibrillation:</p> <p>Assum, I., Krause, J., Scheinhardt, M.O.&nbsp;<em>et al.</em>&nbsp;Tissue-specific multi-omics analysis of atrial fibrillation.&nbsp;<em>Nat Commun&nbsp;</em><strong>13,&nbsp;</strong>441 (2022). https://doi.org/10.1038/s41467-022-27953-1</p> <p>For the related source code, see https://doi.org/https://doi.org/10.5281/zenodo.5094276 or&nbsp;https://github.com/heiniglab/symatrial.</p> <p>Ines Assum<sup>1,2,&dagger;</sup>, Julia Krause<sup>3,4,&dagger;</sup>, Markus O. Scheinhardt<sup>5</sup>, Christian M&uuml;ller<sup>3,4</sup>, Elke Hammer<sup>6,7</sup>, Christin S. B&ouml;rschel<sup>4,8</sup>, Uwe V&ouml;ker<sup>6,7</sup>, Lenard Conradi<sup>9</sup>, Bastiaan Geelhoed<sup>4,8,10</sup>, Tanja Zeller<sup>3,4,</sup>*, Renate B. Schnabel<sup>4,8,</sup>*, Matthias Heinig<sup>1,2,11,</sup>*</p> <p><sup>&dagger; </sup>,*&nbsp;These authors contributed equally.</p> <p><sup>&nbsp;1</sup> Computational Health Center, Helmholtz Zentrum M&uuml;nchen Deutsches Forschungszentrum f&uuml;r Gesundheit und Umwelt (GmbH), Neuherberg, Germany.<br> <sup>&nbsp;2</sup> Department of Informatics, Technical University Munich, M&uuml;nchen, Germany.<br> <sup>&nbsp;3</sup> University Center of Cardiovascular Science, University Heart and Vascular Center Hamburg, Hamburg, Germany.<br> <sup>&nbsp;4</sup> Partner site Hamburg/Kiel/L&uuml;beck, DZHK (German Center for Cardiovascular Research), Hamburg, Germany.<br> <sup>&nbsp;5</sup> Institute of Medical Biometry and Statistics, University of L&uuml;beck, L&uuml;beck, Germany.<br> <sup>&nbsp;6</sup> Interfaculty Institute for Genetics and Functional Genomics, University Medicine Greifswald, Greifswald, Germany.<br> <sup>&nbsp;7</sup> Partner site Greifswald, DZHK (German Center for Cardiovascular Research), Greifswald, Germany.<br> <sup>&nbsp;8</sup> Department of Cardiology, University Heart and Vascular Center Hamburg, Hamburg, Germany.<br> <sup>&nbsp;9</sup> Department of Cardiovascular Surgery, University Heart and Vascular Center Hamburg, Hamburg, Germany.<br> <sup>10&nbsp;</sup>Department of Cardiology, University of Groningen, University Medical Center Groningen, Groningen, Netherlands.<br> <sup>11</sup>Partner site Munich, DZHK (German Center for Cardiovascular Research), Munich, Germany.</p> <p>&nbsp;</p> <p>ABSTRACT:</p> <p>Genome-wide association studies (GWAS) for atrial fibrillation (AF) have uncovered numerous disease-associated variants. Their underlying molecular mechanisms, especially consequences for mRNA and protein expression remain largely elusive. Thus, refined multi-omics approaches are needed for deciphering the underlying molecular networks. Here, we integrate genomics, transcriptomics, and proteomics of human atrial tissue in a cross-sectional study to identify widespread effects of genetic variants on both transcript (cis-eQTL) and protein (cis-pQTL) abundance. We further establish a novel targeted transQTL approach based on polygenic risk scores to determine candidates for AF core genes. Using this approach, we identify two trans-eQTLs and five trans-pQTLs for AF GWAS hits, and elucidate the role of the transcription factor NKX2-5 as a link between the GWAS SNP rs9481842 and AF. Altogether, we present an integrative multi-omics method to uncover trans-acting networks in small datasets and provide a rich resource of atrial tissue-specific regulatory variants for transcript and protein levels for cardiovascular disease gene prioritization.</p> <p>This version contains a reference file identifying effect alleles for all QTL results and adds additional genotype and allele frequency information for all QTL SNPs.&nbsp;</p> <p>TABLE OF CONTENTS:</p> <ul> <li>Reference for effect alleles<br> <em>map_AFHRI_B_effect_alleles.txt</em></li> <li>Reference for genotype and allele frequencies (derived using PLINK) <ul> <li><em>genotype_allele_frequencies_eQTL_SNPs.txt</em></li> <li><em>genotype_allele_frequencies_pQTL_SNPs.txt</em></li> <li><em>genotype_allele_frequencies_resQTL_SNPs.txt</em></li> </ul> </li> <li>Single-omic <em>cis</em>-QTL results <ul> <li><em>cis</em>-eQTLs (all pairs, incl. LD clump info)<br> <em>eQTL_right_atrial_appendage_allpairs_clump.txt</em></li> <li><em>cis</em>-pQTLs (all pairs, incl. LD clump info)<br> <em>pQTL_right_atrial_appendage_allpairs_clump.txt</em></li> <li><em>cis</em>-res eQTLs (all pairs, incl. LD clump info)<br> <em>res_eQTL_right_atrial_appendage_allpairs_clump.txt</em></li> <li><em>cis</em>-res pQTLs (all pairs, incl. LD clump info)<br> <em>res_pQTL_right_atrial_appendage_allpairs_clump.txt</em></li> <li><em>cis</em>-ratioQTLs (all pairs, incl. LD clump info)<br> <em>ratioQTL_right_atrial_appendage_allpairs_clump.txt</em></li> </ul> </li> <li>Functional <em>cis</em>-QTL categories and eQTL/pQTL overlap: <ul> <li>All eQTLs, pQTLs, res eQTLs, res pQTLs and ratioQTLs for all SNP-gene pairs with a significant eQTL and pQTL (FDR&lt;0.05)<br> <em>Fig2a_source_data_Shared_eQTL_pQTL_clump.txt</em></li> <li>All eQTLs, pQTLs, res eQTLs, res pQTLs and ratioQTLs for all SNP-gene pairs with a significant eQTL but no pQTL (FDR&lt;0.05)<br> <em>Fig2b_source_data_Independent_eQTL_clump.txt</em></li> <li>All eQTLs, pQTLs, res eQTLs, res pQTLs and ratioQTLs for all SNP-gene pairs with no eQTL but a significant&nbsp;pQTL (FDR&lt;0.05)<br> <em>Fig2c_source_data_Independent_pQTL_clump.txt</em><span>&nbsp;</span></li> </ul> </li> <li>QTS rankings and enrichment results <ul> <li>eQTS rankings and enrichments<br> <em>TableS6_source_data_eQTS_ranking.txt<br> TableS7_source_data_eQTS_GSEA_results.txt</em></li> <li>pQTS rankings and enrichments<br> <em>TableS8_source_data_pQTS_ranking.txt<br> TableS9_source_data_pQTS_GSEA_results.txt</em></li> </ul> </li> <li><em>Trans</em>-QTLs<br> all tested pairs including <em>trans</em>-pQTLs for <em>trans</em>-eQTLs and <em>trans</em>-eQTLs for <em>trans</em>-pQTLs<br> <em>Table2_source_data_Trans-QTL_results.txt</em></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Paired omics Data Platform projects

<p>The Paired Omics Data Platform is a community-based initiative standardizing links between genomic and metabolomics data in a computer readable format to further the field of natural products discovery. The goals are to link molecules to their producers, find large scale genome-metabolome associations, use genomic data to assist in structural elucidation of molecules, and provide a centralized database for paired datasets. This dataset contains the&nbsp;projects in&nbsp;<a href="http://pairedomicsdata.bioinformatics.nl/">http://pairedomicsdata.bioinformatics.nl/</a>.</p> <p>The JSON documents adhere to the&nbsp;<a href="http://pairedomicsdata.bioinformatics.nl/schema.json">http://pairedomicsdata.bioinformatics.nl/schema.json</a>&nbsp;JSON schema.</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Unsupervised neural network for single cell Multi-omics INTegration (UMINT): An application to health and disease

<p>This dataset repository corresponds to the project&nbsp;Unsupervised neural network for single cell Multi-omics INTegration (UMINT): An application to health and disease.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Rare Variants Prioritized by the Multio-omic Watershed Model

<p>These datasets include genomic annotations included in the Multi-omic Watershed model, trained using data from 1,319 individuals from the Multi-Ethnic Study of Atherosclerosis (MESA) cohort. The resulting rare genetic variants prioritized by the model were provided with their posteriors in each omic dimension (RNA expression, methylation, splicing, and protein expression).&nbsp;</p> <p>For details, please refer to&nbsp;https://www.biorxiv.org/content/10.1101/2022.09.07.507008v1.abstract&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Multi-Omics Visible Drug Activity Prediction with a Biologically Informed Neural Network Model

<p>Drug discovery is a challenging task, it takes several years for a drug to be introduced on the market, with most of<br> the studied drugs not even passing the first phase. The understanding of the mechanisms influencing response to drugs<br> can reduce failures and accelerate drug development. Virtual drug screening, based on Machine Learning models, is a<br> promising field for the prediction of the outcome of a treatment. However, the complex relationships between the features<br> learned by these models are still poorly understood and not easy to interpret.<br> We have designed a Neural Network model for drug sensitivity prediction that leverages a Visible Neural Network, an<br> easily interpretable model, due to its biologically informed nature. The trained model can be inspected to study which<br> biological processes were fundamental for the prediction and to identify the drug properties that affect sensitivity. It<br> combines multi-omics data from various types of tumor tissues and drug representations based on molecular descriptors.<br> The mechanisms learned from the network can also be exploited to find candidate drugs for synergy to predict the effect<br> of combined therapies. We consider the unbalanced nature of public drug screening datasets and show that our model<br> outperforms state-of-the-art visible machine learning models.</p>

opencc-by-4.0Dec 2022View details →
dryad40/100

Discovery of sparse, reliable omic biomarkers with Stabl

<p><span>Adoption of high-content omic technologies in clinical studies, coupled with computational </span><span>methods, have yielded an abundance of candidate biomarkers. However, translating such find</span><span>ings into bona fide clinical biomarkers remains challenging.</span> <span>To facilitate this process, we </span><span>introduce Stabl, a general machine learning framework that identifies a sparse, reliable set </span><span>of biomarkers by integrating noise injection and a data-driven signal-to-noise threshold into </span><span>multivariable predictive modeling.</span> <span>Evaluation of Stabl on synthetic datasets and five inde</span><span>pendent clinical studies demonstrates improved biomarker sparsity and reliability compared to </span><span>commonly used sparsity-promoting regularization methods while maintaining predictive per</span><span>formance; it distills datasets containing 1,400 to 35,000 features down to 4 to 34 candidate </span><span>biomarkers. Stabl extends to multi-omic integration tasks, enabling biological interpretation of </span><span>complex predictive models, as it hones in on a shortlist of proteomic, metabolomic, and cyto</span><span>metric events predicting labor onset, microbial biomarkers of preterm birth, and a pre-operative </span><span>immune signature of post-surgical infections.</span></p>

opencc-zeroOct 2023View details →
dryad40/100

Data for: Multi-omics analysis identifies symbionts and pathogens of blacklegged ticks (Ixodes scapularis) from a Lyme disease hotspot in southeastern Ontario, Canada

Open the record for dataset details and reuse information.

publicNov 2022View details →
dryad40/100

Discovery of sparse, reliable omic biomarkers with Stabl

Open the record for dataset details and reuse information.

publicOct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record