Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

650

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

650 results for “multiomics”

Learn how ShareScore rates datasets ↗
zenodo44/100

Development of Multiomics in situ Pairwise Sequencing (MiP-Seq) for Single-cell Resolution Multidimensional Spatial Omics

<p>The original data used in the article:&nbsp;Development of Multiomics in situ Pairwise Sequencing (MiP-Seq) for Single-cell Resolution Multidimensional Spatial Omics</p> <p>Delineating the spatial multiomics landscape will pave the way to understanding the molecular basis of physiology and pathology. However, current spatial omics technology development is still in its infancy. Here, we developed a high-throughput targeted in situ sequencing strategy, multiomics in situ pairwise sequencing (MiP-Seq), to efficiently decipher multiplexed DNAs, RNAs, proteins, and small biomolecules at subcellular resolution. MiP-Seq simultaneously sequenced the dual barcode base of padlock probes, dramatically increasing the detection capacity to 10N by N rounds of sequencing. We delineated spatial gene profiles in the hypothalamus using MiP-Seq. Moreover, MiP-Seq was unitized to detect tumor gene mutations and allele-specific expression of parental genes and to differentiate sites with and without the m6A RNA modification at specific sites. MiP-Seq was combined with in vivo Ca2+ imaging and Raman imaging to obtain a spatial multiomics atlas correlated to neuronal activity and cellular biochemical fingerprints. Importantly, we proposed a &ldquo;signal dilution strategy&rdquo; to resolve the crowded signals that challenge the applicability of in situ sequencing. Together, our method improves spatial multiomics and precision diagnostics, and facilitates analyzing cell function in connection with gene profiles.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Additional data: Longitudinal single-cell multiomic atlas of high-risk neuroblastoma reveals chemotherapy-induced tumor microenvironment rewiring

<p>This repository provides additional data for the manuscript titled "Longitudinal single-cell multiomic atlas of high-risk neuroblastoma reveals chemotherapy-induced tumor microenvironment rewiring", currently under revision at Nature Genetics. The primary data cohort has been deposited in the HTAN data portal. This repository includes processed 10x Xenium spatial transcriptomic data for six TH-MYCN mice (three chemotherapy-treated and three treatment-naive) as well as processed scRNA-seq data for CHLA15 and CHLA20 neuroblastoma (NBL) cells. The scRNA-seq data includes mono-cultured, co-cultured cells with THP-1 macrophages, and co-culture cells treated with Afatinib/CRM197.&nbsp;&nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo44/100

PBMC CITE-seq, 10x Multiome, and TEA-seq multiomic datasets from Swanson, et al. eLife (2021).

<p>Assembled multiomic datasets from Swanson, et al.&nbsp;<em>Simultaneous trimodal single-cell measurement of transcripts, epitopes, and chromatin accessibility using TEA-seq</em>.&nbsp;eLife 2021;10:e63632&nbsp;DOI:&nbsp;<a href="https://doi.org/10.7554/eLife.63632">10.7554/eLife.63632</a>&nbsp;</p> <p>Datasets were assembled from the GEO repository:&nbsp;<a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE158013">GSE158013</a>&nbsp;</p> <p>Data are provided as SeuratObjects stored in separate .rds files using the saveRDS() function in R.</p> <p>MuData files for use with Python tools (muon, scanpy, scvitools, et al.) will be added soon.</p> <p>Contact Lucas Graybuck (lucasg at alleninstitute dot org) if there are problems with these datasets.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Two metabolomics data sets (mouse kidney, mouse plasma), generated for the publication Bignon et al., 2023: "Multiomics reveals multilevel control of renal and systemic metabolism by the renal tubular circadian clock".

<p><strong>Publication: </strong>Bignon Y, Wigger L, Ansermet C, Weger BD, Lagarrigue S, Centeno G, Durussel F, G&ouml;tz L, Ibberson M, Pradervand S, Quadroni M, Weger M, Amati F, Gachon F, Firsov D. Multiomics reveals multilevel control of renal and systemic metabolism by the renal tubular circadian clock. J Clin Invest. 2023 Mar 2:e167133. doi: 10.1172/JCI167133. Epub ahead of print. PMID: 36862511.</p> <p>&nbsp;</p> <p><strong>Abstract: </strong> Circadian rhythmicity in renal function suggests rhythmic adaptations in renal metabolism. To decipher the role of the circadian clock in renal metabolism, we studied diurnal changes in renal metabolic pathways using integrated transcriptomic, proteomic, and metabolomic analysis performed on control mice and mice with inducible deletion of the circadian clock regulator Bmal1 in the renal tubule (cKOt). With this unique resource, we demonstrated that ~30% RNAs, ~20% proteins and ~20% metabolites are rhythmic in kidneys of control mice. Several key metabolic pathways including NAD+ biosynthesis, fatty acid transport, carnitine shuttle,and b-oxidation displayed impairments in kidneys of cKOt, resulting in a perturbed mitochondrial activity. Carnitine reabsorption from the primary urine was one of the most impacted processes with a ~50% reduction in plasma carnitine levels and a parallel systemic decrease in tissues carnitine content. This suggests that the circadian clock in the renal tubule controls both kidney and systemic physiology.</p> <p>&nbsp;</p> <p><strong>This record contains two separate mass-spectrometry metabolomics data sets associated with this study:</strong></p> <ol> <li>Metabolic profile of renal tubules, MS/MS data, Metabolon, Morrisville, NC (N=60)</li> <li>Metabolic profile of blood plasma, MS/MS data, Biocrates, Innsbruck, Austria (N=60)</li> </ol> <p>For each data set, original data as received from the platforms and processed data as used in the data analysis are provided. Preprocessing of kidney data included removal of metabolites with more than 80% missing data values, median normalization, imputation and glog2 transformation. Preprocessing of plasma data included filtering of metabolites with any missing data and log2 transformation. Details of data processing are available in the STAR*methods of the publication.</p> <p>&nbsp;</p> <p><strong>Data sets in other repositories associated with the same study:</strong></p> <p>Additional data sets (transcriptomics, proteomics) pertaining to the same&nbsp;study have been deposited in public repositories:</p> <ul> <li>Gene Expression Omnibus (NCBI GEO), GSE216252</li> <li>PRIDE Archive (EMBL-EBI), PXD036803</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Data for: Heat induces multiomic and phenotypic stress propagation in zebrafish embryos

<p>This contains the data for the manuscript Feugere et al., &quot;Heat induces multiomic and phenotypic stress propagation in zebrafish embryos&quot; (2023). Zebrafish embryos were exposed to thermal stress (&quot;TS&quot;) and stress metabolites (&quot;SM&quot;) released by heat-stressed conspecifics in a two-way factorial design (&quot;TSxSM&quot;). The folder includes raw molecular data (cortisol levels, HSP70 protein levels, and gene expression acquired with LAMP and RNA-seq) and raw phenotypic data (morphology, hatching, survival, and behaviour) of zebrafish <em>Danio rerio&nbsp;</em>at 1 day and 4 days of development.</p> <p>The .csv files contain all quantitative data, whilst the .tab files contain the gene count data required for gene expression analysis. The data were analysed in R using the code shared in the &quot;TSxSM2.stats.Rmd&quot; file. The &quot;Metadata&quot; document provides the reader with an extensive description of each file.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Linking regulatory variants to target genes by integrating single-cell multiome methods and genomic distance

<p>The below data are associated with our paper entitled "Linking regulatory variants to target genes by integrating single-cell multiome methods and genomic distance."</p> <p>1) SNP-gene link predictions generated by pgBoost and existing methods SCENT (Sakaue et al. 2024 <em>Nat Genet</em>), Signac (Stuart et al. 2021 <em>Nat Methods</em>), ArchR (Granja et al. 2021 <em>Nat Genet</em>), and Cicero (Pliner et al. 2018 <em>Mol Cell</em>).</p> <p><strong>pgBoost_scores.tsv.gz </strong>contains linking predictions made by pgBoost.</p> <p><strong>constituent_method_scores.tsv.gz</strong> contains linking predictions made by constituent methods.</p> <p><em><span>**NOTE: promoters (+/- 1kb from TSS) and candidate links &gt;500kb are excluded from linking predictions (see manuscript)**</span></em></p> <p>Linking scores and percentiles are reported for each method (pgBoost score, SCENT FDR, Signac correlation, ArchR correlation, Cicero co-accessibility). Rank percentiles are computed as: 1 - (rank / n). When multiple links receive the same score, they are assigned the percentile of the top rank. Links unscored by each method (denoted by zeros* in the linking score column) are assigned a percentile equivalent to the percent of links unscored by the focal method. See the Methods section of the paper for further details on computing linking scores and summarizing scores across cell types and data sets.</p> <p>*Candidate links tested and assigned a co-accessibility of zero by the Cicero method are given a score of 1e-100 in the "Cicero" column to distinguish between unscored candidate links and candidate links assigned a partial correlation of zero (see Pliner et al. 2018 <em>Mol Cell</em>).</p> <p><em>NOTE: The predictions associated with this release (version 2) were generated using an expanded set of data sets, an expanded training set, and corrected TSS coordinates.</em></p> <p>2) GWAS-derived evaluation SNP-gene link evaluation set.</p> <p><strong>gwas_evaluation.tsv</strong>: GWAS-derived evaluation SNP-gene link evaluation set. Column 1 provides SNP coordinates in the format &lt;chr-start-end&gt;. This evaluation framework was proposed by Weeks et al. 2024 <em>Nature Genetics</em> based on fine-mapping results from Kanai et al. <em>medrxiv</em>&nbsp;(see Methods: <em>Evaluation data sets</em> of Dorans et al.). "True" links (gold = 1) are non-coding variants fine-mapped to a focal trait (PIP &gt; 0.1) with a coding variant for exactly one candidate gene within 1 Mb&nbsp;attaining PIP &gt; 0.5 for the same trait. "False" links (gold = 0) are candidate SNP-gene pairs involving a SNP with a "true" link. This file of SNP-gene links was adapted from credible set-gene links <a href="https://github.com/Deylab999/GWAS_benchmark_IGVF/blob/bb91d08cc02d59cdd829eb1430569057ac26c5fe/V2G/ENCODE_E2G_2023/UKBiobank.ABCGene.anyabc.tsv">here</a> (the "truth" column defines true/false links) by identifying SNPs with PIP &gt; 0.1 within each credible set-gene link.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Accompanying dataset for: A Multi-scale, Multiomic Atlas of Human Normal and Follicular Lymphoma Lymph Nodes

<p>This dataset accompanies the manuscript titled &ldquo;A Multi-scale, Multiomic Atlas of Human Normal and Follicular Lymphoma Lymph Nodes&rdquo;, A. Radtke et al., bioRxiv, 2022. [<a href="https://doi.org/10.1101/2022.06.03.494716">doi: 10.1101/2022.06.03.494716</a>]</p> <p>The&nbsp;dataset contains the&nbsp;processed scRNA-seq information from human lymph nodes,&nbsp;both normal and from Follicular Lymphoma (FL) patients&nbsp;analyzed in this work as a Seurat object. The scRNA-seq information was saved in the rds format for viewing and analysis using the R programming language (to load it in R: <em>scrna_seq_data &lt;- readRDS(&quot;scRNA_seq_data_object.rds&quot;)</em>).</p> <p>Additionally, the dataset contains comma-separated-value tables describing human lymph nodes, both normal and from Follicular Lymphoma (FL) patients. The files are formatted using the anatomical structures (AS), cell types (CT), and biomarkers (B), ASCT+B format defined by the Human BioMolecular Atlas Program (HuBMAP) for use with the <a href="http://hubmapconsortium.github.io/ccf-asct-reporter/">Reporter visualization tool</a>.&nbsp;Details on the structure of ASCT+B tables and the Reporter tool can be found in the <a href="https://doi.org/10.5281/zenodo.5944386">standard operating procedure</a> authored by the ASCT+B working group.&nbsp;</p> <p><strong>ASCT+B Table Details</strong></p> <p>In support of a human reference atlas (Regev et al., 2017; Snyder et al., 2019), the Human BioMolecular Atlas Program (HuBMAP) is creating machine readable tables that catalog the anatomical structures (AS), cell types (CT), and biomarkers (B) found in human organs (B&ouml;rner et al., 2021). ASCT+B tables facilitate data integration across multimodal assays and support comparisons between normal and diseased tissues. In addition, they are readily visualized with the <a href="http://hubmapconsortium.github.io/ccf-asct-reporter/">ASCT+B Reporter</a>, a web based tool.</p> <p><br> For these reasons, we created 10 ASCT+B tables from the datasets included in our study. To construct these tables, we used the <a href="https://doi.org/10.48539/HBM573.SHCQ.259">Lymph Node v1.1 ASCT+B table</a> as a starting point. The presence or absence of anatomical structures was determined by visual inspection of images and quantitative image analysis of cellular communities. Certain anatomical structures were absent from the excisional biopsies of FL patients e.g., capsule, medulla, hilum, etc. In contrast, the lack of primary follicles, mantle zones, polarized germinal centers (GC), and negligible interfollicular cortex and paracortex in FL LNs reflects changes arising from malignancy. Cell types were defined based on gene biomarkers from bulk and single cell RNA sequencing (RNA-seq) and protein biomarkers from the highly multiplexed imaging method, IBEX (Radtke et al., 2022; Radtke et al., 2020). Whenever possible, cell types captured across assays were defined by both gene and protein biomarkers. However, several cell types were only profiled by bulk RNA-seq, scRNA-seq, or IBEX imaging. In these instances, only assay-specific biomarkers are included in the ASCT+B tables. Whenever possible, we used agreed upon ontology terms to define cell types; however, our study identified several unique cell types not included in ontology databases such as DC-SIGN+ follicular dendritic cells (FDCs). Furthermore, the Reporter does not allow visualization of similar cell types (DC-SIGN- FDCs versus DC-SIGN+ FDCs) in the same anatomical structure if a shared Cell Ontology (CL) identifier is used (FDC: CL:0000442). In these instances, we removed the CL term to allow the Reporter to display the various subpopulations discovered in this study. Cell types were placed in their respective anatomical structures using domain knowledge, visual inspection of images, and quantitative image analysis.</p> <p><strong>Reporter Usage Instructions</strong></p> <ul> <li>Visualizing an individual ASCT+B table: <ol> <li>Go to <a href="https://hubmapconsortium.github.io/ccf-asct-reporter/">Reporter</a></li> <li>Launch Playground</li> <li>Click on Upload tab</li> <li>Attach CSV final of ASCT+B table</li> <li>Use the toolbars on the left to adjust display. Typical parameters include: Tree Height (1400),Tree width (1000), Bimodal Distance X (500), and Bimodal Distance Y (50).&nbsp;</li> <li>Toggle between gene and protein biomarkers by clicking drop down menu under Biomarkers tab on left-side of screen.</li> </ol> </li> <li>Comparing non-FL and FL tables to the Lymph Node v1.1 ASCT+B table: <ol> <li>Go to <a href="https://hubmapconsortium.github.io/ccf-asct-reporter/">Reporter</a></li> <li>Select &ldquo;go to visualization&rdquo; to compare new tables to a master table for lymph node</li> <li>Click check box next to lymph node and select version of published master table v1.1</li> <li>Click submit</li> <li>Click compare button at top right tool bar</li> <li>Attach CSV file of non-FL and FL ASCT+B tables&nbsp;</li> <li>Pick colors&nbsp;</li> <li>Go to bottom of panel and click add</li> <li>Click compare</li> <li>Adjust settings for tree height, tree width, bimodal distance x, bimodal distance y, ontology ID on or off, biomarker type (gene or protein), etc.</li> </ol> </li> </ul> <p><strong>References</strong></p> <ul> <li>B&ouml;rner, K., Teichmann, S.A., Quardokus, E.M., Gee, J.C., Browne, K., Osumi-Sutherland, D., Herr, B.W., Bueckle, A., Paul, H., Haniffa, M., et al. (2021). Anatomical structures, cell types and biomarkers of the Human Reference Atlas. Nature Cell Biology 23, 1117-1128.</li> <li>Radtke, A.J., Chu, C.J., Yaniv, Z., Yao, L., Marr, J., Beuschel, R.T., Ichise, H., Gola, A., Kabat, J., Lowekamp, B., et al. (2022). IBEX: an iterative immunolabeling and chemical bleaching method for high-content imaging of diverse tissues. Nature Protocols.</li> <li>Radtke, A.J., Kandov, E., Lowekamp, B., Speranza, E., Chu, C.J., Gola, A., Thakur, N., Shih, R., Yao, L., Yaniv, Z.R., et al. (2020). IBEX: A versatile multiplex optical imaging approach for deep phenotyping and spatial analysis of cells in complex tissues. Proc Natl Acad Sci U S A 117, 33455-33465.</li> <li>Regev, A., Teichmann, S.A., Lander, E.S., Amit, I., Benoist, C., Birney, E., Bodenmiller, B., Campbell, P., Carninci, P., Clatworthy, M., et al. (2017). The Human Cell Atlas. Elife 6.</li> <li>Snyder, M.P., Lin, S., Posgai, A., Atkinson, M., Regev, A., Rood, J., Rozenblatt-Rosen, O., Gaffney, L., Hupalowska, A., Satija, R., et al. (2019). The human body at cellular resolution: the NIH Human Biomolecular Atlas Program. Nature 574, 187-192.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Multiomics Analysis of the mdx/mTR Mouse Model of Duchenne Muscular Dystrophy

<p>This dataset contains RNA sequencing, proteomics, metabolomics, lipidomics,&nbsp;primer sequence, and Ingenuity Pathway Analysis&nbsp;data for a study evaluating differences between the transcriptome, proteome, metabolome, and lipidome of lower limb muscles of mdx/mTR (Jackson Labs strain 023535) and wild-type C57BL/6J (Jackson Labs strain 000664) mice. Updated from previous submission,&nbsp;DOI: 10.5281/zenodo.3370799.</p>

opencc-by-nc-4.0Aug 2019View details →
dryad40/100

The multiomic landscape of epidemiological factors contributing to preterm birth in low- and middle-income countries

<p><span></span></p> <p><span></span></p> <p><span>Preterm birth (PTB) is the leading cause of death in children under five, yet comprehensive studies are hindered by its multiple complex etiologies. Epidemiological associations between PTB and maternal characteristics have been previously described. This work employed multiomic profiling and multivariate modeling to investigate the biological signatures of these characteristics. Maternal covariates were collected during pregnancy from 13,841 pregnant women across five sites. Plasma samples from 231 participants were analyzed to generate proteomic, metabolomic, and lipidomic datasets. Machine learning models showed robust performance for the prediction of PTB (AUROC=0.70), time-to-delivery (<em>r</em>=0.65), maternal age (<em>r</em>=0.59), gravidity (<em>r</em>=0.56), and BMI (<em>r</em>=0.81). Time-to-delivery biological correlates included fetal-associated proteins (e.g., ALPP, AFP, PGF) and immune proteins (e.g., PD-L1, CCL28, LIFR). Maternal age negatively correlated collagen COL9A1; gravidity with endothelial NOS and inflammatory chemokine CXCL13; and BMI with leptin and structural protein FABP4. These results provide an integrated view of epidemiological factors associated with PTB and identify biological signatures of clinical covariates impacting this disease. </span></p>

opencc-zeroMay 2023View details →
zenodo40/100

Source data for paper "Mapping disease regulatory circuits at cell-type resolution from single-cell multiomics data"

<p>Sample-paired scRNA-seq and scATAC-seq data collected from human&nbsp;peripheral blood mononuclear cells&nbsp;with&nbsp;<em>Staphylococcus aureus</em><em> </em>infection.&nbsp;ScATAC-seq data collected from human&nbsp;peripheral blood mononuclear cells&nbsp;with <em>COVID-19</em>&nbsp;infection.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
dryad40/100

The multiomic landscape of epidemiological factors contributing to preterm birth in low- and middle-income countries

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad40/100

Data from: Multiomics inform invasion risks under global climate change

Open the record for dataset details and reuse information.

publicNov 2024View details →
zenodo36/100

Multiomics analyses reveal a Type III-associated immune response in immunotherapy-induced toxicity in melanoma and identifies potential new therapeutic targets

<p>Immune checkpoint inhibitors (ICIs) are standard-of-care for the treatment of advanced melanoma, but their use is limited by immune-related adverse events (irAEs). Proteomic analyses and multiplex cytokine/chemokine assays from serum at baseline and at irAEs onset in melanoma patients indicated aberrant T-cell activity with differential expression of Type I and III immune signatures. This was in line with an increase in the proportions of monocytes and decrease of IL-17A-producing CD4+ T-cells in the peripheral blood using single cell RNA sequencing. Multiplex immunohistochemistry and spatial transcriptomics on ICI-induced skin rash and colitis showed an increase in the proportion of CD4+ T-cells with IL-17A expression.&nbsp;&nbsp;</p> <p>Anti-IL-17A mAbs were administered in two patients with myocarditis, colitis and skin rash with resolution of the irAE. This study highlights the potential role of Type III CD4+ T-cells in irAEs development and provides proof-of-principle evidence for a corresponding clinical trial using anti-IL17A for treating irAEs.</p> <p>Code to generate the figures are located at https://github.com/pcheng84/AE_analysis/</p> <p>&nbsp;</p>

opengpl-3.0-or-laterNov 2023View details →
zenodo36/100

A multiomic characterization of the leukemia cell line REH using short- and long-read sequencing

<p>This is a public repository containing secondary datasets described in the publication <a href="https://doi.org/10.26508/lsa.202302481">"A multiomic characterization of the leukemia cell line REH using short- and long-read sequencing"</a>. Primary data for this project are available at NCBI/SRA under the BioProject accession numbers PRJNA600820 and PRJNA834955, and include the following sequencing datasets:</p> <p>REH cell line:</p> <ul> <li>PacBio WGS</li> <li>ONT Ultralong WGS</li> <li>Illumina short-read PCR-free WGS</li> <li>IsoSeq RNA-seq</li> <li>Illumina short-read RNA-seq</li> </ul> <p>GM12878 cell line:</p> <ul> <li>Illumina short-read RNA-seq</li> </ul> <p>This dataset includes the following files:</p> <p><strong>Depth of Coverage analysis</strong></p> <ul> <li>Output from `samtools coverage`: <em>samtools.coverage.illumina.txt, samtools.coverage.ont.txt, samtools.coverage.pb.txt</em></li> <li>Output from `copycat` (binned coverage):&nbsp;<em>copycat.ont.coverage.10kb.csv, copycat.pb.coverage.10kb.csv, copycat.pcrfree.coverage.10kb.csv</em></li> </ul> <p><strong>Structural Variant (SV) callsets</strong></p> <ul> <li><strong>Raw:</strong>&nbsp;<em>illumina.tiddit.vcf, ont.sniffles.vcf, pb.sniffles.vcf</em></li> <li><strong>Filtered: </strong><em>REH.svs.filtered.csv</em></li> </ul> <p><strong>SNV callsets</strong></p> <ul> <li><strong>Filtered and annotated:</strong> <em>REH.mutect.filtered.ann.vcf.gz</em></li> </ul> <p><strong>Fusion gene callsets</strong></p> <ul> <li><strong>Short-read:&nbsp;</strong><em>GM12878.fusionreport.txt, illumina.all.txt, illumina.filtered.csv, REH.arriba.fusions.tsv, REH.fusioncatcher.fusion-genes.txt, REH.pizzly.txt, REH.squid.fusions.annotated.txt, REH.starfusion.abridged.tsv, REH.pdf</em></li> <li><strong>Long-read:&nbsp;</strong><em>cupcake.long.csv, cupcake.std.csv, jaffa_results.csv</em></li> <li><strong>Filtered: </strong><em>REH.fusions.filtered.csv</em><br>&nbsp;</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

CT dataset for "Integrated multiomics signatures to optimize the accurate diagnosis of lung cancer" Part-I

<p>To develop and validate a radiomics-based method for lung cancer detection, chest CT images of patients with pulmonary nodules from clinical 5 centers were collected. Due to size limit of zenodo.org, we split the whole dataset into 2 parts, and this is the Part 1.</p> <p>We hope this large-scale dataset could facilitate both clinical research for automatic lung cancer detection and diagnoses, and engineering research for 3D detection, segmentation and classification. This dataset is a research effort of thousands of hours by experienced thoracic surgeons, and radiologists. We kindly ask you to respect our effort by appropriate citation and keeping data license.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Proccessed Data for the Pipelines of the Project "Multiomics and quantitative modelling disentangle diet, host, and microbiota contributions to the host metabolome"

<p><strong>Proccessed and Input Data for the Pipelines of the Project &quot;Multiomics and quantitative modelling disentangle diet, host, and microbiota contributions to the host metabolome&quot;</strong></p> <p>-----------------------------------------------------------------------------------------------------</p> <p>Contents:</p> <p>-----------------------------------------------------------------------------------------------------</p> <p>Folder /ProcessedData/metabolomics/ contains processed metabolomics data from the project:</p> <p>/metabolomics/metabolites_allions_combined_norm_intensity.csv - file containing normalized intensities of ions detected across tissues with six measurement methods.<br> /metabolomics/metabolites_allions_combined_formulas_with_metabolite_filters_spatial100clusters_with_mean.csv - file containing metabolite attribution to spatial clusters and mean intensity values across tissues and conditions.</p> <p>Other files are described in README_ProcessedData.md.</p> <p>-----------------------------------------------------------------------------------------------------</p> <p>Folder /ProcessedData/sequencing/ contains raw and normalized counts of metagenomics and metatransriptomics data mapped to bacterial genomes.</p> <p>Folder /ProccessedData/util/ contains files used for data preprocessing and attribution to chemical classes and pathways.</p> <p>Folder /ProcessedData/example_output/ contains example output of the pipelines:</p> <p>/output/model_results_SMOOTH_raw_2LIcoefHost1LIcoefbact_allions.csv - file containing estimated model parameters (intestinal flux and metabolic flux values) for the forward problem for metabolomics measurements in the GIT.<br> /output/model_results_SMOOTH_normbyabsmax_reciprocal_problem_allions.csv - file containing estimated model parameters for the reverse problem (metabolite intensities) for the parameters estimated with the forward problem.<br> /output/model_results_SMOOTH_normbyabsmax_2LIcoefHost1LIcoefbact_allions.csv - file containing estimated model parameters (intestinal flux and metabolic flux values) for the forward problem for metabolomics measurements in the GIT, normalized by absolute maximum value.<br> /output/model_results_SMOOTH_normbyabsmax_ONLYMETCOEF_2LIcoefHost1LIcoefbact_allions.csv - file containing estimated model parameters (only metabolic flux values) for the forward problem for metabolomics measurements in the GIT, normalized by absolute maximum value.<br> /output/table_hierarchical_clustering_groups.csv - file containing attribution of the annotated metabolites to groups according to hierarchical clustering of the normalized model parameters.<br> /output/cgo_clustergrams_of_model_coefficients.mat - matlab object containing clustergram of the normalized model parameters and manually derived sub-clustergrams corresponding to different largest parameter values.</p> <p>Description of other files is provided in the file README_ProcessedData.md.</p> <p>-----------------------------------------------------------------------------------------------------</p> <p>Folder /InputData/ contains HMDB and KEGG tables used for metabolite annotations and chemical group analysis.</p> <p>Folder InputData_KEGGreaction_path contains matlab files with metabolite-metabolite paths calculated from KEGG reaction-pair information (Each matrix contains a subset of paths). These files are used by the script workflow_extract_keggECpathes_for_SPpairs_final.m.</p> <p>Folder InputData_metabolomics_data contains raw metabolomics data from six methods (three LC columns: C08, C18 and HILIC, and positive and negative acquisition modes) and file tissue_weights.txt with tissue weight information used for normalization.</p> <p>Folder InputData_sequencing_data contains folders ballgown_DNA and ballgown_RNA with results of metagenomic and metatranscriptomic data analysis (raw counts, GetMM normalized counts, EdgeR and DeSeq2 analysis).&nbsp; &nbsp;</p> <p>Description of folders is provided in the file readme_InputData.md.</p> <p>-----------------------------------------------------------------------------------------------------</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Spatial multiomic landscape of the human placenta at molecular resolution

<p>Successful pregnancy and healthy human embryo development rely directly on the placenta&rsquo;s complex, dynamic gene regulatory networks, both within placental subtypes and at the maternal-fetal interface (MFI), that underlie stemness, proliferation, differentiation, invasion, immune tolerance, and communication. These cellular and molecular mechanisms are notoriously challenging to elucidate and make this organ arguably the least understood in the human body. Additionally, disruption of this vast collection of intercellular and intracellular programs and pathways leads to pregnancy complications and developmental defects. Here, we generated a comprehensive spatially resolved multi-modal cell census elucidating the molecular architecture of the first trimester human placenta. We utilized paired single-cell ATAC and RNA sequencing, spatial single-cell ATAC and RNA sequencing (Slide-tags), and in situ sequencing and hybridization mapping of transcriptomes at molecular resolution (STARmap-ISS and STARmap-ISH) using 922,961 cells to construct a spatial single-cell atlas outlining joint epigenomic and transcriptomic regulatory dynamics. Paired analyses unraveled intricate tumor-like gene expression and transcription factor motif programs sustaining the placenta in a hostile uterine environment; further investigation of gene-linked cis-regulatory elements revealed heightened regulatory complexity governing trophoblast differentiation and placental disease risk. Complementary spatial mapping techniques decoded these programs within the placental villous core and extravillous trophoblast (EVT) cell column architecture while simultaneously revealing niche-establishing transcriptional elements and cell-cell communication. To unify our datasets, we computationally imputed 33,357-gene multiomic single-cell profiles and spatially characterized the placental chromatin accessibility landscape. This spatially resolved single-cell multiomic framework of the first trimester human placenta at molecular resolution serves as a blueprint for future studies investigating cellular and molecular programs regulating early placental development and pregnancy.</p>

opencc-by-4.0Apr 2024View details →
dryad36/100

Single cell multiomic analysis identifies key genes differentially expressed in innate lymphoid cells from COVID-19 patients

<p>Innate lymphoid cells (ILCs) are enriched at mucosal surfaces where they respond rapidly to environmental stimuli and contribute to both tissue inflammation and healing. To gain insight into the role of ILCs in the pathology and recovery from COVID-19 infection, we employed a multi-omic approach consisting of Abseq and targeted mRNA sequencing to respectively probe the surface marker expression, transcriptional profile and heterogeneity of ILCs in peripheral blood of patients with COVID-19 compared with healthy controls.  We found that the frequency of ILC1 and ILC2 cells was significantly increased in COVID-19 patients.  Moreover, all ILC subsets displayed a significantly higher frequency of CD69-expressing cells, indicating a heightened state of activation.  ILC2s from COVID-19 patients had the highest number of significantly differentially expressed (DE) genes. The most notable genes DE in COVID-19 vs healthy participants included a) genes associated with responses to virus infections and b) genes that support ILC self-proliferation, activation and homeostasis. In addition, differential gene regulatory network analysis revealed ILC-specific regulons and their interactions driving the differential gene expression in each ILC. Overall, this study provides mechanistic insights into the characteristics of ILC subsets activated during COVID-19 infection.</p>

opencc-zeroJul 2024View details →
zenodo36/100

SCID Multiomics Post-Processed Data and Analysis

<p>In this repository are the post-processed datasets and analytical code for the SCID Multiomics&nbsp;paper.&nbsp;The repository is structured as an installable R package for dependency management and dataset&nbsp;loading; it does not export any functions.</p> <p><strong>Installation</strong></p> <p>The easiest way to install this is to download the repository and install using `devtools::install()`.&nbsp;This will allow the import of various datasets using the `data()` function, upon which many of the&nbsp;analysis scripts depend.</p> <p><strong>Datasets</strong></p> <p>In no particular order, the important datasets are described below:</p> <p>- <strong>intsites</strong>: summary statistics from (Wang et al, Blood, 2010) for timepoints used in this study<br> - <strong>tcr</strong>: Aggregate TCR data from Adaptive Biotechnology&#39;s ImmunoSeq pipeline.<br> - <strong>mb</strong>: Metadata for the microbiome sampling timepoints, as well as species data from Metaphlan (not used)<br> -&nbsp;<strong>agg.mb.kz</strong>: Kraken species data for the microbiome samples, after low-complexity filtering<br> - <strong>agg.vp.kz</strong>: Kraken species data for the virome samples, after low-complexity filtering<br> - <strong>card</strong>: Antibiotic resistance gene data from CARD<br> - <strong>subject_ids.csv</strong>: Provides a mapping from the original sample IDs used in the datasets to the ones used in the manuscript.</p> <p>The code for creating these datasets from the original data files are in the `data-raw` directory.</p> <p><strong>Analysis/Figures</strong></p> <p>The analysis code is broken apart by subject and is largely concerned with figure generation. The R<br> scripts are all located in the `inst` folder. To generate all figures, you should run each script in the<br> order specified by the `GenerateFigures.R` file.</p> <p>Figures are output to the `figures` directory, while tables are output to the `tables` directory.</p> <p>Please note: many of the figures used in the manuscript were aesthetically modified after generation (text size, color palette, orientation), precluding exact figure replication</p>

opencc-by-4.0May 2018View details →
zenodo36/100

Metascape Results for Prostate Cancer Multiomics Data

<p><strong>ABSTRACT&nbsp;</strong></p> <p>Large&nbsp;<em>p</em>&nbsp;small&nbsp;<em>n</em>&nbsp;problem is a challenging problem in big data analytics. There are no de facto standard methods available to it. In this study, we propose a tensor decomposition (TD) based unsupervised feature extraction (FE) formalism applied to multiomics datasets, where the number of features is more than 100000 while the number of instances is as small as about 100. The proposed TD based unsupervised FE outperformed other conventional supervised feature selection methods, such as random forest, categorical regression (also known as analysis of variance, ANOVA), and penalized linear discriminant analysis when they are applied to not only multiomics datasets but also synthetic datasets. Genes selected by TD based unsupervised FE were biologically reliable. TD based unsupervised FE turned out to be not only the superior feature selection method but also the method that can select biologically reliable genes.&nbsp;&nbsp;</p> <p><strong>Instructions:&nbsp;</strong></p> <p>This is a supplementary file of paper submitted to bigdata2020</p> <p>&nbsp;</p> <p><strong>Inspiration:</strong></p> <p>This dataset uploaded to U-BRITE for &quot;AI against CANCER DATA SCIENCE HACKATHON&quot;</p> <p>https://cancer.ubrite.org/hackathon-2021/</p> <p><strong>Acknowledgements</strong></p> <p>Y-h. Taguchi, July 17, 2020, &quot;Metascape results for Prostate cancer multiomics data&quot;, IEEE Dataport, doi: https://dx.doi.org/10.21227/rdmb-jm40.</p> <p>https://ieee-dataport.org/documents/metascape-results-prostate-cancer-multiomics-data</p> <p><strong>U-BRITE last update date:</strong>&nbsp;07/21/2021</p>

opencc-by-4.0Jul 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record