Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

111,170

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

111,170 results for “cell”

Learn how ShareScore rates datasets ↗
edi60/100

Lab disease outcomes data evaluating how antibiotic tolerant vs. non-tolerant cell-free supernatant from Pseudomonas aeruginosa affects the interaction between a fungal pathogen (Batrachochytrium dendrobatidis) and amphibian (Rana sylvaticus), 2022.

Microbes living on hosts and in the environment can play a key role in helping hosts to combat pathogens. However, antibiotic-induced alterations to microbial metabolite production could disrupt this dynamic. Here, we investigated whether antibiotic tolerance influences the anti-pathogenic properties of host-associated (living on the host; biofilms) and environmental (living in the soil of water column; planktonic) microbes in vitro and in vivo. For our model host and pathogen, we used the amphibian (Rana sylvatica)-Batrachochytrium dendrobatidis (Bd) system. For our model host-associated (biofilm) and environmental (planktonic) microbes, we used four strains of Pseudomonas aeruginosa that vary in their tolerance to antibiotics and their biofilm-forming capabilities: Planktonic, non-antibiotic tolerant (ΔsagS/VC); Planktonic, antibiotic tolerant (ΔsagS::sagS_L154A); Biofilm, non-antibiotic tolerant (ΔsagS::sagS_D105A); Biofilm, antibiotic tolerant (ΔsagS::sagS). We collected cell-free supernatants (CFS) from each strain to examine the effects of metabolites. We conducted four experiments. In our pathogen-only exposures to test direct effects of metabolites on Bd, we exposed Bd zoospores to each P. aeruginosa CFS at six concentrations. After 11 days of growth, we measured relative abundance of Bd across each treatment. In our host-only exposures to test effects of metabolites on host disease outcomes, we placed R. sylvatica tadpoles in individual units containing each P. aeruginosa CFS. After 48 hours, water was changed into clean well water (no CFS). Bd zoospores were immediately added to each experimental unit following the water change. After 5 days of Bd exposure, we measured snout-vent length (SVL), mass, developmental stage, and Bd quantification in the mouthparts using qPCR for each tadpole. In our host-pathogen exposures to test interactive effects of metabolites on hosts in the presence of the pathogen, we conducted the same experiment as above. However, ins

openCC (other)May 2025View details →
zenodo56/100

Conserved regulation of RNA processing in somatic cell reprogramming

<p><strong>Data set 1. Transcript expression across human RNA-Seq samples: estimated read counts. </strong>The file contains estimated read counts, generated by kallisto (<a href="https://pachterlab.github.io/kallisto/">https://pachterlab.github.io/kallisto/</a>), for human transcripts and RNA-Seq samples used in this study (see Additional file 2 of the accompanying publication). The format is a compressed (GZIP) tab-separated transcript-by-sample matrix. Ensembl transcript identifiers and a combined Sequence Read Archive study/sample name identifier serve as row and column names, respectively.</p> <p><strong>Data set 2. Transcript expression across murine RNA-Seq samples: estimated read counts. </strong>As in Data set 1, but for mouse transcripts.</p> <p><strong>Data set 3. Transcript expression across simian RNA-Seq samples: estimated read counts. </strong>As in Data set 1, but for chimpanzee transcripts.</p> <p><strong>Data set 4. Transcript expression across across human RNA-Seq samples: estimated transcript abundances. </strong>As in Data set 1, but instead of read counts, transcript abundances in transcripts per million (TPM), as estimated by kallisto (<a href="https://pachterlab.github.io/kallisto/">https://pachterlab.github.io/kallisto/</a>), are listed. Format, column and row names as in Data set 1.</p> <p><strong>Data set 5. Transcript expression across murine RNA-Seq samples: estimated transcript abundances. </strong>As in Data set 4, but for mouse transcripts.</p> <p><strong>Data set 6. Transcript expression across simian RNA-Seq samples: estimated transcript abundances. </strong>As in Data set 4, but for chimpanzee transcripts.</p> <p><strong>Data set 7. Differential expression analyses across human RNA-Seq sample groups: log fold changes. </strong>The file contains log fold changes, inferred by edgeR (<a href="http://bioconductor.org/packages/release/bioc/html/edgeR.html">http://bioconductor.org/packages/release/bioc/html/edgeR.html</a>), for human genes and the RNA-Seq sample group contrasts listed in Additional file 3 of the accompanying publication in a compressed (GZIP) TSV gene-by-comparison matrix. Ensembl gene identifiers and a descriptive contrast identifier serve as row and column names, respectively.</p> <p><strong>Data set 8. Differential expression analyses across murine RNA-Seq sample groups: log fold changes. </strong>As in Data set 7, but for mouse genes.</p> <p><strong>Data set 9. Differential expression analyses across simian RNA-Seq sample groups: log fold changes. </strong>As in Data set 7, but for chimpanzee genes.</p> <p><strong>Data set 10. Differential expression analyses across human RNA-Seq sample groups: false discovery rates. </strong>The file contains false discovery rates (FDR) for the differential expression analyses summarized in Data set 7. Format, column and row names as in Data set 7.</p> <p><strong>Data set 11. Differential expression analyses across murine RNA-Seq sample groups: false discovery rates. </strong>As in Data set 10, but for mouse genes.</p> <p><strong>Data set 12. Differential expression analyses across simian RNA-Seq sample groups: false discovery rates. </strong>As in Data set 10, but for chimpanzee genes.</p> <p><strong>Data set 13. Quantification of alternative splicing events across human RNA-Seq samples. </strong>The file contains &lsquo;percent spliced in&rsquo; (PSI) values computed by SUPPA (<a href="https://github.com/comprna/SUPPA">https://github.com/comprna/SUPPA</a>) for annotated alternative splicing events (inferred from the transcript annotation of the human genome, Ensembl release 84; <a href="http://www.ensembl.org/">http://www.ensembl.org/</a>). The format is a compressed (GZIP) tab-separated transcript-by-sample matrix. SUPPA-provided event identifiers and a combined Sequence Read Archive study/sample name identifier serve as row and column names, respectively.</p> <p><strong>Data set 14. Quantification of alternative splicing events across murine RNA-Seq samples. </strong>As in Data set 13, but for mouse alternative splicing events.</p> <p><strong>Data set 15. Differential splicing analyses across human RNA-Seq sample groups: differences in &lsquo;percent spliced in&rsquo; (&Delta;PSI). </strong>The file contains &Delta;PSI values for human alternative splicing events (as in Data set 13). The RNA-Seq sample group contrasts are listed in Additional file 3 of the accompanying publication. Values were inferred by SUPPA&rsquo;s diffSplice functionality (<a href="https://github.com/comprna/SUPPA">https://github.com/comprna/SUPPA</a>). The format is a compressed (GZIP) tab-separated gene-by-comparison matrix. SUPPA event identifiers and a descriptive contrast identifier serve as row and column names, respectively.</p> <p><strong>Data set 16. Differential splicing analyses across murine RNA-Seq sample groups: differences in &lsquo;percent spliced in&rsquo; (&Delta;PSI). </strong>As in Data set 15, but for mouse alternative splicing events.</p> <p><strong>Data set 17. Differential splicing analyses across human RNA-Seq sample groups: P values. </strong>The file contains P values for the differential splicing analysis of human alternative splicing events summarized in Data set 15. Format, column and row names as in Data set 15.</p> <p><strong>Data set 18. Differential splicing analyses across murine RNA-Seq sample groups: P values. </strong>The file contains P values for the differential splicing analysis of mouse alternative splicing events summarized in Data set 16. Format, column and row names as in Data set 15.</p> <p><strong>Data set 19. Transcript expression across murine RNA-Seq time course data: estimated read counts. </strong>As in Data set 2, but for the time course data generated for the accompanying publication.</p> <p><strong>Data set 20. Transcript expression across murine RNA-Seq time course data: estimated transcript abundances. </strong>As in Data set 5, but for the time course data generated for the accompanying publication.</p> <p><strong>Data set 21. Quantification of alternative splicing events across murine RNA-Seq time course data. </strong>As in Data set 14, but for the time course data generated for the accompanying publication.</p>

opencc-by-4.0Mar 2018View details →
zenodo52/100

Transcriptomic response of human cells to SARS-CoV-2, RSV and H1N1 (STAR + StringTie)

<p>These data represent results from:</p> <ol> <li>Processing reads from 20 experiments (part of GSE147507) by following a standard approach, which includes using STAR to align the reads to GRCh38 and StringTie to calculate the (raw) counts per experiment. These results depict the transcriptomic response&nbsp;of human cells to SARS-CoV-2, RSV and H1N1, and enrichment analyses based on genes differentially expressed in&nbsp;SARS-CoV-2 but not in RSV or H1N1. (Authors: V.A.-P., M.G.F. and A.G.)</li> <li>Aligning to SARS-CoV-2 and quantifying reads&nbsp;by using HISAT2 and StringTie. (Author: C.R.-A.)</li> </ol> <p>Disclaimer: These results were obtained during the virtual BioHackathon 2020. As such, they&nbsp;are subject to ongoing research and have thus NOT yet undergone any scientific peer-review. That is, none of the contents can be considered to be free of errors and must be taken with caution!</p>

opencc-zeroApr 2020View details →
zenodo52/100

Dataset of "Tuning the morphology and energy levels in organic solar cells with metal- organic framework nanosheets"

<p>Metal-organic framework nanosheets (MONs) have proved themselves to be useful<br>additives for enhancing the performance of a variety of thin film solar cell devices. However,<br>to date only isolated examples have been reported. In this work we take advantage of the<br>modular structure of MONs in order to resolve the effect of their different structural and<br>optoelectronic features on the performance of organic photovoltaic (OPV) devices. Three<br>different MONs were synthesized using different combinations of two porphyrin-based ligands<br>meso-tetracarboxyphenyl porphyrin (TCPP) or tetrapyridyl-porphyrin (TPyP) with either zinc<br>and/or copper ions and the effect of their addition to polythiophene-fullerene (P3HT-PCBM)<br>OPV devices was investigated. The power conversion efficiency (PCE) of devices was found to<br>approximately double with the addition of MONs of Zn2(ZnTCPP), but was unchanged with<br>the addition of Cu2(ZnTPyP) and halved upon the addition of Cu2(CuTCPP) compared to<br>devices without nanosheets. Our analysis indicates that there are three different mechanisms<br>by which MONs can influence the photoactive layer &ndash; light absorption, energy level alignment,<br>and morphological changes. Analysis of external quantum efficiency, UV-vis photoelectron<br>spectroscopy data found that MONs have similar effects on light absorption and energy level<br>alignment. However, atomic force and Raman microscopy studies revealed that the nanosheet<br>thickness and lateral size are crucial parameters in enabling the MONs to act as beneficial<br>additives resulting in an improvement of the OPV device performance. We anticipate this<br>study will aid in the design of MONs and other 2D materials for future use in other light<br>harvesting and emitting devices.</p>

opencc-by-4.0Jul 2024View details →
zenodo52/100

Dataset for "An Alternative Chlorine-Assisted Optimization of CdS/Sb2Se3 Solar Cells: Towards Understanding of Chlorine Incorporation Mechanism"

<p>The current strategies in the development of Sb2Se3 thin film solar cells involve fabrication and optimization of<br>superstrate and substrate device architectures, with the preferable choice for TiO2 and CdS heterojunction layers.<br>For CdS-based superstrate cells, several studies reported the necessity to apply CdCl2 or other metal halide-based<br>post-deposition treatment (PDT), highlighting improvement of CdS/Sb2Se3 device efficiency. However, the need,<br>effect, and mechanism of such PDT are very often not described. Additionally, the fact that many groups have not<br>succeeded in demonstrating its benefits suggests that this strategy is not straightforward, requiring a deeper<br>understanding towards a more unified concept. The present study proposes an alternative approach to the<br>challenging CdCl2 PDT of CdS in CdS/Sb2Se3 device, involving controllable Cl incorporation in CdS films by<br>systematically varying the concentration of NH4Cl in the CBD precursor solution from 1 to 8 mM. Structural and<br>electrical characterizations are correlated with advanced measurements of Scanning Kelvin Probe, surface<br>photovoltage, and atomic force microscopy to understand the impact of Cl incorporation on the properties of CdS<br>films and CdS/Sb2Se3 devices. The validity of Cl incorporation in the CdS lattice and interdiffusion processes at<br>the CdS-Sb2Se3 interface is confirmed by secondary ion mass spectrometry analysis. It is demonstrated that<br>incorporation of 1 mM of NH4Cl, as a Cl source in CBD CdS, can boost the PCE of CdS/Sb2Se3 by ~20 %. With this<br>approach, we offer new perspectives on the optimization methodology for Cl-based CdS/Sb2Se3 device processing<br>and complementary understanding of the physiochemistry behind these processes.</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

Dataset for: Associating Mechano-electrochemical Phenomena to Stochastic Current Events in Micro-Electrochemical Cells Containing TiNb2O7 Particles

<p>This is the raw data used in a manuscript that will be submitted to ChemElectroChem. If you have any questions, please email the creators.</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

Dataset of "Activity-stability relationship in magnetron co-sputtered bimetallic catalysts for proton exchange membrane fuel cells"

<p>In the present study, magnetron sputtered PtxM100-x (M = Co, Cu, Y; x = 25, 50, 75 and 100) bimetallic alloys were investigated as PEMFC cathodes. &nbsp;Accurate composition control enabled a systematic study of the correlation between alloy composition, activity, and stability. The catalysts underwent thorough characterization, employing a diverse portfolio of characterization techniques such as scanning electron microscopy, energy-dispersive X-ray spectroscopy, X-ray photoelectron spectroscopy and cyclic voltammetry. The activity of all investigated alloys was tested directly in a fuel cell device, while stability was assessed through potentiodynamic cycling in a half-cell.&nbsp;<br>The activity-stability index, considering experimental results for both activity and stability, was calculated and compared for all investigated catalysts. All alloys exhibited a volcano-type trend in activity-stability index as a function of the concentration of alloying element with peaks observed at Pt50Co50, Pt50Cu50 and Pt75Y25 for respective alloys, surpassing that of monometallic platinum. Overall, Pt50Co50 emerged as a catalyst with the highest activity-stability ratio.</p>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Figure 2a - Rossi et al. Design of Highly Efficient Semitransparent Perovskite/Organic Tandem Solar Cells RRL Solar (2022)

<p>The Data set is related to the<strong> figure 2a</strong> of the paper&nbsp;</p> <p>Design of Highly Efficient Semitransparent Perovskite/Organic Tandem Solar Cells by Daniele Rossi,Karen Forberich,Fabio Matteocci,Matthias Auf der Maur,Hans-Joachim Egelhaaf,Christoph J. Brabec,Aldo Di Carlo, Rapid&nbsp;Research Letter (2022)&nbsp; https://doi.org/10.1002/solr.202200242</p>

opencc-by-4.0Jul 2022View details →
zenodo52/100

Dataset for Towards improved online dissolution evaluation of Pt-alloy PEMFC electrocatalysts via electrochemical flow cell - ICP-MS setup upgrades

<p>Experimental data comprises raw data from ICP-MS (Inductively coupled plasma mass spectrometry) (i.e. time dependence of signal intensity for Co59 and Pt195) for different cell geometry and operating parameters. &nbsp;<br>Model data comprise of time- and space-dependent values of Pt ions concentration in the modelling cell and local velocity vectors.</p>

opencc-by-4.0Jan 2024View details →
zenodo52/100

Dataset for "Impact of the flow-field distribution channel cross-section geometry on PEM fuel cell performance: stamped vs. milled channel"

<p>Experimental data comprises raw data from load curve characterisation of a PEM fuel cell used for the validation of the mathematical model. Model data comprise of space-dependent values of hydrogen and oxygen concentration, local current densities, gas pressures and gas velocities in the modelled cell. These data were used for the investigation of the effect of different geometric parameters of flow-field channels on the performance of a PEM fuel cell.</p>

opencc-by-4.0May 2024View details →
zenodo52/100

Antigen-specific CD4+ T cells exhibit distinct transcriptional phenotypes in the lymph node and blood following vaccination in humans

<p><strong>Abstract:&nbsp;</strong><br>SARS-CoV-2 infection and mRNA vaccination induce robust CD4+ T cell responses that are critical for the development of protective immunity. Here, we evaluated spike-specific CD4+ T cells in the blood and draining lymph node (dLN) of human subjects following BNT162b2 mRNA vaccination using single-cell transcriptomics. We analyze multiple spike-specific CD4+ T cell clonotypes, including novel clonotypes we define here using Trex, a new deep learning-based reverse epitope mapping method integrating single-cell T cell receptor (TCR) sequencing and transcriptomics to predict antigen-specificity. Human dLN spike-specific T follicular helper cells (TFH) exhibited distinct phenotypes, including germinal center (GC)-TFH and IL-10+ TFH, that varied over time during the GC response. Paired TCR clonotype analysis revealed tissue-specific segregation of circulating and dLN clonotypes, despite numerous spike-specific clonotypes in each compartment. Analysis of a separate SARS-CoV-2 infection cohort revealed circulating spike-specific CD4+ T cell profiles distinct from those found following BNT162b2 vaccination. Our findings provide an atlas of human antigen-specific CD4+ T cell transcriptional phenotypes in the dLN and blood following vaccination or infection.</p> <p><strong>More Information:</strong></p> <ul> <li><strong>Preprint:</strong> <a href="https://www.researchsquare.com/article/rs-3304466/v1">Research Square.</a></li> <li><strong>Sample information</strong>: data_inventory.csv file.</li> <li><strong>Code</strong> code_github_repo.zip or at the <a href="https://github.com/ncborcherding/COVID_TCR">original github repo</a></li> <li><strong>Interactive Portal</strong>: <a href="https://cellpilot.emed.wustl.edu/">CellPilot</a></li> </ul>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Data for "Profiling the transcriptomic age of single-cells in humans"

<p>This is a supplementary data for the article titled "Profiling transcriptomic age of human single-cells". Data created in this project is shared here for the scientific community.&nbsp;</p> <p>Here we used available scRNA-seq data of 1,058,909 blood cells of 508 healthy, human donors, for developing cell-type-specific single-cell transcriptomic clocks and predicting the age of human blood cells. &nbsp;We also applied our clocks to different external datasets and evaluated the age of single cells originated from COVID-19 patients and human embryos.</p> <p>For the description of the content of the dataset see the ReadMe file.</p>

opencc-by-4.0Jun 2024View details →
zenodo52/100

Fast and long-term super-resolution imaging of ER nano-structural dynamics in living cells using a neural network

<p>Datasets acquired and generated for the manuscript "Fast and long-term super-resolution imaging of ER nano-structural dynamics in living cells using a neural network". The datasets include test, training and time series datasets each containing the raw data and the predicted data where it applies.&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo52/100

Data for "Measurement of the atom-surface van der Waals interaction by transmission spectroscopy in a wedged nano-cell"

<p>The data presented in publication <a href="http://arxiv.org/abs/1905.02783">&quot;Measurement of the atom-surface van der Waals interaction by transmission spectroscopy in a wedged nano-cell&quot;</a> .&nbsp; Published version: <a href="https://doi.org/10.1103/PhysRevA.100.022503">https://doi.org/10.1103/PhysRevA.100.022503</a></p> <p>The data are in HDF5 format, with associated metadata.</p> <p>To see examples of how to use the data, and the theoretical model for analysis, see <a href="https://github.com/thermal-vapours/TAS-Transmission-Atom-Surface">https://github.com/thermal-vapours/TAS-Transmission-Atom-Surface </a></p>

opencc-by-4.0Apr 2019View details →
zenodo52/100

Data set of the manuscript titled: Follicular Immune Landscaping Reveals a distinct profile of FOXP3hi CD4+ T cells in Treated compared to Untreated HIV

<p>Multiplex imaging data were collected using a scanning confocal system (STELARIS, Leica) and proccessed with the Imaris and Fiji imaging programs. csv files incuding the position identifiers and intensities for each fluorochrome used were generated and data were further analysed using the FlowJo10 program. Neighboring analysis was performed using the G function and mean of minimum distances of relevant cell type pairs.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo52/100

Dataset for "Study of Rapid Capacity Fade in Prismatic Li-ion Cells with Flexible Packaging"

<p>Prismatic lithium-ion batteries (LIBs) are considered promising electric energy sources in electromobility applications due to their cell to pack density. However, their sensitivity to external and internal influences, and reduced durability lead to inflation risk and potential explosions throughout their lifecycle. These critical processes are strongly influenced by the inner construction of the cell, especially concerning the coating and mechanical fixation. This study subjects a commercially available prismatic LIB cell to comprehensive, correlative analysis employing various imaging techniques. The inner structure of the entire cell is visualized non-destructively by X-ray computed tomography (CT), enabling the identification of critical design flaws prior to electrochemical cycling. Electrochemical cycling simulates the battery lifecycle, and the cell is subsequently disassembled in the fully charged state. The usage of the inert-gas transfer system allowed the preparation of Broad Ion Beam (BIB) electrodes cross-sections in a fully native state and for the first time to observe the tearing of graphite particles due to over-lithiation. Established region labeling system allowed to use CT and scanning electron microscopy (SEM) correlatively to identify critical regions. After 100 cycles, a 40% capacity loss was observed and event diagram describing deagradation mechanisms, related both to the cell design and to the processes occurring at high load, was created.</p>

opencc-by-4.0Aug 2024View details →
zenodo52/100

Nucleus and cell segmentations for data in the mudRapp-seq paper

<p>Segmentation masks for images published with the paper describing&nbsp;</p> <p>"<em>Multiple direct RNA padlock probing in combination with in-situ sequencing (mudRapp-seq)</em>":</p> <blockquote> <p>Ahmad S, Gribling-Burrer AS, Schaust J, Fischer SC, Ambil UB, Ankenbrand MJ, Smyth RP. <em>Visualizing the transcription and replication of influenza A viral RNAs in cells by multiple direct RNA padlock probing and in-situ sequencing (mudRapp-seq)</em> (in review)</p> </blockquote> <p>Raw images are published in the <a href="https://www.ebi.ac.uk/bioimage-archive/">Bioimage Archive</a> (identifier pending). To use these masks, run the data formatting code in the accompanying code repository to get the raw data in the correct structure and extract this zip archive into the repository root (the folder structure in the archive matches the folder structure of the repository).</p> <p>Filenames in `analysis/segmentation` contain a hint about how they were created:</p> <ul> <li>cp: direct segmentation with a cellpose model (<a href="https://github.com/BioMeDS/mudRapp-seq/blob/main/models/cellpose/nuclei">nuclei</a>, <a href="https://github.com/BioMeDS/mudRapp-seq/blob/main/models/cellpose/cells">cells</a>)</li> <li>cpws: cell segmentation through watershed with nucleus masks as seeds</li> <li>cpmc: manually corrected cellpose segmentations</li> </ul> <p>Besides the final segmentation masks, the training data are included in `data/training` and the models in `models/cellpose`.</p> <p>Changes:</p> <ul> <li>v1.1 training data and models added</li> </ul>

opencc-by-4.0Sep 2024View details →
zenodo52/100

Chemical structures, Cell Painting and transcriptional profiles for compound bioactivity prediction.

<p>This is the related data, both input and produced for the paper <a href="https://doi.org/10.1101/2020.12.15.422887">&quot;Predicting compound activity from phenotypic profiles and chemical structures&quot;</a>.</p> <p>This data can be merged with <a href="https://github.com/CaicedoLab/2023_Moshkov_NatComm">paper&#39;s GitHub repository</a>&nbsp;for reproduction.</p> <p>Folders and files&nbsp;and are described&nbsp;below:</p> <pre><code>├── assay_data ├── assay_matrix_discrete_270_assays.csv Assay matrix with hits for assays (270) and compounds (16170). Note that this is the final file that we used to produce splits. ├── assay_metadata.csv Assay metadata ├── broad_ids.txt List of broad ids used in this study. That is an unfiltered list of compounds required by some analysis scripts. ├── smiles.txt Same as broad_ids.txt, but SMILES strings. ├── feature_data (for 16978 compounds, can be masked with ./misc/compounds16978to16170.npy) ├── cp.npz Classical chemical features ├── ge.npz Gene expression features ├── ge_scale.npz Gene expression scaled features ├── mo.npz Morphology features (not batch corrected) ├── mobc.npz Morphology features (batch corrected) ├── misc ├── compound_analysis.npz Compounds in the dataset identified as PAINS ├── compounds16978to16170.npy Used to filter features from the bigger set of compounds to the final one ├── fingerprints.npz Calculated fingerprints of compounds, those were then used to calculate similarity ├── similarity_fingerprints.npz Similarity matrix for compounds (16978) ├── population_normalized.csv.gz Well-level morphological profiles that were used for batch-correction ├── Table for PUMA Excel file with additional data and plots ├── predictions ├── scaffold_median(mean)_AUC.csv Aggregated median(mean) AUC scores over scaffold-based cross-validation splits. In the paper, median results were reported. ├── scaffold_median(mean)_EF.csv Aggregated median(mean) enrichment factor (EF) over scaffold-based cross-validation splits. In the paper, median results were reported. ├── toprank_chemical_cv{}_hitsnorm.csv Those files are needed to create enrichment plots and contain hit rate and top rank hit rate. ├── Each folder here stands for an experiment type, the number in the folder name is a number of the split. Inside each folder there are the following elements: ├── predictions Folder with predictions for each assay-compound pair for each modality ├── 2022_01_evaluation_all_data.csv File with AUC scores for each assay for the test set in the split ├── 2022_01_evaluation_all_data_EF.csv File with enrichment factor (EF) values for each assay for the test set in the split. Those files exist only for *chemical* folders. ├── assay_matrix_discrete_train(test)_old_scaff.csv Training and test subsets of data for the split. The first column contains broad_id. ├── assay_matrix_discrete_train(test)_old_scaff.csv Same, but SMILES strings in the first column. Those files are used as input to ChemProp! Experiments in this folder are the following: - chemical Scaffold-based 5-fold cross-validation splits, the main results in the paper are reported with this series of experiments. - chemical_bal Same splits as in chemical, but training were run with ChemProp built-in data balancing. - chemical_st Same splits as in chemical, but separate models were trained for each assay. - CV Random 5-fold cross-validation splits. - GE 5-fold cross-validation splits based on same-size clustering of gene expression features. - MOBC 5-fold cross-validation splits based on same-size clustering of batch-corrected morphology features. - random 10 random splits, ~80% of compounds in the training set and the rest in the test set. ├── splitting This folder contains numpy files which help to match compounds and features to create training and test sets for a split, which can be reused in the analysis notebook for data preparation. ├── scaffold_based_split.npz Splitting for scaffold-based splits. ├── random_split_{}.npz Random split indices of test set compounds (10 files). ├── cross_validation_indicies.npz Indices for random cross-validation splits ├── GE_clusters_size_constrained.npz Indicies of clusters of same-size clustering for gene-expression features. ├── MOBC_clusters_size_constrained.npz Indices of clusters of same-size clustering for batch-corrected morphology features.</code></pre> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo48/100

Raw data to accompany the manuscript 'Data for Engineering Lipid Metabolism of Chinese Hamster Ovary (CHO) Cells for Enhanced Recombinant Protein Production' published in the Journal Data in Brief

<p>This repository consists of the raw western blot, microscopy and mass spectrometry data to accompany the manuscript &#39;Data for Engineering Lipid Metabolism of Chinese Hamster Ovary (CHO) Cells for Enhanced Recombinant Protein Production&#39; published in the Journal Data in Brief and associated with the article &#39;<a href="https://www.ncbi.nlm.nih.gov/pubmed/31805379">Engineering of Chinese hamster ovary cell lipid metabolism results in an expanded ER and enhanced recombinant biotherapeutic protein production</a>&#39; published in the journal Metabolic Engineering (see DOI:&nbsp;10.1016/j.ymben.2019.11.007).&nbsp;</p> <p>The western blot raw file is associated with Figure 1a and 1b of the Data in Brief manuscript.</p> <p>The confocal microscopy raw image files (x3) are associated with Figure 1c&nbsp;of the Data in Brief manuscript.</p> <p>The mass spectrometry files are the raw data that refers to the samples presented in Figure 5 of the Data in Brief manuscript. Files are labelled as in the Data in Brief and Metabolic Engineering manuscripts. The file name structures is as follows;</p> <p>CHO-Controlpoolai</p> <p>Where &#39;a&#39; represents replicate &#39;a&#39; of three biological replicates and &#39;i&#39; refers to mass spectrometry technical analysis 1 of 3 technical analyses of each replicate (thus for each cell pool or line there are three biological replicates that are each analysed in triplicate such that there are 9 raw mass spectrometry files for each cell pool or line).</p> <p>All the mass spectrometry files are found in the compressed (zip) file named mass_spectrometry_raw_files_archive.zip</p>

opencc-by-4.0Jan 2020View details →
zenodo48/100

Synthetic images of cell nuclei in widefield microscopy

<p>The images were generated by&nbsp;<a href="http://www.cs.tut.fi/sgn/csb/simcep/tool.html">SIMCEP</a>, a widefield fluorescence microscopy biological images simulator.</p> <p>The dataset is used to demonstrate the execution of image analysis workflows with BIAFLOWS on a local machine from a jupyter notebook.</p>

opencc-by-4.0Jan 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record