Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13
datasets available to search
ShareScore release 0.9.0
Dataset results
13 results for “Computational Workflows”
A Resilient Workflow to Control a Biomedical HPC Simulation in an Urgent Computing Setting
<p><span><span><span><span>We demonstrate a resilient workflow enabled by the LEXIS Platform, running a time- and safety-critical biomedical simulation of virtual stent placement in intracranial arteries using the HemoFlow application. The workflow, as captured on the video, gracefully handles failures of single computing steps or entire computing systems and thus lends itself to urgent computing applications. <br><br><span><span>The concept of this workflow has potential for realising ab-initio computational biomedical simulations which can provide live, targeted guidance to surgeons.</span></span></span></span></span></span></p>
Additional Artifacts - Supplements to: A Resilient Workflow to Control a Biomedical HPC Simulation in an Urgent Computing Setting
<p>In this dataset, we have collected supplementary artifacts to support an understanding of the workflow presented in the submission cited (see related identifiers).</p> <p>These artifacts are (cf. README.md in the main folder of the tar.gz archive):</p> <p>A1: modified HemoFlow code (cf. https://github.com/gzavo/hemoflow) for our workflow experiments (subfolder "hemoflowcfd");<br>A2: workflow descriptions in python for Apache Airflow (subfolder "workflow");<br>A3: inputs (.xml/.npz) and output (.txt) for the example (subfolder "case").</p> <p> </p>
Cloud-Repro: Reproducible Workflow on a Public Cloud for Computational Fluid Dynamics
<p>In a new effort to make our research transparent and reproducible by others, we developed a workflow to run computational studies on a public cloud. It uses Docker containers to create an image of the application software stack. We also adopt several tools that facilitate creating and managing virtual machines on compute nodes and submitting jobs to these nodes. The configuration files for these tools are part of an expanded "reproducibility package" that includes workflow definitions for cloud computing, in addition to input files and instructions. This facilitates re-creating the cloud environment to re-run the computations under the same conditions.</p> <p>The present Zenodo dataset contains all secondary data required to reproduce the figures of the manuscript ("Reproducible Workflow on a Public Cloud for Computational Fluid Dynamics") without running the CFD simulations again.</p>
Workshop Material - 3D-e-Chem Structural Cheminformatics Workflows for Computer-Aided Drug Discovery
<p>The workshop at the KNIME user meeting (Berlin 9th of March 2018) is set up to stimulate participants with varying degrees of experience in cheminformatics to learn and apply the different structural cheminformatics tools and workflows developed within the context of the 3D-e-Chem project. You will learn how to construct and apply integrated cheminformatics workflows using the 3D-e-Chem KNIME nodes for the exploitation of G protein-coupled receptor and kinase data (two important pharmaceutical target classes) to obtain useful information for drug discovery.</p> <p>Information on the 3D-e-Chem KNIME nodes and workflows can be found online:</p> <p>3D-e-Chem GitHub website: <a href="http://3d-e-chem.github.io/">http://3d-e-chem.github.io/</a></p>
A computational workflow for cell line profiling by Imaging Mass Cytometry.
<p>Imaging Mass Cytometry Data as 32-bit single TIFF with computational analysis from the manuscript: <strong>A computational workflow for cell line profiling by Imaging Mass Cytometry.</strong></p> <p><strong><span lang="EN-US">Breast cancer cell lines SKBR3 MCF7 HCC1143 IMC data and CellProfiler pipelines.zip</span></strong></p> <p><strong><span lang="EN-US">Elongated cell lines HeLa SKOV3 BJ IMC data and CellProfiler pipelines.zip:</span></strong></p> <p><strong><span lang="EN-US">Small cell lines A431 HT29 BxPC3 IMC data and CellProfiler pipelines.zip</span></strong></p> <p><strong><span lang="EN-US">U937 PMA-differentiated cells IMC data and CellProfiler pipeline.zip</span></strong></p> <p><strong><span lang="EN-US">A431 Cisplatin Study IMC data and CellProfiler pipeline.zip</span></strong></p> <p><span lang="EN-US">Contains 1 folder per cell line or drug treatment of single TIFF 32-bit markers exported from MCD/txt files (including Xe131 channel) and their respective cpproj. pipeline file for IMC Cell Line Profiler workstream reproducible analysis</span></p> <p><strong><span lang="EN-US">IMC Cell Line Profiler high dimensional and correlation analysis R scripts.zip</span></strong></p> <p><span lang="EN-US">Contains three adaptable R scripts for high dimensional analysis, correlation analysis and combination of both scripts for Machine Learning classified datasets.</span></p> <p><strong><span lang="EN-US">Breast cancer cell lines nuclear state classification by CellProfiler Analyst MLs.zip</span></strong></p> <p><span lang="EN-US">Contains SQLite databases, properties files, training datasets, nuclear classes visual rendering, and classifier model files with outputs for two machine learning classifiers (Random Forest and Fast Gentle Boosting) per breast cancer cell line for CellProfiler Analyst workflow reproducibility.</span></p> <p><strong><span lang="EN-US">A431 Cisplatin Study IMC data nuclear state classification by CellProfiler Analyst MLs.zip</span></strong></p> <p><span lang="EN-US">Contains SQLite databases, properties files, training datasets, classifier model with outputs for Fast Gentle Boosting and Random Forest per treatment for CellProfiler Analyst workflow reproducibility.</span></p> <p><strong><span lang="EN-US">IMC Cell Line Profiler pseudo-color images with Ki-67 marker Cytoplasm marker and Cell-ID nuclei (Fig2 Fig3), visual nuclei and whole-cell segmentation contours rendered images (Fig4).</span></strong></p> <p><strong><span lang="EN-US">Non-compensated and compensated multiTIFF 32-bit cells lines with Cellprofiler masks SCE objects and FCS files and Datatables.zip</span></strong></p> <p>Contains publicly available compensation matrix (<a href="https://zenodo.org/records/7575859">https://zenodo.org/records/7575859</a>) , R compensation script (<strong>Compensation IMC data with CATALYST.R)</strong>, compensated and non-compensated multiTIFF stacks 32-bit per cell line experiment, exported CellProfiler 16-bit masks per cell line dataset, R single cell experiment script (<strong>Conversion IMC data to Single Cell Experiments Objects and FCS.R)</strong> with inputs and outputs (fcs files, sce files, panel files, metadata files),R<strong> </strong>conversion single cell experiment to datatable script<strong> (Conversion SCE to Datatable and analysis.R)</strong>.</p> <p><strong><span lang="EN-US">Step-by-step guide to assist users with the IMC Cell Line Profiler computational workflow.</span></strong></p>
Computational workflows for perovskites: Case study for lanthanide manganites
<p>Supplemental material for the above manuscript. Revised version (v2.0).</p> <p>Includes the complete code archive, including all Quantum ESPRESSO calculation input and output files, as well as postprocessing scripts used to generate the figures in this manuscript.</p>
Datasets for the computational workflow of multidimensional photoemission spectroscopy
<p>Recorded single-electron event data of bulk 2H-WSe<sub>2</sub> photoemission from a commercial momentum microscope (SPECS METIS 1000).These data are used for demonstration of the computational workflow explained in the following publication.<br> <br> <strong>R. P. Xian, Y. Acremann, S. Y. Agustsson, M. Dendzik, K. Bühlmann, D. Curcio, D. Kutnyakhov, F. Pressacco, M. Heber, S. Dong, T. Pincelli, J. Demsar, Wilfried Wurth, Ph. Hofmann, M. Wolf, M. Scheidgen, L. Rettig, R. Ernstorfer, An open-source, end-to-end workflow for multidimensional photoemission spectroscopy, Scientific Data 7, 442 (2020). DOI: 10.1038/s41597-020-00769-8</strong><br> <br> The zip files are not directly usable for running the computational workflow, but requires first to unzip into HDF5 format (.h5).</p>
A computational workflow for binding free energies in Python
<p>Dataset of distances between a host and six different ligands. The host was beta-cyclodextrin (bCD), while the ligands were phenol, benzene, aspirin, toluene, chlorobenzene and 1,3-dichlorobenzene. No bonds were frozen. </p> <p>The ligand were set to move with a step of 0.25 angstrom from -26 to 26 relative to the bCD (a total of 208 distances). At each distance, a energy biasing potential <span class="math-tex">\(E_{bias}\)</span> was applied the keep two molecules in place. </p> <p><span class="math-tex">\(E_{bias} = \frac{1}{2}\cdot K \cdot (R - R_0)^2\)</span></p> <p>The parameters of the ligands were taken from OpenFF while GLYCAM were used for the host bCD. All of it were applied in Python and the OpenMM framework. Starting parameters, pdb-, and sdf-files can be found in the start folder.</p>
Evaluation of Head-Mounted Spatial Computing and Three-Dimensional (3D) Visualization in Ocular Microsurgery: A Safety and Workflow Study
ClinicalTrials.gov study NCT07301385. IPD Sharing: NO. Countries: 1. Publications: 3.
Computational Artifacts for Performance Feedback Autoscaling Experiments with Workloads of Workflows in Apache Airflow
<p>These computational artifacts are related to the software artifacts DOI:10.5281/zenodo.2635571</p> <p><strong>The content of the computational artifacts:</strong></p> <ul> <li><strong>experiments.pdf</strong> contains the list of all the conducted experiments with the Airflow system. Experiment IDs are not sequential since some experiments required rerunning, etc., we report only successful results. The file lists different experiment configurations, e.g., the number of processed workflows, the name of the used workload, the user budgets, and PFA settings.</li> <li><strong>db.tar.gz</strong> contains directories with Airflow database snapshots and autoscaler logs. The names of the directories correspond to those listed in `experiments.pdf`. Each experiment directory contains an autoscaler log and a full copy of a PostgreSQL database directory just after each experiment finished. The database name is `airflow`, the user name is `ailyushk`. Within each database, most of the paper-related data are stored in the `stat_log` table. The scripts for extracting data from these databases are available as software artifacts in `tools/analysis`.</li> <li><strong>gurobi.tar.gz</strong> contains the results obtained from the Gurobi solver when solving the MIP model.</li> <li><strong>pdf.tar.gz</strong> contains all the figures in pdf format, also those that were not included neither in the paper nor in the technical report. The scripts for creating this plots are delivered as software artifacts.</li> <li><strong>csv.tar.gz</strong> contains the analysis results extracted from Airflow database snapshots. These files are used to create the plots in the `pdf` directory. The scripts for doing this are delivered as software artifacts.</li> <li><strong>wl1.tar.gz</strong> is the first synthetic realistic workload (WL I) with three subsets of 200 workflows each (`1_0`, `1_1`, `1_2`). Each directory contains the file with interarrivals `interarrivals.txt`, and the file with workflow IDs `workload.txt` in the subset. The `dags` directory contains Python-based Airflow descriptors and CSV files that summarise the same descriptors in CSV format for simpler analysis. The scripts for extracting workload statistics from these CSV files are available in the software artifacts in `tools/analysis`. The `inputs` directory contains initial input files for each worfklow. The `dax` contains original DAX files obtained from the generator: <a href="https://github.com/pegasus-isi/WorkflowGenerator/tree/master/bharathi/src/simulation/generator">https://github.com/pegasus-isi/WorkflowGenerator/tree/master/bharathi/src/simulation/generator</a></li> <li><strong>wl2.tar.gz</strong> is the second synthetic realistic workload (WL II) with three subsets of 200 workflows each (`4_0`, `4_1`, `4_2`). Has similar structure as `wl1.tar.gz`, except that `dax` directory is omitted, as WL II uses the same DAX structures as WL I.</li> <li><strong>wl3.tar.gz</strong> is the small synthetic workload based on WL I for the experiment with the MIP solver, contains three subsets with 5 workflows in each, all in the `3_0` directory (thus, the structure differs from the WL I and WL II). The input data files are empty. The identifiers of workflows forming each subset are stored in the `workload_1.txt`, `workload_2.txt`, and `workload_3.txt` files.</li> </ul>
Daily Imaging, Target Identification, and Simulated Computed Tomography-Based Stereotactic Adaptive Radiotherapy Workflow in a Novel Ring Gantry Radiotherapy Device
ClinicalTrials.gov study NCT04008537. IPD Sharing: NO. Countries: 1. Publications: 0.
Full-3D Computer-Assisted Workflow for the Diagnosis and Correction of Deformities? Dentofacial
ClinicalTrials.gov study NCT06806605. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
A computational workflow for the analysis of 3’ Tag-Seq data
GEO Series GSE200778. Candida albicans. 4 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.