Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
650
datasets available to search
ShareScore release 0.9.0
Dataset results
650 results for “Workflow”
Digitization Workflow for Data Mining in Production Technology applied to a Feed Axis of a CNC Milling Machine
<p>Dataset accompanying the publication "Digitization Workflow for Data Mining in Production Technology<br>applied to a Feed Axis of a CNC Milling Machine" (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.procs.2024.01.017" target="_blank" rel="noreferrer noopener"><span>https://doi.org/10.1016/j.procs.2024.01.017</span></a>).</p>
OrgaMapper: A robust and easy-to-use workflow for analyzing organelle positioning
<p><span>Eukaryotic cells are highly compartmentalized by a variety of organelles that carry out specific cellular processes. The position of these organelles within the cell is elaborately regulated and vital for their function. For instance, the position of lysosomes relative to the nucleus controls their degradative capacity and is altered in pathophysiological conditions. The molecular components orchestrating the precise localization of organelles remain incompletely understood. A confounding factor in these studies is the fact that organelle positioning is surprisingly non-trivial to address. E.g., perturbations that affect the localization of organelles often lead to secondary phenotypes such as changes in cell or organelle size. These phenotypes could potentially mask effects or lead to the identification of false positive hits. To uncover and test potential molecular components at scale, accurate and easy-to-use analysis tools are required that allow robust measurements of organelle positioning. </span></p> <h2><span>Results</span></h2> <p><span>Here, we present an analysis workflow for the faithful, robust, and quantitative analysis of organelle positioning phenotypes. Our workflow consists of an easy-to-use Fiji plugin and an R Shiny App. These tools enable users without background in image or data analysis to (1) segment single cells and nuclei and to detect organelles, (2) to measure cell size and the distance between detected organelles and the nucleus, (3) to measure intensities in the organelle channel plus one additional channel, (4) to measure radial intensity profiles of organellar markers, and (5) to plot the results in informative graphs. Using simulated data and immunofluorescent images of cells in which the function of known factors for lysosome positioning has been perturbed, we show that the workflow is robust against common problems for the accurate assessment of organelle positioning such as changes of cell shape and size, organelle size and background.</span></p> <h2><span>Conclusion</span></h2> <p><span>OrgaMapper is a versatile, robust and easy-to-use automated image analysis workflow that can be utilized in microscopy-based hypothesis testing and screens. It effectively allows for the mapping of the intracellular space and thereby enables the discovery of novel regulators of organelle positioning. </span></p>
Visualizing Workflows with the Dragon Telemetry Service
<p>The Dragon telemetry service is an easy-to-use, scalable means for users to visualize both hardware and custom metrics for complex workflows. Dragon is a high-performance distributed runtime for managing processes and data at-scale. It utilizes high-performance communication objects to enable efficient and transparent management of memory and movement of data. We demonstrate the use of the telemetry service for a multi-language AI-in-the-loop workflow where both built-in hardware metrics and custom user metrics are visualized in a Grafana dashboard.</p>
msiFlow: Automated Workflows for Reproducible and Scalable Multimodal Mass Spectrometry Imaging and Immunofluorescence Microscopy Data Processing and Analysis
<p>This record contains example and result data of msiFlow.</p> <p>msiFlow is a collection of automated workflows for reproducible and scalable multimodal mass spectrometry imaging (MSI) and immunofluorescence microscopy (IFM) data processing and analysis. Using an experimental mouse model for urinary tract infection, induced by uropathogenic E.coli (UPEC), we generated data by</p> <ul> <li>matrix-assisted laser desorption ionisation mass spectrometry imaging with laser-induced postionisation (MALDI-2 MSI) using the Bruker timsTOFfleX instrument</li> <li>transmission-mode MALDI-2 MSI (t-MALDI-2)</li> <li>immunofluorescence microscopy (IFM) using the MACSima system from Miltenyi </li> </ul> <p>msiFlow was tested on MALDI-2 MSI, t-MALDI-2 MSI and IFM data of control and UPEC-infected mouse bladder sections. In IFM we used Ly6G and actin for staining neutrophils and the muscle layer. We validated msiFlow on MALDI MSI data of bone marrow (BM)-derived neutrophils. Tentative lipid annotations were validated by MALDI DDA MSI and MALDI MS/MS. All data used and results generated by msiFlow are included in this dataset (besides the intermediate results of the MALDI-2 preprocessing due to data size).</p> <p>The dataset contains the following zip files:</p> <table> <tbody> <tr> <td><strong>zip file</strong></td> <td><strong>description</strong></td> </tr> <tr> <td>ly6g_heterogeneity.zip</td> <td>example and result data (Ly6G clusters) for molecular_heterogeneity_flow</td> </tr> <tr> <td>if_segmentation.zip</td> <td>example and result data (Ly6G segmentation) for if_segmentation_flow</td> </tr> <tr> <td>ly6g_heterogeneity_signatures.zip</td> <td>example and result data (lipids for Ly6G clusters) for molecular_signatures_flow</td> </tr> <tr> <td>ly6g_molecular_signatures.zip</td> <td>example and result data (lipids for Ly6G) for molecular_signatures_flow</td> </tr> <tr> <td>msi_if_registration.zip</td> <td>example and result data for msi_if_registration_flow</td> </tr> <tr> <td>msi_segmentation.zip</td> <td>example and result data (segmented MSI bladder data) for msi_segmentation_flow</td> </tr> <tr> <td>region_group_analysis.zip</td> <td>example and result data (regulated lipids in different bladder tissue regions) for region_group_analysis_flow</td> </tr> <tr> <td>macsima.zip</td> <td>raw IFM data of UPEC-infected bladders containing Ly6G, actin and autofluorescence images</td> </tr> <tr> <td>maldi-bm-neutrophils.zip</td> <td>raw and pre-processed MALDI MSI data of BM-derived neutrophils</td> </tr> <tr> <td>t-maldi-2.zip</td> <td>raw t-MALDI-2 MSI data of a UPEC-infected bladder section</td> </tr> <tr> <td>maldi-2-<em>group-sampleno</em>.zip</td> <td>raw MALDI-2 MSI data of a control/UPEC bladder section</td> </tr> <tr> <td>MALDI_DDA_MSI.zip</td> <td>raw MALDI MSI data acquired in DDA mode</td> </tr> <tr> <td>TIMS_MS_MS.zip</td> <td>raw MALDI TIMS MS/MS data</td> </tr> </tbody> </table> <p> </p>
Optimized Analytical Workflow for Single-Nucleus Transcriptomics in Main Metabolic Tissues
<p><span>Single-nucleus RNA sequencing (snRNA-seq) has emerged as a powerful approach for studying cellular heterogeneity in metabolic tissues. However, snRNA-seq analysis remains challenging due to low gene expression and data complexity. Here, we introduce an optimized analytical workflow for snRNA-seq data from 67 samples across four main metabolic tissues white adipose tissue, hypothalamus, muscle and liver. We emphasized the importance of key steps including ambient RNA removal, doublet identification, normalization and data integration to ensure accurate downstream analysis. </span><span>This workflow </span><span>offers a valuable resource for researchers in metabolism, facilitating deeper insights into cellular diversity and metabolic function through rigorous snRNA-seq analysis.</span></p>
Digital Library Mnemosine. Workflow for creating new collections
<p><span>Flow chart for semi-automatic creation of new collections in Mnemosine.</span></p>
Optimized Analytical Workflow for Single-Nucleus Transcriptomics in Main Metabolic Tissues
<p>Single-nucleus RNA sequencing (snRNA-seq) has emerged as a powerful approach for studying cellular heterogeneity in metabolic tissues. However, snRNA-seq analysis remains challenging due to low gene expression and data complexity. Here, we introduce an optimized analytical workflow for snRNA-seq data from 67 samples across four main metabolic tissues white adipose tissue, hypothalamus, muscle and liver. We emphasized the importance of key steps including ambient RNA removal, doublet identification, normalization and data integration to ensure accurate downstream analysis. This workflow offers a valuable resource for researchers in metabolism, facilitating deeper insights into cellular diversity and metabolic function through rigorous snRNA-seq analysis.</p>
IWC : Test Data For VGP Decontamination Workflow
<p>Dataset used to test the decontamination workflow published with the VGP assembly pipeline in Galaxy. </p>
Acute pseudo-landmarking and Constellation homologies: A generalized workflow to identify and track segmented structures in plant time series images
<p>Assessing plant phenotypes throughout the lifecycle is integral to exploring the development, genetics, and evolution of morphology, and can be critical for agronomic and basic research studies. Although various automated or semi-automated phenomic approaches have been developed, it has been challenging to analyze differential growth because of difficulties in segmenting and annotating specific structures or positions in the plant body and maintaining their identities throughout time-series data. To address this gap, we have developed a generalized workflow linking our previously published function, <i>Acute</i>, with a companion homology workflow, <i>Constellation</i>, in the PlantCV environment. <i>Acute</i> identifies acute shapes (pseudo-landmarks) in the plant body, most often corresponding to leaf tips and ligular regions. <i>Constellation</i> uses a strategy of dimensionality reduction via <i>starscape</i> followed by hierarchical clustering through <i>constella </i>to identify 'constellations' of segments in eigenspace that represent the same landmark in consecutive images of a time-series. We devised a quality control function, <i>constellaQC</i>, to test the accuracy of the clustering approach, and use it to show that the approach appropriately clusters the pseudo-landmarks derived from <i>Acute</i>, with 80-90% accuracy. We discuss the reasons for and consequences of this lack of 100% accuracy in automated workflows and suggest how to develop these functions for other phenomics datasets that may vary in dimensional complexity.</p>
Why Workflows
<p>An open seminar about Workflow Management Systems in the context of Open Science, happened on 28/01/2022 and organized by the <a href="https://bio-it.embl.de/">Bio-IT</a> project.</p> <p>Speakers:</p> <ul> <li>Victoria Tianjing Yan</li> <li>Bastian Drees</li> <li>Georg Zeller</li> <li>Jean-Karim Heriche</li> </ul> <p>Organizers:</p> <ul> <li>Lisanna Paladin</li> <li>Renato Alves</li> </ul>
A formative usability study of workflow management systems in label-free digital pathology - Data and Code
<p>This repository holds the necessary data and code as well as a descriptive Readme file that was used for our publication "A formative usability study of workflow management systems in label-free digital pathology" by Markus Jelonek et al. (2022), submitted to F1000Research.</p> <p> </p> <p>Abstract:</p> <p>We present a formative usability study that investigates the usability of different<br> workflow management systems in the field of biomedical data analysis. Specifically, we study a task in the field of so-called label-free digital pathology and investigate one graphical user interface based workflow and one script-based workflow to solve the task. Our main intention is to gain first insights into the systematic study of usability in the context of biomedical image analysis, and formulate experiences and guidelines for future usability studies dealing with workflow management systems. Embedded in a specific setup dealing with label-free digital pathology, the core question behind our contribution is how usability studies for scientific workflow management can be conducted, and how they can be used systematically to improve such tools. Further, we address specific questions about the resource utilisation and management of usability studies, including the recruitment of participants as well as the design of specific workflows to be investigated.</p>
A formative usability study of workflow management systems in label-free digital pathology - Questionnaires
<p>This repository holds the necessary questionnaires, participant data, interview questions and data, as well as a descriptive Readme file that was used for our publication "A formative usability study of workflow management systems in label-free digital pathology" by Markus Jelonek et al. (2022), submitted to F1000Research.</p> <p> </p> <p>Abstract:</p> <p>We present a formative usability study that investigates the usability of different<br> workflow management systems in the field of biomedical data analysis. Specifically, we study a task in the field of so-called label-free digital pathology and investigate one graphical user interface based workflow and one script-based workflow to solve the task. Our main intention is to gain first insights into the systematic study of usability in the context of biomedical image analysis, and formulate experiences and guidelines for future usability studies dealing with workflow management systems. Embedded in a specific setup dealing with label-free digital pathology, the core question behind our contribution is how usability studies for scientific workflow management can be conducted, and how they can be used systematically to improve such tools. Further, we address specific questions about the resource utilisation and management of usability studies, including the recruitment of participants as well as the design of specific workflows to be investigated.</p>
An automated workflow for parallel processing of large multiview SPIM recordings
<p>Selective Plane Illumination Microscopy (SPIM) allows to image developing organisms in 3D at unprecedented temporal resolution over long periods of time. The resulting massive amounts of raw image data requires extensive processing interactively via dedicated graphical user interface (GUI) applications. The consecutive processing steps can be easily automated and the individual time points can be processed independently, which lends itself to trivial parallelization on a high performance computing (HPC) cluster. Here, we introduce an automated workflow for processing large multiview, multichannel, multiillumination time-lapse SPIM data on a single workstation or in parallel on a HPC cluster. The pipeline relies on <em>snakemake</em> to resolve dependencies among consecutive processing steps and can be easily adapted to any cluster environment for processing SPIM data in a fraction of the time required to collect it.</p>
101K workflow jobs dataset (1.2 million tasks); a composition of Epigenomics and Montage workflows
<p>This dataset contains detailed information about more than 101 thousand jobs with 1.2 million tasks. This dataset can be used in various simulation, training, and modelling processes. various information about each job is recorded and the underlying environment was based on IoT/Fog/Cloud in which IoT nodes only generate requests and Fog and Cloud nodes process these requests.</p>
Non-informative dataset to be used with GA-VirReport workflow
<p>Input dataset of plant chloroplast, mitochondria and rRNA</p>
CWL run of Somatic Variant Calling Workflow (CWLProv 0.5.0 Research Object)
<p>The somatic variant calling workflow included in this case study is designed by <a href="http://bcb.io/">Blue Collar Bioinformatics (bcbio)</a>, a community-driven initiative to develop best-practice pipelines for variant calling, RNA-seq and small RNA analysis workflows. According to the documentation, the goal of this project is to facilitate the automated analysis of high throughput data by making the resources quantifiable, analyzable, scalable, accessible and reproducible.</p> <p>All the underlying tools are containerized, facilitating software use in the workflow. The somatic variant calling workflow defined in CWL is available on GitHub and equipped with a well defined test dataset.</p> <p>This dataset folder is a CWLProv Research Object that captures the Common Workflow Language execution provenance, see <a href="https://w3id.org/cwl/prov/0.5.0">https://w3id.org/cwl/prov/0.5.0</a> or use <a href="https://pypi.org/project/cwlprov/">https://pypi.org/project/cwlprov/</a> to explore</p> <p><strong>Steps to reproduce</strong></p> <p>To build the research object again, use Python 3 on macOS. Built on:</p> <ul> <li>Processor 2.8GHz Intel Core i7</li> <li>Memory: 16GB</li> <li>OS: macOS High Sierra, Version 10.13.3</li> <li>Storage: 250GB</li> </ul> <p>To run the workflow:<br> </p> <pre><code class="language-bash">pip3 install cwltool==1.0.20180912090223 git clone https://github.com/FarahZKhan/bcbio_test_cwlprov cd bcbio_test_cwlprov/somatic/somatic-workflow/ cwltool --provenance somaticwf_0.5.0_mac main-somatic.cwl main-somatic-samples.json</code></pre> <p>To package the research object:<br> </p> <pre><code class="language-bash">zip -r somaticwf_0.5.0_mac.zip somaticwf_0.5.0_mac/ sha256sum somaticwf_0.5.0_mac.zip > somaticwf_0.5.0_mac.zip.sha256</code></pre> <p>The <a href="https://github.com/FarahZKhan/bcbio_test_cwlprov">cloned git repository</a> is a fork of <a href="https://github.com/bcbio/test_bcbio_cwl">https://github.com/bcbio/test_bcbio_cwl</a>. It was obtained using:</p> <pre><code class="language-bash">wget -O test_bcbio_cwl.tar.gz https://github.com/bcbio/test_bcbio_cwl/archive/master.tar.gz</code></pre> <p>The content is from an archived version from the documentation here: <a href="https://bcbio-nextgen.readthedocs.io/en/latest/contents/cwl.html#install-bcbio-vm-with-containers">https://bcbio-nextgen.readthedocs.io/en/latest/contents/cwl.html#install-bcbio-vm-with-containers</a></p>
Sample data for "Live Cell Fluorescence Microscopy – An End-to-End Workflow for High-Throughput Image and Data Analysis"
<p>This repository contains:</p> <ul> <li> <p>Sample data for the "Live Cell Fluorescence Microscopy – From Sample Preparation to Numbers and Plots" methodology paper by Zahumensky & Malinsky. The paper describes the preparation of live yeast cell samples for microscopy, the subsequent semi-automatic analysis of the microscopy images using our custom-written Fiji macros, and automatic processing of the output (Results table) from the image analys using custom-written R scripts. The data provided here are real experimental data from two publications of our group: Zahumensky et al., 2022 and Vesela et al., 2023</p> </li> <li> <p>"Results tables" from the Fiji based analysis</p> </li> <li> <p>Outputs of the processing of these Results tables using our R scripts, in the form of summary tables, graphs, and statistical analyses</p> </li> </ul>
Supplemental Data for "Adaptive Container Service: a New Paradigm for Robust and Optimized Bioinformatics Workflow Deployment in the Cloud."
<p>All supplemental data for "Adaptive Container Service: a New Paradigm for Robust and Optimized Bioinformatics Workflow Deployment in the Cloud."<br><br>Abstract:<br>We propose Adaptive Container Service (ACS), a new paradigm for deploying bioinformatics workflows in cloud computing environments. By encapsulating the entire workflow within a single virtual container, combined with automatic workflow checkpointing and dynamic migration to appropriately scaled containers, ACS-based deployment demonstrates several key advantages over alternative strategies: it enables optimal resource provision to any workflow that comprise of multiple applications with diverse computing needs; it provides protection against application-agnostic out-of-memory (OOM) errors or spot instance interruptions; and it reduces efforts required for workflow development, optimization, and management because it runs workflows with minimal or no code modifications. Proof-of-concept experiments show that ACS avoided both under- and over-provisioning in monolithic single-container deployment. Despite being deployed as a single container, it achieved comparable resource utilization efficiency as optimized Nextflow-managed, multi-modular workflows. Analysis of over 18,000 workflow runs demonstrated that ACS can effectively reduce workflow failures by two-thirds. These findings suggest that ACS frees developers from navigating the complexity of deploying robust workflows and rightsizing compute resources in the cloud, leading to significant reduction in workflow development time and savings in cloud computing costs.<br><br>Contains the following directories:<br>Fig2-bbtools: running metrics for BBTools<br>Fig3-rna-seq: running metrics for RNA-Seq<br>Fig4-Job_records: meta data and running metrics of 18,000+ jobs</p>
Demo dataset for: SPACEc, a streamlined, interactive Python workflow for multiplexed image processing and analysis
<p>Multiplexed imaging technologies provide insights into complex tissue architectures. However, challenges arise due to software fragmentation with cumbersome data handoffs, inefficiencies in processing large images (8 to 40 gigabytes per image), and limited spatial analysis capabilities. To efficiently analyze multiplexed imaging data, we developed SPACEc, a scalable end-to-end Python solution, that handles image extraction, cell segmentation, and data preprocessing and incorporates machine-learning-enabled, multi-scaled, spatial analysis, operated through a user-friendly and interactive interface.</p> <p>The demonstration dataset was derived from a previous analysis and contains TMA cores from a human tonsil and tonsillitis sample that were acquired with the Akoya PhenocyclerFusion platform. The dataset can be used to test the workflow and establish it on a user's system or to familiarize oneself with the pipeline.</p>
FAIRmat Tutorial 7: Molecular Dynamics Trajectories and Workflows in NOMAD
<p>The FAIRmat consortium is committed to extending the NOMAD infrastructure to a wide variety of materials science data. To support soft matter simulations (e.g., atomistic molecular dynamics simulations), a number of challenges arise, primarily due to the volume and variety of data. The FAIRmat team is working to overcome these challenges, and the NOMAD infrastructure is now equipped with new metadata, features, and tools specifically designed to ease the FAIR treatment of trajectory data and workflows. Parsers have been implemented for two of the most popular molecular dynamics codes (Gromacs and Lammps), with plans for quick expansion to additional codes within the next year. The NOMAD Metainfo now describes the system’s hierarchical structure (in terms of bond topology) through the concept of fixed chemical bonds defined within classical force fields. The NOMAD GUI provides a bespoke overview page for molecular dynamics data, which includes tools that ease visualization of the system topology and automatically displays structural, dynamic, and thermodynamic observables that can assist in a fast assessment of system equilibration. Additionally, a native workflow visualizer allows the user to connect individual simulation entries into complex workflows. Finally, the NOMAD Python module facilitates custom trajectory analysis, for instance in a Jupyter notebook, with functions that convert a NOMAD archive entry to an instance of the MDAnalysis data class.</p> <p>This tutorial invites both experienced and completely novice NOMAD users to learn about these new features for molecular dynamics trajectories. A brief introduction to the FAIRmat consortium and the NOMAD infrastructure will be given, followed by guided and interactive tutorials highlighting the various features described above</p> <p><strong>Disclaimer:</strong> NOMAD is being continuously developed based on input and feedback from the scientific community. Hence the features, services or interface may have changed since the time of recording of this video. For up-to-date information please consult our latest tutorials and the NOMAD documentation <a href="https://nomad-lab.eu/prod/v1/docs/">https://nomad-lab.eu/prod/v1/docs/</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.