Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

650

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

650 results for “Workflow”

Learn how ShareScore rates datasets ↗
zenodo32/100

Test data for running snakePipes : noncoding-RNA-seq workflow

<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a>&nbsp;for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the noncoding-RNA-seq workflow under snakePipes. To test the workflow, follow the following steps :&nbsp;</p> <ul> <li>Download or prepare genome fasta, indices and annotations for mouse (<strong>GRCm38</strong>) genome. Repeat Masker file is required for this workflow.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a>&nbsp;with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Figure S1: Tissue detection in StrataQuest software via creation of a digital mask; Figure S2: StrataQuest workflow for detection of SOX2+ nuclei and thresholding.

<p><strong>Figure S1: Tissue detection in StrataQuest software via creation of a digital mask</strong><strong>.</strong> A digital &lsquo;mask&rsquo; is created to measure the area of tissue for quantification of cell densities. The tissue mask is automatically generated by StrataQuest software via conversion of the scanned image from RGB to grayscale and then application of an intensity threshold.&nbsp; Manual adjustments are made to the tissue mask to remove necrotic areas and/or staining artefacts. Only nuclei present within the tissue mask are quantified. Representative tumour cores with low (A) and high (B) SOX2 densities, respectively, with corresponding overlaid tissue masks are shown (C-D; purple colour).&nbsp; Tissue mask generation is based on haematoxylin staining so is not affected by the level of DAB staining. Scale bars 200&micro;m.</p> <p><strong>Figure S2: StrataQuest workflow for detection of SOX2+ nuclei and thresholding. </strong>Inbuilt colour deconvolution algorithms within StrataQuest software separate SOX2 DAB (brown) staining from haematoxylin (blue) staining to produce grayscale images for each channel. Nuclear segmentation was then performed on the resulting grayscale DAB image to detect brown-stained nuclei.&nbsp; This is possible as SOX2 expression is localised to the nucleus. Segmented nuclear masks overlaid onto the DAB (brown) channel (A) and the original colour image (B) allow visualisation of the segmentation algorithm. The software then calculates parameters such as nuclear size, haematoxylin intensity and DAB intensity for each nuclear mask, which are reported on a scattergram, with each dot representing a single nuclear mask. A scattergram displaying DAB intensity vs haematoxylin intensity is used to threshold and accurately detect SOX2+ stained nuclei (C). Gated nuclei from the scattergram (C) can be visualised on an image of segmented nuclear masks overlaid onto the RGB image (D). Red nuclei represent SOX2+ cells (red gate on C) and green nuclei are SOX2- cells (green gate on C). Scale bars 50&mu;m.</p>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Datasets for the computational workflow of multidimensional photoemission spectroscopy

<p>Recorded single-electron event data of bulk 2H-WSe<sub>2</sub>&nbsp;photoemission from a commercial momentum microscope (SPECS METIS 1000).These data are used for demonstration of the computational workflow explained in the following publication.<br> <br> <strong>R. P.&nbsp;Xian, Y.&nbsp;Acremann, S. Y. Agustsson, M. Dendzik, K. B&uuml;hlmann, D.&nbsp;Curcio, D.&nbsp;Kutnyakhov, F. Pressacco, M.&nbsp;Heber, S.&nbsp;Dong, T.&nbsp;Pincelli, J.&nbsp;Demsar, Wilfried Wurth, Ph. Hofmann, M. Wolf, M.&nbsp;Scheidgen, L.&nbsp;Rettig,&nbsp;R.&nbsp;Ernstorfer,&nbsp;An open-source, end-to-end workflow for multidimensional photoemission spectroscopy, Scientific Data 7, 442 (2020). DOI:&nbsp;10.1038/s41597-020-00769-8</strong><br> <br> The zip files are not directly usable for running the computational workflow, but requires first to unzip into HDF5 format (.h5).</p>

opencc-by-4.0Aug 2020View details →
zenodo32/100

Workflow Trace Archive Galaxy trace

Traces from different biomedical research workflows, executed on the public Galaxy server in Europe.

opencc-zeroNov 2020View details →
dryad32/100

Data from: A from-benchtop-to-desktop workflow for validating HTS data and for taxonomic identification in diet metabarcoding studies

The main objective of this work was to develop and validate a robust and reliable 'from benchtop-to-desktop' metabarcoding workflow to investigate the diet of invertebrate-eaters. We applied our workflow to fecal DNA samples of an invertebrate-eating fish species. A fragment of the COI gene was amplified by combining two minibarcoding primer sets to maximize the taxonomic coverage. Amplicons were sequenced by an Illumina MiSeq platform. We developed a filtering approach based on a series of non-arbitrary thresholds established from control samples and from molecular replicates in order to address the elimination of cross-contamination, PCR/sequencing errors and mistagging artifacts. This resulted in a conservative and informative metabarcoding dataset. We developed a taxonomic assignment procedure that combines different approaches and that allowed the identification of ~75% of invertebrate COI variants to the species level. Moreover, based on the diversity of the variants, we introduced a semi-quantitative statistic in our diet study, the Minimum Number of Individuals (MNI), which is based on the number of distinct variants in each sample. The metabarcoding approach described in this paper may guide future diet studies that aim to produce robust datasets associated with a fine and accurate identification of prey items.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Expanding the described metabolome of the marine cyanobacterium Moorea producens JHB through orthogonal natural products workflows

Moorea producens JHB, a Jamaican strain of tropical filamentous marine cyanobacteria, has been extensively studied by traditional natural products techniques. These previous bioassay and structure guided isolations led to the discovery of two exciting classes of natural products, hectochlorin (1) and jamaicamides A (2) and B (3). In the current study, mass spectrometry-based 'molecular networking' was used to visualize the metabolome of Moorea producens JHB, and both guided and enhanced the isolation workflow, revealing additional metabolites in these compound classes. Further, we developed additional insight into the metabolic capabilities of this strain by genome sequencing analysis, which subsequently led to the isolation of a compound unrelated to the jamaicamide and hectochlorin families. Another approach involved stimulation of the biosynthesis of a minor jamaicamide metabolite by cultivation in modified media, and provided insights about the underlying biosynthetic machinery as well as preliminary structure-activity information within this structure class. This study demonstrated that these orthogonal approaches are complementary and enrich secondary metabolomic coverage even in an extensively studied bacterial strain.

opencc-zeroDec 2014View details →
zenodo32/100

Data Processing for a Small-Scale Long-Term Coastal Ocean Observing System Near Mobile Bay, Alabama: A Geoscience Papers of the Future (GPF) Workflow Diagram

<p>The Dauphin Island Sea Lab (DISL) has been operating a permanent moored oceanographic station at 30 05.410&#39;N, 88 12.694&#39;W, 25 km southwest of the entrance to Mobile Bay, Alabama, since 2004. It collects hydrographic and current velocity data.</p> <p>This diagram shows the processing steps for data from the instruments at this mooring, from initial download to initial scientific analysis. The accompanying text explains how to apply the workflow to the example dataset (10.5281/zenodo.18943) using the provided software (10.5281/zenodo.32741).&nbsp;</p> <p>The files&nbsp;have&nbsp;been prepared&nbsp;as supplementary material for a Geoscience Paper of the Future (GPF) in prep for publication at Earth and Space Science, as part of the OntoSoft GPF Initiative.</p>

opencc-by-nc-sa-4.0Nov 2015View details →
zenodo32/100

Music Workflow used in the evaluation of the VFramework

<p>Music Workflow used in the evaluation of the VFramework</p> <p>Original execution Windows, re-execution Linux.</p>

opencc-by-4.0Oct 2016View details →
zenodo32/100

The weather workflow used in demonstration of the VFramework

<p>The weather workflow used in demonstration of the VFramework.</p> <p>The package contains context model for two executions.</p> <p>The workflow is also included.</p>

opencc-by-4.0Oct 2016View details →
zenodo32/100

Artifact of the paper "An Empirical Investigation on the Challenges in Scientific Workflow Systems Development"

<p><strong><span>Scientific Workflow Systems (SWSs)</span></strong><span> play a critical role in the contemporary scientific landscape, significantly enriching research endeavors by augmenting productivity and fostering collaboration</span><span>, SWSs</span><span> elevate the standard of scholarly inquiry, fortifying its pillars of reproducibility and ethical adherence. Essentially, they </span><span>serve as</span><span> the bedrock upon which efficient, transparent, and impactful research </span><span>is built</span><span>, propelling knowledge and innovation across diverse fields. SWSs accomplish mundane yet essential tasks intrinsic to scientific inquiry&mdash;ranging from data acquisition to analysis and reporting. By liberating researchers from the shackles of manual labor, SWSs enable them to channel their energies toward more intellectually demanding pursuits, thereby enhancing the pace and quality of research outcomes. Moreover, SWSs wield a formidable influence in standardizing workflows across research cohorts, instilling a sense of uniformity in experimental methodologies and data-handling practices. This standardization not only cultivates a culture of rigor and coherence but also fosters cross-disciplinary dialogue and collaboration.</span></p> <p>Integral to the operation of SWSs is their capacity to integrate diverse tools, software, and data sources, effectively functioning as centralized hubs for research management. This integration expedites the research process and facilitates seamless data exchange and interoperability&mdash;a pivotal asset in an era characterized by the deluge of data and the imperative of interdisciplinary collaboration. Furthermore, SWSs afford researchers and project managers real-time insights into the progress of research endeavors, empowering them to identify bottlenecks, allocate resources judiciously, and optimize workflow execution. This granular oversight enhances project transparency and accountability and serves as a catalyst for informed decision-making.</p> <p>Crucially, SWSs are engineered to accommodate the complexities inherent in scientific inquiry, adeptly handling vast volumes of data and supporting parallel processing to meet the evolving demands of research projects. This scalability underscores their adaptability to diverse research paradigms, ensuring their relevance across a spectrum of scientific disciplines. Facilitating collaboration across geographic and temporal divides, SWSs offer a suite of collaborative features&mdash;including version control, shared workspaces, and communication tools&mdash;that transcend the constraints of physical proximity. By fostering a culture of inclusivity and knowledge exchange, SWSs catalyze innovation and synergy among distributed research teams.</p> <p>Moreover, SWSs serve as custodians of reproducibility, meticulously documenting each facet of the research workflow&mdash;from data sources to analysis methods&mdash;thus safeguarding the integrity of scientific inquiry. This commitment to transparency and methodological rigor underpins the credibility of research findings, engendering trust within the scientific community and beyond. The customizable nature of SWSs empowers research teams to tailor their workflows to suit their unique needs and preferences, further amplifying their utility and versatility. In essence, SWSs emerge not merely as tools of convenience but as indispensable allies in the relentless pursuit of scientific excellence.</p> <p>Numerous developers actively participate in the advancement of SWSs through diverse roles, including designing system architectures to ensure flexibility and performance, developing algorithms for data processing and analysis, crafting user-friendly interfaces, handling backend logic, integrating with external tools, and ensuring quality, security, and compliance. They address challenges such as optimizing performance and scalability by leveraging parallel processing and distributed computing techniques. To tackle these diverse tasks, developers encounter numerous challenges, often turning to crowd-sourced platforms like Stack Overflow and GitHub to discuss and address them. Stack Overflow serves as a vital resource for developers to seek solutions, learn new technologies, validate best practices, and engage with the programming community. Similarly, GitHub facilitates collaborative development by allowing developers to report problems, propose enhancements, and contribute to open-source projects. Our research draws insights from Stack Overflow discussions, GitHub issues, and pull request reports related to SWSs, reflecting the dynamic and collaborative nature of software development in this domain.</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Dataset for "NanoVar: a Comprehensive Workflow for Structural Variant Detection to uncover the Genome's Hidden Patterns"

<h2><strong>Output Files for Long-Read Structural Variant and Repeat Analysis in Colorectal Cancer Samples (HRR698464, HRR698460, C586, C588)</strong></h2> <h3>Description:</h3> <p>This Zenodo dataset includes comprehensive output files generated during the application of a long-read sequencing analysis protocol for structural variant (SV) detection and repeat element characterization in colorectal cancer samples. The dataset is organized into two main directories:</p> <p><strong>1. HRR698464_MSI-H_Tumor</strong><br>This directory contains all primary output files generated from the analysis pipeline applied to the MSI-H tumor sample HRR698464 (also referred to as patient C586.T). Each subdirectory corresponds to a specific stage in the protocol:</p> <ul> <li>NanoPlot_output<br>Output from Stage 1 &ndash; Quality assessment of raw reads using NanoPlot.</li> <li>SAMtools_output<br>BAM file processing outputs from Stage 2 &ndash; Alignment of long reads to the reference genome using SAMtools.</li> <li>NanoVar_output<br>Output from Stage 3 &ndash; Structural variant calling using NanoVar.</li> <li>VCF_filtering_output<br>Output from Stage 4 &ndash; Filtering of structural variants using SURVIVOR and BCFtools; includes the filtered VCF files.</li> <li>NanoINSight_output<br>Output from Stage 5 &ndash; Characterization of repeat elements using NanoINSight.</li> <li>VEP_output<br>Output from Stage 6 &ndash; Annotation of structural variants using Ensembl Variant Effect Predictor (VEP).</li> </ul> <p>&nbsp;</p> <p><strong>2. Additional_output_files</strong><br>This directory contains supplementary output files used for comparison and visualization in Figures 4&ndash;7 of the associated publication. These include:</p> <ul> <li>HRR698460.NanoPlot.report.html<br>NanoPlot quality summary of a lower-quality tumor sample (HRR698460), used in Figure 4 for comparison with HRR698464.</li> <li>C586.N.nanovar.pass.vcf<br>NanoVar VCF output for the matched normal sample of patient C586, used to filter somatic calls in Stage 4.</li> <li>C586.N.nanovar.pass.report.html<br>NanoVar summary report of the normal sample of C586; used in Figure 5a.</li> <li>C588.N.nanovar.pass.vcf<br>NanoVar VCF output of the MSS normal sample (C588) for comparison with the MSI-H patient (C586).</li> <li>C588.N.nanovar.pass.report.html<br>NanoVar summary report of the MSS normal sample; used in Figure 5b.</li> <li>C588.T.nanovar.pass.vcf<br>NanoVar VCF output of the MSS tumor sample (C588); used in comparative analyses with the MSI-H sample.</li> <li>C588.T.nanovar.pass.report.html<br>NanoVar summary report of the MSS tumor sample; used in Figure 5b.</li> <li>MSS.tumor.unique.vcf<br>VCF file of somatic SVs in the MSS sample, generated by comparing matched tumor and normal pairs.</li> <li>MSS.tumor.unique.RepeatMasker.tbl<br>RepeatMasker output annotating somatic insertions in the MSS tumor sample; used in Figure 6.</li> <li>MSS.tumor.unique.vep.html<br>Ensembl VEP annotation report of somatic SVs in the MSS patient; used in Figures 7a and 7b.</li> <li>This dataset supports reproducibility and transparency of the protocol and offers a valuable resource for researchers interested in long-read-based SV detection, repeat annotation, and comparative cancer genomics.</li> </ul>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Predictive Design of Ultrastretchable Electrodes with Strain-Insensitive Performance via Robotics- and Machine Learning-Integrated Workflow

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Snakemake workflow for Illumina bacterial assembly and QC v1.0.1

<p>This Illumina bacterial assembly snakemake repo is a snakemake workflow to assemble bacterial genomes from Illumina paired-end short-reads (typically 150bp); It also performs several QC steps on the reads and the resulting assemblies. Everything needed can be found in this repository (except for the GTDB reference database for taxonomic assembly classification and the Kraken2 database for taxonomical read classification, installation instructions provided in the README). The latest version of the scripts can be found at https://gitlab.ilvo.be/genomics/wgs/illumina-bacterial-assembly-snakemake and was updated to v1.0.1 with interactive and printer-friendly HTML and PDF reports summarising all QC metrics of reads and assemblies, tool versions used and a graphical workflow overview.</p>

opencc-by-4.0Sep 2025View details →
zenodo32/100

Test datasets for the iwc workflow : Workflow VGP8 - HiC scaffolding

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
zenodo32/100

Test data for iwc workflow : Purgedups VGP6

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
zenodo32/100

Test Datasets for iwc workflow VGP5 : Trio Genome Assembly

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
zenodo32/100

Test datasets for the iwc workflow : Scaffolding with Bionano VGP7

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
zenodo32/100

Test datasets for iwc workflow : Assembly with HiC phasing VGP4

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo32/100

Maximizing Data Utility for HPC Python Workflow Execution

<p>Large-scale HPC workflows are increasingly implemented in dynamic languages such as Python, which allow for more rapid development than traditional techniques. However, the cost of executing Python applications at scale is often dominated by the distribution of common datasets and complex software dependencies. As the application scales up, data distribution becomes a limiting factor that prevents scaling beyond a few hundred nodes. To address this problem, we present the integration of Parsl (a Python-native parallel programming library) with TaskVine (a data-intensive workflow execution engine). Instead of relying on a shared filesystem to provide data to tasks on demand, Parsl is able to express advance data needs to TaskVine, which then performs efficient data distribution at runtime. This combination provides a performance speedup of 1.48x over the typical method of on-demand paging from the shared filesystem, while also providing an average task speedup of 1.79x with 2048 tasks and 256 nodes.</p>

opencc-by-4.0Sep 2023View details →
zenodo32/100

Dragon Proxy Runtimes and Multi-system Workflows

<p>We present a novel method for obtaining proxy access to remote instances of the Dragon distributed runtime. Dragon is a composable distributed runtime for managing dynamic processes, high-performance communication objects, memory and data at scale that is based on an abstraction of a distributed system. Proxy access to a remote instance of the Dragon runtime allows the client Dragon runtime to run any command that could be run directly by the remote Dragon runtime, but executes the command on the remote runtime. Commands to be run on a remote Dragon runtime are mediated by a Python object that acts as a proxy for the remote runtime, which we call a \textit{proxy runtime}. These proxy runtimes, combined with the ability to start and tear down remote Dragon runtimes both programmatically and via the command line interface, make a number of challenging workflows simple to program. Such workflows include edge-to-cloud scientific workflows, batch services and scientific applications based on Python multiprocessing. The ability to program complex workflows on systems that span clusters, scientific instruments and cloud resources is critical to the development of post-exascale applications, infrastuctures and frameworks.</p>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record