Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
650
datasets available to search
ShareScore release 0.9.0
Dataset results
650 results for “Workflow”
Workflow Trace Archive askalon-new_ee65 trace
BWA (short for Burroughs-Wheeler Alignment tool) is a genomics analysis workflow, courtesy of Scott Emrich and Notre Dame Bioinformatics Laboratory. It maps low-divergent sequences against a large reference genome, such as the human genome.
Workflow Trace Archive LANL_Mustang trace
This workload was published by Amvrosiadis et al. as part of their ATC 2018 paper titled "On the diversity of cluster workloads and its impact on research results".
Workflow Trace Archive askalon-new_ee17 trace
Wien2k uses a full-potential Linearized Augmented Plane Wave (LAPW) approach for the computation of crystalline solids.
Workflow Trace Archive askalon-new_ee53 trace
Wien2k uses a full-potential Linearized Augmented Plane Wave (LAPW) approach for the computation of crystalline solids.
Workflow Trace Archive askalon-new_ee54 trace
Wien2k uses a full-potential Linearized Augmented Plane Wave (LAPW) approach for the computation of crystalline solids.
Workflow Trace Archive askalon-new_ee21 trace
Wien2k uses a full-potential Linearized Augmented Plane Wave (LAPW) approach for the computation of crystalline solids.
Workflow Trace Archive askalon-new_ee18 trace
Wien2k uses a full-potential Linearized Augmented Plane Wave (LAPW) approach for the computation of crystalline solids.
Workflow Trace Archive askalon-new_ee63 trace
BWA (short for Burroughs-Wheeler Alignment tool) is a genomics analysis workflow, courtesy of Scott Emrich and Notre Dame Bioinformatics Laboratory. It maps low-divergent sequences against a large reference genome, such as the human genome.
Workflow Trace Archive askalon-new_ee13 trace
Wien2k uses a full-potential Linearized Augmented Plane Wave (LAPW) approach for the computation of crystalline solids.
Workflow Trace Archive askalon-new_ee44 trace
Wien2k uses a full-potential Linearized Augmented Plane Wave (LAPW) approach for the computation of crystalline solids.
Workflow Trace Archive askalon-new_ee61 trace
BWA (short for Burroughs-Wheeler Alignment tool) is a genomics analysis workflow, courtesy of Scott Emrich and Notre Dame Bioinformatics Laboratory. It maps low-divergent sequences against a large reference genome, such as the human genome.
Workflow Trace Archive Google trace
<p>This workload contains the popular Google cluster trace (2014) in the workflow trace archive format.</p>
reproducible geospatial scientific workflows
<p>Additional material for submission to JSS.</p> <p>Manuscript ID: TJSS-2018-0191.r1</p>
CLIJ: Benchmarking workflow results
<p>CLIJ Workflow Benchmarking leads to a folder of resulting images. As two GPUs and two CPUs were tested, four of these result folders exist. Additionally, statistics CSV files are stored in this data set.</p> <p>Read more in <a href="https://www.biorxiv.org/content/10.1101/660704v1">https://www.biorxiv.org/content/10.1101/660704v1</a></p>
A Graph Neural Network Based Workflow for Real-time Lightning Location with Continuous Waveforms
<p>The dataset for "A Graph Neural Network Based Workflow for Real-time Lightning Location with Continuous Waveforms" can be divided into training and validation sets at any desired ratio.</p> <p> </p> <p>The code has been published on GitHub: <a href="https://github.com/cqtian-kk/Lightning_Detection_Location">Lightning_Detection_Location</a> or <a href="https://zenodo.org/records/14048427">DOI 10.5281/zenodo.13350849</a></p>
Analyzing Scientific Workflow Management Systems
<p>Contains data and scripts to reproduce the test runs that were used in the thesis "Analyzing Scientific Workflow Management Systems", and the complete result data.</p>
FASTA or Tabular Feature Retriever galaxy workflow example output
<p>FASTA Feature Retriever and Tabular Feature Retriever are galaxy workflow that retrieves features(like genes) in fasta or tabular format using .bed and genome .fasta as input</p> <p>input data derived from:</p> <ol> <li><a href="https://solgenomics.net/" target="_blank" rel="noopener">solgenomics.net</a></li> <li><a title="mycocosm" href="https://mycocosm.jgi.doe.gov/Fusso1/Fusso1.home.html" target="_blank" rel="noopener">mycocosm</a></li> </ol>
A Unified Multi-Wavelength Data Analysis Workflow with gammapy.
<p>Repository containing multi-wavelength (MWL) data of the distant quasar OP 313 (z=0.997) taken during the nights of MJD 60373 (3-4 March 2024) and MJD 60384 (14-15 March 2024) using <em>Fermi</em>-LAT, <em>Swift</em>-XRT, <em>Swift</em>-UVOT, and the Liverpool IO:O photometric data.</p> <div> <h3>Directory structure:</h3> <a href="https://github.com/mireianievas/gammapy_mwl_workflow/tree/main#directory-structure"></a></div> <ul> <li>Notebooks/DatasetGenerator: Notebooks (one per instrument) that summarize the steps to generate gammapy-compliant 1D and 3D binned datasets.</li> <li>Notebooks/DatasetAnalysis: Notebooks (one per instrument) to analyse each dataset independently and one notebook to perform the MWL joint analysis.</li> <li>Helpers: Auxiliary functions to generate native multiplicative models for dust extinction (out of xspec's redden), neutral hydrogen (out of xspec's tbabs), EBL absorption, and utility functions for file handling and plotting.</li> <li>Models: multiplicative models in tabular format for easy 2d interpolation.</li> <li>Figures: collection of figures for the paper.</li> </ul> <p>Link: <a href="https://github.com/mireianievas/gammapy_mwl_workflow">https://github.com/mireianievas/gammapy_mwl_workflow/ </a></p> <p>Based on the original work from <a href="https://github.com/luca-giunti/gammapyXray">https://github.com/luca-giunti/gammapyXray</a></p>
The AstroPath Image Acquisition and Segmentation Workflow
<p>Multidimensional, spatially resolved analyses of cells from pathology slides are of great diagnostic and prognostic interest. New multispectral, multiplex immunofluorescence microscopy platforms have the potential to facilitate such analyses, and here, we further improve and standardize the image acquisition and cell classification workflow. Studies to date on this emerging technology have typically assessed ~10 operator-dependent high power fields (HPFs) per slide, which represents a fraction of the tissue available for study. Standard cell segmentation and classification algorithms often oversegment larger cells, when they are segmented at the same time as smaller cells. Here we describe our AstroPath imaging platform, which addresses each of these considerations. In our study, slides from formalin-fixed paraffin embedded tissue specimens were stained with an optimized 6-plex multiplex immunofluorescence (mIF) assay. The slides were then scanned at 35 unique wavelengths using a multispectral microscope (Vectra 3.0 or Vectra Polaris) with 20% overlap of HPFs in an operator-independent fashion. An average of 1300 HPFs per slide was required to image the entire tissue, and each microscope scanned between 2 to 3 slides per day with this approach. After the images were captured and organized, overlaps were used to measure, quantify and correct systematics in the imagery (see Eminizer abstract). The central parts of the images were used to create a set of seamless “primary” tiles, similar to the strategy of the Sloan Digital Sky Survey, for a statistically fair pixel coverage of the whole tissue area (see Roskes abstract). Images were then linearly unmixed from the 35 wavelengths to 8 component layers (DAPI, tissue auto-fluorescence, and the 6 added fluorescent dyes) using inForm Cell Analysis©. We then employed a bespoke method for ‘multi-pass’ classification of cells wherein each marker was segmented and classified separately from the other markers, then merged into a single plane using a unique set of rules and predefined cell hierarchy. We showed that our segmentation and classification method reduced error in over-counting larger cells, e.g. tumor cells, by 25% and increased the specificity and sensitivity in each classification algorithm. Due to the amount of data, each algorithm was run automatically through one of 20 virtual machines housed on a set of servers in the Physics and Astronomy Department. Following the methodology developed during the SDSS project, image data was stored in a well-defined file system structure that facilitated further automatic processing and ingestion into a SQL Server database. Raw data for each slide was 200-300 GBs, which is on par with a full scale (30x) human genome. In summary, we have developed a unique facility and workflow that generates whole slide multispectral imagery with high-fidelity, single cell resolution. Our facility houses five multispectral microscopes (2 Vectra 3.0 and 3 Vectra Polaris) allowing us to collect a petabyte of raw data per year, on scale of the largest sky survey.</p>
Satellite Datasets used for MIRA Workflows
<p>CANGA Remapping Intercomparison Satellite Data Set</p> <p>Included is the data used to generate sampling fields. Included is the script for generating spatial power spectra and fitting later used in reconstruction over any unstructured spherical grid. This data set is included for reproducibility of results provided in a journal article submission.</p> <p>REQUIRES:</p> <ol> <li><a href="http://code.google.com/p/netcdf4-python/">http://code.google.com/p/netcdf4-python/</a> Python NetCDF IO modules: "pip install netcdf4"</li> <li><a href="https://shtools.oca.eu/shtools/">https://shtools.oca.eu/shtools/</a> Python spherical harmonic tools package: "pip install pyshtools"</li> <li>Numpy</li> <li>Scipy (KDTree search)</li> <li><a href="https://plot.ly/python/">https://plot.ly/python/</a> Plotly (Fancy, web-based plotting)</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.