Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
650
datasets available to search
ShareScore release 0.9.0
Dataset results
650 results for “Workflow”
Data of "A workflow to study the microbiota profile of piglet's umbilical cord blood: from sampling to data analysis".
<p>The present study proposes a workflow – from the sampling method to DNA extraction, bioinformatics and data analysis – that characterises the bacterial profile of umbilical cord blood samples, taking into account the contaminants found throughout the procedure of bacterial DNA extraction and amplification.</p> <p>Ps_umbilical.rds: A phyloseq object file of data containing the amplicon sequences variants (ASVs) of thirteen umbilical cord samples and two negative control samples, created by DADA2.</p> <p>R-script.doc: A word document containing the scripts used to characterize the taxonomical composition of the fifteen umbilical cord samples and two negative control samples before and after the application of Decontam R-package (Davis et al., 2018).</p> <p>metadata.docx: meta data for R-script.doc</p>
Triple-resolution of spectral phases via semi-relativistic ab-initio RABBITT simulations DATA & WORKFLOW
<p>This dataset contains the necessary atomic structure and input files to use the <a href="https://gitlab.com/Uk-amor/RMT/rmt">R-Matrix with Time-dependence code suite</a> (open source and freely available) to replicate the results presented in "Triple-resolution of spectral phases via semi-relativistic ab-initio RABBITT simulations".</p> <p>Additionally, the output photoelectron momentum spectra data output from the RMT simulations are provided, to allow replication of the post-processing and spectral phase extraction processes in the absence of access to a large HPC cluster.</p> <p>Finally, a link is provided to a <a href="https://gitlab.com/lukeroantree/argon_rabbitt_scripts">git repository</a> hosted on gitlab.com where post-processing, spectral phase extraction, and visualisation tools are available to operate on these momentum spectra, and an interactive example is provided via a webhosted (via mybinder) Python Jupyter notebook.</p>
MALDI-MS dataset for use with open-source untargeted metabolomic workflow for complex biological samples
<p class="MsoNormal">Untargeted metabolomics is a powerful tool for measuring and understanding complex biological chemistries. However, employment, bioinformatics and downstream analysis of mass spectrometry (MS) data can be daunting for inexperienced users. Numerous open-source and free to-use data processing and analysis tools exist for various untargeted MS approaches, but choosing the 'correct' pipeline isn't straight-forward. This data set can be used in conjunction with a user-friendly online guide which presents a workflow for connecting these tools to process, analyse and annotate various untargeted MS datasets. The workflow is intended to guide exploratory analysis in order to inform decision-making regarding costly and time-consuming downstream targeted MS approaches. The workflow provides practical advice concerning experimental design, organisation of data and downstream analysis, and offers details on sharing and storing valuable MS data for posterity. The workflow is editable and modular, allowing flexibility for updated/ changing methodologies and increased clarity and detail as user participation becomes more common allowing contributions and improvements to the workflow via the online repository. </p>
The digital workflow and the IIIF initiatives at the Vatican Library
<p>The presentation focuses on:<br> 1) the digital workflow for long-term preservation and the web dissemination of the Vatican Digital Library.<br> 2) the Library's IIIF initiatives and the implementation of the "Thematic Pathways on the Web: IIIF annotations of manuscripts from the Vatican collections".</p>
Run of digital pathology tissue/tumor prediction workflow
<p>This dataset is an <a href="https://www.researchobject.org/ro-crate/">RO-Crate</a> representation of an execution of the tissue/tumor prediction workflow for digital pathology from <a href="https://github.com/crs4/deephealth-pipelines/tree/c54840df08742e3aa454394e0e74d15fbd640f07">crs4/deephealth-pipelines</a>. It follows the <a href="https://w3id.org/ro/wfrun/provenance/0.1">Provenance Run Crate</a> profile. The workflow has been run with <a href="https://github.com/common-workflow-language/cwltool/tree/3.1.20230213100550">cwltool</a>, using the --provenance option to generate a <a href="https://doi.org/10.1093/gigascience/giz095">CWLProv</a> RO bundle, and then converted to an RO-Crate using <a href="https://github.com/ResearchObject/runcrate/tree/755fb7f0a8ba6fc238a2cb7a3218175644eb78b5">runcrate</a>. The input dataset is <a href="https://openslide.cs.cmu.edu/download/openslide-testdata/Mirax/Mirax2-Fluorescence-2.zip">Mirax2-Fluorescence-2</a> by Yves Sucaet, from the <a href="https://openslide.cs.cmu.edu/download/openslide-testdata/Mirax/">MIRAX test data</a>.<br> </p>
Run of an example Galaxy collection workflow
<p>This dataset is an <a href="https://www.researchobject.org/ro-crate/">RO-Crate</a> representation of an execution of an example Galaxy workflow, making use of some of Galaxy's platform specific features. It follows the <a href="https://w3id.org/ro/wfrun/workflow/0.1">Workflow Run Crate</a> profile. The workflow has been run with <a href="https://docs.galaxyproject.org/en/latest/index.html">Galaxy version 23.0</a> and exported using the implemented <a href="https://galaxyproject.org/news/2023-02-23-structured-data-exports-ro-bco/">export invocation to RO-crate feature</a>.</p>
Workflow for Remote Sensing for Forest Dynamics and Its Implications for Tree Outside Forest over Maryland, U.S.A.
<p>The workflow shows the process of using the data to plot the figures in the paper "Remote Sensing for Forest Dynamics and Its Implications for Tree Outside Forest over Maryland, U.S.A."</p>
Comparison between the results from JGA analysis somatic short variant discovery workflow and those from the compatible Terra workflow
<p>Files starting from <code>HCC1143.somatic</code> are the results from <a href="https://github.com/ddbj/jga-analysis/tree/main/somatic-short-variant">JGA analysis somatic short variant discovery workflow</a>. Files starting from <code>submissions_</code> are the results from the compatible Terra workflow.</p> <p>VCFs are identical between two workflows except for the header lines. MAFs are also identical except for the header lines.</p>
WRF/EMEP UK Workflow Inputs
<p>Example input data for a linear WRF/EMEP workflow.</p> <p>Setup is a small (60x90 grid cells, 30 vertical levels) domain over the UK, to be run for 4.75 days. Intermediate inputs are included so that different stages of the workflow can be run in isolation.</p> <p>The workflow scripts are available on WorkflowHub: <a href="https://workflowhub.eu/workflows/455">https://workflowhub.eu/workflows/455</a></p>
Comparison between the results from JGA analysis mitochondrial short variant discovery workflow and those from the compatible Terra workflow
<p>Files starting from <code>NA12878.chrM</code> are the results from <a href="https://github.com/ddbj/jga-analysis/tree/mitocondrial-variant">JGA analysis mitochondrial short variant discovery workflow</a>. Files starting from <code>submissions_</code> are the results from the compatible Terra workflow.</p> <p>VCFs are identical between two workflows except for the header lines.</p>
Protein Structure Files and Galaxy Workflows for Conducting Molecular Dynamics Simulations of Flavivirus Helicases -- Output Files
<p>These are the output files generated using the input files and Galaxy workflows for flavivirus helicase simulations, from: </p> <pre>https://doi.org/10.5281/zenodo.7493015</pre>
Tryps-IN: A streamlined palaeoproteomics workflow enables ZooMS analysis of 10,000-year-old petrous bones from Jordan rift-valley
<p>Poor preservation of collagen in dry and/or arid environments has hindered the application of Zooarchaeology by mass spectrometry (ZooMS) analysis in many regions of the world, and as a result many zooarchaeological investigations have relied exclusively on the morphological assessment of fragmentary remains, due to the inadequate preservation of biomolecules. The climatic conditions of Southwest Asia include extreme temperature fluctuations unconducive to preservation of proteins and DNA. We performed zooarchaeological analysis of remains from the 10,000-year-old site of Shkārat Msaied in Jordan and sub-sampled twenty-eight petrous bones, the hardest bone in the mammalian skeleton, for species identification by ZooMS. Using an unconventional and simplified extraction protocol we call Tryps-IN, in which digestion was performed without removal of the demineralising EDTA, we taxonomically identified several fragments, outperforming the established ZooMS work-flow. A subset of identifications was subsequently confirmed using liquid chromatography coupled to tandem mass spectrometry (LC-MS/MS) protein sequencing. The new methodology presented here opens the possibility of further bioarchaeological investigation of other fragmentary faunal assemblages within this region of archaeological significance. </p>
Workflow for structured literature reviews using the Open Research Knowledge Graph (ORKG)
<p>Figure showing a workflow of making a structured literature review using the core features of the Open Research Knowledge Graph (ORKG). </p>
Computational workflows for perovskites: Case study for lanthanide manganites
<p>Supplemental material for the above manuscript. Revised version (v2.0).</p> <p>Includes the complete code archive, including all Quantum ESPRESSO calculation input and output files, as well as postprocessing scripts used to generate the figures in this manuscript.</p>
Exploratory Search Workflows (ESW) collection
<p>This Zenodo digital object represents the dataset of the Exploratory Search Worklflows (ESW) collection. It contains the ontology (owl file) and a folder with the exploratory workflows divided by "track". Each track contains the query logs, the exploratory workflows execution, the exploratory workflows evaluated, the ground truths and the serialized turtle files.</p>
StreamFlow run of digital pathology tissue/tumor prediction workflow
<p>This dataset is an <a href="https://www.researchobject.org/ro-crate/">RO-Crate</a> representation of an execution of the tissue/tumor prediction workflow for digital pathology from <a href="https://github.com/crs4/deephealth-pipelines/tree/c54840df08742e3aa454394e0e74d15fbd640f07">crs4/deephealth-pipelines</a>. It follows the <a href="https://w3id.org/ro/wfrun/provenance/0.1">Provenance Run Crate</a> profile. The workflow has been run with <a href="http://streamflow.di.unito.it">StreamFlow</a>, using the following commands to produce an RO-Crate bundle:</p> <pre><code class="language-bash"># Run the pipeline streamflow run \ --name ml-predict-pipeline-streamflow \ streamflow.yml # Generate the RO-Crate bundle streamflow prov \ --add-file src=README.md,dst=/README.md,about="{\"@id\":\"./\"}",encodingFormat=text/markdown \ --add-property \./.license=https://spdx.org/licenses/MIT \ --add-property \./.name="DeepHealth Pipeline" \ --add-property \./.description="Run of digital pathology tissue/tumor prediction workflow" \ --file streamflow.yml \ ml-predict-pipeline-streamflow</code></pre> <p>The input dataset is <a href="https://openslide.cs.cmu.edu/download/openslide-testdata/Mirax/Mirax2-Fluorescence-2.zip">Mirax2-Fluorescence-2</a> by Yves Sucaet, from the <a href="https://openslide.cs.cmu.edu/download/openslide-testdata/Mirax/">MIRAX test data</a>.</p>
Workflow-Based Spatio-Temporal Data Analytics
<p>In the biodiversity domain, researchers often have to combine a large variety of heterogeneous spatio-temporal data sources. For example, the loss of biodiversity can be quantified by analyzing occurrence observations of various species across time. To find the root cause of that loss, occurrence data may need to be combined with satellite images to find possible correlations with climate variables. To facilitate an exploratory approach for this combination of data sources, it is essential to provide researchers with workflow-based tools such that each step during the formulation of a research hypothesis can be tracked. In this presentation, we will discuss Geo Engine, a workflow-based analysis platform for spatio-temporal data analytics, and its place within FAIR data spaces.</p>
Data Management Workflows with CaosDB
<p>A figure illustrating how the open source research data management system CaosDB can be integrated into data management workflows. It is shown that data is typically integrated into the system using a file crawler. Afterwards data can be accessed using e.g. the web frontend or other client interfaces.</p>
Open‐source workflow approaches to passive acoustic monitoring of bats
<ol> <li>The affordability, storage, and power capacity of compact modern recording hardware has evolved passive acoustic monitoring (PAM) of animals and soundscapes into a non-invasive, cost-effective tool for research and ecological management and is particularly effective for bats and toothed whales that consistently echolocate. The use of PAM at large scales hinges on effective automated detectors and species classifiers which, combined with distance sampling approaches, have enabled species abundance estimation of toothed whales. But standardized, user-friendly, and open-access automated detection and classification workflows are in demand for this key conservation metric to be realized for bats.</li> <li>We used the PAMGuard toolbox including its new deep learning classification module to test the performance of four open-source workflows for automated analyses of acoustic datasets from bats. Each workflow used a different initial detection algorithm followed by the same deep learning classification algorithm and was evaluated against the performance of an expert manual analyst.</li> <li>Workflow performance depended strongly on the signal-to-noise ratio and detection algorithm used: the full deep learning workflow had the best classification accuracy (≤67%) but was computationally too slow for practical large-scale bat PAM. Workflows using PAMGuard's detection module or triggers onboard an SM4BAT or AudioMoth accurately classified up to 47%, 59% and 34%, respectively, of calls to species. Not all workflows included noise sampling critical to estimating changes in detection probability over time, a vital parameter for abundance estimation. The workflow using PAMGuard's detection module was 40 times faster than the full deep learning workflow and missed as few calls (recall for both ~0.6), thus balancing computational speed and performance. </li> <li>We show that complete acoustic detection and classification workflows for bat PAM data can be efficiently automated using open-source software such as PAMGuard and exemplify how detection choices, whether pre- or post-deployment, hardware or software-driven, affect the performance of deep learning classification and <span>the downstream ecological information that can be extracted from acoustic recordings. In particular, understanding, and quantifying detection/classification accuracy and the probability of detection are key to avoid introducing biases that may ultimately affect the quality of data for ecological management. </span> </li> </ol>
Supporting data for "Software pipelines for RNA-Seq, ChIP-Seq and Germline Variant calling analyses in Common Workflow Language (CWL)"
<p>Datasets produced during the validation of CWL-based pipelines, designed for the analysis of data from RNA-Seq, ChIP-Seq and germline variant calling experiments. Specifically, the workflows were tested using publicly available High-throughput (HTS) data from published studies on Chronic Lymphocytic Leukemia (CLL) (accession numbers: E-MTAB-6962, GSE115772) and Genome in a Bottle (GIAB) project samples (accession numbers: SRR6794144, SRR22476789, SRR22476790, SRR22476791).</p> <p>The supporting data include:</p> <ul> <li>Differential transcript and gene expression results produced during the analysis with the CWL-based RNA-Seq pipeline</li> <li>Bigwig and narrowPeak files, differential binding results, table of consensus peaks and read counts of EZH2 and H3K27me3, produced during the analysis with the CWL-based ChIP-Seq pipeline</li> <li>VCF files containing the detected and filtered variants, along with the respective hap.py () results regarding comparisons against the GIAB golden standard truth sets for both CWL-based germline variant calling pipelines</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.