Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
650
datasets available to search
ShareScore release 0.9.0
Dataset results
650 results for “Workflow”
Example workflow using KLIFS nodes in KNIME - identifying structures with similar molecules
<p>This is KNIME workflow created during the recording of the following <a href="https://www.youtube.com/channel/UCzSo1G_wyTv1vp42AhFDT8A">YouTube demonstration video</a>.</p> <p>Using this workflow, the user can draw a molecule and compare this molecule to all ligands from the KLIFS database (<a href="https://klifs.net">https://klifs.net</a>) to identify structures with molecules that are similar to the drawn molecule.</p> <p>In this workflow the follow steps are performed:</p> <ul> <li>Draw a molecule</li> <li>Collect all kinase ligands with a known structures from KLIFS</li> <li>Compare all KLIFS ligands to the drawn molecule using the ECFP-4 fingerprint and calculate a Tanimoto score </li> <li>Select the highest scoring three ligands and search for their PDB structures</li> <li>Collect the MOL2 files of the ligands as observed while binding in the PDB structures (note: all the PDB structures were first aligned by KLIFS)</li> </ul>
Test datasets for iwc workflows : Assembly decontamination VGP9
Open the record for dataset details and reuse information.
Surfalex HF formability study - Workflow 4 - Estimate hardening curves
<p>This MatFlow workflow is the fourth in a set of eight workflows developed during our formability study of the Surfalex HF (AA6016A) material. In this workflow, we estimated the plastic stress-strain curves of the material during different loading conditions, using crystal plasticity simulations (via DAMASK). In turn, this table was used a "plastic table" Abaqus input for the fifth and final workflow.</p> <p>This workflow can be downloaded and explored in a Jupyter notebook, as explained in the <a href="https://github.com/LightForm-group/surfalex_data_explorer">GitHub repository here</a>.</p>
Source data for "Benchmarking commonly used software suites and analysis workflows for DIA proteomics and phosphoproteomics"
<p>Source data of "Benchmarking commonly used software suites and analysis workflows for DIA proteomics and phosphoproteomics".</p> <p>For reproduction of main and supplementary figures.</p> <p>MS raw files, spectral libraries, and MS data search results are stored in iProX with identifier IPX0004576001.</p> <p> </p>
CCTBX.XFEL Workflow Manager Demo
<p>A demo showcasing the CCTBX.XFEL Computational Crystallography workflow tool. It is being narrated by Aaron Brewster as he demonstrates the tool running on a High-Performance Computing (HPC) system. The demo shows:</p> <ol> <li>Starting the CCTBX.XFEL tool on a pre-configured HPC system.</li> <li>Configuring CCTBX.XFEL to look for new data</li> <li>Configuring data analysis parameters</li> <li>Automatically submitting jobs to the batch queue</li> <li>Watching the data analysis in real time</li> </ol>
An ArcGIS Pro workflow to extract vegetation indices from aerial imagery of small‐plot turfgrass research
<p>Collection of multispectral imagery from an aerial sensor is a means to obtain plot-level vegetation index (VI) values; however, post-capture image processing and analysis remain a challenge for small-plot researchers. An ArcGIS Pro workflow of two task items was developed with established routines and commands to extract plot-level VI values (Normalized Difference VI, Ratio VI, and Chlorophyll Index-Red Edge) from multispectral aerial imagery of small-plot turfgrass experiments. Users can access and download task item(s) from the ArcGIS Online platform for use in ArcGIS Pro. The workflow standardizes the processing of aerial imagery to ensure repeatability between sampling dates and across site locations. A guided workflow saves time with assigned commands, ultimately allowing users to obtain a table with plot descriptions and index values within a .csv file for statistical analysis. The workflow was used to analyze aerial imagery from a small-plot turfgrass research study evaluating herbicide effects on St. Augustinegrass [<em>Stenotaphrum secundatum</em> (Walt.) Kuntze] grow-in. To compare methods, index values were extracted from the same aerial imagery by TurfScout, LLC and were obtained by handheld sensor. Index values from the three methods were correlated with visual percentage cover to determine the sensitivity (i.e., the ability to detect differences) of the different methodologies.</p>
Common Workflow Scheduler Evaluation with Nextflow and Kubernetes
<ul> <li>Setup scripts to test Nextflow with CWS on Kubernetes</li> <li>Traces and logs of 990 workflow executions</li> </ul>
Resources for "Function-as-a-Service Performance Evaluation with Application-level Workflows"
<p><strong>Thesis</strong></p> <p>"Function-as-a-Service Performance Evaluation with Application-level Workflows"</p> <p><strong>Dataset</strong></p> <p>All extracted data from the benchmark runs in each CSP are available as machine-readable CSV.</p>
Dataset for surgical workflow and context recognition
<p>Dataset based on 6 procedures of radical prostatectomy with lymphadenectomy using da Vinci Xi surgical system (Intuitive Surgical Inc., USA) at European Institute of Oncology (IEO, Milan, Italy). <br> <br> It is divided into 6 folders concerning the step of the procedure: Clips, Dissection, Irrigation, Suction, Suturing, Traction</p> <p>Each step is then divided into the four phases of the procedure: Collapse of the peritoneum, Prostate removal, Lymphadenectomy, Anastomosis</p> <p>Each frame is labeled as follows: PatientNumber_camera_step_NumberOfTheFrame</p> <p>for example, for the first frame of patient 1, left lens of the Da Vinci endoscope, in the dissection step we have:<br> pa1_L_dissection_000000</p> <p>In the folders concerning the step there are also collected the text files created on anvil (ground truth labels)</p> <p>We appreciate Alice Pierini, Eleonora Pollini, and Raffaella Salama's efforts in making this dataset.</p>
Raw data of workflow execution results used in tonkaz's experiments
<p>This constitutes the raw data of workflow execution results employed in <a href="http://github.com/sapporo-wes/tonkaz">Tonkaz</a>'s experiments.</p> <p>Further information regarding the generation methods and additional details can be found at <a href="https://github.com/sapporo-wes/tonkaz/tree/main/tests">https://github.com/sapporo-wes/tonkaz/tree/main/tests</a>."</p> <p>The contents of this deposit are basically licensed under <a href="https://spdx.org/licenses/CC0-1.0.html">the Creative Commons Zero v1.0 Universal</a>.<br> However, there are files that could be licensed under other licenses, such as the nf-core workflow and its dependencies.<br> Because Zenodo does not provide the capability to attach licenses to individual files, we have described the licenses for these workflows in license.txt and ro-crate-metadata.json.<br> Please check them.</p>
A computational workflow for binding free energies in Python
<p>Dataset of distances between a host and six different ligands. The host was beta-cyclodextrin (bCD), while the ligands were phenol, benzene, aspirin, toluene, chlorobenzene and 1,3-dichlorobenzene. No bonds were frozen. </p> <p>The ligand were set to move with a step of 0.25 angstrom from -26 to 26 relative to the bCD (a total of 208 distances). At each distance, a energy biasing potential <span class="math-tex">\(E_{bias}\)</span> was applied the keep two molecules in place. </p> <p><span class="math-tex">\(E_{bias} = \frac{1}{2}\cdot K \cdot (R - R_0)^2\)</span></p> <p>The parameters of the ligands were taken from OpenFF while GLYCAM were used for the host bCD. All of it were applied in Python and the OpenMM framework. Starting parameters, pdb-, and sdf-files can be found in the start folder.</p>
Staged Snakemake Workflow: stage
This is a a snapshot of the outputs of a Snakemake workflow
On the Outdatedness of Workflows in the GitHub Actions Ecosystem
<p>This replication package contains all the material required to replicate the analyses we made for the paper entitled "On the Outdatedness of Workflows in the GitHub Actions Ecosystem" for Journal of Systems and Software.</p> <p>Refer to the README file(s) for more details on the content of this replication package.</p>
Supplementary Figure Legend S1: An overview of DNA samples included in the NGS based chimerism validation process and AlloSeq HCT library preparation workflow.
<p><strong>Supplementary Figure </strong><strong>Legend S1: </strong>An overview of DNA samples included in the NGS based chimerism validation process and AlloSeq HCT library preparation workflow. <strong>S1a: </strong>An overview of clinical and artificial DNA included in the NGS based chimerism validation process (Clinical samples, Artificial construct samples and ASHI EMO Proficiency samples). <strong>S1b:</strong> Schematic presentation of time required for the NGS Chimerism assay. <strong>S1c.</strong> AlloSeq HCT library preparation workflow.</p>
Human Disease Ontology 2018 update: classification, content and workflow expansion
<p><strong>ABSTRACT:</strong></p> <p>The Human Disease Ontology (DO) (http://www.disease-ontology.org), database has undergone significant expansion in the past three years. The DO disease classification includes specific formal semantic rules to express meaningful disease models and has expanded from a single asserted classification to include multiple-inferred mechanistic disease classifications, thus providing novel perspectives on related diseases. Expansion of disease terms, alternative anatomy, cell type and genetic disease classifications and workflow automation highlight the updates for the DO since 2015. The enhanced breadth and depth of the DO's knowledgebase has expanded the DO's utility for exploring the multi-etiology of human disease, thus improving the capture and communication of health-related data across biomedical databases, bioinformatics tools, genomic and cancer resources and demonstrated by a 6.6× growth in DO's user community since 2015. The DO's continual integration of human disease knowledge, evidenced by the more than 200 SVN/GitHub releases/revisions, since previously reported in our DO 2015 NAR paper, includes the addition of 2650 new disease terms, a 30% increase of textual definitions, and an expanding suite of disease classification hierarchies constructed through defined logical axioms.</p> <p><strong>Instructions:</strong></p> <p>Data was cleaned. Duplicates and unnecessary columns were removed. Title of columns were changed.</p> <p><strong>Inspiration:</strong></p> <p>This dataset uploaded to U-BRITE for "DRG_DEPOT" summer 2023 team project.</p> <p><strong>Acknowledgements:</strong></p> <p>Schriml, L. M., Mitraka, E., Munro, J., Tauber, B., Schor, M., Nickle, L., Felix, V., Jeng, L., Bearer, C., Lichenstein, R., Bisordi, K., Campion, N., Hyman, B., Kurland, D., Oates, C. P., Kibbey, S., Sreekumar, P., Le, C., Giglio, M., & Greene, C. </p> <p>Human Disease Ontology 2018 update: classification, content and workflow expansion</p> <p><a href="https://doi.org/10.1093/nar/gky1032">Nucleic Acids Research</a> 2019; <em>47</em>(D1), D955–D962;PMID:<a href="https://pubmed.ncbi.nlm.nih.gov/30407550/">30407550</a>;DOI:<a href="http:// https://doi.org/10.1093/nar/gky1032">https://doi.org/10.1093/nar/gky1032</a></p> <p><strong>U-BRITE last update data: </strong>06/28/2023</p>
Analysis datasets for NEMO_validation workflow Byrne et al 2023 GMD. "Using the COAsT Python package to develop a standardised validation workflow for ocean physics models"
<p>Analysis datasets in support of Byrne et al. (2023) "Using the COAsT Python package to develop a standardised validation workflow for ocean physics models", <em>Geoscientific Model Development</em>.</p> <p> </p> <p>The datasets are from a comparative analysis of two versions of the European shelf sea AMM15 (Atlantic Margin Model at 1.5km horizontal resolution) configuration. These are NEMO ocean model configurations with different code base versions. The configurations are CO7, which is based on NEMOv3.6, and CO9p0 (also referred to as P0.0), which is based on NEMOv4.0.4.</p>
Fig. 5. A in Assessing specialized metabolite diversity of Alnus species by a digitized LC-MS/MS data analysis workflow
Fig. 5. A boxplot showing the ion intensities of MS/MS feature 10 (gallic acid) in extracts which were active (IC50 <30 μg/mL) and inactive against α-glucosidase.
Fig. 4 in Assessing specialized metabolite diversity of Alnus species by a digitized LC-MS/MS data analysis workflow
Fig. 4. Discrimination of the analyzed Alnus extracts into chemogroups. The analyzed extracts can be discriminated into three chemogroups by visualizing the CSCS distance metric between samples as PCoA plot (A) and chemical dendrogram (B). On the other hand, conventional methods such as PCA score plot (C) or hierarchical clustering analysis (HCA) using the Euclidean distance (D; chemogroups 1–3 are visualized with same colors used in B to make it easy to be compared) could not discriminate the samples into the same chemotypes. By mapping the chemogrouping of samples on the molecular network, it could be visualized that the three chemogroups were rich in diarylheptanoid, flavonoid, and tannins, respectively (E). (For interpretation of the references to color in this figure legend, the reader is referred to the Web version of this article.)
Fig. 3 in Assessing specialized metabolite diversity of Alnus species by a digitized LC-MS/MS data analysis workflow
Fig. 3. MS2LDA-driven substructural annotation of diarylheptanoids of Alnus species. Integrated with GNPS library matching and NAP in silico annotation, diarylheptanoid-related Mass2Motifs 41, 49, 72, and 81 could be characterized and correlated with specific substructures of diarylheptanoid aglycones. Scaffold diversity within diarylheptanoid molecular families A, D, and I were revealed by mapping these Mass2Motifs on the molecular network with different colors. (For interpretation of the references to color in this figure legend, the reader is referred to the Web version of this article.)
Fig. 1 in Assessing specialized metabolite diversity of Alnus species by a digitized LC-MS/MS data analysis workflow
Fig. 1. LC–MS base peak ion (BPI) chromatograms of 15 Alnus extracts. Gaps between chromatogram were added to visualize their difference, so y-axis values do not equal to the absolute intensities.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.