Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
650
datasets available to search
ShareScore release 0.9.0
Dataset results
650 results for “workflow”
Open Research Skills Workshops - GitHub collaborative workflows
<p>This is the fourth workshop on GitHub collaborative workflows in a series of workshop about Open Research Skills.</p><p>This workshop covers:</p><p>- Introduction to version control</p><p>- How to fork a repository</p><p>- Forking exercises</p><p>- How to work in a team and create and merge branches</p><p>- Branching exercises</p><p><strong>List of training workshops in Open Research Skills:</strong></p><ul><li>24th February 2023 - Open access publishing</li><li>24th March 2023 - Using repositories</li><li>21st April 2023 - GitHub basics</li><li><strong>28th April 2023 - GitHub collaborative workflows</strong></li><li>26th May 2023 - Standard vocabularies and ontologies</li><li>30th June 2023 - FAIR data</li></ul><p><strong>Project overview:</strong></p><p>Our project aims to upskill participants in open research skills to increase the quality and reusability of phytolith research and related disciplines such as archaeology, palaeosciences and plant sciences. We will run six hands-on training workshops on open access publishing and research outputs, using repositories, ontologies and standard vocabularies, implementation of FAIR Guidelines for phytolith research, and two workshops on Github basic and advanced skills. The materials from all workshops will be archived as self-study courses on our website (<a href="https://open-phytoliths.netlify.app/">https://open-phytoliths.netlify.app/</a>). We will also provide translation during workshops and training materials into multiple languages. </p><p>Github is a collaborative, project management tool used to run reproducible research projects with version control. In these videos, you will learn how to use version control, how to branch and fork a repository, how to pull a request and how to collaborate as part of a team on GitHub.</p>
Genome annotation workflow for Effrenium voratum RCC1521
<p>Scripts of complete genome annotation workflow for Effrenium voratum RCC1521, associated with the key genome paper (Shah et al., 2024, Massive genome reduction predates the divergence of Symbiodiniaceae dinoflagellates, under review in <em>ISME Journal</em>). An earlier preprint of this manuscript is available at <em>bioRxiv</em>: <a href="https://doi.org/10.1101/2023.03.24.534093" target="_blank" rel="noopener">https://doi.org/10.1101/2023.03.24.534093</a>.</p> <p>See <strong>README_EvRCC1521.txt</strong> for more detail.</p>
A Resilient Workflow to Control a Biomedical HPC Simulation in an Urgent Computing Setting
<p><span><span><span><span>We demonstrate a resilient workflow enabled by the LEXIS Platform, running a time- and safety-critical biomedical simulation of virtual stent placement in intracranial arteries using the HemoFlow application. The workflow, as captured on the video, gracefully handles failures of single computing steps or entire computing systems and thus lends itself to urgent computing applications. <br><br><span><span>The concept of this workflow has potential for realising ab-initio computational biomedical simulations which can provide live, targeted guidance to surgeons.</span></span></span></span></span></span></p>
Additional Artifacts - Supplements to: A Resilient Workflow to Control a Biomedical HPC Simulation in an Urgent Computing Setting
<p>In this dataset, we have collected supplementary artifacts to support an understanding of the workflow presented in the submission cited (see related identifiers).</p> <p>These artifacts are (cf. README.md in the main folder of the tar.gz archive):</p> <p>A1: modified HemoFlow code (cf. https://github.com/gzavo/hemoflow) for our workflow experiments (subfolder "hemoflowcfd");<br>A2: workflow descriptions in python for Apache Airflow (subfolder "workflow");<br>A3: inputs (.xml/.npz) and output (.txt) for the example (subfolder "case").</p> <p> </p>
Datasets of synthetic workflows for evaluating a multi-objective and multi-constrained scheduling approach for cyber-physical applications
<p>These datasets of synthetic workflows (task graphs) were generated to evaluate the performance and scalability of a multi-objective and multi-constrained scheduling approach for workflow applications of various structures, sizes, and sensing/actuating requirements in a cyber-physical system (CPS) based on the edge-hub-cloud paradigm. The examined CPS comprised four edge devices (i.e., single-board computers, each attached to an unmanned aerial vehicle (UAV) equipped with sensors/actuators) interacting with a hub device (e.g., a laptop), which in turn communicated with a more computationally capable cloud server. All system devices featured heterogeneous multicore processors with different processing core failure rates and varied sensing/actuating or other specialized capabilities. Our objectives were the minimization of the overall latency, the minimization of the overall energy consumption, and the maximization of the overall reliability of the workflow application in the specific CPS, under deadline, reliability, memory, storage, energy, capability, and task precedence constraints.</p> <p>We generated 25 random task graphs with 10, 20, 30, 40, and 50 nodes (5 task graphs for each size), utilizing the Task Graphs For Free (TGFF) random task graph generator [1],[2]. Additional task parameters (e.g., execution time, power consumption, memory, storage, output data size, capability, reliability threshold) were included post-generation, using appropriate values. More details are provided in README.txt.<br><br>References:<br>[1] R. P. Dick, D. L. Rhodes, and W. Wolf, "TGFF: Task graphs for free," Proceedings of the Sixth International Workshop on Hardware/Software Codesign (CODES/CASHE), 1998, pp. 97-101, doi: 10.1109/HSC.1998.666245.<br>[2] R. P. Dick, D. L. Rhodes, and K. Vallerio, "TGFF," https://robertdick.org/projects/tgff/.</p>
Supplemental Data from the article "The SmARTR pipeline: a modular workflow for the cinematic rendering of 3D scientific imaging data"
<h1><strong>Please, refer to <a href="https://github.com/MeVisLab/SmARTR-Networks">this GitHub repository</a> for additional info, updates, issue reports, and discussion<br></strong></h1> <p><strong>A collection of configuration files (SmARTR networks) published in "<a href="https://doi.org/10.1016/j.isci.2024.111475">The SmARTR Pipeline: a modular workflow for the cinematic rendering of 3D scientific imaging data</a>", enabling the creation of cinematic (photorealistic) renderings of 3D data in the FREE software <a href="https://www.mevislab.de/download">MeVisLab</a><br></strong></p> <ul> <li>Each folder in the archive contains one or more SmARTR network files, the scan and mask files required for the practical examples detailed in the <a href="https://www.cell.com/cms/10.1016/j.isci.2024.111475/attachment/8d79036b-acb6-4cda-a5ff-f56317691ebc/mmc1.pdf">Supplemental Data</a> of the article, and an additional folder with LUT presets.</li> </ul>
Data for Publication: "Automated Investigation of Metal-Ligand Interactions by a Newly Established Robotic Workflow for Titrations"
<p>This dataset contains the whole primary and raw (original) data for the manuscript "Automated investigation of metal-ligand interactions by a newly established robotic workflow for titrations".</p>
A High-Performance Data Processing Workflow to Incorporate Effect-Directed Analysis in Suspect and Nontarget Screening [Feature Tables]
<p>This repository is supplementary to the manuscript "High-Performance Data Processing Workflow Incorporating Effect-Directed Analysis for Feature Prioritization in Suspect and Nontarget Screening" (DOI: 10.1021/acs.est.1c04168) and includes an overview of all measured chemical features and annotations in a waste water treatment plant (WWTP) effluent, dust standard reference material (SRM) 2585 and fetal calf serum (FCS) sample.</p> <p>Samples were measured using liquid chromatography - high resolution mass spectrometry (LC-HRMS) and fractionated into 80 micro-fractions encompassing a couple of seconds from the chromatographic run. The fractions were tested for their bioactivity in the antibiotics and the TTR-binding assay. The samples were processed separately using one, two, and three technical replicates in positive and negative ion mode. The first excel sheet includes all measured chemical features, suspect screening annotation, and corresponding bioassay responses. The second sheet includes all possible isomer annotations from the CECscreen database (DOI: <a href="https://doi.org/10.5281/zenodo.3956586">10.5281/zenodo.3956586</a>) for the annotated features. </p>
VEP_workflow_simulatedData
<p>This dataset is used for the paper: </p> <p>Wang, H. E., Woodman, M., Triebkorn, P., Lemarechal, J.-D., Jha, J., Dollomaja, B., Vattikonda, A. N., Sip, V., Medina Villalon, S., Hashemi, M., Guye, M., Makhalova, J., Bartolomei, F., & Jirsa, V. (2023). Delineating epileptogenic networks using brain imaging data and personalized modeling in drug-resistant epilepsy. <em>Science Translational Medicine</em>, <em>15</em>(680). <a href="https://doi.org/10.1126/scitranslmed.abp8982">https://doi.org/10.1126/scitranslmed.abp8982</a></p> <p>Code to read and process this dataset: <a href="http://zenodo.org/record/7573382#.Y9PN9C8w3RI">https://zenodo.org/record/7573382#.Y9PN9C8w3RI</a></p> <p>It includes both anatomical and simulated data for one example patient. The anatomical data includes the personalized brain surface meshes, gain matrix from SEEG electrodes and source brain regions, and global structural connectivity matrices. The functional data includes the neural field simulation on both vertices and SEEG sensors. The datasets are used in BIDS format. The codes to read and use this dataset are available in https://github.com/HuifangWang/VEP_INS_workflow. </p> <p> </p>
A workflow for exploring ligand dissociation from a macromolecule: Efficient random acceleration molecular dynamics simulation and interaction fingerprint analysis of ligand trajectories
<p>Containes input data for MD simulations of 3 HSP90- small compound complexes from the paper</p> <p>A workflow for exploring ligand dissociation from a macromolecule: Efficient random acceleration molecular dynamics simulation and interaction fingerprint analysis of ligand trajectories" from Daria B. Kokh, Bernd Doser , Stefan Richter , Fabian Ormersbach , Xingyi Cheng, Rebecca C. Wade, publishe in J. Chem. Phys. <strong>153</strong>, 125102 (2020); <a href="https://doi.org/10.1063/5.0019088">https://doi.org/10.1063/5.0019088</a></p> <ul> <li>ref.pdb - structure of the complex in PDB format</li> <li>ref.prmtop - topology file in AMBER</li> <li>ref-equal-NTP.pdb - structure after NTP equilibration </li> <li>ref-equal-NTP.rst7 - coordinates after NTP equilibration</li> <li>ref-equal-NTP.crd - coordinates after NTP equilibration </li> <li>gromacs.gro - coordinates in Gromacs format (after NTP equalibration)</li> <li>gromacs.top - Gromacs topology </li> </ul> <p> </p>
ESCALATOR - Stakeholder map data workflow
<p>The stakeholder map project aims to collect and share data on Digital Humanities (DH), Computational Social Sciences (CSS) and related activities and initiatives in South Africa. This data includes information about South African researchers, projects, publications, tools, datasets, academic programmes, training events, learning materials, and more. The aim is to provide deeper insight into the breadth of activities in this area, facilitate enhanced networking and collaboration, and support the optimal use of resources. The stakeholder map will, for example, support researchers looking for collaborators, help potential students to identify undergraduate and postgraduate training programmes, and highlight gaps and opportunities to funders and institutions.</p> <p>The initial design of the data pipeline and workflow for data visualisation has been completed. The pipeline is primarily based on open-source software and platforms often used in the open science community. Development is currently under way. Data will be captured via Google Forms and manipulated using R scripts, available on GitHub and archived in Zenodo. Interactive visualisations will be published on the ESCALATOR website. These visualisations include a [Shiny app](https://shiny.rstudio.com/) that will allow the community to explore data through a web interface and a [Kumu network visualisation](https://kumu.io/). Research articles can be added to an [open collection in Zotero](https://www.zotero.org/groups/3866799/dhcssza) to facilitate easy access to publications from the South African community.</p> <p><br> This diagramme shows the high-level workflow. We anticipate the diagramme will be updated as design and development progresses to incorporate lessons learned and feedback from the community.</p>
Molecule dataset used in workflow memoization experiments
<p>This is a collection of Simplified Molecular Input Line Entry System (SMILES) strings that we used to evaluate our workflow memoization system in:</p> <p>> Vassiliadis, V., Johnston, A. M., McDonagh, L. J. "Fast, Transparent, and High-Fidelity Memoization Cache-Keys for Computational Workflows." 2022 IEEE International Conference on Services Computing (SCC). IEEE, 2022.</p>
Reproducible and Attributable Materials Science Workflows
<p>This set includes the deidentified data, reproducible analysis and research report of the project on Reproducible and Attributable Materials Science Workflows.</p>
Spectral Libraries for Metabolome Annotation Workflow (MAW)
<p>MassBank saved at 2022-09-12 10:28:52 with release version 2022.06 as mbankNIST.rda (MsBackendMsp)<br> GNPS saved at 2022-09-12 13:37:42 as gnps.rda (MsBackendMsp)<br> HMDB saved with the release version 4 as hmdb.rda (MsBackendHmdb)</p> <p>All .rda files can be reloaded into R session using the respective Backends. These databases were created for MAW version 1.</p> <p>hmdb_dframe_str.csv is downloaded from HMDB Downloads for structural information on HMDB IDs present in the HMDB version 4 spectral data.</p>
myExperiment Workflows, "abstracted" (all non-analytical nodes removed)
<p>To do the SCOFF analysis (detecting highly similar workflow fragments) we took all bioinformatics-related workflows from myExperiment and removed all non-analytical nodes. These included nodes referred-to as "shims" - those that do data structure/type transformations, but not any "semantic" transformation. This deposit contains all such abstracted workflows.</p>
Datasets of synthetic workflows for cyber-physical edge-hub-cloud systems
<p>These datasets of synthetic workflows were generated to evaluate the performance and scalability of a multi-constrained scheduling approach for workflow applications of various structures, sizes, and sensing/actuating requirements in a cyber-physical system (CPS) following the edge-hub-cloud paradigm. The examined CPS comprised four edge devices (i.e., single-board computers, each attached to an unmanned aerial vehicle (UAV) equipped with sensors/actuators) interacting with a hub device (e.g., a laptop), which in turn communicated with a more computationally capable cloud server. All system devices featured heterogeneous multicore processors and varied sensing/actuating or other specialized capabilities. The problem objective was the minimization of the overall latency of the application under deadline, memory, storage, energy, capability, and task precedence constraints.</p> <p>We generated 25 random workflows (task graphs) with 10, 20, 30, 40, and 50 nodes (5 task graphs for each size), utilizing the Task Graphs For Free (TGFF) random task graph generator [1],[2]. Additional task parameters (e.g., execution time, power consumption, memory, storage, output data size, capability) were included post-generation, using appropriate values. More details are provided in README.txt and in [3].<br><br>References:<br>[1] R. P. Dick, D. L. Rhodes, and W. Wolf, "TGFF: Task graphs for free," in Proc. Sixth International Workshop on Hardware/Software Codesign (CODES/CASHE), 1998, pp. 97-101, doi: 10.1109/HSC.1998.666245.</p> <p>[2] R. P. Dick, D. L. Rhodes, and K. Vallerio, "TGFF," https://robertdick.org/projects/tgff/.</p> <p>[3] A. Kouloumpris, G. L. Stavrinides, M. K. Michael, and T. Theocharides, “Optimal multi-constrained workflow scheduling for cyber-physical systems in the edge-cloud continuum,” in Proc. 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC), Jul. 2024, pp. 483-492, doi: 10.1109/COMPSAC61105.2024.00072.</p>
Binding Affinity Prediction Workflow - Simulation Input Files and Absolute Binding Free Energies
<p>The Binding Affinity Prediction (BAP) workflow calculates absolute binding free energies for protein-ligand complexes by taking their crystal structures, converting them into input files for molecular dynamics (MD) simulations with GROMACS after they have passed extensive quality checks, and analysing the resulting trajectories with the Generalised Born model of implicit solvation as implemented in gmx_MMPBSA to obtain the free-energy estimates. The workflow was designed for soluble proteins without post-translational modifications, co-factors and non-standard amino acids, and it has limited support for coordinated ions.</p> <p>For the dataset published here, the BAP workflow was run on the PDBbind 2020 (http://www.pdbbind.org.cn/index.php) refined set. This entry contains the MD simulation input files (BAPSimulationInputFiles.tar.gz) and the ABFE estimates (BAPBindingFreeEnergyEstimates.csv) obtained from four 250 ns trajectories for each complex. The MD simulations for more than 4000 complexes were run on the Leonardo supercomputer while the implicit-solvent calculations were carried out on Galileo, both operated by Cineca (Italy). The MD trajectories will be stored at Cineca for approx. 1 year after publication of this entry; contact Cineca's user support if you are interested in the trajectories.</p> <p>The README file describes how to reproduce the MD trajectories and the subsequent implicit-solvent calculations yielding the free-energy estimates. The workflow scripts can be downloaded from GitHub (https://github.com/LigateProject/Binding-Affinity-Prediction-workflow). The MD simulations were run with GROMACS 2023.2 (https://manual.gromacs.org/2023.2/index.html), and the implicit-solvent calculations were carried out with gmx_MMPBSA 1.6.1 (https://valdes-tresanco-ms.github.io/gmx_MMPBSA/v1.6.1/).</p>
Digitization Workflow: Talk with Joana Meier
<p>Jane Haller, a sociologist, digital project manager, and president of the Digitales Schaudepot, is in conversation with Joana Meier. Joana holds a BA in Sociology and English Literature, is a Master's student in Digital Humanities, and works in museum education and digitization.</p> <ul> <li>As a Digital Humanities Master’s student and museum education expert, talking about experimenting and hands on exercises with the digital</li> <li>What does "curating data stories" mean from a technical and academic perspective?</li> </ul> <p>As a winning project of the Dariah Theme Call 2022-2024 on Workflows, we evaluated a showcase project called “<a href="https://curiositas.digitalesschaudepot.ch/en/">curiositas5.0</a>”, initiated by the <a href="https://www.digitalesschaudepot.ch/">Digitales Schaudepot Association</a> (DSD) it is intended to assist in planning projects and, above all, avoiding unwanted missteps. Keep in mind that each project is distinct and may necessitate alternative measures. </p> <p>To capture voices from the community and provide insight into the different working methods, expertise, and backgrounds of the people collaborating on the curiositas5.0 project, we conducted 3 interviews.</p> <p> </p> <p> </p>
Digital Repository of Ireland Member Digitisation Workflows for 2D Image Files: Survey Questions and Dataset
<p>The Digital Repository of Ireland (DRI) issued a survey to its membership, <strong>DRI Member Digitisation Workflows for 2D Images</strong>, which ran from December 7, 2023–January 31, 2024. The survey was conducted to improve the DRI’s understanding of the technical processes and metadata workflows that our members use to digitise and share images in the Repository, in order to better tailor our support for this work and deliver the most complete information about digital images files available to our users. </p> <p>The survey informed the actions taken in WorldFAIR Project WP13 deliverable <a href="https://doi.org/10.5281/zenodo.10850009" target="_blank" rel="noopener">13.3 Implementing and Testing the Cultural Heritage Image Sharing Recommendations: DRI Case Study Report</a>. The data will inform ongoing work at DRI aimed at improving the transparency of technical information associated with digital assets accessed through the Repository.</p> <p>Read more about the Cultural Heritage Image Sharing Case Study DRI on our website: <a href="https://dri.ie/the-worldfair-project/">https://dri.ie/the-worldfair-project/</a>. </p> <p>Summary: DRI is Ireland's national repository for the arts, humanities, and social sciences data, and operates on a membership scheme. There were 20 respondents to the survey, giving us a response rate of about 35% of DRI's membership. Representation from professional fields of work across the cultural heritage sector was captured in the results (note that some institutions gave multiple responses): 17 Archives, 12 Libraries, 5 Museums and 11 Higher Education Institutions. </p>
A Bioconductor workflow for processing, evaluating and interpreting expression proteomics data
<p>Files for users of the workflow "A Bioconductor workflow for processing, evaluating and interpreting expression proteomics data". Files include Proteome Discoverer (v2.5) processing and consensus workflows for both TMT and LFQ expression proteomics data. Also provided are the output .txt files of a corresponding Proteome Discoverer identification search, as required for users to follow the workflow themselves. For raw data please refer to PRIDE. Appendix is provided as a PDF.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.