Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

650

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

650 results for “workflows”

Learn how ShareScore rates datasets ↗
zenodo40/100

Data for FEgrow: An Open-Source Molecular Builder and Free Energy Preparation Workflow

<p>Data illustrating the use of de novo design in building and scoring protein-ligand complexes.</p> <p>This is relationship to the FEgrow publication with the intiial preprint here:&nbsp;<br> https://chemrxiv.org/engage/chemrxiv/article-details/6287bb98a42e9c78d34769f6<br> &nbsp;</p> <p>The FEgrow software snapshot used can be found here:&nbsp;https://zenodo.org/record/7105647#.YzFwINLMIUE</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Data sets for the Simulated AMPI (SAMPI) load balancing simulation workflow and Ondes3D performance analysis (Companion to CCPE - Euro-Par 2017 special issue)

<p>This package contains data sets and scripts (in&nbsp;an Org-mode file) related to our submission to the special Euro-Par 2017 issue of the&nbsp;&nbsp;journal &quot;Concurrency and Computation: Practice and Experience&quot;, under the title&nbsp;&quot;Performance Modeling of a Geophysics Application to Accelerate Over-decomposition Parameter Tuning through Simulation&quot;.</p>

opencc-by-sa-4.0Nov 2017View details →
zenodo40/100

Figure 2. Workflow in Collaborative Portals-Questions Regarding Alterity in Social Collaborative Networks

<p>Despite the divergence of each members individual interests and structure, the media as well<br> as the Internet become new forms of social binding, and we believe they do not exclude the<br> traditional ways of communication and interaction, but help people in their duties, in addition to the<br> traditional, formal ways. In Figure 2 [7]. there is a clear image about how each of the benefficiaries<br> interact with the educational portal.</p>

opencc-by-4.0Jan 2010View details →
zenodo40/100

OPMW Formatted Workflow Fragments from myExperiment Workflows

<p>The bioinformatics-associated workflows from myExperiment.org, fragmented randomly into 2 or 3-node (overlapping) fragments.&nbsp; Each fragment is then represented as an OPMW ontology instance (opmw:WorkflowTemplateProcess), which is annotated with as many ontological terms as could be discovered based on the kinds of services in each node of that fragment.</p>

opencc-by-4.0Jan 2018View details →
zenodo40/100

Output from the SCOFF clustering of myExperiment Workflow Fragments

<p>Using a semantic-similarity approach, we clustered the annotated, abstracted workflow fragments (see https://doi.org/10.5281/zenodo.1147545&nbsp; and https://doi.org/10.5281/zenodo.1147485).&nbsp; This zip file contains the results of that clustering.</p>

opencc-by-4.0Jan 2018View details →
zenodo40/100

Development of a desorption electrospray ionization –multiple-reaction-monitoring mass spectrometry (DESI-MRM) workflow for spatially mapping oxylipins in pulmonary tissue

<p>Data from desorption electrospray ionization mass spectrometry &ndash; multiple-reaction-monitoring mass spectrometry (DESI-MRM) analysis of oxylipins in guinea pig lung tissue following<em> in vivo</em> exposure to house dust mite extract.</p> <p>Data are provided as Waters *.raw data folders, each incuding an 'Analyte .txt' file, which is generated from processing within MassLynx (Waters). The 'ion_library.txt' file includes details about the MRM transitions and is required for processing the data with quantMSImageR (<span><a href="https://github.com/targeted-lipidomics/quantMSImageR"><span>https://github.com/targeted-lipidomics/quantMSImageR</span></a></span><span>).</span></p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Pose Selector Workflow - Structure Input Files for Machine Learning (Set 2)

<p>Second set of structure input files for the docking poses of the remaining 2022 protein-ligand complexes. Together with the structure files and absolute binding free energy (ABFE) estimates shared in 10.5281/zenodo.11397017, this data can be used to train a machine-learning (ML) model predicting the ABFE of binding poses of protein-ligand complexes.</p> <p>More background on the workflow generating the structure files and ABFE estimates is provided in 10.5281/zenodo.11397017.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Pose Selector Workflow - Docking Poses, Absolute Binding Free Energy Estimates and Structure Input Files for Machine Learning

<p>The Pose Selector (PS) workflow calculates absolute binding free energies (ABFEs) for binding poses of protein-ligand complexes. First, it converts the binding poses (both docking poses as well as experimentally observed ligand binding poses), which are provided as a combination of protein PDB file and ligand MOL2 file, into input files for molecular dynamics (MD) simulations with GROMACS after they have passed extensive quality checks and repair steps. Next, the PS workflow post-processes and analyses the last frame of the resulting eight 100 ps trajectories per binding pose with the Generalised Born model of implicit solvation as implemented in gmx_MMPBSA to obtain the ABFE estimates. The workflow was designed for soluble proteins without post-translational modifications, co-factors and non-standard amino acids, and it has limited support for coordinated ions.</p> <p>For the dataset published here, the PS workflow was run on docking poses generated for the PDBbind 2020 dataset (http://www.pdbbind.org.cn/index.php), shared in dockingPosesPDBBind2020.tar.gz. This entry and its partner entry 10.5281/zenodo.11397486 also share the intial coordinates used in the MD simulations of &gt;800,000 docking poses of 4022 protein-ligand complexes (structureFiles_dockingPoses1.tar.gz in this entry and structureFiles_dockingPoses2.tar.gz in 10.5281/zenodo.11397486) and of the experimental ligand binding pose of 4549 complexes (structureFiles_experimentalStructures.tar.gz) as well as the corresponding ABFE estimates (absoluteBindingFreeEnergyEstimates.tar.gz). The MD simulations were run on the LUMI and MeluXina supercomputers while the implicit-solvent calculations were carried out on Galileo (Cineca).</p> <p>The README file describes the structure of the shared data in more detail and points out how to reproduce the MD trajectories and the subsequent implicit-solvent calculations yielding the free-energy estimates as well as how to use the data provided in this entry to train a machine-learning model predicting the ABFE of binding poses of protein-ligand complexes. The workflow scripts can be downloaded from GitHub (https://github.com/LigateProject/Pose-Selector-workflow). The MD simulations were run with GROMACS 2023.2 (https://manual.gromacs.org/2023.2/index.html), and the implicit-solvent calculations were carried out with gmx_MMPBSA 1.6.1 (https://valdes-tresanco-ms.github.io/gmx_MMPBSA/v1.6.1/).</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Dataset for "Quality and Contamination control" workflow

<p>This dataset is associated with the workflow "Quality and Contamination control for paired end data".</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Workflow Run Crate specification

<p><strong>Web version:</strong>&nbsp;<a href="https://w3id.org/ro/wfrun/workflow/0.5">https://w3id.org/ro/wfrun/workflow/0.5</a></p> <p>This specification is part of a collection of&nbsp;<a href="https://w3id.org/ro/crate/">RO-Crate</a>&nbsp;profiles&nbsp;for capturing the provenance of an execution of a computational workflow. The Workflow Run Crate profile can be used to describe the execution of a&nbsp; <em>workflow</em>, i.e., a computational tool that has orchestrated the execution of other tools.</p>

openapache2.0Feb 2023View details →
zenodo40/100

Replication Package for "Catching Smells in the Act: A GitHub Actions Workflow Investigation" (SCAM 2024)

<p>Welcome to our artifact! In here we provide additional information on how to retrace our steps performed during the research. We have split up our content into four sections based on the RQ's we have answered. Below you can find a quick summary of the contents of each folder, each folder also contains additional information regarding any data and scripts present.</p> <ul> <li>RQ1 + 2: Contains excel files with the commits we have analyzed and the scripts we have used to automate this process.</li> <li>RQ3: Contains our smell detector and evaluation of the detector</li> <li>RQ4: Contains the data on our contribution study</li> </ul>

opencc-by-4.0Jun 2024View details →
zenodo40/100

"Wissen schaffen (lassen!?)". Workflows mit Generativer KI in den Digital Humanities

<p>Die rasante Entwicklung von generativen KI-Technologien stellt eine bedeutende Ver&auml;nderung f&uuml;r die Forschungspraxis nicht nur in den Digital Humanities dar. Dieser Vortrag untersucht den Einsatz von GPT-4-Tier LLM (Gemini Advanced und Claude 3) sowie deren M&ouml;glichkeiten und Grenzen in verschiedenen Forschungsprojekten der Digital Humanities. Der Fokus liegt dabei auf Workflows wie der Datenerfassung, Transkription, &Uuml;bersetzung, Datenmodellierung, Datengenerierung oder -analyse sowie der Visualisierung geisteswissenschaftlicher Daten. Anhand ausgew&auml;hlter Fallstudien wird die Integration von generativer KI in diese Prozesse dargestellt, wobei sowohl die Automatisierung von Standardaufgaben als auch die Unterst&uuml;tzung komplexerer, analytischer und anspruchsvoller T&auml;tigkeiten wie Datenmodellierung thematisiert werden. Haben generative KI-Modelle, wenn sie im Einklang mit menschlicher Expertise und komplement&auml;ren Systemen eingesetzt werden, das Potenzial, die Effizienz und Tiefe (digitaler) geisteswissenschaftlicher Forschung zu steigern? Die Studie betont auch die Notwendigkeit, die Grenzen und Herausforderungen, wie die Abh&auml;ngigkeit von gro&szlig;en Technologieunternehmen beim Einsatz von generativer KI in den Digital Humanities kritisch zu hinterfragen.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Dataset of "Key Aspects in Designing High-Throughput Workflows in Electrocatalysis Research: A Case Study on IrCo Mixed-Metal Oxidese"

<p>With the growing interest of the electrochemical community in high-throughput (HT) experimentation as a powerful tool in accelerating materials discovery, the implementation of HT methodologies and the design of HT workflows has gained traction. We identify 6 aspects essential to HT workflow design in electrochemistry and beyond to ease the incorporation of HT methods in the community&rsquo;s research and to assist in their improvement. We study IrCo mixed-metal oxides (MMOs) for the oxygen evolution reaction (OER) in acidic media using the mentioned aspects to provide a practical example of possible workflow design pitfalls and strategies to counteract them.&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Figure 1. General Workflow of Algorithm-.Data Conflict Resolution among Same Entities in Web of Data

<p>The Page Rank algorithm which is widely used in most search engines such as Google could<br> be easily used to rank linked data. By starting from a point and random surfing, this algorithm<br> evaluates the probability of finding any given page. The algorithm assumes a link between a page i<br> to a page j demonstrates the importance of page j. In addition, the importance of page j is associated<br> to the importance of page i itself and inversely proportional to the number of pages i point to. To<br> adapt this algorithm to web of data, any page considered as a dataset and links between pages<br> considered as links between datasets.</p>

opencc-by-4.0Jun 2012View details →
zenodo40/100

Figure 1. Workflow of approach-Hybridization of Fuzzy Clustering and Hierarchical Method for Link Discovery

<p>In this section, we present our model in more detail. Figure 1 gives an overview of the<br> workflow. Our approach is organized in two phases: first, the division of data in two clusters; thenthe determination of the worst cluster and splitting. The number of clusters is unknown, but our<br> algorithms can find this parameter based on the complexity of cluster structure.</p>

opencc-by-4.0Jun 2012View details →
zenodo40/100

Figure 1. The AGLO Workflow-Generative Learning Objects Instantiated with Random Numbers Based Expressions

<p>E-learning is a key area of research with a great influence on the developments of several<br> industries. For example, the nowadays ITC industry is in a continuous growth because of its<br> applications in almost all industrial domains. Companies tend to lack qualified human resources and<br> because of that they reject high economical value projects. In response to this lack of human<br> resource problem, universities started to develop several alternative study programs, many of them<br> are based on e-learning technology and namely on electronic learning materials. Learning objects<br> (LO) are considered to be digital resources that support learning and can be delivered across<br> networks in large or small sizes (Wiley, 2000). In order to increase the reusability and<br> interoperability of LOs, standards were developed by several organizations (IEEE Learning<br> Technology Standards Committee) (e.g. LOM http://ltsc.ieee.org/doc/wg12/LOMv4.1.htm.)</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

NEUBIAS TS7 - data used in the workflow deconstruction session on quantifying monolayer cell migration

<p>Training&nbsp;session details (including slides):&nbsp;https://github.com/miura/NEUBIAS_AnalystSchool2018/tree/master/Assaf</p> <p>Matlab source code:&nbsp;https://github.com/assafzar/MonolayerKymographs</p>

opencc-by-sa-4.0Feb 2018View details →
zenodo40/100

Data sets for the Simulated AMPI (SAMPI) load balancing simulation workflow and Ondes3D performance analysis (Companion to CCPE paper)

<p>This package contains data sets and scripts (in&nbsp;an Org-mode file) related to our submission to the&nbsp; journal &quot;Concurrency and Computation: Practice and Experience&quot;, under the title&nbsp;<em>&quot;Performance Modeling of a Geophysics Application to Accelerate the Tuning of Over-decomposition Parameters through Simulation&quot;</em>.</p>

opencc-by-sa-4.0Jun 2018View details →
zenodo40/100

K-mer databases of plant virus sequences for use with the Kodoja workflow

<p><strong>Details</strong></p> <p>This is a gzipped tar file that includes the plant virus database files required to run the Kodoja workflow (https://github.com/abaizan/kodoja)[1]. Kodoja is a workflow for the detection of plant virus sequences in RNA-seq data files that uses two previoulsy published tools Kraken[2] and Kaiju[3].</p> <p>This file contains databases for Kraken [2] and Kaiju [3]. The file includes the kraken database files: database.idx, database.kdb, nodes.dmp, names.dmp and the kaiju database file kaij_library.fmi.</p> <p>These k-mer databases are based on virus sequences in RefSeq [4] (ttps://www.ncbi.nlm.nih.gov/refseq/) with plant hosts as defined in the Virus-Host Database [5] (https://www.genome.jp/virushostdb/).</p> <p><strong>Version</strong><strong> 1.0</strong></p> <p>kodojaDB_v1.0 is based on RefSeq v89 and the Virus-Host Database (accessed 03/09/2018 which is based on RefSeq 89 and Genbank 226.0). The viral partition of RefSeq v89 genome comprises 7946 viruses (ftp://ftp.ncbi.nlm.nih.gov/genomes/refseq/viral/assembly_summary.txt).</p> <p>kodojaDB_v1.0 was created using kodoja_retrieve.py which is part of the kodoja workflow (v0.05) (https://github.com/abaizan/kodoja).</p> <p><strong>References</strong></p> <p>[1] Baizan-Edge, A, Cock, P, MacFarlane, S, McGavin, W, Torrance, T, Jones, S. Kodoja: A workflow for virus detection in plants using k-mer analysis of RNA-sequencing data (under review Nucleic Acids Research).&nbsp;</p> <p>[2] Wood,D.E. and Salzberg,S.L. (2014) Kraken: ultrafast metagenomic sequence classification using exact alignments. <em>Genome Biol.</em>, <strong>15</strong>, R46</p> <p>[3] Menzel,P., Ng,K.L. and Krogh,A. (2016) Fast and sensitive taxonomic classification for metagenomics with Kaiju. <em>Nat. Commun.</em>, <strong>7</strong>, 1&ndash;9.</p> <p>[4] O&rsquo;Leary,N.A., Wright,M.W., Brister,J.R., Ciufo,S., Haddad,D., McVeigh,R., Rajput,B., Robbertse,B., Smith-White,B., Ako-Adjei,D., <em>et al.</em> (2016) Reference sequence (RefSeq) database at NCBI: Current status, taxonomic expansion, and functional annotation. <em>Nucleic Acids Res.</em>, <strong>44</strong>, D733&ndash;D745.</p> <p>[5] Mihara,T., Nishimura,Y., Shimizu,Y., Nishiyama,H., Yoshikawa,G., Uehara,H., Hingamp,P., Goto,S. and Ogata,H. (2016) Linking virus genomes with host taxonomy. <em>Viruses</em>, <strong>8</strong>, 10&ndash;15</p>

opencc-by-4.0Sep 2018View details →
zenodo40/100

W2Share Case Study: Workflow Research Object (WRO)

<p>Case Study - Molecular Dynamics</p> <p>Our case study is based on a molecular dynamics simulation defined in the following article:</p> <p>Silveira, R.L. and Skaf, M. S. Molecular Dynamics Simulations of Family 7 Cellobiohydrolase Mutants Aimed at Reducing Product Inhibition. J. Phys. Chem. B 119, 9295-9303 (2015). DOI: <a href="https://doi.org/10.1021/jp509911m">https://doi.org/10.1021/jp509911m</a></p>

opencc-by-4.0Oct 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record