Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
650
datasets available to search
ShareScore release 0.9.0
Dataset results
650 results for “Workflow”
Market Research - What's become... of new entrants in research workflows and scholarly communication? ('long list')
<p>Market Research - What’s become... of new entrants in research workflows and scholarly communication?</p> <p>Full report to be published as preprint on OSF Preprints shortly.</p> <p>To systematically create a list of the new entrants in research workflows and scholarly communication that I’ve seen over the years, I used a range of resources, listed hereunder:</p> <ul> <li>400+ Tools and innovations in scholarly communication (charting the <a href="http://bit.ly/innoscholcomm-list">creation and availability</a> (supply side))<a href="#_edn1">[i]</a>. This database continues to build on the 2015-2016 survey by Bianca Kramer & Jeroen Bosman (both at Utrecht University Library), who are interested in the way information is created, shared, and processed in academia<a href="#_edn2">[ii]</a>. From this source I pulled 683 players. From the answers to the open and closed questions in the 2015-2016 survey I found 225 players.</li> <li>Presenters at STM events<a href="#_edn3">[iii]</a> from 2010-2017. This contributed 53 players.</li> <li>Outsell<a href="#_edn4">[iv]</a> ‘Companies to watch’, as per their annual Information Industry Outlook (2009-2017). This source provided 29 players.</li> <li>Pitches at APE (Academic Publishing in Europe) Conferences<a href="#_edn5">[v]</a> between 2011-2017. Here I found 44 players.</li> <li>Nominees for the ALPSP Award for Innovation in Publishing. This source contributed 10 players from 2014-2017.<a href="#_edn6">[vi]</a>.</li> </ul> <p>This file contains the 'long list' per 31 December 2018.</p> <p><a href="#_ednref1">[i]</a> https://docs.google.com/spreadsheets/d/1KUMSeq_Pzp4KveZ7pb5rddcssk1XBTiLHniD0d3nDqo/edit#gid=0</p> <p><a href="#_ednref2">[ii]</a> https://101innovations.wordpress.com/</p> <p><a href="#_ednref3">[iii]</a> https://www.stm-assoc.org/events/?previous</p> <p><a href="#_ednref4">[iv]</a> https://www.outsellinc.com/</p> <p><a href="#_ednref5">[v]</a> https://www.ape2019.eu/ape-literature</p> <p><a href="#_ednref6">[vi]</a> https://www.alpsp.org/Awards</p>
Market Research - What's become... of new entrants in research workflows and scholarly communication? ('sample V2')
<p>Market Research - What’s become... of new entrants in research workflows and scholarly communication?</p> <p>Full report to be published as preprint on OSF Preprints shortly.</p> <p>To systematically create a list of the new entrants, a range of approaches and sources were used. This resulted in a long list (see: <a href="https://zenodo.org/record/2530048#.XDnfnFxKjIV">https://zenodo.org/record/2530048#.XDnfnFxKjIV</a>).</p> <p>The ‘long list’ was shortened to a sample of 120 independent for-profit startups through various filtering exercises, described in the full report. For the sample, three questions were investigated: 1. Did they still exist (independently) in 2018? 2. If so, how were they funded and how were they doing? 3. If they were acquired by 2018, by whom and when were they taken over?</p> <p>This file contains the ‘sample’ and answers to the research questions per 31 December 2018.</p>
Method to Improve Workflow Net Decomposition for Process Model Repair - experiments
<p>This repository contains data that was used to carry out experiments as well as the results of these experiments.</p> <p>File names are presented in the following format: LM2-[repair method]-[b/f]-[number of experiment], where:<br> - repair method may take values "Greedy" or "Smart". "Greedy" means that the repair method of the model was greedy algorithm working with maximal decomposition. "Smart" means that the repair method was greedy algorithm as well, with the difference of decomposition method being the developed one.<br> - "b" (broken) means that the model has not undergone repair. "f" (fixed) means that the file presents a repaired model (the one which fits the initial log perfectly).<br> - number of experiment ranging from 1 to 10.</p> <p>This repository also contains the following files:<br> - LM2-CL.xes - the initial LM2 model log;<br> - LM2-CM.pnml - the initial LM2 model;<br> - LM2-Greedy-data.txt - auto-generated measurements of greedy approach performance;<br> - LM2-Smart-data.txt - auto-generated measurements of smart (developed) approach perfomance.</p>
LIBER 2019 Workshop. Open Access books in academic libraries – how can we adapt workflows and cost management to an open scholarly communications landscape?
<p>This dataset includes all the results from a workshop held at the LIBER Annual Conference 2019 on June 26, 2019, in Dublin, Ireland. The workshop aimed at collecting and discussing current library practices related to open access books.</p> <p>The dataset includes information from a survey made in preparation for the conference, where 67 European libraries responded to a questionnaire based on activities or workflows in libraries related to open access books. Both the survey questionnaire and the results from the survey are uploaded as separate files.</p> <p>The dataset also includes the presentation made by keynote speaker Eelco Ferwerda from the OAPEN Foundation. He presented results based on the 2017 landscape study report on open access monographs with some new results from more recent studies by Springer Nature and a follow-up report by Knowledge Exchange.</p> <p>Olaf Siegert from ZBW - Leibniz Information Centre for Economics presented a brief overview of what libraries can do to promote OA books in terms of collection management, publication services and development of staff and organisation. The conclusion is that it is not necessarily big changes that are needed.</p> <p>Sofie Wennström presented results from a survey aimed at European research libraries on behalf of the LIBER Open Access Working Group. The survey reveals that many libraries are already working with processes to promote OA books. This is done by libraries organising publishing services or inhouse publishing, by including OA books in discovery services and repositories and by supporting authors to learn more about open access and open licensing.</p> <p>Finally, the LIBER Open Access Working Group shares a report from the workshop providing some quick takeaways and some good examples brought up during the breakout session with the workshop participants.</p>
The demo data set for the meta16S-Seq workflow using Qiime2
<p>The demo data used in the introduction of meta16S-Seq workflow using Qiime2. The original manuscript of the workflow introduction is written by Yuh Shiwa. The workflow is translated in Common Workflow Language by Tazro Ohta.</p>
Cloud-Repro: Reproducible Workflow on a Public Cloud for Computational Fluid Dynamics
<p>In a new effort to make our research transparent and reproducible by others, we developed a workflow to run computational studies on a public cloud. It uses Docker containers to create an image of the application software stack. We also adopt several tools that facilitate creating and managing virtual machines on compute nodes and submitting jobs to these nodes. The configuration files for these tools are part of an expanded "reproducibility package" that includes workflow definitions for cloud computing, in addition to input files and instructions. This facilitates re-creating the cloud environment to re-run the computations under the same conditions.</p> <p>The present Zenodo dataset contains all secondary data required to reproduce the figures of the manuscript ("Reproducible Workflow on a Public Cloud for Computational Fluid Dynamics") without running the CFD simulations again.</p>
ROC_all SINTEF workflow results with consensus steady state from CCLE
<p>This dataset has the input data + result dataset that was the product of using the <strong>DrugLogics</strong> computational pipeline with the <strong>rbbt</strong> workflow system to predict synergistic drug combinations across 8 cell lines that were also tested in SINTEF. The <strong>atopo topology</strong> was used and the logical models were trained to a <strong>consensus steady state derived from ~1000 cell lines from CCLE</strong>.</p>
ROC_all SINTEF workflow results with Atopo topology
<p>This dataset has the input data + result dataset that was the product of using the <strong>DrugLogics</strong> computational pipeline with the <strong>rbbt</strong> workflow system to predict synergistic drug combinations across 8 cell lines that were also tested in SINTEF. <strong>An automated generated Signor-based topology</strong> was used and the logical models were trained to a <strong>steady state activity profile </strong>that was derived using the <strong>PARADIGM tool</strong> and input from the<strong> CCLE</strong>.</p>
ROC_all SINTEF workflow results with CASCADE topology
<p>This dataset has the input data + result dataset that was the product of using the <strong>DrugLogics</strong> computational pipeline with the <strong>rbbt</strong> workflow system to predict synergistic drug combinations across 8 cell lines that were also tested in SINTEF. The <strong>CASCADE topology</strong> was used and the logical models were trained to a <strong>steady state activity profile </strong>that was derived using the <strong>PARADIGM tool</strong> and input from the<strong> CCLE</strong>.</p>
A small dataset for demonstrating the benchmarking of spot-detection/spot-counting workflows with BIAFLOWS
<p>The images were generated by <a href="http://www.cs.tut.fi/sgn/csb/simcep/tool.html">SIMCEP</a>, a widefield fluorescence microscopy biological images simulator.</p> <p>The dataset contains 5 input images and 5 ground-truth images with the suffix _lbl.</p>
Fig. 2 in Workflow of Lotmaria passim isolation: Experimental infection with a low-passage strain causes higher honeybee mortality rates than the PRA-403 reference strain
Fig. 2. Kaplan-Meier survival curves for the experimental groups (C1, control and PRA-403), showing the cumulative mortality over time. Vertical ticks indicate censored observations.
Fig. 1 in Workflow of Lotmaria passim isolation: Experimental infection with a low-passage strain causes higher honeybee mortality rates than the PRA-403 reference strain
Fig. 1. Workflow for the isolation of bee-infecting trypanosomatid parasites from honeybee guts. A. Dissection and tissue processing (steps 1–4), and trypanosomatid culture and expansion (step 5) in liquid or Solid Cultures. B. Growth curve of L. passim PRA-403 strain in decreasing concentrations of 5-Fluorocytosine (1 × 106 μg/ mL-100 μg/mL) to determine the maximum dose for parasite survival. C. Giemsa staining of L. passim C1 (CCP 1). D. Hoescht DNA staining of live L. passim C1 (CCP 1): N, Nucleus; K, Kinetoplast; E, Scanning Electron Microscopy of L. passim C1 (CCP 1) grown in Agar Solid cultures 20 days post-inoculation.
PySCENIC reduced test dataset for workflow testing
<p>Reduced size test dataset for workflow components of PySCENIC, including expected results (inputs for subsequent steps), directly collected or derived from https://github.com/aertslab/SCENICprotocol/tree/master/example .</p>
UniSpec: Deep Learning for Predicting the Full Range of Peptide Fragment Ion Series to Enhance the Proteomics Data Analysis Workflow
<p>UniSpec is a comprehensive DL spectrum predictor that can predict the intensity of the entire HCD MS/MS fragment ion series, going beyond existing tools limited to b/y ion series. </p> <p>All datasets developed for UniSpec model are shared on Zenodo as part of the UniSpec publication, "UniSpec: Deep Learning for Predicting Comprehensive Peptide Fragment Ion Series to Improve Peptide-Spectrum Matches from Shotgun Proteomics Experiments".</p> <p>This includes UniSpec datasets, downstream evaluation and analysis, and application case studies.</p> <p>1. pre-processed training, evaluation and testing data for machine learning;</p> <p> UniSpec-Datasets.7z, Readme_UniSpecDatasets.txt</p> <p>2. Streamlined input datasets based on the fragmentation dictionary;</p> <p> Streamlined_inputdatasets.7z, Readme_Streamlined_inputdatasets.txt</p> <p>3. Predictions on the validation and test sets;</p> <p> UniSpecPred_Validation-Test.7z, Readme_Predictons_ValidationTest.txt</p> <p>4. Evaluation by comparison with Prosit;</p> <p> a. Predictions: prosit_and_unispec_predictions.7z, Readme_prosit_and_unispec_predictions.txt</p> <p> b. Cosine similarity scores: prosit_vs_unispec_CS.7z, Readme_prosit_vs_unispec_CS.txt</p> <p>5. CSS for Different HCD Fragment Ion Series;</p> <p> CS_for_ion_splits.tsv</p> <p>6. Application 1: PSM rescoring;</p> <p> PSM rescoring_zipfiles.7z, PSM rescoring_readme.txt</p> <p>7. Application 2: In-silico spectral library search </p> <p> in-silico_librarysearch.7z, in-silico_librarysearch_readme.txt</p> <p> </p>
Galaxy workflow from Galaxy 101 for everyone
<p>Galaxy workflow from Galaxy 101 for everyone. This workflow is used in the training "How to reproduce published Galaxy analyses" to learn how to run a published Galaxy workflow.</p>
SAAG Workflow Evaluation Results for Anirudh Prabhu's PhD Dissertation
<p>Part of Anirudh Prabhu's PhD Dissertation. Contains Interpretability, Coverage and Normalized Coverage scores for 65 stories used in the SAAG workflow. More details about the evaluation can be found in Chapter 6 and Appendix F of the dissertation document. </p>
Protein Structure Files and Galaxy Workflows for Conducting Molecular Dynamics Simulations of Coronavirus Helicases
<p>The files included here are a set of Galaxy workflows, starting structure files (PDB, mol2, and frcmod), and specialized force field files (ZAFF) for the simulation of coronavirus helicases in the apo and drug-bound state. The inhibitor molecules include those from virtual screening (FCID1 and thioguanine), as well as experimentally validated candidates (Lumacaftor and SSYA10-001).</p>
Protein Structure Files and Galaxy Workflows for Conducting Molecular Dynamics Simulations of Flavivirus Helicases
<p>The files included here are a set of Galaxy workflows and starting structure files (PDB, mol2, and frcmod) for the simulation of flavivirus helicases in the apo and drug-bound state. The inhibitors include the 4th highest ranking compound from a virtual screening of more than 12.7 million drug-like molecules.</p>
Datset workflow
<p>Dataset usado para la practica final.</p>
WhereWulff: A semi-autonomous workflow for systematic catalyst surface reactivity under reaction conditions
<p>This repository houses electronic structure data and metadata generated as part of a computational chemistry case study, enabling full analysis of the paper "WhereWulff: A semi-autonomous workflow for systematic catalyst surface reactivity under reaction conditions" by Rohan Yuri Sanspeur, Javier Heras-Domingo, John R. Kitchin and Zachary Ulissi.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.