Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
650
datasets available to search
ShareScore release 0.9.0
Dataset results
650 results for “Workflow”
Supplementary material 1 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Appendices
Figure 1e from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 1e A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Liquid preserved specimen (Natural History Museum 2010)
Supplementary material 4 from: Seebens H, Clarke DA, Groom Q, Wilson JRU, García-Berthou E, Kühn I, Roigé M, Pagad S, Essl F, Vicente J, Winter M, McGeoch M (2020) A workflow for standardising and integrating alien species distribution data. NeoBiota 59: 39-59. https://doi.org/10.3897/neobiota.59.53578
Intermediate output files of the case study applying the SInAS workflow
Supplementary material 3 from: Seebens H, Clarke DA, Groom Q, Wilson JRU, García-Berthou E, Kühn I, Roigé M, Pagad S, Essl F, Vicente J, Winter M, McGeoch M (2020) A workflow for standardising and integrating alien species distribution data. NeoBiota 59: 39-59. https://doi.org/10.3897/neobiota.59.53578
Final output files of the case study applying the SInAS workflow
Supplementary material 2 from: Seebens H, Clarke DA, Groom Q, Wilson JRU, García-Berthou E, Kühn I, Roigé M, Pagad S, Essl F, Vicente J, Winter M, McGeoch M (2020) A workflow for standardising and integrating alien species distribution data. NeoBiota 59: 39-59. https://doi.org/10.3897/neobiota.59.53578
Supplementary Tables S1–S4
Supplementary material 5 from: Seebens H, Clarke DA, Groom Q, Wilson JRU, García-Berthou E, Kühn I, Roigé M, Pagad S, Essl F, Vicente J, Winter M, McGeoch M (2020) A workflow for standardising and integrating alien species distribution data. NeoBiota 59: 39-59. https://doi.org/10.3897/neobiota.59.53578
Unresolved entries of the case study applying the SInAS workflow
Supplementary material 1 from: Seebens H, Clarke DA, Groom Q, Wilson JRU, García-Berthou E, Kühn I, Roigé M, Pagad S, Essl F, Vicente J, Winter M, McGeoch M (2020) A workflow for standardising and integrating alien species distribution data. NeoBiota 59: 39-59. https://doi.org/10.3897/neobiota.59.53578
Technical description and manual of the SInAS workflow implementation in R
Figure 2 from: Owen D, Groom Q, Hardisty A, Leegwater T, Livermore L, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e58030. https://doi.org/10.3897/rio.6.e58030
Figure 2 A possible semi-automatic digitisation workflow to extract data from the labels of collection specimens.
KNIME workflows for the evaluation of neurotoxic effects in zebrafish embryos by automatic measurement of early motor behaviours
<p>Zebrafish (<em>Danio rerio</em>) has rapidly become a popular model species for behavioural studies that may be relevant to drug screening and safety toxicology. Zebrafish embryos show a complex behavioural repertoire already a few hours after fertilization. Particularly, early stage zebrafish show characteristic behavioural features such as spontaneous tail coiling (STC) or induced movements when exposed to a short and bright light flash (called photomotor response -PMR-). In this chapter, we provide the methods for assessing STC and PMR in zebrafish embryos and to detect changes provoked by chemicals. One of the protocols uses video analysis suitable for automated high-throughput screening. Moreover, both protocols describe the use of automated video analysis by using an open-source integration platform (KNIME analytics platform), providing a flexible workflow system that can be adapted to a diversity of video recordings. We also provide a toxicological validation of this assay and show that these protocols can be used to provide an automated, high data-content readout for zebrafish behavioural responses. </p>
From leaf to label: A robust automated workflow for stomata detection: Light microscope images of stomata.
<p>All light microscope images used for training and testing of the deep learning model developed in the study: Meeus S., Van den Bulcke J., wyffels F. (2020) From leaf to label: A robust automated workflow for stomata detection. Ecol. Evol. <a href="https://doi.org/10.1002/ece3.6571">https://doi.org/10.1002/ece3.6571</a>.</p>
Data from: Setup in a clinical workflow and impact on radiotherapy routine of an in vivo dosimetry procedure with an electronic portal imaging device
High conformal techniques such as intensity-modulated radiation therapy and volumetric-modulated arc therapy are widely used in overloaded radiotherapy departments. In vivo dosimetric screening is essential in this environment to avoid important dosimetric errors. This work examines the feasibility of introducing in vivo dosimetry (IVD) checks in a radiotherapy routine. The causes of dosimetric disagreements between delivered and planned treatments were identified and corrected during the course of treatment. The efficiency of the corrections performed and the added workload needed for the entire procedure were evaluated. The IVD procedure was based on an electronic portal imaging device. A total of 3682 IVD tests were performed for 147 patients who underwent head and neck, abdomen, pelvis, breast, and thorax radiotherapy treatments. Two types of indices were evaluated and used to determine if the IVD tests were within tolerance levels: the ratio R between the reconstructed and planned isocentre doses and a transit dosimetry based on the γ-analysis of the electronic portal images. The causes of test outside tolerance level was investigated and corrected and IVD test was repeated during subsequent fraction. The time needed for each step of the IVD procedure was registered. Pelvis, abdomen, and head and neck treatments had 10% of tests out of tolerance whereas breast and thorax treatments accounted for up to 25%. The patient setup was the main cause of 90% of the IVD tests out of tolerance and the remaining 10% was due to patient morphological changes. An average time of 42 min per day was sufficient to monitor a daily workload of 60 patients in treatment. This work shows that IVD performed with an electronic portal imaging device is feasible in an overloaded department and enables the timely realignment of the treatment quality indices in order to achieve a patient's final treatment compliant with the one prescribed.
Data from: From benchtop to desktop: important considerations when designing amplicon sequencing workflows
Amplicon sequencing has been the method of choice in many high-throughput DNA sequencing (HTS) applications. To date there has been a heavy focus on the means by which to analyse the burgeoning amount of data afforded by HTS. In contrast, there has been a distinct lack of attention paid to considerations surrounding the importance of sample preparation and the fidelity of library generation. No amount of high-end bioinformatics can compensate for poorly prepared samples and it is therefore imperative that careful attention is given to sample preparation and library generation within workflows, especially those involving multiple PCR steps. This paper redresses this imbalance by focusing on aspects pertaining to the benchtop within typical amplicon workflows: sample screening, the target region, and library generation. Empirical data is provided to illustrate the scope of the problem. Lastly, the impact of various data analysis parameters is also investigated in the context of how the data was initially generated. It is hoped this paper may serve to highlight the importance of pre-analysis workflows in achieving meaningful, future-proof data that can be analysed appropriately. As amplicon sequencing gains traction in a variety of diagnostic applications from forensics to environmental DNA (eDNA) it is paramount workflows and analytics are both fit for purpose.
Workflow and data for: Elevated temperature decreases stony coral tissue loss disease (SCTLD) transmission rate, with little effect of nutrients V1.1
<p>Changes for review 1</p>
Dataset for EBAII practical session (ChIP-seq workflow)
<p>Dataset for EBAII practical session (ChIP-seq workflow)</p>
Supplementary material 1 from: Borisenko A, Young R, Hanner R (2024) A lab-centric, workflow-based data management system for environmental DNA research. Research Ideas and Outcomes 10: e120483. https://doi.org/10.3897/rio.10.e120483
eDNA Laboratory Database Schema Outline
BioToFlow: a corpus annotated with bioinformatics workflows information
<div> <div><em>BioToFlow</em> is a corpus describing bioinformatics workflows in English publications. These annotations are available in the BRAT Rapid Annotation Tool (BRAT) standoff format (https://brat.nlplab.org/standoff.html).</div> <br> <div>This corpus is composed of 52 articles (26 articles related to Nextflow workflows and 26 on Snakemake workflows, randomly selected from PubMed) with a total of 78 419 tokens 27 786 annotated tokens.</div> <div> </div> <h2>Repository organisation</h2> <div> <div> <div> <div>The articles are separated into two directories:</div> <div> <ul> <li>one containing all the articles for the training phases (39) and</li> <li>the other with 13 articles for test.</li> </ul> </div> <br> <h2>Papers</h2> <br>Please cite <em>BioToFlow</em> in any research that uses or extends it :<br> <div> </div> <div> <ul> <li>Sebe, C., Cohen-Boulakia, S., Ferret, O., Névéol, A.: Extracting information in a low-resource setting: Case study on bioinformatics workflows (2024), https://arxiv.org/abs/2411.19295</li> </ul> </div> <div>In this article accepted to IDA 2025 (in English), we present *BioToFlow* and experiments with few shot named entity recognition (NER) using an autoregressive language model, we also use a pre-existing corpus and test integration of workflow knowledge in NER models.</div> <br> <div> </div> <div> <ul> <li>Clémence Sebe, Sarah Cohen-Boulakia, Olivier Ferret, Aurélie Névéol. Extraction d’entités nommées décrivant des chaînes de traitement bioinformatiques dans des articles scientifiques en anglais. 35emes Journées d’Études sur la Parole (JEP 2024) 31eme Conférence sur le Traitement Automatique des Langues Naturelles (TALN 2024) 26eme Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RECITAL 2024), Jul 2024, Toulouse, France. pp.422-434. hal-04623033.</li> </ul> </div> <div>In this article (in French), we present the second version of *BioToFlow*. The new articles are annotated with entities and attributes. We conduct preliminary experiments with NER with a specific focus on the memorization vs. generalization abilities of statistical and rule-based methods.</div> <br><br> <div> <ul> <li>Sebe C., Névéol A., Cohen-Boulakia S. & Gaignard A. (2023). Extraction d’informations sur les workflows scientifiques à partir de la littérature. volume Extraction et Gestion des Connaissances, RNTI-E-39, p. 313.</li> </ul> </div> <div>In this article (in French), we present the first version of the corpus with 24 articles annotated with entities and relations. We show the feasibility of the task of NER.</div> </div> <div> </div> <h2><br>Contact</h2> <div> <div> <ul> <li>Clémence Sebe, clemence.sebe@universite-paris-saclay.fr</li> </ul> </div> <br> <h2>Funding</h2> This work received support from the National Research Agency under the France 2030 program, with reference to ANR-22-PESN-0007.</div> </div> </div> </div>
Supplementary material 1 from: Seebens H, Kaplan E (2022) DASCO: A workflow to downscale alien species checklists using occurrence records and to re-allocate species distributions across realms. NeoBiota 74: 75-91. https://doi.org/10.3897/neobiota.74.81082
Manual of DASCO
Supplementary material 1 from: Niehues A, de Visser C, Hagenbeek FA, Karu N, Kindt ASD, Kulkarni P, Pool R, Boomsma DI, van Dongen J, van Gool AJ, `t Hoen PAC (2022) A Multi-omics Data Analysis Workflow Packaged as a FAIR Digital Object. Research Ideas and Outcomes 8: e94042. https://doi.org/10.3897/rio.8.e94042
Members of the ACTION Consortium
Supplementary material 1 from: Vohland K, Hoffmann A, Underwood E, Weatherdon L, Bonet F, Häuser C, Wetzel F (2016) 3rd EU BON Stakeholder Roundtable (Granada, Spain): Biodiversity data workflow from data mobilization to practice. Research Ideas and Outcomes 2: e8622. https://doi.org/10.3897/rio.2.e8622
3rd EU BON Stakeholder Roundtable – Acronyms
Reciprocal best hits BLAST files for "EXCRETE workflow enables deep proteomics of the microbial extracellular environment"
<p>.fasta file input and tabular output from a BLAST reciprocal best hits analysis on https://usegalaxy.eu/</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.