Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
146
datasets available to search
ShareScore release 0.9.0
Dataset results
146 results for “data workflow”
Data from: A from-benchtop-to-desktop workflow for validating HTS data and for taxonomic identification in diet metabarcoding studies
Open the record for dataset details and reuse information.
Data from: From population genomics to conservation and management: a workflow for targeted analysis of markers identified using genome-wide approaches in Atlantic salmon Salmo salar
Open the record for dataset details and reuse information.
Overview of XCT data processing workflow for ammonium nitrate prills quantitative analysis
<p>This video presents the data processing workflow that was developped to perform the quantitative structureal and morphological analysis of ammonium nitrate prills by X-ray computed tomography.. </p>
Test data for running snakePipes : DNA-mapping workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run DNA-mapping workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for human (<strong>hg38</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>
MePPi: A complete and flexible workflow for metaproteomics data analyses
<p>Data for an examplary metaproteomics data analysis with the <a href="https://github.com/compomics/meta-proteome-analyzer">MetaProteomeAnalyzer</a> (MPA) and <a href="https://gitlab.com/s.fuchs/prophane/">Prophane</a> software tools. Data is from the PRIDE dataset <a href="https://www.ebi.ac.uk/pride/archive/projects/PXD010550/">PXD010550</a>.</p> <p>Files include:</p> <ul> <li>protein databases (FASTA) : <ol> <li>UniProt Swiss-Prot: <a href="https://zenodo.org/record/3727600/files/UniprotSwP-2020_03.fasta">UniprotSwP-2020_03.fasta</a></li> <li>Metagenome (+ Swiss-Prot): <a href="https://zenodo.org/record/3727600/files/MG_BG__UPSP-sp_2020_03.fasta">MG_BG__UPSP.fasta</a></li> </ol> </li> <li>MS Datasets (MGF): <ol> <li>FASP digest: <a href="https://zenodo.org/record/3727600/files/FASP_BGP_A.mgf">FASP_BGP_A.mgf</a></li> <li>In-gel digest: <a href="https://zenodo.org/record/3727600/files/InGel_BGP_A.mgf">InGel_BGP_A.mgf</a></li> </ol> </li> <li>Example results for a single experiment analysis (Sample A, based on: MS data: FASP digest, FASTA: UniProt Swiss-Prot): <ul> <li>MPA results: <a href="https://zenodo.org/record/3727600/files/mpa_result-sample_a-fdr_0.05-single_exp.csv">mpa_result-sample_a-fdr_0.05-single_exp.csv</a></li> <li>Prophane results: <a href="https://zenodo.org/record/3727600/files/prophane_result-sample_a.zip">prophane_result-sample_a.zip</a></li> </ul> </li> <li>Example results for a multi-experiment analysis (Sample B, based on: MS data: FASP + in-gel digest, FASTA: Metagenome): <ul> <li>MPA results: <a href="https://zenodo.org/record/3727600/files/mpa_result-sample_b-fdr_0.01-multi_exp.csv">mpa_result-sample_b-fdr_0.01-multi_exp.csv</a></li> <li>Prophane results: <a href="https://zenodo.org/record/3727600/files/prophane_result-sample_b.zip">prophane_result-sample_b.zip</a></li> </ul> </li> <li><a href="https://zenodo.org/record/3727600/files/mpa_ressources_incl_swissprot_03-2020.zip">MPA data dump</a> including preprocessed UniProt Swiss-Prot FASTA (optionally used by <a href="https://anaconda.org/bioconda/mpa-server">conda mpa-server package</a>)</li> </ul>
Supplementary material 4 from: Seebens H, Clarke DA, Groom Q, Wilson JRU, García-Berthou E, Kühn I, Roigé M, Pagad S, Essl F, Vicente J, Winter M, McGeoch M (2020) A workflow for standardising and integrating alien species distribution data. NeoBiota 59: 39-59. https://doi.org/10.3897/neobiota.59.53578
Intermediate output files of the case study applying the SInAS workflow
Supplementary material 3 from: Seebens H, Clarke DA, Groom Q, Wilson JRU, García-Berthou E, Kühn I, Roigé M, Pagad S, Essl F, Vicente J, Winter M, McGeoch M (2020) A workflow for standardising and integrating alien species distribution data. NeoBiota 59: 39-59. https://doi.org/10.3897/neobiota.59.53578
Final output files of the case study applying the SInAS workflow
Supplementary material 2 from: Seebens H, Clarke DA, Groom Q, Wilson JRU, García-Berthou E, Kühn I, Roigé M, Pagad S, Essl F, Vicente J, Winter M, McGeoch M (2020) A workflow for standardising and integrating alien species distribution data. NeoBiota 59: 39-59. https://doi.org/10.3897/neobiota.59.53578
Supplementary Tables S1–S4
Supplementary material 5 from: Seebens H, Clarke DA, Groom Q, Wilson JRU, García-Berthou E, Kühn I, Roigé M, Pagad S, Essl F, Vicente J, Winter M, McGeoch M (2020) A workflow for standardising and integrating alien species distribution data. NeoBiota 59: 39-59. https://doi.org/10.3897/neobiota.59.53578
Unresolved entries of the case study applying the SInAS workflow
Supplementary material 1 from: Seebens H, Clarke DA, Groom Q, Wilson JRU, García-Berthou E, Kühn I, Roigé M, Pagad S, Essl F, Vicente J, Winter M, McGeoch M (2020) A workflow for standardising and integrating alien species distribution data. NeoBiota 59: 39-59. https://doi.org/10.3897/neobiota.59.53578
Technical description and manual of the SInAS workflow implementation in R
Data from: Setup in a clinical workflow and impact on radiotherapy routine of an in vivo dosimetry procedure with an electronic portal imaging device
High conformal techniques such as intensity-modulated radiation therapy and volumetric-modulated arc therapy are widely used in overloaded radiotherapy departments. In vivo dosimetric screening is essential in this environment to avoid important dosimetric errors. This work examines the feasibility of introducing in vivo dosimetry (IVD) checks in a radiotherapy routine. The causes of dosimetric disagreements between delivered and planned treatments were identified and corrected during the course of treatment. The efficiency of the corrections performed and the added workload needed for the entire procedure were evaluated. The IVD procedure was based on an electronic portal imaging device. A total of 3682 IVD tests were performed for 147 patients who underwent head and neck, abdomen, pelvis, breast, and thorax radiotherapy treatments. Two types of indices were evaluated and used to determine if the IVD tests were within tolerance levels: the ratio R between the reconstructed and planned isocentre doses and a transit dosimetry based on the γ-analysis of the electronic portal images. The causes of test outside tolerance level was investigated and corrected and IVD test was repeated during subsequent fraction. The time needed for each step of the IVD procedure was registered. Pelvis, abdomen, and head and neck treatments had 10% of tests out of tolerance whereas breast and thorax treatments accounted for up to 25%. The patient setup was the main cause of 90% of the IVD tests out of tolerance and the remaining 10% was due to patient morphological changes. An average time of 42 min per day was sufficient to monitor a daily workload of 60 patients in treatment. This work shows that IVD performed with an electronic portal imaging device is feasible in an overloaded department and enables the timely realignment of the treatment quality indices in order to achieve a patient's final treatment compliant with the one prescribed.
Data from: From benchtop to desktop: important considerations when designing amplicon sequencing workflows
Amplicon sequencing has been the method of choice in many high-throughput DNA sequencing (HTS) applications. To date there has been a heavy focus on the means by which to analyse the burgeoning amount of data afforded by HTS. In contrast, there has been a distinct lack of attention paid to considerations surrounding the importance of sample preparation and the fidelity of library generation. No amount of high-end bioinformatics can compensate for poorly prepared samples and it is therefore imperative that careful attention is given to sample preparation and library generation within workflows, especially those involving multiple PCR steps. This paper redresses this imbalance by focusing on aspects pertaining to the benchtop within typical amplicon workflows: sample screening, the target region, and library generation. Empirical data is provided to illustrate the scope of the problem. Lastly, the impact of various data analysis parameters is also investigated in the context of how the data was initially generated. It is hoped this paper may serve to highlight the importance of pre-analysis workflows in achieving meaningful, future-proof data that can be analysed appropriately. As amplicon sequencing gains traction in a variety of diagnostic applications from forensics to environmental DNA (eDNA) it is paramount workflows and analytics are both fit for purpose.
Workflow and data for: Elevated temperature decreases stony coral tissue loss disease (SCTLD) transmission rate, with little effect of nutrients V1.1
<p>Changes for review 1</p>
Supplementary material 1 from: Borisenko A, Young R, Hanner R (2024) A lab-centric, workflow-based data management system for environmental DNA research. Research Ideas and Outcomes 10: e120483. https://doi.org/10.3897/rio.10.e120483
eDNA Laboratory Database Schema Outline
Supplementary material 1 from: Niehues A, de Visser C, Hagenbeek FA, Karu N, Kindt ASD, Kulkarni P, Pool R, Boomsma DI, van Dongen J, van Gool AJ, `t Hoen PAC (2022) A Multi-omics Data Analysis Workflow Packaged as a FAIR Digital Object. Research Ideas and Outcomes 8: e94042. https://doi.org/10.3897/rio.8.e94042
Members of the ACTION Consortium
Supplementary material 1 from: Vohland K, Hoffmann A, Underwood E, Weatherdon L, Bonet F, Häuser C, Wetzel F (2016) 3rd EU BON Stakeholder Roundtable (Granada, Spain): Biodiversity data workflow from data mobilization to practice. Research Ideas and Outcomes 2: e8622. https://doi.org/10.3897/rio.2.e8622
3rd EU BON Stakeholder Roundtable – Acronyms
BlockClust workflow-testing data
<p>Data is taken from https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM450239.</p>
Code to reproduce the data analysis performed in the study "EXCRETE workflow enables deep proteomics of the microbial extracellular environment"
Open the record for dataset details and reuse information.
Data supporting publication: MiFoDB, a workflow for microbial food metagenomic characterization, enables high-resolution analysis of fermented food microbial dynamics
<p>MiFoDB (Microbial Foods Database) is a workflow and primary reference database which includes 675 assembled MAGs and RefSeq bacterial, yeast, fungal, and substrate genomes from fermented foods.</p>
Figure 6 from: Vohland K, Hoffmann A, Underwood E, Weatherdon L, Bonet F, Häuser C, Wetzel F (2016) 3rd EU BON Stakeholder Roundtable (Granada, Spain): Biodiversity data workflow from data mobilization to practice. Research Ideas and Outcomes 2: e8622. https://doi.org/10.3897/rio.2.e8622
Figure 6 - Participants of the 3rd EU BON Stakeholder Roundtable discussing details of the workflow (credits: Katrin Vohland).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.