Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
650
datasets available to search
ShareScore release 0.9.0
Dataset results
650 results for “Workflow”
data for BioImage.IO workflows
<p>Test data and cover images for workflows implemented in https://github.com/bioimage-io/workflows-bioimage-io-python</p> <p><br> References:<br> stardist_chatty_frog.npy derived from https://doi.org/10.5281/zenodo.6338615<br> dask_inference_cover.svg/png: raw data from https://cremi.org/, predictions from https://bioimage.io/#/?id=10.5281%2Fzenodo.5874741, dask icon: https://dask.org</p> <p> </p>
Supplementary materials of the manuscript "Establishing a new workflow in the study of core reduction intensity and distribution".
<p>This repository hosts the R code scripts and datasets that allow reproducibility and replicability of the statistical analyses implemented in the paper: Lombao et al. (2022). Establishing a new workflow in the study of core reduction distribution. Journal of Lithic Studies.</p> <p> </p> <p>The analytical and statistical protocols applied for this study were implemented in in R (version 3.6.3) (R Core Team, 2021).</p> <p>Contents:<br> 1. Script.txt: The R- Script with all the packages and functions used and all the steps followed in the statistical analyses.<br> 2. vrm_experiment_database.xlsx: the database with the data used in this manuscript. The data comes from an experiment presented in Lombao et al., 2020. A new approach to measure reduction intensity on cores and tools on cobbles: the Volumetric Reconstruction Method. Archaeological and Anthropological Sciences 12(9)<br> DOI: 10.1007/s12520-020-01154-7<br> 3. Supplementary_table_S1.docx<br> 4. Supplementary_table_S2.docx<br> 5. Supplementary_table_s3.xlsx This database is used in the R-script. </p>
Earthquake catalogs for: A specific earthquake processing workflow for studying long-lived explosive volcanic eruptions with application to the 2008 Okmok eruption
<p>Repository for the seismic catalogs from Garza-Giron et al. (2023a,b). These include the catalog with absolute locations using NonLinLoc (Lomax et al., 2001; Lomax and Curtis, 2001), and the relocated catalogs using hypoDD (Waldhauser and Ellsworth, 2000) and GrowClust (Trugman and Shearer, 2017).</p> <p>The header of the CSV files is as follows:</p> <p><strong>Date</strong> (year/month/day), <strong>Time</strong> (hr:min:sec:msec), <strong>Latitude</strong> (decimal degrees), <strong>Longitude</strong> (decimal degrees), <strong>Depth</strong> (km), <strong>Magnitude</strong> (Ml calculated for this study), <strong>Event_type</strong> (VT:vulcano-tectonic;LP:long-period), <strong>Number of stations</strong> where the event was detected, <strong>ID</strong></p> <p>References:</p> <div>Garza‐Giron, R., Brodsky, E. E., Spica, Z. J., Haney, M. M., & Webley, P. W. (2023a). A specific earthquake processing workflow for studying long‐lived, explosive volcanic eruptions with application to the 2008 Okmok Volcano, Alaska, eruption. <em>Journal of Geophysical Research: Solid Earth</em>, e2022JB025882.</div> <div> </div> <div> <div>Garza‐Girón, R., Brodsky, E. E., Spica, Z. J., Haney, M. M., & Webley, P. W. (2023b). Earthquakes record cycles of opening and closing in the enhanced seismic catalog of the 2008 Okmok Volcano, Alaska, eruption. <em>Journal of Geophysical Research: Solid Earth</em>, <em>128</em>(7), e2023JB026893.</div> <div> </div> </div> <p>Lomax A, Curtis A (2001) Fast, probabilistic earthquake location in 3-D models using oct-tree importance sampling. Geophys Res Abstracts, 3:955.</p> <p>Lomax, A., Zollo, A., Capuano, P., and Virieux, J. (2001). Precise, absolute earthquake location under Somma‐Vesuvius volcano using a new 3D velocity model.Geophysical Journal International,146, 313–331.</p> <p>Trugman, D. T., and Shearer, P. M. (2017). GrowClust: A hierarchical clustering algorithm for relative earthquake relocation, with application to the Spanish Springs and Sheldon, Nevada, earthquake sequences. Seismological Research Letters, 88(2A), 379-391.</p> <p>Waldhauser, F., and Ellsworth, W. L. (2000). A double-difference earthquake location algorithm: Method and application to the northern Hayward fault, California. Bulletin of the Seismological Society of America, 90(6), 1353-1368.</p>
Data for manuscript, "An optimized workflow for MS-based quantitative proteomics of challenging clinical bronchoalveolar lavage fluid (BALF) samples"
<p>Clinical BALF samples are rich in biomolecules, including proteins, and useful for molecular studies of lung health and disease. However, MS based proteomic analysis of BALF is impeded by the dynamic range of protein abundance, and potential for interfering contaminants. We have developed a workflow that eliminates these challenges. By combining high abundance protein depletion, protein trapping, clean-up, and in-situ tryptic digestion, our workflow is compatible with both qualitative and quantitative MS-based proteomic analysis. The workflow includes collection of endogenous peptides for peptidomic analysis of BALF, if desired, as well as amenability to offline semi-preparative or microscale fractionation of peptide mixtures prior to LC-MS/MS analysis, for increased depth of analysis. We show the effectiveness of this workflow on BALF samples from COPD patients. Overall, our workflow should allow MS-based proteomics to be applied to a wide variety of studies focused on BALF clinical samples. </p> <p>Note: Due to the nature of some of the files, file <em>wendt005_ostr0103_18260_20210831_BALF_FAIMS_MS2_TMT16.msf, wendt005_ostr0103_18976_20230202_quantReport.msf, cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_1R.raw, cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_2R.raw, cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_3R.raw and cmsptc_higgi022_18988_20230203_18976DW_Eclipse_noFAIMS_quantReport.msf</em> were zipped into compressed folders before uploading.</p>
Protein Structure Files and Galaxy Workflows for Conducting Molecular Dynamics Simulations of Coronavirus Helicases -- Output Files
<p>These are the output files generated using the input files and Galaxy workflows for coronavirus helicase simulations, from: </p> <pre>https://doi.org/10.5281/zenodo.7492987</pre>
Supplementary materials to: Nano-Strainer: a workflow for identification of single-copy nuclear loci for plant systematic studies, using target capture kits and Oxford Nanopore long reads
<p>In the paper associated with this dataset, a workflow is presented which enables the identification of single-/low-copy nuclear molecular markers for a plant group of interest, by mining data from a small representative target capture experiment done using a commercial probe kit and Oxford Nanopore long-read sequencing. The proposed pipeline first assesses sequence variability contained in the data from targeted loci and assigns reads to their respective genes, via a combined BLAST/clustering procedure. Cluster consensus sequences are then examined based on four pre-defined criteria presumably indicative for absence of paralogy. This is done by calculating four specialized indices; loci are ranked according to their performance in these indices, and top-scoring loci are considered putatively single- or low-copy. The approach can be applied to any probe set. As it relies on long reads, the contribution also provides template workflows for processing Nanopore-based target capture data. Identified loci can be used for NGS amplicon sequencing. For detection of possibly remaining paralogy in these data, which might occur in groups with rampant paralogy, the long-read assembly tool CANU is employed. The presented workflow can be useful for researchers dealing with reticulate or polyploidization phylogenetic histories in plants.</p> <p>The present dataset contains several documents supplementing the original paper. Its most important elements are a detailed description (alongside two graphical workflow figures) of all methods employed in the study, suitable for reproducing the steps of the workflow and also the wet-lab work. The workflow employs a collection of BASH, Python and R scripts which is available here, together with a detailed account on command line use in Linux. Also, reference sequences for the identified markers can be found as well as sequence alignments derived from the amplicon sequencing.</p>
metaGOflow: a workflow for the analysis of marine Genomic Observatories shotgun metagenomics data - use case
<p>Data products returned by <a href="https://github.com/emo-bon/MetaGOflow">metaGOflow</a> (<a href="https://github.com/emo-bon/MetaGOflow/releases/tag/v1.0.0">v1.0.0</a>) and packed as a Research Object (RO) Crate, when performed with:</p> <ul> <li>a <strong>seawater metagenomic sample </strong>(TARA OCEAN, <a href="https://www.ebi.ac.uk/ena/browser/view/ERR599171">ERR599171</a>)</li> <li>a <strong>fish gut </strong>sample (<a href="https://www.ebi.ac.uk/ena/browser/view/ERR4765907">ERR4765907</a>)</li> <li>a<strong> human gut </strong>sample (<a href="https://www.ebi.ac.uk/ena/browser/view/SRR9654976">SRR9654976</a>)</li> </ul> <p>This Zenodo repo accompanies the metaGOflow paper and more about the analysis of this sample can be found there.</p> <p>You can also have a look at some visual components of the workflow at this <a href="https://data.emobon.embrc.eu/MetaGOflow/">GitHub page</a>. </p> <p>The source code of metaGOflow is available through <a href="http://github.com/emo-bon/MetaGOflow">GitHub</a>.</p>
FIGURE 1 The Senckenberg model for a streamlined taxonomy workflow involving a in Why is there no service to support taxonomy?
FIGURE 1 The Senckenberg model for a streamlined taxonomy workflow involving a commercial service that covers transferable technical aspects of species descriptions. This figure was designed using icons from Flaticon.com.
Supplementary Datasets for: 'A processing and analytics system for microscopy data workflows: the Pycroscopy ecosystem of packages'
<p>The repository contains four independent datasets that are a part of the publication (<a href="https://arxiv.org/abs/2302.14629">arXiv:2302.14629</a>), which delineates the capabilities of the Pycroscopy ecosystem of packages. The details of the individual datasets can be found below. </p> <p>1) bfo_iv_final.hf5: Dataset of I-V curves captured by conductive atomic force microscopy on a BiFeO3 sample. The data has been transformed so that we plot not the log of the current density (J) as a function of the square root of the electric field. The dataset was originally presented in the paper 10.1038/s41467-017-01334-5 </p> <p>2) bto_atomic.dm3: Atomically resolved data BaTiO3 thin film acquired with scanning transmission electron microscopy. These were originally captured in the dm3 file format. This dataset was a part of the publication: doi.org/10.1002/adma.202106426</p> <p>3) EELS_STO.dm3: Scanning transmission electron microscope (STEM)-Electron energy loss spectroscopy (EELS) dataset of SrTiO3.</p> <p>4) STO-stack.h5: High-angle annular dark-field imaging (HAADF) scanning transmission electron microscope (STEM) image stack of SrTiO3. This image stack contains 25 images.</p> <p> </p>
Test datasets for the IWC workflow Assembly-Hifi-only-VGP3
<p>Test datasets for the workflow Assembly-Hifi-only-VGP3</p> <p>Datasets : </p> <ul> <li>Hifi reads</li> <li>Meryl database</li> <li>Genomescope profile summary</li> </ul> <p> </p>
Predicting time of failure of Internet of Things devices using Bayesian workflow
<p>Repository includes the environment details in ```energies_iot_env.yml``` file.</p> <p>All code for model analysis is included in the ```iot_tests_refactor.ipynb``` notebook.</p> <p>Code for computing simulation based calibration is in the ```compute_sbc.py``` file, and can be run by ```just_csv.ipynb``` notebook.</p> <p> </p> <p>The data set was created in the project NCN OPUS "Process Fault Prediction and Detection" (UMO-2021/41/B/ST7/03851)</p>
Spheroids workflow: KapoorLabs
<p>Presentation material for workflow developed for cell segmentation, cell action type classification, tracking and auto track correction with track analysis and track classification for the group of Prof. Chris Bakal at ICR, London.</p>
Supplementary materials to: Nano-Strainer: a workflow for identification of single-copy nuclear loci for plant systematic studies, using target capture kits and Oxford Nanopore long reads
Open the record for dataset details and reuse information.
Tracking small animals in complex landscapes: a comparison of localisation workflows for automated radio telemetry systems
Open the record for dataset details and reuse information.
specleanr: An R package for automated flagging of environmental outliers in ecological data for modeling workflows
Open the record for dataset details and reuse information.
Data from: Automated workflow for the cell cycle analysis of (non-)adherent cells using a machine learning approach
Open the record for dataset details and reuse information.
LipidQuant 1.0: Automated data processing in lipid class separation - mass spectrometry quantitative workflows
Open the record for dataset details and reuse information.
The FloRes Database: A floral resources trait database for pollinator habitat-assessment generated by a multistep workflow
Open the record for dataset details and reuse information.
Bat-aggregated time series workflow
Open the record for dataset details and reuse information.
BIDS Data for "An Optimized Registration Workflow and Standard Geometric Space for Small Animal Brain Imaging"
<p>Base data package for the “An Optimized Registration Workflow and Standard Geometric Space for Small Animal Brain Imaging” article, formatted corresponding to the Brain Imaging Data Structure.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.