Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

146

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

146 results for “data workflow”

Learn how ShareScore rates datasets ↗
zenodo52/100

Input data for MFAssignR Galaxy workflow tutorial

<p>This is the input dataset for the MFAssignR Galaxy training workflow. The input dataset corresponds to the model data of MFAssignR (<a title="Raw_Neg_ML" href="https://github.com/skschum/MFAssignR/tree/master/MFAssignR/data" target="_blank" rel="noopener">Raw_Neg_ML</a>), containing a raw mass list, measured in a negative ESI mode.</p>

openmit-licenseSep 2024View details →
zenodo48/100

Data from: A FAIR and modular image-based workflow for knowledge discovery in the emerging field of imageomics

<p>Data and results from the Imageomics Workflow. These include data files from the Fish-AIR repository (https://fishair.org/) for purposes of reproducibility and outputs from the application-specific imageomics workflow contained in the Minnow_Segmented_Traits repository (https://github.com/hdr-bgnn/Minnow_Segmented_Traits).</p> <p>Fish-AIR:<br> This is the dataset downloaded from&nbsp;Fish-AIR, filtering for Cyprinidae and the Great Lakes Invasive Network (GLIN) from the Illinois Natural History Survey (INHS) dataset. These files contain information about fish images, fish image quality, and path for downloading the images. The data download ARK ID is dtspz368c00q. (2023-04-05). The following files are unaltered from the Fish-AIR download. We use the following files:</p> <p>extendedImageMetadata.csv: A CSV file containing information about each image file. It has the following columns: ARKID, fileNameAsDelivered, format, createDate, metadataDate, size, width, height, license, publisher, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>imageQualityMetadata.csv: A CSV file containing information about the quality of each image. It has the following columns: ARKID, license, publisher, ownerInstitutionCode, createDate, metadataDate, specimenQuantity, containsScaleBar, containsLabel, accessionNumberValidity, containsBarcode, containsColorBar, nonSpecimenObjects, partsOverlapping, specimenAngle, specimenView, specimenCurved, partsMissing, allPartsVisible, partsFolded, brightness,&nbsp;<br> uniformBackground, onFocus, colorIssue, quality, resourceCreationTechnique. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>multimedia.csv: A CSV file containing information about image downloads. It has the following columns: ARKID, parentARKID, accessURI, createDate, modifyDate, fileNameAsDelivered, format, scientificName, genus, family, batchARKID, batchName, license, source, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>meta.xml: A XML file with the metadata about the column indices and URIs for each file contained in the original downloaded zip file. This file is used in the fish-air.R script to extract the indices for column headers.</p> <p>The outputs from the Minnow_Segmented_Traits workflow are:</p> <p>sampling.df.seg.csv: Table with tallies of the sampling of image data per species during the data cleaning and data analysis. This is used in Table S1 in Balk et al.&nbsp;</p> <p>presence.absence.matrix.csv: The Presence-Absence matrix from segmentation, not cleaned. This is the result of the combined outputs from the presence.json files created by the rule &ldquo;create_morphological_analysis&rdquo;. The cleaned version of this matrix is shown as Table S3 in Balk et al.</p> <p>heatmap.avg.blob.png and heatmap.sd.blob.png: Heatmaps of average area of biggest blob per trait (heatmap.avg.blob.png) and standard deviation of area of biggest blob per trait (heatmap.sd.blob.png). These images are also in Figure S3 of Balk et al.</p> <p>minnow.filtered.from.iqm.csv: Filtered fish image data set after filtering (see methods in Balk et al. for filter categories).</p> <p>burress.minnow.sp.filtered.from.iqm.csv: Fish image data set after filtering and selecting species from&nbsp;Burress et al. 2017.</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Associated data from: An end-to-end workflow to study newly synthesized mRNA following rapid protein depletion in Saccharomyces cerevisiae

<p>This dataset includes two custom BED files described in "An end-to-end workflow to study newly synthesized mRNA following rapid protein depletion in&nbsp;<em>Saccharomyces cerevisiae</em>" (Ridenour and Donczew, submitted), which were used to define counting windows for processing SLAM-seq data in SLAM-DUNK (version 0.4.3) [1]. The BED files contain all annotated open reading frames (ORFs) in the<em> Saccharomyces cerevisiae</em> genome or the <em>Schizosaccharomyces</em><em>&nbsp;pombe</em> genome and were created using BEDOPS (version 2.4.3) [2]. All ORFs were then extended 250 bp beyond their stop position to capture 3&prime; untranslated regions (UTRs) using SAMtools (version 1.14) [3] and BEDTools (version 2.30.0) [4]. The reference genome annotations for <em>S. cerevisiae</em> strain S288C (version R64-3-1, RefSeq Assembly GCF_000146045.2) and <em>S. pombe</em> strain 972h- (version ASM294v2, RefSeq Assembly GCF_000002945.1) were retrieved from the NCBI Datasets repository. The <em>S. cerevisiae </em>chromosome names were modified to reflect standard nomenclature (https://www.yeastgenome.org/).</p>

opencc-by-4.0Feb 2024View details →
zenodo48/100

Data and Workflow to: Three-dimensional buoyant hydraulic fracture growth: constant release from a point source (Möri and Lecampion, (2022))

<p>This upload contains the relevant scripts, notebooks, and datasets to reproduce the numerically obtained results of the Journal article &quot;Three-dimensional buoyant hydraulic fracture growth: constant release from a point source&quot; by M&ouml;ri and Lecampion, (2022).</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Test data for running snakePipes : ATAC-seq workflow

<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a>&nbsp;for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the ATAC-seq workflow under snakePipes. To test the workflow, follow the following steps :&nbsp;</p> <ul> <li>Download or prepare genome fasta, indices and annotations for fruit fly (<strong>dm6</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a>&nbsp;with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo44/100

Test data for running snakePipes : mRNA-seq workflow

<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a>&nbsp;for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the mRNA-seq workflow under snakePipes. To test the workflow, follow the following steps :&nbsp;</p> <ul> <li>Download or prepare genome fasta, indices and annotations for mouse (<strong>GRCm38</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a>&nbsp;with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo44/100

Data and code to perform the"Target deformation" workflow in R: virtual reconstruction of the Equus stenonis holotype skulll

<p>Data and code to perform the&quot;Target deformation&quot; workflow in R: virtual reconstruction of the Equus stenonis holotype skulll.</p> <p>TargetDeformation.R: R code with for the Target Deformation procedure.<br> IGF560.ply: 3D mesh of the holotype IGF560 in ply extension.<br> IGF560_set.txt: landmark set of the holotype IGF560.<br> Dm. 5/154.3/4.A4.5.ply: 3D mesh of Dm 5/154.3/4.A4.5 in .ply extension.<br> Dm_set.txt: landmark set of the Dm 5/154.3/4.A4.5 sample.<br> IGF11023: 3D mesh of IGF11023 in.ply extension.<br> IGF11023_set.txt: landmark set on the IGF11023 sample.<br> IGF560R: 3D mesh of IGF560R in.ply extension.<br> IGF560W: 3D mesh of IGF560W in.ply extension.<br> IGF560R-s: 3D mesh of IGF560R-s in.ply extension.<br> IGF560W-s: 3D mesh of IGF560W-s in.ply extension.<br> IGF560_IGF560R_IGF560W.html: file that contain WebGL code to reproduce the 3D meshes of IGF560, IGF560R and IGF560W in a browser.<br> IGF560Rs_IGF560Ws.html: file that contain WebGL code to reproduce the 3D meshes of IGF560R-S and IGF560W-S in a browser.<br> IGF560W Mesh area variation.html: file that contain WebGL code to reproduce two 3d meshes of IGF560W using localmeshDist() and meshdist() in a browser.<br> &nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Coordinates and checklists of alien species populations as obtained from the DASCO workflow and the SInAS data set

<p>This data set contains coordinate records of alien (i.e., non-native) species populations worldwide and aggregated checklists of alien species for individual regions. The regions consists of non-overlapping polygons representing countries, sub-national or coastal marine ecoregions.&nbsp;</p><p>The data set was produced by applying the DASCO workflow (https://doi.org/10.5281/zenodo.5841930) using the SInAS database (version 2.5; https://doi.org/10.5281/zenodo.10038256). The workflow imports checklists of alien species such as those stored in SInAS, and extracts coordinates for the alien regions (according to SInAS) from GBIF and OBIS. After cleaning and thinning the coordinates, the workflow exports a list of coordinates of alien populations for all species included in SInAS and with records on GBIF or OBIS.</p><p>These files are part of a manuscript published in the journal Neobiota, where the workflow is described in detail (Seebens &amp; Kaplan 2022, https://doi.org/10.3897/neobiota.74.81082).</p><p>DASCO_AlienCoordinates_SInAS_2.5.gz contains the coordinates of alien populations.</p><p>DASCO_AlienRegions_SInAS_2.5.csv contains the checklists of alien species per region. Note that this only includes species with GBIF and OBIS records. For more comprehensive checklists, other databases such as those listed here (https://doi.org/10.5281/zenodo.10038256) should be consulted.</p><p>OBIS_SpeciesKeys_SInAS_2.5.csv contains the species keys from OBIS.</p><p>GBIF_SpeciesKeys_SInAS_2.5.csv contains the species keys from GBIF.</p><p>DASCO_TaxonHabitats_SInAS_2.5.csv contains habitat information for individual species if available from WoRMS, Fishbase or Sealifebase (used to identify marine species).</p><p>The file DASCO_ListOriginalGBIFData_keys_SInAS_2.5.csv contains the DOIs of the originally downloaded files from GBIF, which provides the basis for the generation of the GBIF part (ie. the DASCO workflow was applied to these data sets from GBIF). Note that OBIS does not provide a DOI for downloads, and thus we cannot provide this.</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Supplemental Data from the article "The SmARTR pipeline: a modular workflow for the cinematic rendering of 3D scientific imaging data"

<h1><strong>Please, refer to <a href="https://github.com/MeVisLab/SmARTR-Networks">this GitHub repository</a>&nbsp; for additional info, updates, issue reports, and discussion<br></strong></h1> <p><strong>A collection of configuration files (SmARTR networks) &nbsp;published in "<a href="https://doi.org/10.1016/j.isci.2024.111475">The SmARTR Pipeline: a modular workflow for the cinematic rendering of 3D scientific imaging data</a>", enabling the&nbsp; creation of cinematic (photorealistic) renderings of 3D data in the FREE software <a href="https://www.mevislab.de/download">MeVisLab</a><br></strong></p> <ul> <li>Each folder in the archive contains one or more SmARTR network files, the scan and mask files required for the practical examples detailed in the <a href="https://www.cell.com/cms/10.1016/j.isci.2024.111475/attachment/8d79036b-acb6-4cda-a5ff-f56317691ebc/mmc1.pdf">Supplemental&nbsp; Data</a> of the article,&nbsp; and an additional folder with LUT presets.</li> </ul>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Data for Publication: "Automated Investigation of Metal-Ligand Interactions by a Newly Established Robotic Workflow for Titrations"

<p>This dataset contains the whole primary and raw (original) data for the manuscript "Automated investigation of metal-ligand interactions by a newly established robotic workflow for titrations".</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

A High-Performance Data Processing Workflow to Incorporate Effect-Directed Analysis in Suspect and Nontarget Screening [Feature Tables]

<p>This repository is supplementary to&nbsp;the manuscript &quot;High-Performance Data Processing Workflow Incorporating Effect-Directed Analysis for Feature Prioritization in Suspect and Nontarget Screening&quot; (DOI: 10.1021/acs.est.1c04168)&nbsp;and&nbsp;includes an overview of all measured chemical features and annotations in a&nbsp;waste water treatment plant (WWTP)&nbsp;effluent, dust standard reference material (SRM) 2585 and fetal calf serum (FCS) sample.</p> <p>Samples were measured using liquid chromatography - high resolution mass spectrometry (LC-HRMS)&nbsp;and fractionated into 80 micro-fractions encompassing a couple of&nbsp;seconds from the chromatographic run. The fractions were tested for their bioactivity in the antibiotics and the TTR-binding assay. The samples were processed separately&nbsp;using one, two, and three technical replicates in positive and negative ion mode. The first excel sheet includes all measured chemical features, suspect screening annotation, and corresponding bioassay responses. The second sheet includes all possible isomer&nbsp;annotations from the CECscreen database (DOI:&nbsp;<a href="https://doi.org/10.5281/zenodo.3956586">10.5281/zenodo.3956586</a>) for the annotated features.&nbsp;&nbsp;&nbsp;</p>

opencc-by-4.0May 2021View details →
zenodo44/100

ESCALATOR - Stakeholder map data workflow

<p>The stakeholder map project aims to collect and share data on Digital Humanities (DH), Computational Social Sciences (CSS) and related activities and initiatives in South Africa. This data includes information about South African researchers, projects, publications, tools, datasets, academic programmes, training events, learning materials, and more. The aim is to provide deeper insight into the breadth of activities in this area, facilitate enhanced networking and collaboration, and support the optimal use of resources. The stakeholder map will, for example, support researchers looking for collaborators, help potential students to identify undergraduate and postgraduate training programmes, and highlight gaps and opportunities to funders and institutions.</p> <p>The initial design of the data pipeline and workflow for data visualisation has been completed. The pipeline is primarily based on open-source software and platforms often used in the open science community. Development is currently under way. Data will be captured via Google Forms and manipulated using R scripts, available on GitHub and archived in Zenodo. Interactive visualisations will be published on the ESCALATOR website. These visualisations include a [Shiny app](https://shiny.rstudio.com/) that will allow the community to explore data through a web interface and a [Kumu network visualisation](https://kumu.io/). Research articles can be added to an [open collection in Zotero](https://www.zotero.org/groups/3866799/dhcssza) to facilitate easy access to publications from the South African community.</p> <p><br> This diagramme shows the high-level workflow. We anticipate the diagramme will be updated as design and development progresses to incorporate lessons learned and feedback from the community.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

A Bioconductor workflow for processing, evaluating and interpreting expression proteomics data

<p>Files for users of the workflow "A Bioconductor workflow for processing, evaluating and interpreting expression proteomics data". Files include Proteome Discoverer (v2.5) processing and consensus workflows for both TMT and LFQ expression proteomics data. Also provided are the output .txt files of a corresponding Proteome Discoverer identification search, as required for users to follow the workflow themselves. For raw data please refer to PRIDE. Appendix is provided as a PDF.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Test data for running snakePipes : scRNA-seq workflow

<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a>&nbsp;for further information on snakePipes.</p> <p>This folder contains test files that can be used to run scRNA-seq workflow under snakePipes. To test the workflow, follow the following steps :&nbsp;</p> <ul> <li>Download or prepare genome fasta, indices and annotations for mouse (<strong>mm10</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a>&nbsp;with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo40/100

Output reports and supplementary data for MTB workflow

<p>The archive contains the following data:</p> <ul> <li>Output of the workflow on all validation samples (tabular summaries and full HTML reports).</li> <li>Database with AMR regions and mutations.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Data and code for the publication of a surge-specific DEM workflow on ASTER DEMs.

<p>This repository is associated to the publication submitted with the title "Glacier surge monitoring from temporally dense elevation time series: application to an ASTER dataset over the Karakoram region".<br>It contains elevation change maps produced by the workflow, the Python script of the workflow, and vector outlines used in the study.</p> <p>_______________________________<br>Content of the data repository:</p> <p>1) Elevation change maps (raster <em>**.tif</em> files)<br>&nbsp; &nbsp; Regional maps of the elevation changes over 3 years periods, interpolated results in metre (m).<br>&nbsp; &nbsp; Name: <em>dh_[date1]_[date2].tif</em><br>&nbsp; &nbsp;&nbsp;<br>2) Surge-affected areas (vector <em>**.gpkg </em>files)<br>&nbsp; &nbsp; Surge-affected areas drawn manually of four selected glacier surges, divided into reservoir and receiving areas.<br>&nbsp; &nbsp;&nbsp;<br>3) Python script <em>dem_processing_publi.py</em><br>&nbsp; &nbsp; Script with the implementation of the workflow presented in this study.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Eddy covariance data processing workflow example utilizing openeddy and REddyProc R packages

<p>The example dataset is provided within the folder structure required by the workflow files (version 2025-04-27; amended on 2025-07-31) related to the R package openeddy version 0.0.0.9009. Only files needed for successful processing are included. It is shared here as part of a data processing example at <a href="https://github.com/lsigut/EC_workflow">https://github.com/lsigut/EC_workflow</a> to overcome the file size limitation of GitHub.</p>

opencc-by-4.0Oct 2018View details →
zenodo40/100

Test data for jga-analysis per-sample workflow

<p>Test data for jga-analysis per-sample workflow.</p> <p>Please see:</p> <p>-&nbsp;<a href="https://github.com/biosciencedbc/jga-analysis">https://github.com/biosciencedbc/jga-analysis</a></p> <p>-&nbsp;<a href="https://github.com/biosciencedbc/jga-analysis/blob/main/per-sample/Workflows/per-sample.cwl">https://github.com/biosciencedbc/jga-analysis/blob/main/per-sample/Workflows/per-sample.cwl</a></p>

openapache2.0May 2022View details →
zenodo40/100

Data for FEgrow: An Open-Source Molecular Builder and Free Energy Preparation Workflow

<p>Data illustrating the use of de novo design in building and scoring protein-ligand complexes.</p> <p>This is relationship to the FEgrow publication with the intiial preprint here:&nbsp;<br> https://chemrxiv.org/engage/chemrxiv/article-details/6287bb98a42e9c78d34769f6<br> &nbsp;</p> <p>The FEgrow software snapshot used can be found here:&nbsp;https://zenodo.org/record/7105647#.YzFwINLMIUE</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Data sets for the Simulated AMPI (SAMPI) load balancing simulation workflow and Ondes3D performance analysis (Companion to CCPE - Euro-Par 2017 special issue)

<p>This package contains data sets and scripts (in&nbsp;an Org-mode file) related to our submission to the special Euro-Par 2017 issue of the&nbsp;&nbsp;journal &quot;Concurrency and Computation: Practice and Experience&quot;, under the title&nbsp;&quot;Performance Modeling of a Geophysics Application to Accelerate Over-decomposition Parameter Tuning through Simulation&quot;.</p>

opencc-by-sa-4.0Nov 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record