Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

650

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

650 results for “workflow”

Learn how ShareScore rates datasets ↗
zenodo52/100

Input data for MFAssignR Galaxy workflow tutorial

<p>This is the input dataset for the MFAssignR Galaxy training workflow. The input dataset corresponds to the model data of MFAssignR (<a title="Raw_Neg_ML" href="https://github.com/skschum/MFAssignR/tree/master/MFAssignR/data" target="_blank" rel="noopener">Raw_Neg_ML</a>), containing a raw mass list, measured in a negative ESI mode.</p>

openmit-licenseSep 2024View details →
zenodo48/100

Multi-faceted analyses of Poland's Bronze and Early Iron Age hoards: Fig.3. Workflow in the Biography of Hoards project

<p>The set contains a figure and editable files associated with the figure.</p> <p>Figure presenting workflow of the project described in the related paper.</p> <p>The paper and data were prepared as part of a project funded by the National Science Centre, Poland: <em>A Biography of Late Bronze and Early Iron Ages Hoards. A Multi-Faceted Analysis of Metal Objects Related to Monumental Constructions in Poland</em> (UMO-2021/41/B/HS3/00038)</p>

opencc-zeroSep 2023View details →
zenodo48/100

Data from: A FAIR and modular image-based workflow for knowledge discovery in the emerging field of imageomics

<p>Data and results from the Imageomics Workflow. These include data files from the Fish-AIR repository (https://fishair.org/) for purposes of reproducibility and outputs from the application-specific imageomics workflow contained in the Minnow_Segmented_Traits repository (https://github.com/hdr-bgnn/Minnow_Segmented_Traits).</p> <p>Fish-AIR:<br> This is the dataset downloaded from&nbsp;Fish-AIR, filtering for Cyprinidae and the Great Lakes Invasive Network (GLIN) from the Illinois Natural History Survey (INHS) dataset. These files contain information about fish images, fish image quality, and path for downloading the images. The data download ARK ID is dtspz368c00q. (2023-04-05). The following files are unaltered from the Fish-AIR download. We use the following files:</p> <p>extendedImageMetadata.csv: A CSV file containing information about each image file. It has the following columns: ARKID, fileNameAsDelivered, format, createDate, metadataDate, size, width, height, license, publisher, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>imageQualityMetadata.csv: A CSV file containing information about the quality of each image. It has the following columns: ARKID, license, publisher, ownerInstitutionCode, createDate, metadataDate, specimenQuantity, containsScaleBar, containsLabel, accessionNumberValidity, containsBarcode, containsColorBar, nonSpecimenObjects, partsOverlapping, specimenAngle, specimenView, specimenCurved, partsMissing, allPartsVisible, partsFolded, brightness,&nbsp;<br> uniformBackground, onFocus, colorIssue, quality, resourceCreationTechnique. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>multimedia.csv: A CSV file containing information about image downloads. It has the following columns: ARKID, parentARKID, accessURI, createDate, modifyDate, fileNameAsDelivered, format, scientificName, genus, family, batchARKID, batchName, license, source, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>meta.xml: A XML file with the metadata about the column indices and URIs for each file contained in the original downloaded zip file. This file is used in the fish-air.R script to extract the indices for column headers.</p> <p>The outputs from the Minnow_Segmented_Traits workflow are:</p> <p>sampling.df.seg.csv: Table with tallies of the sampling of image data per species during the data cleaning and data analysis. This is used in Table S1 in Balk et al.&nbsp;</p> <p>presence.absence.matrix.csv: The Presence-Absence matrix from segmentation, not cleaned. This is the result of the combined outputs from the presence.json files created by the rule &ldquo;create_morphological_analysis&rdquo;. The cleaned version of this matrix is shown as Table S3 in Balk et al.</p> <p>heatmap.avg.blob.png and heatmap.sd.blob.png: Heatmaps of average area of biggest blob per trait (heatmap.avg.blob.png) and standard deviation of area of biggest blob per trait (heatmap.sd.blob.png). These images are also in Figure S3 of Balk et al.</p> <p>minnow.filtered.from.iqm.csv: Filtered fish image data set after filtering (see methods in Balk et al. for filter categories).</p> <p>burress.minnow.sp.filtered.from.iqm.csv: Fish image data set after filtering and selecting species from&nbsp;Burress et al. 2017.</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Associated data from: An end-to-end workflow to study newly synthesized mRNA following rapid protein depletion in Saccharomyces cerevisiae

<p>This dataset includes two custom BED files described in "An end-to-end workflow to study newly synthesized mRNA following rapid protein depletion in&nbsp;<em>Saccharomyces cerevisiae</em>" (Ridenour and Donczew, submitted), which were used to define counting windows for processing SLAM-seq data in SLAM-DUNK (version 0.4.3) [1]. The BED files contain all annotated open reading frames (ORFs) in the<em> Saccharomyces cerevisiae</em> genome or the <em>Schizosaccharomyces</em><em>&nbsp;pombe</em> genome and were created using BEDOPS (version 2.4.3) [2]. All ORFs were then extended 250 bp beyond their stop position to capture 3&prime; untranslated regions (UTRs) using SAMtools (version 1.14) [3] and BEDTools (version 2.30.0) [4]. The reference genome annotations for <em>S. cerevisiae</em> strain S288C (version R64-3-1, RefSeq Assembly GCF_000146045.2) and <em>S. pombe</em> strain 972h- (version ASM294v2, RefSeq Assembly GCF_000002945.1) were retrieved from the NCBI Datasets repository. The <em>S. cerevisiae </em>chromosome names were modified to reflect standard nomenclature (https://www.yeastgenome.org/).</p>

opencc-by-4.0Feb 2024View details →
zenodo48/100

Data and Workflow to: Three-dimensional buoyant hydraulic fracture growth: constant release from a point source (Möri and Lecampion, (2022))

<p>This upload contains the relevant scripts, notebooks, and datasets to reproduce the numerically obtained results of the Journal article &quot;Three-dimensional buoyant hydraulic fracture growth: constant release from a point source&quot; by M&ouml;ri and Lecampion, (2022).</p>

opencc-by-4.0May 2022View details →
zenodo48/100

GTN_PAR-CLIP_workflow

<p>Data from&nbsp;https://trace.ncbi.nlm.nih.gov/Traces/sra/sra.cgi?exp=SRX105188&amp;cmd=search&amp;m=downloads&amp;s=seq and ftp://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/000/002/985/GCA_000002985.3_WBcel235</p> <p>For Galaxy Training https://rna.usegalaxy.eu/workflows/run?id=a108b575b16e6cb9</p>

opencc-by-4.0Jan 2019View details →
zenodo48/100

Dataset for Training Material - Galaxy Workflow - Analyse unaligned ncRNAs

<p>Input dataset for Galaxy Training Material for the Analyze unaligned ncRNAs workflow.</p> <p>See https://github.com/galaxyproject/training-material for more information.</p>

opencc-by-4.0Oct 2019View details →
zenodo48/100

Workflow for detecting biomedical articles with openly available underlying datasets - Datasets and extraction forms

<p>The open data screening datasets contain both automatically detected (TRUE) Open Data statements by <a href="https://github.com/quest-bih/oddpub">ODDPub</a>, and its manual validation using <a href="https://github.com/bgcarlisle/Numbat">Numbat</a> extraction tool. Furthermore, extraction forms for both screenings &ndash; 2020 and 2021 &ndash; are included. The manually processed dataset for the calculation of the inter-rater reliability of manual validation can be also found here.&nbsp;&nbsp;</p> <p>(i) Data from articles published in 2020 (file &lsquo;<em>charite_open_data_2020.csv</em>&rsquo;) have been collected applying a slightly different sequence of questions in the extraction workflow than the articles published in 2021 (file &lsquo;<em>charite_open_data_2021.csv</em>&rsquo;). Both datasets were cleaned for any personal data or internal comments. Thus, they do not contain the default columns which in the raw export from Numbat contained commentaries regarding different question. Also, in another regard these files do not represent raw outputs of the Numbat extraction tool, but a processed version. This means that articles validated by more than two raters were first reconciled in Numbat, resulting in one final decision (output of extractions <strong>after reconciliation</strong>). Then from the output of extractions <strong>before reconciliation</strong> those articles validated by only 1 rater (and thus not part of the inter-rater reliability calculation) were selected, which were afterwards joined with the already reconciled dataset.&nbsp;&nbsp;</p> <p>The actual decision about Openness of validated dataset can be analysed in various ways:&nbsp;</p> <ol> <li>Column &lsquo;<em>open_data_assessment</em>&rsquo;/&rsquo;<em>assessment</em>&rsquo; shows a binary decision between Open Data TRUE and FALSE.&nbsp;</li> <li>If that column indicates &lsquo;<em>NULL</em>&rsquo;, the dataset was classified into &lsquo;non&rsquo;-open category, and the result can be found on one of the following ways:&nbsp; <ul> <li>Column &lsquo;<em>reference_to_data</em>&rsquo; as &lsquo;<em>n_a</em>&rsquo; for excluded articles, e.g. not producing any data.</li> <li>Column &lsquo;<em>data_access</em>&rsquo; as &lsquo;<em>restricted</em>&rsquo;.&nbsp;</li> <li>Column &lsquo;<em>own_or_reuse_data</em>&rsquo; as &lsquo;<em>open_data_reuse</em>&rsquo;.&nbsp;</li> </ul> </li> </ol> <p>The original extraction form contains an option &lsquo;unsure_open_data&rsquo; besides &lsquo;<em>open_data</em>&rsquo;/&rsquo;<em>no_open_data</em>&rsquo; which was resolved either during reconciliation between multiple raters or by case-related consultation with a second rater in case of doubt, and is not included here.&nbsp;</p> <p>(ii) The inter-rater reliability calculation was made on randomly selected 100 articles for 2 raters. The third rater screened 20 articles sample, which is part of 100 sample. The tables provided here include both article-level data, and dataset-level data.&nbsp;</p> <p>(iii) The Numbat extarction forms used for the screenings in 2020 and 2021 are included in two formats - JSON and Markdown.</p> <p>(iv) &lsquo;<em>data_dictionary_open_data.csv</em>&rsquo; table documents all variables of each data file containing here.&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Test data for running snakePipes : ATAC-seq workflow

<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a>&nbsp;for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the ATAC-seq workflow under snakePipes. To test the workflow, follow the following steps :&nbsp;</p> <ul> <li>Download or prepare genome fasta, indices and annotations for fruit fly (<strong>dm6</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a>&nbsp;with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo44/100

Test data for running snakePipes : mRNA-seq workflow

<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a>&nbsp;for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the mRNA-seq workflow under snakePipes. To test the workflow, follow the following steps :&nbsp;</p> <ul> <li>Download or prepare genome fasta, indices and annotations for mouse (<strong>GRCm38</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a>&nbsp;with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo44/100

Digitaler Workflow in der Zusammenarbeit mit Externen beim Leihverkehr

<p>Abb. 4: Digitaler Workflow in der Zusammenarbeit mit Externen beim Leihverkehr</p> <p>from Gasser, Sonja. &bdquo;Das digital transformierte Museum.&ldquo; Online-Erweiterung zur Museumskunde Band 84/2019.</p> <p>Find all figures from the article:<br> Abb. 1: <a href="https://doi.org/10.5281/zenodo.3752544">10.5281/zenodo.3752544</a><br> Abb. 2: <a href="https://doi.org/10.5281/zenodo.3752586">10.5281/zenodo.3752586</a><br> Abb. 3: <a href="https://doi.org/10.5281/zenodo.3752590">10.5281/zenodo.3752590</a><br> Abb. 4: <a href="https://doi.org/10.5281/zenodo.3752597">10.5281/zenodo.3752597</a></p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Data and code to perform the"Target deformation" workflow in R: virtual reconstruction of the Equus stenonis holotype skulll

<p>Data and code to perform the&quot;Target deformation&quot; workflow in R: virtual reconstruction of the Equus stenonis holotype skulll.</p> <p>TargetDeformation.R: R code with for the Target Deformation procedure.<br> IGF560.ply: 3D mesh of the holotype IGF560 in ply extension.<br> IGF560_set.txt: landmark set of the holotype IGF560.<br> Dm. 5/154.3/4.A4.5.ply: 3D mesh of Dm 5/154.3/4.A4.5 in .ply extension.<br> Dm_set.txt: landmark set of the Dm 5/154.3/4.A4.5 sample.<br> IGF11023: 3D mesh of IGF11023 in.ply extension.<br> IGF11023_set.txt: landmark set on the IGF11023 sample.<br> IGF560R: 3D mesh of IGF560R in.ply extension.<br> IGF560W: 3D mesh of IGF560W in.ply extension.<br> IGF560R-s: 3D mesh of IGF560R-s in.ply extension.<br> IGF560W-s: 3D mesh of IGF560W-s in.ply extension.<br> IGF560_IGF560R_IGF560W.html: file that contain WebGL code to reproduce the 3D meshes of IGF560, IGF560R and IGF560W in a browser.<br> IGF560Rs_IGF560Ws.html: file that contain WebGL code to reproduce the 3D meshes of IGF560R-S and IGF560W-S in a browser.<br> IGF560W Mesh area variation.html: file that contain WebGL code to reproduce two 3d meshes of IGF560W using localmeshDist() and meshdist() in a browser.<br> &nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Rbbt SINTEF workflow bliss and hsa task results

<p>The two datasets&nbsp;have the results of <strong>rbbt&#39;s</strong>&nbsp;<strong>SINTEF</strong>&nbsp;workflow tasks <strong>bliss</strong> and <strong>hsa </strong>and the rbbt docker image used for the analysis.</p> <p>See the analysis results&nbsp;of the SINTEF dataset (Flobak et. al 2019) in:&nbsp;<a href="https://druglogics.github.io/sintef-obs-synergies/">https://druglogics.github.io/sintef-obs-synergies/</a></p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

PROGRAMS project. PRM workflows from IDEKO S. Coop. on 2020-09 sample 1

<p>These data is the PRM workflow elaboration of the ones collected from FIDIA machine tool controller during milling operation.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

PROGRAMS project. PRM workflows from IDEKO S. Coop. on 2020-09 sample 3

<p>These data is the PRM workflow elaboration of the ones collected from FIDIA machine tool controller during milling operation.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

PROGRAMS project. PRM workflows from IDEKO S. Coop. on 2020-09 sample 2

<p>These data is the PRM workflow elaboration of the ones collected from FIDIA machine tool controller during milling operation.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Supplementary Datasets for dadasnake workflow

<p>This dataset contains configuration and results files for the proof-of-principle of the dadasnake pipeline. Includes dadasnake output and tables with the composition of ground-truth data or mock-communities.</p> <p>dadasnake is a user-friendly, one-command Snakemake pipeline that wraps the pre-processing of sequencing reads and the delineation of exact sequence variants by using the favorably benchmarked and widely-used DADA2 algorithm with a taxonomic classification and the post-processing of the resultant tables, including hand-off in standard formats. The suitability of the provided default configurations is demonstrated using mock-community data from bacteria and archaea, as well as fungi. By use of Snakemake, dadasnake makes efficient use of high-performance computing infrastructures. Easy user configuration guarantees flexibility of all steps, including the processing of data from multiple sequencing platforms. dadasnake facilitates easy installation via conda environments. dadasnake is available at <a href="https://github.com/a-h-b/dadasnake">https://github.com/a-h-b/dadasnake</a> .</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Coordinates and checklists of alien species populations as obtained from the DASCO workflow and the SInAS data set

<p>This data set contains coordinate records of alien (i.e., non-native) species populations worldwide and aggregated checklists of alien species for individual regions. The regions consists of non-overlapping polygons representing countries, sub-national or coastal marine ecoregions.&nbsp;</p><p>The data set was produced by applying the DASCO workflow (https://doi.org/10.5281/zenodo.5841930) using the SInAS database (version 2.5; https://doi.org/10.5281/zenodo.10038256). The workflow imports checklists of alien species such as those stored in SInAS, and extracts coordinates for the alien regions (according to SInAS) from GBIF and OBIS. After cleaning and thinning the coordinates, the workflow exports a list of coordinates of alien populations for all species included in SInAS and with records on GBIF or OBIS.</p><p>These files are part of a manuscript published in the journal Neobiota, where the workflow is described in detail (Seebens &amp; Kaplan 2022, https://doi.org/10.3897/neobiota.74.81082).</p><p>DASCO_AlienCoordinates_SInAS_2.5.gz contains the coordinates of alien populations.</p><p>DASCO_AlienRegions_SInAS_2.5.csv contains the checklists of alien species per region. Note that this only includes species with GBIF and OBIS records. For more comprehensive checklists, other databases such as those listed here (https://doi.org/10.5281/zenodo.10038256) should be consulted.</p><p>OBIS_SpeciesKeys_SInAS_2.5.csv contains the species keys from OBIS.</p><p>GBIF_SpeciesKeys_SInAS_2.5.csv contains the species keys from GBIF.</p><p>DASCO_TaxonHabitats_SInAS_2.5.csv contains habitat information for individual species if available from WoRMS, Fishbase or Sealifebase (used to identify marine species).</p><p>The file DASCO_ListOriginalGBIFData_keys_SInAS_2.5.csv contains the DOIs of the originally downloaded files from GBIF, which provides the basis for the generation of the GBIF part (ie. the DASCO workflow was applied to these data sets from GBIF). Note that OBIS does not provide a DOI for downloads, and thus we cannot provide this.</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

RO-Crate created using Autosubmit version 4.0.100 workflow running kinow/auto-mhm-test-domains

<p>This dataset is an <a href="https://www.researchobject.org/ro-crate/">RO-Crate</a> representation of an execution of an example Autosubmit workflow that executes <a href="https://mhm-ufz.org/">mHM</a>, the UFZ mesoscale Hydrologic Model. The example was created with Autosubmit version 4.0.100. The source code of the workflow can be found at <a href="https://github.com/kinow/auto-mhm-test-domains">https://github.com/kinow/auto-mhm-test-domains</a>, commit <a href="https://github.com/kinow/auto-mhm-test-domains/commit/c2672568fb14706c6e66481bc7ceac467db8261e">c267256</a>. This dataset was <a href="https://github.com/ResearchObject/workflow-run-crate/pull/61">validated by the RO-Crate community</a>.</p>

openapache2.0Jul 2023View details →
zenodo44/100

SQUID manuscript workflow with outputs

<p>This repository reproduces the analysis performed in for our method <strong>SQUID</strong> (<strong>S</strong>urrogate <strong>Qu</strong>antitative <strong>I</strong>nterpretability for <strong>D</strong>eepnets; Seitz, McCandlish, Kinney and Koo). It contains tools to apply SQUID on several previously-published genomic models, and compare its results to existing attribution methods.</p><p>The scripts contained in this release are a replica of those found in the current release of our GitHub repository (https://github.com/evanseitz/squid-manuscript as of commit 5b47a7d), with the addition that all intermediate and final outputs are provided here.</p>

openmit-licenseOct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record