Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
650
datasets available to search
ShareScore release 0.9.0
Dataset results
650 results for “Workflow”
GeneSeqToFamily: a Galaxy workflow to find gene families based on the Ensembl Compara GeneTrees Pipeline.
<p>Gene duplication is a major factor contributing to evolutionary novelty, and the contraction or expansion of gene families has often been associated with morphological, physiological and environmental adaptations. The study of homologous genes helps us to understand the evolution of gene families. It plays a vital role in finding ancestral gene duplication events as well as identifying genes that have diverged from a common ancestor under positive selection. There are various tools available, such as MSOAR, OrthoMCL and HomoloGene, to identify gene families and visualise syntenic information between species, providing an overview of syntenic regions evolution at the family level. Unfortunately, none of them provide information about structural changes within genes, such as the conservation of ancestral exon boundaries amongst multiple genomes. The Ensembl GeneTrees computational pipeline generates gene trees based on coding sequences and provides details about exon conservation, and is used in the Ensembl Compara project to discover gene families. </p>
BlockClust workflow-testing data
<p>Data is taken from https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM450239.</p>
Digital Archaeological Workflow
<p>The archaeological workflow in terms of a computational pipeline from data acquisition to unpublished data, and re-use. Problems of quality impact data in each stage. Black boxes in this workflow occur wherever archaeologists employ software and tools whose code are unavailable to review and modify and that do not enable documentation of transformations</p>
Computational Artifacts for Performance Feedback Autoscaling Experiments with Workloads of Workflows in Apache Airflow
<p>These computational artifacts are related to the software artifacts DOI:10.5281/zenodo.2635571</p> <p><strong>The content of the computational artifacts:</strong></p> <ul> <li><strong>experiments.pdf</strong> contains the list of all the conducted experiments with the Airflow system. Experiment IDs are not sequential since some experiments required rerunning, etc., we report only successful results. The file lists different experiment configurations, e.g., the number of processed workflows, the name of the used workload, the user budgets, and PFA settings.</li> <li><strong>db.tar.gz</strong> contains directories with Airflow database snapshots and autoscaler logs. The names of the directories correspond to those listed in `experiments.pdf`. Each experiment directory contains an autoscaler log and a full copy of a PostgreSQL database directory just after each experiment finished. The database name is `airflow`, the user name is `ailyushk`. Within each database, most of the paper-related data are stored in the `stat_log` table. The scripts for extracting data from these databases are available as software artifacts in `tools/analysis`.</li> <li><strong>gurobi.tar.gz</strong> contains the results obtained from the Gurobi solver when solving the MIP model.</li> <li><strong>pdf.tar.gz</strong> contains all the figures in pdf format, also those that were not included neither in the paper nor in the technical report. The scripts for creating this plots are delivered as software artifacts.</li> <li><strong>csv.tar.gz</strong> contains the analysis results extracted from Airflow database snapshots. These files are used to create the plots in the `pdf` directory. The scripts for doing this are delivered as software artifacts.</li> <li><strong>wl1.tar.gz</strong> is the first synthetic realistic workload (WL I) with three subsets of 200 workflows each (`1_0`, `1_1`, `1_2`). Each directory contains the file with interarrivals `interarrivals.txt`, and the file with workflow IDs `workload.txt` in the subset. The `dags` directory contains Python-based Airflow descriptors and CSV files that summarise the same descriptors in CSV format for simpler analysis. The scripts for extracting workload statistics from these CSV files are available in the software artifacts in `tools/analysis`. The `inputs` directory contains initial input files for each worfklow. The `dax` contains original DAX files obtained from the generator: <a href="https://github.com/pegasus-isi/WorkflowGenerator/tree/master/bharathi/src/simulation/generator">https://github.com/pegasus-isi/WorkflowGenerator/tree/master/bharathi/src/simulation/generator</a></li> <li><strong>wl2.tar.gz</strong> is the second synthetic realistic workload (WL II) with three subsets of 200 workflows each (`4_0`, `4_1`, `4_2`). Has similar structure as `wl1.tar.gz`, except that `dax` directory is omitted, as WL II uses the same DAX structures as WL I.</li> <li><strong>wl3.tar.gz</strong> is the small synthetic workload based on WL I for the experiment with the MIP solver, contains three subsets with 5 workflows in each, all in the `3_0` directory (thus, the structure differs from the WL I and WL II). The input data files are empty. The identifiers of workflows forming each subset are stored in the `workload_1.txt`, `workload_2.txt`, and `workload_3.txt` files.</li> </ul>
Workflow Trace Archive askalon_ee2 trace
Trace description unavailable.
Workflow Trace Archive Pegasus_P5 trace
Trace description unavailable.
Workflow Trace Archive workflowhub_epigenomics_dataset-hep_grid5000_schema-0-2_epigenomics-hep-g5k-run001 trace
Workload downloaded from WorkflowHub, see http://workflowhub.isi.edu/.
Workflow Trace Archive workflowhub_epigenomics_dataset-ilmn_chameleon-cloud_schema-0-2_epigenomics-ilmn-100000-cc-run004 trace
Workload downloaded from WorkflowHub, see http://workflowhub.isi.edu/.
Workflow Trace Archive Pegasus_P8 trace
Trace description unavailable.
Workflow Trace Archive Pegasus_P1 trace
Trace description unavailable.
Workflow Trace Archive workflowhub_soykb_grid5000_schema-0-2_soykb-g5k-run002 trace
Workload downloaded from WorkflowHub, see http://workflowhub.isi.edu/.
Workflow Trace Archive Pegasus_P6b trace
Trace description unavailable.
Workflow Trace Archive workflowhub_montage_ti01-971107n_degree-2-0_osg_schema-0-2_montage-2-0-osg-run007 trace
Workload downloaded from WorkflowHub, see http://workflowhub.isi.edu/.
Workflow Trace Archive Pegasus_P6a trace
Trace description unavailable.
Workflow Trace Archive Pegasus_P3 trace
Trace description unavailable.
Workflow Trace Archive Pegasus_P7 trace
Trace description unavailable.
Workflow Trace Archive workflowhub_epigenomics_dataset-hep_futuregrid_schema-0-2_epigenomics-hep-fg-run001 trace
Workload downloaded from WorkflowHub, see http://workflowhub.isi.edu/.
Workflow Trace Archive workflowhub_epigenomics_dataset-hep_chameleon-cloud_schema-0-2_epigenomics-hep-100000-cc-run005 trace
Workload downloaded from WorkflowHub, see http://workflowhub.isi.edu/.
Workflow Trace Archive workflowhub_epigenomics_dataset-taq_chameleon-cloud_schema-0-2_epigenomics-taq-100000-cc-run002 trace
Workload downloaded from WorkflowHub, see http://workflowhub.isi.edu/.
Workflow Trace Archive askalon_ee trace
Trace description unavailable.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.