Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
146
datasets available to search
ShareScore release 0.9.0
Dataset results
146 results for “data workflow”
Knime workflow and data files to calculate a lead-likeness evaluation function 'Λ' for small molecule compounds
<p class="RSCB01COMAbstract"><span class="06CHeading"><span>A simple and rational method to rank molecules' lead-likeness using continuous evaluation functions was developed to support epigenetic drug discovery. This strategy proved to be highly effective on model chemical libraries and finally helped driving synthetic efforts towards candidates of interest for epigenetic applications towards HDAC6, BRD4 and EZH2.</span></span></p> <p>Here we provide sample data and the knime workflow for generatating lead-like chemical probes for epigenetics. This tool and example data set will enable researchers to establish this workflow in their own laboratories and apply it to new synthetic scaffolds for medicinal chemistry.</p>
DLC networks from: Application of a novel deep learning based 3D videography workflow to bat flight data
Open the record for dataset details and reuse information.
Data from: Detection of the endangered European weather loach (Misgurnus fossilis) via water and sediment samples: testing multiple eDNA workflows.
Open the record for dataset details and reuse information.
Knime workflow and data files to calculate a lead-likeness evaluation function 'Λ' for small molecule compounds
Open the record for dataset details and reuse information.
Data from: Ortho2Web: A workflow for disentangling the roles of hybridization and allopolyploidization in reticulation within Campanulaceae
Open the record for dataset details and reuse information.
Data from: Evaluating culture-free targeted next-generation sequencing for diagnosing drug-resistant tuberculosis: A multicentre clinical study of two end-to-end commercial workflows
Open the record for dataset details and reuse information.
Test data for running snakePipes : WGBS workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the WGBS workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for the mouse (<strong>GRCm38</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>
Test data for running snakePipes : noncoding-RNA-seq workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the noncoding-RNA-seq workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for mouse (<strong>GRCm38</strong>) genome. Repeat Masker file is required for this workflow.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>
Data from: A from-benchtop-to-desktop workflow for validating HTS data and for taxonomic identification in diet metabarcoding studies
The main objective of this work was to develop and validate a robust and reliable 'from benchtop-to-desktop' metabarcoding workflow to investigate the diet of invertebrate-eaters. We applied our workflow to fecal DNA samples of an invertebrate-eating fish species. A fragment of the COI gene was amplified by combining two minibarcoding primer sets to maximize the taxonomic coverage. Amplicons were sequenced by an Illumina MiSeq platform. We developed a filtering approach based on a series of non-arbitrary thresholds established from control samples and from molecular replicates in order to address the elimination of cross-contamination, PCR/sequencing errors and mistagging artifacts. This resulted in a conservative and informative metabarcoding dataset. We developed a taxonomic assignment procedure that combines different approaches and that allowed the identification of ~75% of invertebrate COI variants to the species level. Moreover, based on the diversity of the variants, we introduced a semi-quantitative statistic in our diet study, the Minimum Number of Individuals (MNI), which is based on the number of distinct variants in each sample. The metabarcoding approach described in this paper may guide future diet studies that aim to produce robust datasets associated with a fine and accurate identification of prey items.
Data from: Expanding the described metabolome of the marine cyanobacterium Moorea producens JHB through orthogonal natural products workflows
Moorea producens JHB, a Jamaican strain of tropical filamentous marine cyanobacteria, has been extensively studied by traditional natural products techniques. These previous bioassay and structure guided isolations led to the discovery of two exciting classes of natural products, hectochlorin (1) and jamaicamides A (2) and B (3). In the current study, mass spectrometry-based 'molecular networking' was used to visualize the metabolome of Moorea producens JHB, and both guided and enhanced the isolation workflow, revealing additional metabolites in these compound classes. Further, we developed additional insight into the metabolic capabilities of this strain by genome sequencing analysis, which subsequently led to the isolation of a compound unrelated to the jamaicamide and hectochlorin families. Another approach involved stimulation of the biosynthesis of a minor jamaicamide metabolite by cultivation in modified media, and provided insights about the underlying biosynthetic machinery as well as preliminary structure-activity information within this structure class. This study demonstrated that these orthogonal approaches are complementary and enrich secondary metabolomic coverage even in an extensively studied bacterial strain.
Data Processing for a Small-Scale Long-Term Coastal Ocean Observing System Near Mobile Bay, Alabama: A Geoscience Papers of the Future (GPF) Workflow Diagram
<p>The Dauphin Island Sea Lab (DISL) has been operating a permanent moored oceanographic station at 30 05.410'N, 88 12.694'W, 25 km southwest of the entrance to Mobile Bay, Alabama, since 2004. It collects hydrographic and current velocity data.</p> <p>This diagram shows the processing steps for data from the instruments at this mooring, from initial download to initial scientific analysis. The accompanying text explains how to apply the workflow to the example dataset (10.5281/zenodo.18943) using the provided software (10.5281/zenodo.32741). </p> <p>The files have been prepared as supplementary material for a Geoscience Paper of the Future (GPF) in prep for publication at Earth and Space Science, as part of the OntoSoft GPF Initiative.</p>
Test data for iwc workflow : Purgedups VGP6
Open the record for dataset details and reuse information.
Maximizing Data Utility for HPC Python Workflow Execution
<p>Large-scale HPC workflows are increasingly implemented in dynamic languages such as Python, which allow for more rapid development than traditional techniques. However, the cost of executing Python applications at scale is often dominated by the distribution of common datasets and complex software dependencies. As the application scales up, data distribution becomes a limiting factor that prevents scaling beyond a few hundred nodes. To address this problem, we present the integration of Parsl (a Python-native parallel programming library) with TaskVine (a data-intensive workflow execution engine). Instead of relying on a shared filesystem to provide data to tasks on demand, Parsl is able to express advance data needs to TaskVine, which then performs efficient data distribution at runtime. This combination provides a performance speedup of 1.48x over the typical method of on-demand paging from the shared filesystem, while also providing an average task speedup of 1.79x with 2048 tasks and 256 nodes.</p>
Data from: "Navigating through Complexity: Optimizing Cathodes for Organic Electrohydrogenation through Coherent Workflows"
<p>The data used in "<strong>Navigating through Complexity: Optimizing Cathodes for Organic Electrohydrogenation through Coherent Workflows</strong>". </p>
Supporting data for "The benefit of in silico predicted spectral libraries in data-independent acquisition data analysis workflows"
Open the record for dataset details and reuse information.
Data for article: Robust characterization of forest structure from airborne laser scanning – a systematic assessment and sample workflow for ecologists
<p><strong>### Update 03/02/2025: the most up to date version of the processing pipeline presented here, also working on Linux, is available on github: https://github.com/fischer-fjd/GCA/tree/main, and a worked example with open data from the Dutch AHN surveys is available on Zenodo: https://zenodo.org/records/14722001 ###<br></strong></p> <p>This is a collection of scripts and research data to assess the robustness of forest structure characterization from airborne laser scanning (ALS). It accompanies the article "Robust characterization of forest structure from airborne laser scanning – a systematic assessment and sample workflow for ecologists" (accepted in Methods in Ecology and Evolution on 25/08/2024). </p> <p>In the article, we assess the derivation of canopy height models (CHMs) from point cloud data, how sensitive CHM algorithms are to point cloud degradation (pulse density thinning, large scan angles, loss of higher-order returns) and how uncertainties and biases propagate to commonly used forest structure metrics. In addition, we provide a standardized processing pipeline in R to convert point clouds into CHMs. </p> <p>The main data source for this study are ALS point clouds from nine Australian research sites belonging to the Terrestrial Ecosystem Research Network (TERN, 5 km x 5 km extent each). The underlying data can be found here: https://portal.tern.org.au/metadata/TERN/4ff0b4c9-cfa0-4d09-9520-b5402adc583f. For one site (Robson Creek), we also used field data to assess the sensitivity of aboveground biomass estimates to ALS point cloud characteristics. Data are available here: https://portal.tern.org.au/metadata/supersite.174. </p> <p>To characterize climatic/environmental differences between sites, we used climatic data from the CHELSA/BIOCLIM+ climatology 1981-2010 (Brun et al. 2022: Global climate-related predictors at kilometer resolution for the past and future. Earth System Science Data, 14(12), 5573–5603. https://doi.org/10.5194/essd-14-5573-2022; Karger et al. 2017: Climatologies at high resolution for the earth's land surface areas. Scientific Data, 4(1), 170122. https://doi.org/10.1038/sdata.2017.122). </p> <p>We note that the enormous size of the full set of manipulated point clouds (original + thinned + individual flightlines: ~400 GB) and the derived raster products (~200 GB) by far exceeds limits on data storage in Zenodo. However, all analyses can be recreated from scratch from the openly available data and the R code in this repository. In addition, we include derived products for the nine study sites that allow to replicate results in the main text without any point cloud processing (CHMs and other rasters across thinned point clouds + summary statistics). </p> <p>The different data layers are:</p> <p><strong>01_rscripts.zip:</strong></p> <ul> <li>contains a sample script to test the processing pipeline (<em>test.processing.R</em>) as well as a collection of helper functions (<em>ALS_processing_helperfunctions_v40.R</em>); the script can be run directly after unzipping the folder, but an installation of LAStools (https://rapidlasso.de) is necessary (path_lastools = "PATH/TO/LASTOOLS/BIN"); we note that the script was developed on Windows PCs, its application with the recent Linux distribution of LAStools has not yet been tested</li> <li>contains the full set of scripts necessary to reproduce the analyses, including point cloud manipulations and derivation of CHMs from the raw data (<em>create.CHMs.R)</em> as well as the overall robustness analysis (<em>analyze.CHMs.R</em>); to replicate the processing of the raw point clouds step by step, .laz files should be downloaded from the TERN repository (cf. citation above) and placed in a "data" folder, with subfolders for each site and with the same naming conventions as in this repository (e.g., "/data/Alice Mulga")</li> </ul> <p><strong>02_reference.zip</strong></p> <ul> <li>contains reference digital surface models (DSMs), canopy height models (CHMs) and digital terrain models (DTMs) for all nine TERN sites, based on the original ALS point clouds</li> <li>note that these reference layers are produced with the "CHMhighest" algorithm, which provides an easily interpretable canopy description as long as pulse densities are high (>= 20 shots per squaremetre)</li> </ul> <p><strong>03_climate.zip</strong></p> <ul> <li>contains site coordinates</li> <li>contains the climate layers from the CHELSA climatology (cf. citation above, only used to evaluate climatic ranges of sites)</li> </ul> <p><strong>04_robson_additional.zip</strong></p> <ul> <li>contains biomass estimates for Robson Creek</li> <li>contains shapefiles for large trees at Robson Creek (only used for visualization purposes)</li> </ul> <p><strong>05_downsampling_pulse_[Site name].zip</strong></p> <ul> <li>[Site name] is a stand-in for the nine TERN sites (e.g., "Alice Mulga.zip", "Credo.zip", etc.)</li> <li>contains the data necessary to reproduce results in the main text of the study, i.e. DSMs, CHMs, and DTMs for all nine TERN sites, and at different pulse density levels (from 16 down to 0.5 laser shots per squaremetre)</li> <li>also contains calculated summary statistics for each site</li> </ul> <p>All zip files should be extracted into the same folder, except for 05_downsampling_pulse_[Site name].zip which should all be moved to a subfolder called "downsampling_pulse".</p>
Test Data for iwc Pre-curation PretextMap generation workflow
Open the record for dataset details and reuse information.
Test data for mapseq-to-ampvis2 workflow on Galaxy iwc
Open the record for dataset details and reuse information.
Test data for GUNC+BUSCO filtering workflow
<p>Test data for running GUNC+BUSCO filtering workflow, as described on <a href="https://github.com/trajkovski-lab/Quality-filtering">this GitHub page.</a></p>
Data from: A data-driven geospatial workflow to map species distributions for conservation assessments
<p>We developed a geospatial workflow that refines the distribution of a species from its extent of occurrence (EOO) to area of habitat (AOH) within the species range map. The range maps are produced with an inverse distance weighted (IDW) interpolation procedure using presence and absence points derived from primary biodiversity data (GBIF and eBird hotspots respectively). Here we provide sample data to run the geospatial workflow for nine forest species across Mexico and Central America.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.