Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
650
datasets available to search
ShareScore release 0.9.0
Dataset results
650 results for “Workflow”
An optimized RNA-seq workflow for isolating single nuclei from clinical biopsies
GEO Series GSE195719. Homo sapiens. 5 samples. Type: Expression profiling by high throughput sequencing.
Optimization of mESC ChIP-re-ChIP workflow
GEO Series GSE242686. Mus musculus. 26 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Systematic assessment of tissue dissociation and storage biases in single-cell and single-nucleus RNA-seq workflows
GEO Series GSE141115. Mus musculus. 54 samples. Type: Expression profiling by high throughput sequencing.
Sensitive and quantitative detection of MHC-I displayed neoepitopes using a semi-automated workflow and TOMAHAQ mass spectrometry
GEO Series GSE163326. Mus musculus. 12 samples. Type: Expression profiling by high throughput sequencing.
Design of an unbiased machine learning workflow to predict Multiple Sclerosis staging from blood transcriptome
GEO Series GSE136411. Homo sapiens. 336 samples. Type: Expression profiling by array.
GO-CRISPR: a highly controlled workflow to improve discovery of gene essentiality
GEO Series GSE150246. Homo sapiens; unidentified plasmid. 13 samples. Type: Expression profiling by high throughput sequencing.
A cost-effective and flexible workflow for high-resolution spatial transcriptomics in fixed tissue
GEO Series GSE292893. Homo sapiens; Mus musculus. 52 samples. Type: Expression profiling by high throughput sequencing.
Workflow for the implementation of precision genomics in healthcare
<p>To enable the implementation of precise genomics in a local healthcare system, we devised a pipeline for filtering and reporting of relevant genetic information to healthy individuals based on exome or genome data. In our analytical pipeline, the first tier of filtering is variant-centric and it is based on the selection of annotated pathogenic, protective, risk factor and drug response variants, and their one-by-one detailed evaluation. This is followed by a second-tier gene-centric deconstruction and filtering of virtual gene lists associated with diseases, and VUS-centric filtering according to ACMG pathogenicity criteria and pre-defined deleteriousness criteria. By applying this filtering protocol, we were able to provide valuable insights regarding the carrier status, pharmacogenetic profile, actionable cardiovascular and cancer predispositions, and potentially pathogenic variants of unknown significance to our patients. Our experience demonstrates that genomic profiling can be implemented into routine healthcare and provide information of medical significance</p>
Figure 13 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 13 Gold standard versus NER output.
Figure 12 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 12 An example of a specimen label.
Figure 10 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 10 Results per field from Google Cloud Vision.
Test data for sv-callers workflow
<p>This distribution includes data analyzed by the <em>sv-callers</em> workflow (v1.1.0) in the single-sample (germline) and paired-sample (somatic) modes:</p> <ul> <li>human reference genomes (in<em> .fa[sta]</em>)</li> <li>excluded genomic regions (in<em> .bed[pe]</em>) <ul> <li><a href="https://identifiers.org/encode/ENCFF001TDO">ENCODE:ENCFF001TDO</a></li> <li><a href="https://doi.org/10.1186/gb-2014-15-6-r84#ref-CR28">CEPH</a> by <a href="https://doi.org/10.1186/gb-2014-15-6-r84">Layer <em>et al</em>. (2014)</a></li> </ul> </li> <li>structural variants (SVs) detected by the workflow (in <em>.vcf</em>)</li> <li>SV <em>truth</em> sets (in<em> .bed[pe] </em>and <em>.vcf.gz</em>) <ul> <li><a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/technical/svclassify_Manuscript/Supplementary_Information/Personalis_1000_Genomes_deduplicated_deletions.bed">Personalis/1000 Genomes Project</a> data by <a href="https://doi.org/10.1186/s12864-016-2366-2">Parikh <em>et al</em>. (2016)</a></li> <li><a href="https://static-content.springer.com/esm/art%3A10.1186%2Fgb-2014-15-6-r84/MediaObjects/13059_2013_3363_MOESM4_ESM.zip">PacBio/Moleculo</a> data by <a href="https://doi.org/10.1186/gb-2014-15-6-r84">Layer <em>et al</em>. (2014)</a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/dbVar/data/Homo_sapiens/by_study/vcf/nstd167.GRCh37.variant_call.vcf.gz">dbVar:nstd167</a> data by <a href="https://doi.org/10.1038/s41587-019-0217-9">Wenger <em>et al</em>. (2019)</a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/dbVar/data/Homo_sapiens/by_study/vcf/nstd137.GRCh37.variant_call.vcf.gz">dbVar:nstd137</a> data by <a href="https://doi.org/10.1101/gr.214007.116">Huddleston <em>et al</em>. (2017)</a></li> </ul> </li> <li>workflow samples (in<em> .csv</em>) and config files (in <em>.yaml</em>)</li> <li>short-read alignments are not included due to large sizes but are freely available for download (in <em>.bam</em>) <ul> <li>NA12878 <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/NA12878/NIST_NA12878_HG001_HiSeq_300x/RMNISTHS_30xdownsample.bam">sample</a></li> <li>NA24385 <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG002_NA24385_son/NIST_Illumina_2x250bps/novoalign_bams/HG002.hs37d5.2x250.bam">sample</a></li> <li>CHM1_CHM13 <a href="https://identifiers.org/ena.embl:ERX1413368">sample</a></li> <li>COLO829 <a href="https://identifiers.org/ena.embl:ERX2765496">tumor sample</a> with matched <a href="https://identifiers.org/ena.embl:ERX2765495">normal sample</a></li> </ul> </li> <li><a href="https://github.com/GooglingTheCancerGenome/notebooks">Jupyter Notebooks</a> to analyze SV callsets (in <em>.ipynb</em>)</li> </ul>
Data for validation STEC workflow
<p>Additional data<br> ========</p> <p>This archive contains additional data for the manuscript "Validation of a bioinformatics workflow for characterization of Shiga Toxin-Producing *Escherichia coli*, applied to a high-quality reference dataset, demonstrates high performance for using WGS for routine pathogen typing"</p> <p># Notes</p> <p>Samples were analyzed with anonymized file names. They can be linked back to the original sample name as indicated in the Excel sheet.</p> <p># Content</p> <p>## Validation results ('all_results.xlsx')</p> <p>This spreadsheet contains detailed results for the validation. It contains all workflow output, corresponding metadata and classification (TP, FN, TN or FP).</p> <p>## KMA output ('kma.tar')</p> <p>This folder contains the output from KMA for the various assays. They were executed in isolation for each of the assays (instead of with the workflow). </p> <p><br> ## Example reports ('report_EH1236_*.zip')</p> <p>Example output reports of the workflow for each of the three detection methods on the same sample (EH1236).</p> <p><br> ## Virulence gene custom database ('virulence_genes_db.fasta')</p> <p>This folder contains a FASTA file with the sequences that were used to evaluate the performance of the virulence gene detection.</p> <p>## Virulence gene detection ('virulence_gene_detection.tar')</p> <p>This folder contains the output of the virulence gene detection for the three detection methods.<br> This was executed separately because the custom virulence gene database is not included in the bioinformatics workflow.</p> <p>## Workflow reports ('workflow_reports_updated.tar')</p> <p>This folder contains the output of the workflow for all of the validation samples with the three detection methods. <br> All runs were executed in August 2020, with database updated to the latest available version. <br> A single archive is created for each sample, containing the output for the three bioinformatics approaches (BLAST+, KMA, SRST2).<br> BAM files were omitted from the archives due to their large sizes.</p> <p>## Workflow reports - validation ('workflow_reports_validation.tar')</p> <p>This folder contains the output reports used for the validation (with older database version, etc).<br> This does not include KMA results because they were validated per-assay (see KMA archive).</p> <p># Contact</p> <p>For further questions you can contact Bert Bogaerts (bert.bogaerts@sciensano.be)<br> </p>
Data from: Sorting specimen-rich invertebrate samples with cost-effective NGS barcodes: validating a reverse workflow for specimen processing
Biologists frequently sort specimen-rich samples to species. This process is daunting when based on morphology, and disadvantageous if performed using molecular methods that destroy vouchers (e.g., metabarcoding). An alternative is barcoding every specimen in a bulk sample and then presorting the specimens using DNA barcodes, thus mitigating downstream morphological work on presorted units. Such a "reverse workflow" is too expensive using Sanger sequencing, but we here demonstrate that is feasible with an NGS barcoding pipeline that allows for cost-effective high throughput generation of short specimen-specific barcodes (313 bp of COI; lab cost <$0.50 per specimen) through Next Generation Sequencing of tagged amplicons. We applied our approach to a large sample of tropical ants, obtaining barcodes for 3290 of 4032 specimens (82%). NGS barcodes and their corresponding specimens were then sorted into molecular operational taxonomic units (mOTUs) based on objective clustering and Automated Barcode Gap Discovery (ABGD). High diversity of 88-90 mOTUs (4% clustering) was found and morphologically validated based on preserved vouchers. The mOTUs were overwhelmingly in agreement with morphospecies (match ratio 0.95 at 4% clustering). Because of lack of coverage in existing barcode databases, only 18 could be accurately identified to named species, but our study yielded new barcodes for 48 species, including 28 that are potentially new to science. With its low cost and technical simplicity, the NGS barcoding pipeline can be implemented by a large range of laboratories. It accelerates invertebrate species discovery, facilitates downstream taxonomic work, helps with building comprehensive barcode databases, and yields precise abundance information.
Figure 1 from: Borisenko A, Young R, Hanner R (2024) A lab-centric, workflow-based data management system for environmental DNA research. Research Ideas and Outcomes 10: e120483. https://doi.org/10.3897/rio.10.e120483
Figure 1 Schematic representation of key ontological entities of an eDNA data management system.
Sample FastQ file for sm-SNIPER workflow
<p>Reference data for sm-SNIPER workflow.</p>
Workflow record:Antique Old Telephone
Maybe you will be interested:https://youtu.be/pLjICQGqkkQ If you like it , put the like button let me know ✧◝(⁰▿⁰)◜✧ thank you~ Actually…I just want to made the wire. Modeling&UV - 3Ds Max Sculpt detail - ZBrush Bake - Marmoset Toolbag Texture - Substance Painter Final in Unity and Shader with Amplify Shader Editor Source: Objaverse 1.0 / Sketchfab
Dataset for workflow Isocor to Tracegroomer to DIMet
<p>These files are input files for workflow <a href="https://workflow4metabolomics.usegalaxy.fr/workflows/list_published" target="_self"><span>Workflow constructed from history 'ISOCOR_TRACEGROOMER_DIMET'</span></a> in W4M galaxy. https://workflow4metabolomics.usegalaxy.fr/workflows/list_published</p>
Bioinformatical workflow and results for Group B Streptococcal transcriptome when interacting with brain endothelial cells
<p>Bioinformatical workflow and results for the manuscript Group B Streptococcal transcriptome when interacting with brain endothelial cells </p>
Dataset for the article "Benchmarks and Workflow for Harmonic IR and Raman Spectra"
<p>This dataset contains the data and scripts used in the article "Benchmarks and Workflow for Harmonic IR and Raman Spectra"</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.