Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9
datasets available to search
ShareScore release 0.9.0
Dataset results
9 results for “Nextflow”
NextClone and CloneDetective: An Integrated Nextflow Pipeline and R Package for Clonal Barcode Extraction and Quantification
<p>Raw FASTQ files for the DNA-seq data required to replicate the analyses presented at: https://phipsonlab.github.io/NextClone-analysis/.</p>
Dataset to test the Nextflow Tutorial
<p>Tutorial: <a href="https://telatin.github.io/microbiome-bioinformatics/Nextflow-start/">https://telatin.github.io/microbiome-bioinformatics/Nextflow-start/</a></p> <p>Repository: <a href="https://github.com/telatin/nextflow-example">https://github.com/telatin/nextflow-example</a></p>
NextClone and CloneDetective: An Integrated Nextflow Pipeline and R Package for Clonal Barcode Extraction and Quantification
<p>BAM file (chunk 26-50 of 50) for the scRNAseq data required to replicate the analyses presented at: https://phipsonlab.github.io/NextClone-analysis/.</p><p>Chunk 1-25 can be downloaded from https://zenodo.org/records/10129134.</p><p>The original BAM file was too big to fit in an entry. Thus it was split into 50 using picard:</p><blockquote><p>SplitSamByNumberOfReads -I possorted_genome_bam.bam -O split_sam/ -N_FILES 50 <i>--CREATE_MD5_FILE</i></p></blockquote><p>Before re-running all the analyses in https://phipsonlab.github.io/NextClone-analysis/, make sure you merge all 50 chunks first using picard (change xxx to point to the directory storing all the 50 chunks you have downloaded):</p><blockquote><p>outdir="xxx"</p><p>args=""</p><p><i># Loop through each BAM file</i></p><p>for file in ${outdir}/*.bam; do</p><p> args+="-I $file "</p><p>done</p><p># Do the actual merging</p><p>MergeSamFiles $args -O $outdir/merged_v2/merged_bam_v2.bam --USE_THREADING --CREATE_MD5_FILE</p></blockquote><p>Picard can be downloaded from: https://github.com/broadinstitute/picard</p>
NextClone and CloneDetective: An Integrated Nextflow Pipeline and R Package for Clonal Barcode Extraction and Quantification
<p>BAM file (chunk 1-25 of 50) for the scRNAseq data required to replicate the analyses presented at: https://phipsonlab.github.io/NextClone-analysis/.</p><p>Chunk 26-50 can be downloaded from https://zenodo.org/uploads/10129625</p><p>The original BAM file was too big to fit in an entry. Thus it was split into 50 using picard:</p><blockquote><p>SplitSamByNumberOfReads -I possorted_genome_bam.bam -O split_sam/ -N_FILES 50 <i>--CREATE_MD5_FILE</i></p></blockquote><p>Before re-running all the analyses in https://phipsonlab.github.io/NextClone-analysis/, make sure you merge all 50 chunks first using picard (change xxx to point to the directory storing all the 50 chunks you have downloaded):</p><blockquote><p>outdir="xxx"</p><p>args=""</p><p><i># Loop through each BAM file</i></p><p>for file in ${outdir}/*.bam; do</p><p> args+="-I $file "</p><p>done</p><p># Do the actual merging</p><p>MergeSamFiles $args -O $outdir/merged_v2/merged_bam_v2.bam --USE_THREADING --CREATE_MD5_FILE</p></blockquote><p>Picard can be downloaded from: https://github.com/broadinstitute/picard</p>
Samples for CNIC-Proteomics Nextflow pipelines: nf-PTM-compass
<p>We present several input sample files for the execution of various Nextflow pipelines developed by the <em>Cardiovascular Proteomics Lab/Proteomics Unit at the National Centre for Cardiovascular Research</em> (CNIC, <a href="https://www.cnic.es/" target="_blank" rel="noopener">https://www.cnic.es</a>).</p> <p>The <strong>nf-PTM-compass</strong> enhances the identification and quantification of Post-Translational Modifications (PTMs) (<a href="https://github.com/CNIC-Proteomics/nf-PTM-compass" target="_blank" rel="noopener">https://github.com/CNIC-Proteomics/nf-PTM-compass</a>).</p> <p>The available sample files are as follows:</p> <ul> <li><strong>heteroplasmic_heart.zip</strong>: Heart tissue. Input files for nf-PTM-compass, derived from the study by Bagwan N, Bonzon-Kulichenko E, Calvo E, et al. (*), with results from an open search conducted using nf-SearchEngine results.</li> <li><strong>heteroplasmic_liver.zip</strong>: Liver tissue. Input files for nf-PTM-compass, derived from the same study (*), with results from an open search conducted using nf-SearchEngine results.</li> <li><strong>heteroplasmic_muscle.zip</strong>: Mucle tissue. Input files for nf-PTM-compass, derived from the same study (*), with results from an open search conducted using nf-SearchEngine results.</li> </ul> <p> </p> <p>(*) Bagwan N, Bonzon-Kulichenko E, Calvo E, et al. Comprehensive Quantification of the Modified Proteome Reveals Oxidative Heart Damage in Mitochondrial Heteroplasmy. <em>Cell Reports</em>. 2018;23(12):3685-3697.e4. doi:10.1016/j.celrep.2018.05.080</p> <p> </p>
Common Workflow Scheduler Evaluation with Nextflow and Kubernetes
<ul> <li>Setup scripts to test Nextflow with CWS on Kubernetes</li> <li>Traces and logs of 990 workflow executions</li> </ul>
Underlying data for: nf-core/clipseq - a robust Nextflow pipeline for comprehensive CLIP data analysis
<p>Underlying data for: nf-core/clipseq - a robust Nextflow pipeline for comprehensive CLIP data analysis</p>
Bioinformatics workflow for the detection of eQTL in the cattle genome using Nextflow DSL2
<p>The <em>in silico</em> detection of expression quantitative trait loci (eQTL) demands high throughput processing from hundreds of samples, which is often a challenge to handle and run such large datasets. In order to focus on the core analysis, it is convenient to have simple coding and hassle-free installation of different software tools required for the bioinformatics workflow. In this context, the newly available technologies like workflow managers and software containers enabled to develop workflows with less complexity. In this study, we developed an eQTL bioinformatics pipeline with the workflow manager Nextflow and docker container software, for coding and installing the required software tools. This workflow can be portable to a different computer environment, and the results are reproducible. We tested the functionality of our workflow with a sample dataset and the runtime estimates from this demo run will provide important information in planning future analyses with much larger datasets.</p>
Nextflow Pipeline for Visium and H&E Data from Patient-Derived Xenograft Samples
GEO Series GSE238004. Mus musculus. 4 samples. Type: Expression profiling by high throughput sequencing; Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.