Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

14

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

14 results for “bioinformatic pipeline”

Learn how ShareScore rates datasets ↗
zenodo44/100

Bioinformatic pipeline: Genomic diversity landscape of the honey bee gut microbiota

<p>This data-set describes the full bioinformatic pipeline used to analyze 54 metagenomic samples of the honey bee gut microbiota. Each sample was isolated from an individual honey bee, and all samples originate from two colonies of the Engel laboratory at the University of Lausanne, Switzerland. The full raw data-set is available from the sequence-read archive: SRP150166.</p> <p>A publication based on this analysis is currently under review, with the title: &quot;Genomic diversity landscape of the honey bee gut microbiota&quot;, and an upload to Biorxiv is also underway.</p> <p>The data-set contains tar-balls for the different main workflows of the analysis. Dowload and unpack to view the contents (tar -zxvf filename.tar.gz). For each workflow, all directories contain README.txt files, describing the contents of the directory. Due to size constraints, some intermediate files have been omitted, and some workflows are demonstrated for a subset of the data. However, the full analysis can be reproduced from the raw data, using the provided scripts.</p> <p>Scripts are included within workflow directories, and are also provided as a separate tar-ball for convenience. All perl-scripts come with documentation, which can be viewed by typing: &quot;perl script_name.pl -h&quot;. For R scripts, the usage is indicated as a comment in the top lines of each script. Note that many of the scripts require specific input-files to be present in the run-directory. Their usage is demonstrated within the workflow directories in bash-scripts (*.sh). Commands used for generating plots and some statistics are given within workflow directories in text-files &quot;R.commands&quot; when applicable.</p> <p>Aside from custom code, the pipeline also utilizes various open-source Software packages, which are detailed in the file &quot;software_dependencies.txt&quot;. Note, while many of the scripts will run fast on any computer, some steps of the pipeline are computationally demanding, and will require significant computing time, as well as storage space. When scripts are known to be time-consuming, this is indicated in the script help message.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2018View details →
dryad40/100

A two-tier bioinformatic pipeline to develop probes for target capture of nuclear loci with applications in Melastomataceae

<p><b><i>Premise of the study</i></b><b>: </b>Putatively single-copy nuclear (SCN) loci, identified using genomic resources of closely related species, are ideal for phylogenomic inference. However, suitable genomic resources are not available for many clades, including Melastomataceae. We introduce a versatile approach to identify SCN loci for clades with few genomic resources and use it to develop probes for target enrichment<i> </i>in the distantly related <i>Memecylon</i> and <i>Tibouchina</i> (Melastomataceae).</p> <p><b><i>Methods</i></b>: We present a two-tiered pipeline. First, we identified putatively SCN loci using MarkerMiner and transcriptomes from distantly related species in Melastomataceae. Published loci and genes of functional significance were added (384 total loci). Second, using HybPiper, we retrieved 689 homologous template sequences for these loci using genome-skimming data from within the focal clades.</p> <p><b><i>Results</i></b>: We sequenced 193 loci from both <i>Memecylon</i> and <i>Tibouchina</i>, with probes designed from 56 template sequences successfully targeting sequences in both clades. Probes designed from genome-skimming data within a focal clade were more successful than probes designed from other sources.</p> <p><b><i>Discussion: </i></b>Our pipeline successfully identified and targeted SCN loci in <i>Memecylon </i>and <i>Tibouchina</i>, enabling phylogenomic studies in both clades and potentially across Melastomataceae. This pipeline could be easily applied to other clades with few genomic resources. </p>

opencc-zeroJan 2021View details →
dryad40/100

A two-tier bioinformatic pipeline to develop probes for target capture of nuclear loci with applications in Melastomataceae

Open the record for dataset details and reuse information.

publicJan 2021View details →
dryad40/100

UnFATE: A comprehensive probe set and bioinformatics pipeline for phylogeny reconstruction and multilocus barcoding of filamentous ascomycetes (Ascomycota, Pezizomycotina)

Open the record for dataset details and reuse information.

publicJan 2025View details →
dryad36/100

Bioinformatic pipeline from: Increasing confidence for discerning species and population compositions from metabarcoding assays of environmental samples: case studies of fishes in the Laurentian Great Lakes and Wabash River

<p>Community composition data are essential for conservation management, facilitating identification of rare native and invasive species, along with abundant ones. However, traditional capture-based morphological surveys require considerable taxonomic expertise, are time consuming and expensive, can kill rare taxa and damage habitats, and often are prone to false negatives. Alternatively, metabarcode assays can be used to assess the genetic identity and compositions of entire communities from environmental samples, comprising a more sensitive, less damaging, and relatively time- and cost-efficient approach. However, there is a trade-off between the stringency of bioinformatic filtering needed to remove false positives and the potential for false negatives. The present investigation thus evaluated use of four mitochondrial (mt) DNA metabarcode assays and a customized bioinformatic pipeline to increase confidence in species identifications by removing false positives, while achieving high detection probability. Positive controls were used to calculate sequencing error, and results that fell below those cutoff values were removed, unless found with multiple assays. The performance of this approach was tested to discern and identify North American freshwater fishes using lab experiments (mock communities and aquarium experiments) and processing of a bulk ichthyoplankton sample. The method then was applied to field environmental (e)DNA water samples taken concomitant with electrofishing surveys and morphological identifications. This protocol detected 100% of species present in concomitant electrofishing surveys in the Wabash River and an additional 21 that were absent from traditional sampling. Using single 1 L water samples collected from just four locations, the metabarcoding assays discerned 73% of the total fish species that were discerned in comparison to four months of an extensive electrofishing river survey in the Maumee River, along with an additional nine species. In both rivers, total fish species diversity was best resolved when all four metabarcode assays were used together, which identified 35 additional species missed by electrofishing. Ecological distinction and diversity levels among the fish communities also were better resolved with the metabarcode assays than with morphological sampling and identifications, especially with the combined assays. At the population-level, metabarcode analyses targeting the invasive round goby <i>Neogobius melanostomus</i> and the silver carp <i>Hypophthalmichthys molitrix</i> identified all population haplotype variants found using Sanger sequencing of morphologically sampled fish, along with additional intra-specific diversity, meriting further investigation. Overall findings demonstrated that the use of multiple metabarcode assays and custom bioinformatics that filter potential error from true positive detections improves confidence in evaluating biodiversity.</p>

opencc-zeroAug 2021View details →
dryad36/100

Bioinformatic pipeline from: Increasing confidence for discerning species and population compositions from metabarcoding assays of environmental samples: case studies of fishes in the Laurentian Great Lakes and Wabash River

Open the record for dataset details and reuse information.

publicJan 2021View details →
zenodo32/100

Bioinformatic pipeline: Vast differences in strain-level diversity in the gut microbiota of two closely related honey bee species

<p>This data-set contains the full bioinformatic pipeline used to analyze metagenomic samples in the study &quot;Vast differences in strain-level diversity in the gut microbiota of two closely related honey bee species&quot; (Ellegaard et al. 2020, Current Biology).&nbsp;</p> <p>New metagenomic samples were generated for the study, for which the raw data is available on the NCBI Sequence Read Achive, under accession: PRJNA59809.</p> <p>The data of this submission consist of 9 tar-balls, as further described here below. Download and unpack to view the contents (tar -zxvf filename.tar.gz). For each tarball, all directories contain README.txt files, describing the contents of the directory. Due to size constraints, some intermediate files have been omitted, and some workflows are demonstrated for a subset of the data. However, the full analysis can be reproduced from the raw data, using the provided scripts.</p> <p>All scripts are included within the directories where they were applied. Perl-scripts contain documentation, which can be viewed by typing: &quot;perl script_name.pl -h&quot;. For R scripts, the usage is indicated as a comment in the top lines of each script. Note that many of the scripts require specific input-files to be present in the run-directory. Their usage is demonstrated within the workflow directories in bash-scripts (*.sh). Commands used for generating plots and some statistics are given within workflow directories in text-files &quot;R.commands&quot; when applicable.</p> <p>Aside from custom code, the pipeline also utilizes various open-source Software packages, which are detailed in the file &quot;software_dependencies.txt&quot;. Note, while many of the scripts will run fast on any computer, some steps of the pipeline are computationally demanding, and will require significant computing time, as well as storage space. When scripts are known to be time-consuming, this is indicated in the script help message.</p> <p>Description of tarballs.</p> <p>raw_data_processing.tar.gz: Describes the quality-control and trimming of raw data, and includes info on the sequencing run.</p> <p>databases.tar.gz: Contains all databases used for analysis, in addition to relevant meta-data.</p> <p>mapping_stats.tar.gz: Contains a file with the number of reads mapped to the honey bee gut microbiota database and the host genomes, for each sample. Bash-scripts are provided, detailing how the mapping was done and quantified.</p> <p>orthologs_phylogenies.tar.gz: Contains the pipeline for inferring orthologous gene-families and core genome phylogenies, as well as scripts for filtering of single-copy core gene families.</p> <p>assemblies.tar.gz: Contains the final de novo metagenome assembly files (contig fasta-files), gener<br> ated for both complete and rarefied read subsets. Bash-scripts detailing the assembly commands are also provided.</p> <p>SDP_validation.tar.gz: Contains the pipeline for metagenomic validation of candidate SDPs. Final output-files, containing the percentage identity of recruited metagenomic ORFs to database core genes, are provided for each candidate SDP. Additionally, a small example dataset is provided, where the intermediate result-files can be viewed.</p> <p>community_profiling.tar.gz: Contains the pipeline for community profiling, i.e. the quantification of individual community members (SDPs) across samples. Final output files are provided, including mapped read coverage on core gene families and corresponding plots. A small bam-file (containing data from a single subset sample), is also provided, in order to demonstrate the pipeline, together with all scripts used.</p> <p>snv_profiling.tar.gz: Contains the pipeline used for SNV profiling, including filtering and analysis. Final filtered vcf-files are provided for each SDP. Analytical output files are also provided, including data on shared SNV fractions, distance matrices, and cumulative curves.</p> <p>metagenomic_ORF_analyses.tar.gz: Contains the pipeline for analysis of metagenomic ORFs. This includes prediction of ORFs, clustering, annotation and functional characterization. ORF sequences, annotation files, and cluster-files are provided.</p>

opencc-by-4.0Apr 2020View details →
zenodo28/100

Raw data for comparison of bioinformatics pipelines for Diatom DNA metabarcoding for ecological assessment

<p>This archive contains the raw .fatsq files for&nbsp;29 samples from&nbsp;water bodies (lakes and rivers) located in Nordic countries (Sweden, Finland, Norway)&nbsp;sequenced on Illumina MiSeq, with the 18S-V4 marker&nbsp;and with the <em>rbc</em>L marker. For both marker two separate datasets are provided, containing the&nbsp;F and R fragments ( R1 and R2). The&nbsp;samples tags and primer sequences are also provided.</p> <p>The archive also contains the custom curated reference database used for the taxonomic identification using bioinformatics pipeline. For both marker, two files are provided: an .rarl file with the sequences and sequence ID and a .tax file with the taxonomic information associated.&nbsp;</p>

opencc-by-4.0Mar 2020View details →
dryad28/100

Semi-artificial datasets as a resource for validation of bioinformatics pipelines for plant virus detection

<p>In the last decade, High-Throughput Sequencing (HTS) has revolutionized biology and medicine. This technology allows the sequencing of huge amount of DNA and RNA fragments at a very low price. In medicine, HTS tests for disease diagnostics are already brought into routine practice. However, the adoption in plant health diagnostics is still limited. One of the main bottlenecks is the lack of expertise and consensus on the standardization of the data analysis. The Plant Health Bioinformatic Network (PHBN) is an Euphresco project aiming to build a community network of bioinformaticians/computational biologists working in plant health. One of the main goals of the project is to develop reference datasets that can be used for validation of bioinformatics pipelines and for standardization purposes.</p> <p>Semi-artificial datasets have been created for this purpose (Datasets 1 to 10). They are composed of a "real" HTS dataset spiked with artificial viral reads. It will allow researchers to adjust their pipeline/parameters as good as possible to approximate the actual viral composition of the semi-artificial datasets. Each semi-artificial dataset allows to test one or several limitations that could prevent virus detection or a correct virus identification from HTS data (<i>i.e.</i> low viral concentration, new viral species, non-complete genome).</p> <p>Eight artificial datasets only composed of viral reads (no background data) have also been created (Datasets 11 to 18). Each dataset consists of a mix of several isolates from the same viral species showing different frequencies. The viral species were selected to be as divergent as possible. These datasets can be used to test haplotype reconstruction software, the goal being to reconstruct all the isolates present in a dataset.</p> <p><span>A GitLab repository (<a href="https://gitlab.com/ilvo/VIROMOCKchallenge">https://gitlab.com/ilvo/VIROMOCKchallenge</a>) is available and provides a complete description of the composition of each dataset, the methods used to create them and their goals.</span></p>

opencc-zeroNov 2021View details →
dryad28/100

Semi-artificial datasets as a resource for validation of bioinformatics pipelines for plant virus detection

Open the record for dataset details and reuse information.

publicNov 2021View details →
dryad28/100

Data from: SSR_pipeline: a bioinformatic infrastructure for identifying microsatellites from paired-end Illumina high-throughput DNA sequencing data

Open the record for dataset details and reuse information.

publicSep 2013View details →
geo24/100

ChEC-seq2: an improved chromatin endogenous cleavage sequencing method and bioinformatic analysis pipeline for mapping in vivo protein–DNA interactions

GEO Series GSE246951. Saccharomyces cerevisiae. 42 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenNov 2023View details →
geo24/100

IS-Seq: a novel bioinformatics pipeline for integration sites analysis

GEO Series GSE203211. Homo sapiens. 3 samples. Type: Other.

openGEO-OpenAug 2023View details →
geo24/100

A bioinformatic pipeline for analysis of M.EcoGII methylation footprint PacBio long-read sequence data

GEO Series GSE243114. Saccharomyces cerevisiae; Escherichia coli. 12 samples. Type: Other.

openGEO-OpenMar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record