Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
250
datasets available to search
ShareScore release 0.7.1
Dataset results
250 results for “Bioinformatics”
A two-tier bioinformatic pipeline to develop probes for target capture of nuclear loci with applications in Melastomataceae
Open the record for dataset details and reuse information.
UnFATE: A comprehensive probe set and bioinformatics pipeline for phylogeny reconstruction and multilocus barcoding of filamentous ascomycetes (Ascomycota, Pezizomycotina)
Open the record for dataset details and reuse information.
CEA301 bioinformatics workshop data
<p>Post-processed FreeBayes variant calls for use in the 2020 bioinformatics workshop for unit CEA301.</p> <p>These are derived from data published previously in <a href="https://zenodo.org/record/3243160">Training material for the course "Exome analysis with GALAXY"</a>. Credit for uploading the original data goes to Paolo Uva and Gianmauro Cuccuru!</p>
Bioinformatic pipeline from: Increasing confidence for discerning species and population compositions from metabarcoding assays of environmental samples: case studies of fishes in the Laurentian Great Lakes and Wabash River
<p>Community composition data are essential for conservation management, facilitating identification of rare native and invasive species, along with abundant ones. However, traditional capture-based morphological surveys require considerable taxonomic expertise, are time consuming and expensive, can kill rare taxa and damage habitats, and often are prone to false negatives. Alternatively, metabarcode assays can be used to assess the genetic identity and compositions of entire communities from environmental samples, comprising a more sensitive, less damaging, and relatively time- and cost-efficient approach. However, there is a trade-off between the stringency of bioinformatic filtering needed to remove false positives and the potential for false negatives. The present investigation thus evaluated use of four mitochondrial (mt) DNA metabarcode assays and a customized bioinformatic pipeline to increase confidence in species identifications by removing false positives, while achieving high detection probability. Positive controls were used to calculate sequencing error, and results that fell below those cutoff values were removed, unless found with multiple assays. The performance of this approach was tested to discern and identify North American freshwater fishes using lab experiments (mock communities and aquarium experiments) and processing of a bulk ichthyoplankton sample. The method then was applied to field environmental (e)DNA water samples taken concomitant with electrofishing surveys and morphological identifications. This protocol detected 100% of species present in concomitant electrofishing surveys in the Wabash River and an additional 21 that were absent from traditional sampling. Using single 1 L water samples collected from just four locations, the metabarcoding assays discerned 73% of the total fish species that were discerned in comparison to four months of an extensive electrofishing river survey in the Maumee River, along with an additional nine species. In both rivers, total fish species diversity was best resolved when all four metabarcode assays were used together, which identified 35 additional species missed by electrofishing. Ecological distinction and diversity levels among the fish communities also were better resolved with the metabarcode assays than with morphological sampling and identifications, especially with the combined assays. At the population-level, metabarcode analyses targeting the invasive round goby <i>Neogobius melanostomus</i> and the silver carp <i>Hypophthalmichthys molitrix</i> identified all population haplotype variants found using Sanger sequencing of morphologically sampled fish, along with additional intra-specific diversity, meriting further investigation. Overall findings demonstrated that the use of multiple metabarcode assays and custom bioinformatics that filter potential error from true positive detections improves confidence in evaluating biodiversity.</p>
Combined bioinformatics and machine learning methodologies reveal prognosis-related ceRNA network and propose ABCA8, CAT, and CXCL12 as independent protective factors against osteosarcoma
<p><strong>Supplementary Table 1</strong>. Basic traits of the seven microarray datasets from the Gene Expression Omnibus and The Cancer Genome Atlas.</p><p><strong>Supplementary Table 2</strong>. Basic characteristics of the nine differentially expressed circRNAs</p><p><strong>Supplementary Table 3</strong>. Index of concordance (C-index) and variance inflation factor (VIF) of ABCA8, CXCL12, and CAT.</p><p><strong>Supplementary Table 4.</strong> LASSO and cox analysis of ceRNA with coef/se(coef) < 0.01 and P-value of proportional hazards assumption (PH) > 0.05.</p><p><strong>Supplementary Table 5</strong>. Robust rank aggregation analysis of ABCA8, CXCL12, and CAT. LogFC in the four datasets and RRA score of the three RNAs.</p><p><strong>Supplementary Figure 1.</strong> Competitive endogenous RNA in osteosarcoma.</p><p><strong>Supplementary Figure 2.</strong> Protein–protein interaction (PPI) network of genes in the competitive endogenous RNA network</p><p><strong>Supplementary Figure 3.</strong> Proportional hazards assumption (left) and linearity assumption (right).</p><p><strong>Supplementary Figure 4.</strong> Survival analysis for competitive endogenous RNA in osteosarcoma.</p>
Impact of open bioinformatics resources on industry (worked example for an ELIXIR deposition database)
Open the record for dataset details and reuse information.
Mind Map of The European Bioinformatic infrastructure for the Life Sciences (ELIXIR) activities
<p>This diagram showcases the organisation of the different components of the ELIXIR network, with brief description of the activities therein.</p>
Pre-Analysis Bioinformatics Files for The Microbiome and Volatile Organic Compounds Reflect the State of Decomposition in an Indoor Environment
<p>Data statistics before and after trimming, FastQC reports before and after trimming, MultiQC reports before and after trimming, commands for the Kraken2-Bracken analysis, and classification reports. Read the read.me file for file names and descriptions. </p>
Screening Rafflesia and Sapria Metabolites Using a Bioinformatics Approach to Assess Their Potential as Drugs
<p>This dataset contains molecular docking results read using LigPlus, referred in the text:</p> <p>Wicaksono et al. (2022) Screening Rafflesia and Sapria Metabolites Using a Bioinformatics Approach to Assess Their Potential as Drugs. Philippine Journal of Science.</p> <p>As Appendix II.</p>
EMBL-EBI scRNA Bioinformatics T cell course 2022 (RESULTS)
<p>Repository: Final results for the projects developed during the 'Bioinformatics for T-Cell immunology' course (11-15/07/2022) at EMBL-EBI: [https://www.ebi.ac.uk/training/events/bioinformatics-t-cell-immunology-2022](https://www.ebi.ac.uk/training/events/bioinformatics-t-cell-immunology-2022)</p> <p>Official website: [https://elolab.github.io/Bioinfo_Tcell_projects_22](https://elolab.github.io/Bioinfo_Tcell_projects_22)</p> <p>Maintainer: [Elo lab](https://elolab.utu.fi)</p> <p>Contact: **António Sousa** ([ENLIGHT-TEN+](http://www.enlight-ten.eu) PhD student at the [Medical Bioinformatics Centre](https://elolab.utu.fi), TBC, University of Turku & Åbo Akademi)</p> <p>Last update: 07/07/2022</p> <p> </p> <p><br></p> <p> </p> <p>---</p> <p> </p> <p><br></p> <p> </p> <p>### Projects</p> <p> </p> <p>This repository hosts the results related with three standalone/independent data analysis projects examples using publicly available data generated and published elsewhere properly referenced below:</p> <p> 1. _Integration of single-cell data from patients developing arthritis arAE under ICI_<br> <br> + _publication_: [Kim et al., 2022](https://www.nature.com/articles/s41467-022-29539-3)<br> <br> + _data_: GEO [GSE173303](https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE173303)<br> <br> + _R markdown notebook_: `01_integration_arthritis_arAE_ICI.Rmd`<br> <br> + _vignette_: [01_integration_arthritis_arAE_ICI.html](https://elolab.github.io/Bioinfo_Tcell_projects_22/pages/01_integration_arthritis_arAE_ICI.html)<br> <br> + _results_ (_in this repository_): GSE173303.tar.gz<br> <br> 2. _Fine-grained clustering of single-cell data of melanoma immune/stroma cells_<br> <br> + _publication_: [Jerby-Arnon et al., 2018](https://www.sciencedirect.com/science/article/pii/S0092867418311784?via%3Dihub)<br> <br> + _data_: GEO [GSE115978](https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE115978)<br> <br> + _R markdown notebook_: `02_clustering_seurat_vs_iloreg.Rmd`<br> <br> + _vignette_: [02_clustering_seurat_vs_iloreg.html](https://elolab.github.io/Bioinfo_Tcell_projects_22/pages/02_clustering_seurat_vs_iloreg.html)</p> <p> + _results_ (_in this repository_): GSE115978.tar.gz</p> <p> 3. _Differential gene expression of stimulated CD4+ T single-cell data with single-cell and pseudobulk methods_<br> <br> + _publication_: [Cano-Gamez et al., 2020](https://www.nature.com/articles/s41467-020-15543-y)<br> <br> + _data_: [www.opentargets.org](https://www.opentargets.org/projects/effectorness)<br> <br> + _R markdown notebook_: `03_pseudobulks_dge_rots_cd4_act.Rmd`<br> <br> + _vignette_: [03_pseudobulks_dge_rots_cd4_act.html](https://elolab.github.io/Bioinfo_Tcell_projects_22/pages/03_pseudobulks_dge_rots_cd4_act.html)</p> <p> + _results_ (_in this repository_): CanoGamez_et_al_2020.tar.gz</p> <p> </p> <p><br></p> <p> </p> <p>---</p> <p> </p> <p><br></p> <p> </p> <p>### Disclaimer</p> <p> </p> <p>>All the data used along each project notebook was made public elsewhere by the respective authors and it has been properly referenced in each project (proper links were provided along each project notebook). The data and tools chosen to address the topic(s) of each project notebook reflect only my personal experience/knowledge and they were chosen to highlight particular aspects that I consider important. The results generated and explored within each project notebook have just the general purpose of give a brief introduction to the topics addressed in each project and do not aim, at any point, to reproduce or question neither the approaches taken nor the main findings published along with the data sets used herein.</p> <p> </p>
Supplemental Data for "Adaptive Container Service: a New Paradigm for Robust and Optimized Bioinformatics Workflow Deployment in the Cloud."
<p>All supplemental data for "Adaptive Container Service: a New Paradigm for Robust and Optimized Bioinformatics Workflow Deployment in the Cloud."<br><br>Abstract:<br>We propose Adaptive Container Service (ACS), a new paradigm for deploying bioinformatics workflows in cloud computing environments. By encapsulating the entire workflow within a single virtual container, combined with automatic workflow checkpointing and dynamic migration to appropriately scaled containers, ACS-based deployment demonstrates several key advantages over alternative strategies: it enables optimal resource provision to any workflow that comprise of multiple applications with diverse computing needs; it provides protection against application-agnostic out-of-memory (OOM) errors or spot instance interruptions; and it reduces efforts required for workflow development, optimization, and management because it runs workflows with minimal or no code modifications. Proof-of-concept experiments show that ACS avoided both under- and over-provisioning in monolithic single-container deployment. Despite being deployed as a single container, it achieved comparable resource utilization efficiency as optimized Nextflow-managed, multi-modular workflows. Analysis of over 18,000 workflow runs demonstrated that ACS can effectively reduce workflow failures by two-thirds. These findings suggest that ACS frees developers from navigating the complexity of deploying robust workflows and rightsizing compute resources in the cloud, leading to significant reduction in workflow development time and savings in cloud computing costs.<br><br>Contains the following directories:<br>Fig2-bbtools: running metrics for BBTools<br>Fig3-rna-seq: running metrics for RNA-Seq<br>Fig4-Job_records: meta data and running metrics of 18,000+ jobs</p>
Identification and Exploration of Immunity-Related Genes and Natural Products for Alzheimer's Disease Based on Bioinformatics, Molecular Docking and Molecular Dynamics
<p>Supplementary material to the article: Identification and Exploration of Immunity-Related Genes and Natural Products for Alzheimer’s Disease Based on Bioinformatics, Molecular Docking and Molecular Dynamics,These data are available to researchers.</p>
Literature consistency of bioinformatics sequence databases is effective for assessing record quality
<p>Bioinformatics sequence databases such as Genbank or UniProt contain hundreds of millions of records of genomic data. These records are derived from direct submissions from individual laboratories, as well as from bulk submissions from large-scale sequencing centres; their diversity and scale means that they suffer from a range of data quality issues including errors, discrepancies, redundancies, ambiguities, incompleteness and inconsistencies with the published literature. In this work, we seek to investigate and analyze the data quality of sequence databases from the perspective of a curator, who must detect anomalous and suspicious records. Specifically, we emphasize the detection of inconsistent records with respect to the literature. Focusing on GenBank, we propose a set of 24 quality indicators, which are based on treating a record as a query into the published literature, and then use query quality predictors. We then carry out an analysis that shows that the proposed quality indicators and the quality of the records have a mutual relationship, in which one depends on the other. We propose to represent record literature consistency as a vector of these quality indicators. By reducing the dimensionality of this representation for visualization purposes using principal component analysis, we show that records which have been reported as inconsistent with the literature fall roughly in the same area, and therefore share similar characteristics. By manually analyzing records not previously known to be erroneous that fall in the same area than records know to be inconsistent, we show that one record out of four is inconsistent with respect to the literature. This high density of inconsistent record opens the way towards the development of automatic methods for the detection of faulty records. We conclude that literature inconsistency is a meaningful strategy for identifying suspicious records.</p>
Data supplementing the article "Diatom DNA metabarcoding for biomonitoring : strategies to avoid major taxonomical and bioinformatical biases limiting molecular indices capacities" K. Tapolczai, F. Keck, A. Bouchez, F. Rimet, M. Kahlert and V. Vasselon submitted to "Frontiers in Ecology and Evolution" journal
<p>These data supplement the article "Diatom DNA metabarcoding for biomonitoring : strategies to avoid major taxonomical and bioinformatical biases limiting molecular indices capacities" K. Tapolczai, F. Keck, A. Bouchez, F. Rimet, M. Kahlert and V. Vasselon submitted to "Frontiers in Ecology and Evolution" journal.</p> <p>The directory contains the following files:</p> <p><strong>464_samples_fastq_files_(mothur).rar </strong>- contains the 464 fastq files proceed together during the Mothur bioinformatics treatments to produce the OTUs and ISUs tables. As the contig and the demultiplexing steps were performed by the sequencing platform, there is 1 fastq file per sample. From this 464 samples OTU/ISU tables, only information regarding 76 samples were used in this study and are listed in the "<strong>76_samples_list_(mothur).xlsx" </strong>file<strong>.</strong></p> <p><strong>76_samples_list_(mothur).xlsx </strong>- contains the information regarding the 76 samples used to create the OTUs and ISUs tables presented in the paper.</p> <p><strong>76_samples_R1_R2_fastq_files(DADA2).rar - </strong>contains the raw demultiplexed fastq files (R1.fastq and R2.fastq) for each of the 76 samples used in this study to produce the ESVs table using the DADA2 bioinformatics pipeline.</p>
Bioinformática para el análisis de SARS-CoV-2 para principiantes / Bioinformatics for SARS-CoV-2 analysis for beginners
<ul> <li>Data for the e-learning tutorial <a href="https://github.com/cabana-online/Course_SARS_CoV-2">Bioinformatics for SARS-CoV-2 analysis for beginners</a></li> <li>Datos para el tutorial e-learning <a href="Bioinformática para el análisis de SARS-CoV-2 para principiantes" target="_blank" rel="noopener">Bioinformática para el análisis de SARS-CoV-2 para principiantes</a></li> </ul>
EvolvingSTEM Bioinformatics mutant colony sequencing
<p>Mutant colony sequencing for EvolvingSTEM Bioinformatics</p>
Supplementary materials for "SPAG9 expression predicts a good prognosis in patients with clear cell renal cell carcinoma: A bioinformatics integrative analysis"
<p>Supplementary materials for "SPAG9 expression predicts a good prognosis in patients with clear cell renal cell carcinoma: A bioinformatics integrative analysis".</p>
Bioinformatic pipeline from: Increasing confidence for discerning species and population compositions from metabarcoding assays of environmental samples: case studies of fishes in the Laurentian Great Lakes and Wabash River
Open the record for dataset details and reuse information.
Data from: Transcriptional regulation of human <em>NMNAT2</em>: Insights from 3D genome sequencing and bioinformatics
Open the record for dataset details and reuse information.
Bioinformatic pipeline: Vast differences in strain-level diversity in the gut microbiota of two closely related honey bee species
<p>This data-set contains the full bioinformatic pipeline used to analyze metagenomic samples in the study "Vast differences in strain-level diversity in the gut microbiota of two closely related honey bee species" (Ellegaard et al. 2020, Current Biology). </p> <p>New metagenomic samples were generated for the study, for which the raw data is available on the NCBI Sequence Read Achive, under accession: PRJNA59809.</p> <p>The data of this submission consist of 9 tar-balls, as further described here below. Download and unpack to view the contents (tar -zxvf filename.tar.gz). For each tarball, all directories contain README.txt files, describing the contents of the directory. Due to size constraints, some intermediate files have been omitted, and some workflows are demonstrated for a subset of the data. However, the full analysis can be reproduced from the raw data, using the provided scripts.</p> <p>All scripts are included within the directories where they were applied. Perl-scripts contain documentation, which can be viewed by typing: "perl script_name.pl -h". For R scripts, the usage is indicated as a comment in the top lines of each script. Note that many of the scripts require specific input-files to be present in the run-directory. Their usage is demonstrated within the workflow directories in bash-scripts (*.sh). Commands used for generating plots and some statistics are given within workflow directories in text-files "R.commands" when applicable.</p> <p>Aside from custom code, the pipeline also utilizes various open-source Software packages, which are detailed in the file "software_dependencies.txt". Note, while many of the scripts will run fast on any computer, some steps of the pipeline are computationally demanding, and will require significant computing time, as well as storage space. When scripts are known to be time-consuming, this is indicated in the script help message.</p> <p>Description of tarballs.</p> <p>raw_data_processing.tar.gz: Describes the quality-control and trimming of raw data, and includes info on the sequencing run.</p> <p>databases.tar.gz: Contains all databases used for analysis, in addition to relevant meta-data.</p> <p>mapping_stats.tar.gz: Contains a file with the number of reads mapped to the honey bee gut microbiota database and the host genomes, for each sample. Bash-scripts are provided, detailing how the mapping was done and quantified.</p> <p>orthologs_phylogenies.tar.gz: Contains the pipeline for inferring orthologous gene-families and core genome phylogenies, as well as scripts for filtering of single-copy core gene families.</p> <p>assemblies.tar.gz: Contains the final de novo metagenome assembly files (contig fasta-files), gener<br> ated for both complete and rarefied read subsets. Bash-scripts detailing the assembly commands are also provided.</p> <p>SDP_validation.tar.gz: Contains the pipeline for metagenomic validation of candidate SDPs. Final output-files, containing the percentage identity of recruited metagenomic ORFs to database core genes, are provided for each candidate SDP. Additionally, a small example dataset is provided, where the intermediate result-files can be viewed.</p> <p>community_profiling.tar.gz: Contains the pipeline for community profiling, i.e. the quantification of individual community members (SDPs) across samples. Final output files are provided, including mapped read coverage on core gene families and corresponding plots. A small bam-file (containing data from a single subset sample), is also provided, in order to demonstrate the pipeline, together with all scripts used.</p> <p>snv_profiling.tar.gz: Contains the pipeline used for SNV profiling, including filtering and analysis. Final filtered vcf-files are provided for each SDP. Analytical output files are also provided, including data on shared SNV fractions, distance matrices, and cumulative curves.</p> <p>metagenomic_ORF_analyses.tar.gz: Contains the pipeline for analysis of metagenomic ORFs. This includes prediction of ORFs, clustering, annotation and functional characterization. ORF sequences, annotation files, and cluster-files are provided.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.