Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

250

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

250 results for “Bioinformatics”

Learn how ShareScore rates datasets ↗
zenodo48/100

Benchmarking bioinformatic tools for amplicon-based sequencing of norovirus

<p>This repository contains associated datasets and accession numbers for a study entitled &#39;<strong>Benchmarking bioinformatic tools for amplicon-based sequencing of norovirus&#39;</strong>. The scripts for this project can be found on the GitHub project<a href="https://github.com/ahfitzpa/Benchmarking-bioinformatics-norovirus-amplicons">&nbsp;page</a>.&nbsp;</p> <p>Expected composition tsv files are the OTU tables for each simulation performed (001-010). OTU IDs in this case are the expected taxonomy with the&nbsp;associated accession numbers. Samples are numbered 1-40, including the simulation number. Expected sequences fasta files contain the sequences used as input for each simulation, without primers or Illumina adapter sequences.</p> <p>Amplicons were generated using the following primers:</p> <p><strong>GI Primers&nbsp;</strong><br> GISKF: CTG CCC GAA TTY GTA AAT GA 4<br> GISKR: CCA ACC CAR CCA TTR TAC A 5<br> <br> <strong>GII Primers&nbsp;</strong><br> G2SKF: CNT GGG AGG GCG ATC GCAA 8<br> G2SKR: CCR CCN GCA TRH CCR TTR TAC AT</p> <p>In this study, three databases and multiple classifiers were compared. Here we include the taxonomy and fasta files for each database; noronet =NoroNet RIVM, calicinet= HuCat CDC and custom, randomly generated database. Fasta files for the classifiers include the GI/GII primers listed above in a 5-3 orientation.&nbsp;</p> <p>The tags.txt file&nbsp;contains the Illumina adapters used for the simulation component of the study.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Dataset for genes linked to gastroschisis along with bioinformatics analysis

<p>The dataset consists of records from genes linked to gastroschisis. Genes displaying statistical significance with gastroschisis (excluding those genes undergoing adjusted calculations with covariates) were selected and manually curated for further bioinformatics analysis.</p> <p>Figure 1 illustrates the systematic review of the literature search strategy and selection criteria, publishing crude (unadjusted) genes linked to gastroschisis from January 1, 1990 to August 2, 2020.&nbsp;&nbsp;</p> <p>The full list of tables is described in the file READ ME and remains available in CSV files.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Dataset for publication: An inter-laboratory study characterizes the impact of bioinformatic approaches on genome-based cluster detection for foodborne bacterial pathogens

<p>This dataset is part of a dry-lab interlaboratory study conducted across Germany, regarding bacterial outbreak detection based on NGS data, with a focus on bioinformatic analysis of four species to identify potential variability caused by different data analysis approaches and human interpretation. Participants were asked to follow their usual in-house protocols while adhering to the general guidelines. A quality assessment (with sample exclusion) was followed by 7-gene Multilocus-Sequence Typing (MLST), core genome Multilocus Sequencing Typing (cgMLST), and SNP calling. The participants were then asked to identify clusters. The study was not intended to resemble a standard proficiency test with a passing/failing grade, but rather to investigate and quantify obvious variability in the results and, where possible, the reasons for it. For this purpose, the datasets included borderline cases in terms of quality.</p>

opencc-by-4.0Oct 2025View details →
zenodo44/100

The use of Foundational Ontologies in Bioinformatics - Supplementary Material

<p>Supplementary material for the paper &quot;The use of Foundational Ontologies in Bioinformatics&quot;.</p>

opencc-byDec 2021View details →
zenodo44/100

Bioinformatics for public health microbiologists: Module 1 dataset (South African Salmonella)

<p>Module 1 (WGS) dataset (paired-end Illumina reads of <em>Salmonella enterica&nbsp;</em>strains isolated from animals and animal products in South Africa)</p> <ul> <li> <p>module1_dataset_ZAsalmonella.tar.gz: raw Illumina paired-end reads</p> </li> <li> <p>trimmed_reads.tar.gz: trimmed Illumina paired-end reads (i.e., trimmed via fastp v0.23.4)</p> </li> <li> <p>contigs.tar.gz: assembled genomes (i.e., trimmed reads assembled into contigs using SKESA v2.5.1)</p> </li> <li> <p>prokka.tar.gz: whole-genome annotation results (i.e., produced via Prokka v1.14.6)</p> </li> <li> <p>enterobase_salmonella.tar.gz: publicly available assembled genomes (downloaded via Enterobase; https://enterobase.warwick.ac.uk/, accessed 1 June 2024)</p> </li> <li> <p>snippy_input.tsv: input file used for Snippy (https://github.com/tseemann/snippy)</p> </li> <li> <p><span>snippy_final.tar.gz: output files produced by Snippy (https://github.com/tseemann/snippy), Gubbins (https://github.com/nickjcroucher/gubbins), and SNP-sites (https://github.com/sanger-pathogens/snp-sites)</span></p> </li> </ul>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Bioinformatic pipeline: Genomic diversity landscape of the honey bee gut microbiota

<p>This data-set describes the full bioinformatic pipeline used to analyze 54 metagenomic samples of the honey bee gut microbiota. Each sample was isolated from an individual honey bee, and all samples originate from two colonies of the Engel laboratory at the University of Lausanne, Switzerland. The full raw data-set is available from the sequence-read archive: SRP150166.</p> <p>A publication based on this analysis is currently under review, with the title: &quot;Genomic diversity landscape of the honey bee gut microbiota&quot;, and an upload to Biorxiv is also underway.</p> <p>The data-set contains tar-balls for the different main workflows of the analysis. Dowload and unpack to view the contents (tar -zxvf filename.tar.gz). For each workflow, all directories contain README.txt files, describing the contents of the directory. Due to size constraints, some intermediate files have been omitted, and some workflows are demonstrated for a subset of the data. However, the full analysis can be reproduced from the raw data, using the provided scripts.</p> <p>Scripts are included within workflow directories, and are also provided as a separate tar-ball for convenience. All perl-scripts come with documentation, which can be viewed by typing: &quot;perl script_name.pl -h&quot;. For R scripts, the usage is indicated as a comment in the top lines of each script. Note that many of the scripts require specific input-files to be present in the run-directory. Their usage is demonstrated within the workflow directories in bash-scripts (*.sh). Commands used for generating plots and some statistics are given within workflow directories in text-files &quot;R.commands&quot; when applicable.</p> <p>Aside from custom code, the pipeline also utilizes various open-source Software packages, which are detailed in the file &quot;software_dependencies.txt&quot;. Note, while many of the scripts will run fast on any computer, some steps of the pipeline are computationally demanding, and will require significant computing time, as well as storage space. When scripts are known to be time-consuming, this is indicated in the script help message.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

PhasAGE Training School 1 -Overview of bioinformatics tools for the life sciences & Classification and evolution of non-globular proteins- LECTUREs

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

EAW FASTQ files for bioinformatic courses (16S rRNA genes, 2018)

<p>Selection of 6 samples, with 6 Forward (R1) and 6 Reverse (R2) files, including primers. R1 and R2 reads are ca. 300 bp long, and were obtained from Illumina MiSeq technologies, at the FEM facility sequencing platform. The files refer to the16S rRNA gene reads obtained from the analyses carried out on the samples collected and filtered (Sterivex<sup>TM</sup> 0.22 &micro;m) in different areas and depths of Lake Garda on September, 2018 (see EAW_2018_FASTQ_16S_description.docx).</p> <p>Sampling and analyses were carried out in the framework of the project Eco-AlpsWater (ASP569), funded by the Interreg Alpine Space program.</p> <p>A bioinformatic protocol for analyzing these&nbsp;data using DADA2 is available in Zenodo:</p> <pre>https://doi.org/10.5281/zenodo.5232772 </pre>

opencc-by-4.0Aug 2021View details →
zenodo44/100

EAW FASTQ files for bioinformatic courses (18S rRNA genes, 2018)

<p>Selection of 6 samples, with 6 Forward (R1) and 6 Reverse (R2) files, including primers. R1 and R2 reads are ca. 300 bp long, and were obtained from Illumina MiSeq technologies, at the FEM facility sequencing platform. The files refer to the18S rRNA gene reads obtained from the analyses carried out on the samples collected and filtered (Sterivex<sup>TM</sup> 0.22 &micro;m) in different areas and depths of Lake Garda on September, 2018 (see EAW_2018_FASTQ_18S_description.docx).</p> <p>Sampling and analyses were carried out in the framework of the project Eco-AlpsWater (ASP569), funded by the Interreg Alpine Space program.</p> <p>A bioinformatic protocol for analyzing these&nbsp;data using DADA2 is available in Zenodo:</p> <pre>https://doi.org/10.5281/zenodo.5233527</pre> <p>The corresponding 16S rRNA gene reads, obtained from the same eDNA extracts, are saved in <a href="https://doi.org/10.5281/zenodo.5215815">https://doi.org/10.5281/zenodo.5215815</a></p>

opencc-by-4.0Aug 2021View details →
dryad40/100

A two-tier bioinformatic pipeline to develop probes for target capture of nuclear loci with applications in Melastomataceae

<p><b><i>Premise of the study</i></b><b>: </b>Putatively single-copy nuclear (SCN) loci, identified using genomic resources of closely related species, are ideal for phylogenomic inference. However, suitable genomic resources are not available for many clades, including Melastomataceae. We introduce a versatile approach to identify SCN loci for clades with few genomic resources and use it to develop probes for target enrichment<i> </i>in the distantly related <i>Memecylon</i> and <i>Tibouchina</i> (Melastomataceae).</p> <p><b><i>Methods</i></b>: We present a two-tiered pipeline. First, we identified putatively SCN loci using MarkerMiner and transcriptomes from distantly related species in Melastomataceae. Published loci and genes of functional significance were added (384 total loci). Second, using HybPiper, we retrieved 689 homologous template sequences for these loci using genome-skimming data from within the focal clades.</p> <p><b><i>Results</i></b>: We sequenced 193 loci from both <i>Memecylon</i> and <i>Tibouchina</i>, with probes designed from 56 template sequences successfully targeting sequences in both clades. Probes designed from genome-skimming data within a focal clade were more successful than probes designed from other sources.</p> <p><b><i>Discussion: </i></b>Our pipeline successfully identified and targeted SCN loci in <i>Memecylon </i>and <i>Tibouchina</i>, enabling phylogenomic studies in both clades and potentially across Melastomataceae. This pipeline could be easily applied to other clades with few genomic resources. </p>

opencc-zeroJan 2021View details →
zenodo40/100

command line interface practical - 2022 bioinformatics summer school

<p>Files and folders for the comman line interface practical of the 2022 bioinformatics summer school. Second release</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Figure 7. T in Bioinformatics and expression analysis of the Xeroderma Pigmentosum complementation group C (XPC) of Trypanosoma evansi in Trypanosoma cruzi cells

Figure 7. T. cruzi growth assessment after cisplatin treatment (300 ΜM). (a) Wild type (WT). (b) TcXPC superexpressor (Tc-TcXPC). (c) TevXPC expressor (Tc-TevXPC). The solid lines represent the untreated cells, while the dotted lines represent the cells treated with cisplatin. Statistical student's t test: (*) On that point, cells treated with cisplatin presented a statistically significant lower growth in relation to untreated cells (p &lt;0.05). Representative results of three independent experiments.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 6 in Bioinformatics and expression analysis of the Xeroderma Pigmentosum complementation group C (XPC) of Trypanosoma evansi in Trypanosoma cruzi cells

Figure 6. Growth assessment of T. cruzi: wild type (WT), TcXPC superexpressor (Tc-TcXPC) and TevXPC expressor (Tc-TevXPC). Statistical student's t test: (*) On that point, only Tc-TevXPC presented a statistically significant lower growth in relation to WT (p &lt;0.05); (**) On that point, both Tc-TcXPC and Tc-TevXPC presented a significant lower growth in relation to WT (p &lt;0.05). All parasites were at same initial concentration, grown on LIT medium and were counted daily. Representative results of three independent experiments.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 5 in Bioinformatics and expression analysis of the Xeroderma Pigmentosum complementation group C (XPC) of Trypanosoma evansi in Trypanosoma cruzi cells

Figure 5. TevXPC amplification by RT-PCR with the cDNA from cell cultures. Lanes: (1) 1Kb DNA Ladder; (2) WT; (3) Tc-TcXPC; (4 and 5) Tc-TevXPC; (6) positive control (DNA from T. evansi); (7) negative control.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 4 in Bioinformatics and expression analysis of the Xeroderma Pigmentosum complementation group C (XPC) of Trypanosoma evansi in Trypanosoma cruzi cells

Figure 4. (a) TcXPB-R protein model. (b) TevXPB-R protein model. (c) TcXPB-R (blue) and TevXPB-R (orange) models overlay.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 2 in Bioinformatics and expression analysis of the Xeroderma Pigmentosum complementation group C (XPC) of Trypanosoma evansi in Trypanosoma cruzi cells

Figure 2. (a) Alignment between TcXPC and TevXPC proteins (mismatches highlighted) and its domains. Green: RAD4/PNGase transglutaminase-like fold. Blue: RAD4 beta-hairpin domain 1. Red: RAD4 beta-hairpin domain 2. Yellow: RAD4 beta-hairpin domain 3. (b) Candidate sequence motif involved in p62 interaction (highlighted by brown rectangle) found in TcXPC and TevXPC. This sequence is suggested based on the sequence motif described for Human XPC and yeast RAD4: D/E-F/W-E-D/E-V.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 1 in Bioinformatics and expression analysis of the Xeroderma Pigmentosum complementation group C (XPC) of Trypanosoma evansi in Trypanosoma cruzi cells

Figure 1. (a) TcXPC protein model. (b) TevXPC protein model. (c) TbXPC protein model (d) Model of TevXPC protein bound to a mismatch DNA. (e) TcXPC (red), TevXPC (blue) and TbXPC (green) models overlay. (f) Crystal structure of Rad4-Rad23 bound to a mismatched DNA performed by Min and Pavletich (2007).

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 3 in Bioinformatics and expression analysis of the Xeroderma Pigmentosum complementation group C (XPC) of Trypanosoma evansi in Trypanosoma cruzi cells

Figure 3. (a) TcXPB protein model. (b) TevXPB protein model. (c) TcXPB (green) and TevXPB (red) models overlay.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Machine learning and bioinformatics analysis of diagnostic biomarkers associated with the occurrence and development of lung adenocarcinoma

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo40/100

Bioinformatic databases survey

<h1>Bioinformatic databases survey</h1> <p>The dataset surveys bioinformatic databases published in the <a href="https://academic.oup.com/nar">NAR database issue</a> from 1995 to 2022. It evaluates the current number of citations and availability of each ressources.</p> <h2>Data content</h2> <p>The dataset is composed of two tables :</p> <p><strong>A. Databases table :</strong> Contains the information of each database published in the NAR database issue.</p> <ul> <li>db_id : Database ID in the dataset</li> <li>resource_name : Name(s) of the database</li> <li>current_access : Latest known web address of the database</li> <li>is_a_pun : The database name is a play on word</li> <li>available_2022 : The database was accessible online during the 2022 survey</li> <li>last_accessible_year : If not accessible, latest point in time where the database was found online (using the Internet web archive snapshots)</li> <li>unavailable_message : If not accessible, the message/error when trying to access the ressource</li> <li>year_first_publication : Year of first publication of the database</li> <li>year_last_publication : Year of latest publication of the database (including database update publications)</li> <li>total_citations_2022 : Cumulative number of citation for all articles of the database</li> <li>nb_authors_max : Maximum number of authors associated to any articles published for that database</li> <li>nb_articles_2022 : Number of articles published for that database in 2022</li> </ul> <p><strong>B. Articles table :</strong> Contains the information collected for the NAR articles</p> <ul> <li>collector : Person who contributed to add this database in the dataset</li> <li>article_global_id : DOI of the article surveyed</li> <li>db_id : Database ID of the ressource described in the article</li> <li>article_id : Article unique ID</li> <li>article_year : Article publication year</li> <li>Authors : list of authors of the article. Separated by ";"</li> <li>Author.ID : list of ORCID of the authors of the article. Separated by ";"</li> <li>Title : Title of the atricle</li> <li>Source.title : Journal name</li> <li>Volume : Volume number</li> <li>Issue : Issue number</li> <li>Funding.Details : Funding information of the article</li> <li>Funding.Text : Funding text provided by the authors</li> <li>PubMed.ID : Pubmed ID of the article</li> <li>citations_2016 : Number of citations of the article in 2016 (if published)</li> <li>citations_2022 : Number of citations of the article in 2022</li> <li>nb_authors : Number of authors in the article</li> <li>Index.Keywords : Keywords associated to the publication</li> </ul> <h2>Data sources</h2> <p>Note that the presented dataset leverage and expand on the dataset gathered and published in Imker, H.J., 2020. Who Bears the Burden of Long-Lived Molecular Biology Databases?. Data Science Journal, 19(1), p.8. The original dataset collected by Dr. Imker is available at : <a href="https://doi.org/10.13012/B2IDB-4311325_V1">https://doi.org/10.13012/B2IDB-4311325_V1</a>&nbsp;</p> <p>The dataset was collected and is maintained by undergraduate students of a CURE class (Course-based Undergraduate Research Experience) held at the University of Arizona. All students of the class have participated to the collection, update and curation the dataset that is available as a database and a web-portal at <a href="https://hurwitzlab.shinyapps.io/DS_Heroes/">https://hurwitzlab.shinyapps.io/DS_Heroes/</a>. Students could elect to be added or not as author to this Zenodo repository.</p> <p>The <a href="https://ur.arizona.edu/content/course-based-research-undergraduate-experiences-cures">CURE class BAT102</a> "<strong>Data Science Heroes: An undergraduate research experience in Open Data Science Practices"</strong> gives the students an opportunity to learn about open science and investigate open data practices in bioinformatics through a survey of the databases published in the NAR database issue.</p>

opencc-by-4.0Jul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record