Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,109

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,109 results for “sequence analysis”

Learn how ShareScore rates datasets ↗
zenodo36/100

Analysis results for association study of long-term kidney transplant rejection using whole-exome sequencing

<p>Association study results for long-term kidney transplant rejection. Single-variant association results are provided as Plink output files. Meta-analysis results are provided as METAL output files. FDR results are sorted lists of the top association result from random sample label permutations and are included with the plink and meta-analysis results. SKAT and GSEA results are provided for gene and pathway level analyses, respectively.</p>

opencc-by-4.0Oct 2018View details →
zenodo36/100

Data for 'Comparative Analysis of Single-Cell RNA Sequencing Methods'

<p>Raw sequencing data to &quot;Comparative Analysis of Single-Cell RNA Sequencing Methods&quot;.&nbsp;</p> <p>https://www.ncbi.nlm.nih.gov/pubmed/28212749</p> <p>&nbsp;</p> <p>In addition to the GEO submission&nbsp;https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE75790, you can find here raw bam files for UMI-methods tagged with cell barcode and UMI sequences.</p> <p>MD5 checksum:&nbsp;f10825509952fffd9c4dc0c1dcb9eb8e</p>

opencc-by-nc-sa-4.0Feb 2017View details →
zenodo36/100

Nanopore sequence analysis - Galaxy Training Material

<p>Twelve MDR plasmids harboring samples were prepared according to the MinION library construction protocols, followed by library sequencing. After 8 hours of sequencing run, a total of 287 725 reads ranging from dozens to tens of thousands of bases in length were obtained, covering a total of 493 Mbp. The raw data were subjected to several stages of processing, including basecalling, de-multiplexing, fasta sequence extraction. For this tutorial one out of the twelve samples is chosen as example.</p> <p>This dataset is extracted of a&nbsp;project studying the Efficient generation of complete sequences of MDR-encoding plasmids by rapid assembly of MinION barcoding sequencing data (<a href="https://doi.org/10.1093/gigascience/gix132">https://doi.org/10.1093/gigascience/gix132</a>)</p>

opencc-by-4.0Oct 2018View details →
zenodo36/100

Human sequence alignment data set used for analysis of SPDI algorithm and tools

<p>Collection of alignment segments produced on October 30, 2019.&nbsp;&nbsp;The ADS currently consists of over 2,680,000 pairwise alignment segments generated from over 350,000 distinct input sequences.&nbsp;&nbsp;</p> <ul> <li> <p>Old assembly to current Genome Reference Consortium (GRC) <a href="http://f1000.com/work/citation?ids=111899&amp;pre=&amp;suf=&amp;sa=0">(Church et al., 2011)</a> primary assemblies (e.g. GRCh36(hg18) or GRCh37(hg19) with GRCh38(hg38))</p> </li> </ul> <ul> <li> <p>Patches, alternative loci, or pseudoautosomal regions (PAR) to GRC primary assembly</p> </li> <li> <p>RefSeq <a href="http://f1000.com/work/citation?ids=2599029&amp;pre=&amp;suf=&amp;sa=0">(O&rsquo;Leary et al., 2016)</a> and select GenBank <a href="http://f1000.com/work/citation?ids=6183037&amp;pre=&amp;suf=&amp;sa=0">(Benson et al., 2018)</a> transcripts to selected RefSeq genomic regions, also known as RefSeqGene (NG), a member of the Locus Reference Genome (LRG) collaboration <a href="http://f1000.com/work/citation?ids=3225699&amp;pre=&amp;suf=&amp;sa=0">(Dalgleish et al., 2010)</a>.</p> </li> <li> <p>Current RefSeq transcripts (NM/NR/XM/XR) and RefSeq genomic regions (NG) to the latest Assembly</p> </li> <li> <p>Previous versions of NG and RefSeq transcripts (NM/NR) to GRC primary assembly</p> </li> </ul>

opencc-by-4.0Oct 2019View details →
zenodo36/100

Data for paper: Analysis and control of untemplated DNA polymerase activity for guided synthesis of kilobase-scale DNA sequences

<p>This deposit contains data for the paper entitled: "<strong>Analysis and control of untemplated DNA polymerase activity for guided synthesis of kilobase-scale DNA sequences</strong>"</p> <p>Contents include:</p> <ul> <li><strong>00_sequencing-metadata.xlsx</strong> - Metadata for each sequencing sample.</li> <li><strong>01_raw-fastqs.zip</strong> - Raw basecalled FASTQ data.</li> <li><strong>02_clean-fastas.zip</strong> - Cleaned FASTA files (removal of adapters and barcode sequences).</li> <li><strong>03_read-statistics.zip</strong> - Read statistics for all samples.</li> <li><strong>04_sequence-analysis.zip</strong> - Sequence analysis output for all samples.</li> <li><strong>05_qpcr-data.xlsx</strong> - qPCR data for all the dNTP mixes studied.</li> <li><strong>06_analysis-scripts.zip</strong> - Analysis scripts.</li> </ul>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Data accompanying "In silico analysis of the profilaggrin sequence indicates alterations in the stability, degradation route, and intracellular protein fate in filaggrin null mutation carriers" article.

<p>This research was supported by the National Science Centre, Poland, grant PRELUDIUM number 2021/41/N/NZ1/03473 to NS, National Science Centre, Poland, grant SONATA BIS number 2019/34/E/NZ6/00354 to DG-O, as well as POIR.04.04.00-00-21FA/16&ndash;00 grant, carried out within the First TEAM programme of the Foundation for Polish Science co-financed by the European Union under the European Regional Development Fund (awarded to DG-O). WP was supported by the National Science Centre, Poland, grant SONATA-BIS number 2021/42/E/NZ1/00190. SB is supported by a Wellcome Trust Senior Research Fellowship (220875/Z/20/Z).</p>

opencc-by-4.0May 2023View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on Human Genome Diversity Project Whole-Genome Sequencing Data

<p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the HGDP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Simons Genome Diversity Project</strong>: Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>2b388c1fa446ecec70e33ea0471e06f8</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>50b80ed32b1ae542c8967cc31986dd19</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>995f30b74c4bb094a674b1a994853246</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Copy Number</td> <td>e79e61efab4c491fa2825b7d1853df58</td> </tr> </tbody> </table> </div> <div>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</div>

opencc-by-4.0Sep 2024View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on Simons Genome Diversity Project Whole-Genome Sequencing Data

<p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the SGDP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Simons Genome Diversity Project</strong>: Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <table> <tbody> <tr> <th>Public Dataset</th> <th>EGP Result File Type</th> <th>MD5</th> </tr> </tbody> <tbody> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>86b09553f80926c1c29c57000ec1a88f</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>010026d77bee81e7b8daf5836bd12da3</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>f1ea3edf4a82b42f2028467fb3544dc4</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Copy Number</td> <td>4872eeb792c214ad49662e98e4b14620</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on 1000 Genomes Project 2504 Whole-Genome Sequencing Data

<p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the 1000 Genomes Project 2504 Dataset. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on 1000 Genomes Project 2504 Dataset</strong>:&nbsp;Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000G_2504_high_coverage/1000G_2504_high_coverage.sequence.index</code>. Please note that the index files are there as well. You just have to append a <code>.crai</code>. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>dbf39d6ff0e4389b900f9d985f2e6c64</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>4d53ef60ec16f3e4b566c45fdf0fb977</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Variant Tables</td> <td>16925b546051d37cce27df8ec57ccc5e</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Copy Number</td> <td>365c1b360ea327795d981356064658a6</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p> </div>

opencc-by-4.0Sep 2024View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on 1000 Genomes Project 698 Related Whole-Genome Sequencing Data

<div> <p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the 1000 Genomes Project 698 Related Dataset. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on 1000 Genomes Project 698 Related Dataset</strong>:&nbsp;Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://ftp-trace.ncbi.nlm.nih.gov/1000genomes/ftp/1000G_2504_high_coverage/additional_698_related/1000G_698_related_high_coverage.sequence.index</code>. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>322038d61b4da2e937b32410613c3532</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>36c782c12245100478903f7fa191a402</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Variant Tables</td> <td>68b2a51361ffae4e7ad9d420b8becd38</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Copy Number</td> <td>1e83c8ae132b0a7ef33b090757b29063</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p> </div> <p>&nbsp;</p> </div> <h2>&nbsp;</h2>

opencc-by-4.0Sep 2024View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on Gambian Genome Variation Project Whole-Genome Sequencing Data

<p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the GGVP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Gambian Genome Variation Project</strong>:&nbsp;Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>d21e1e91e8b4c00627171fae79a1f54d</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>b359d1068d4f84f7746d1ebde82df29a</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>ee2b93aa93d2177d92ec0f8f308b43ed</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome Copy Number</td> <td>fda509ba1d2bf33fd2d6b77b92e76c03</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p> </div>

opencc-by-4.0Sep 2024View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on GIAB Whole-Genome Sequencing Data

<div> <p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the GIAB. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on GIAB</strong>: Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://raw.githubusercontent.com/genome-in-a-bottle/giab_data_indexes/refs/heads/master/AshkenazimTrio/alignment.index.AJtrio_Illumina300X_wgs_novoalign_GRCh37_GRCh38_NHGRI_07282015</code></p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>5eac6ec7d36307aa401fd5441b38a506</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>1151ae74c8e515f4f39bef816bb55d6a</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Variant Tables</td> <td>b3e342fe9827df2e399f5685f84cd4dc</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Copy Number</td> <td>3c45c19f76f71b3ccaad155565ead5e4</td> </tr> </tbody> </table> </div> <div>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</div> <p>&nbsp;</p> <p>&nbsp;</p> </div> <h2>&nbsp;</h2>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Additional files for Horvath et al., 2024. Detection and classification of long terminal repeat sequences in plant LTR-retrotransposons and their analysis using explainable machine learning.

<p>Additional data for Horvath et al., 2024 (source code freeze, models, data, supplementary figures, tables and files(.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

NPIP: A Comprehensive Analysis Pipeline for Rapid Pathogen Detection in Clinical Samples Based on Nanopore Sequencing

<p>Background: Rapid and accurate pathogen detection is important for effective control of infectious diseases. However, traditional pathogen culture methods have a very long detection time, as well as high rates of false-positive and false-negative results. Third generation sequencing (TGS) technology brings the new possibility of being used as a pathogen detection method. However, the practicability of a pathogen detection report based on TGS is still lacking. There is also a lack of professional and accurate report interpretation.</p> <p>Results: Here, we report on the development of a pathogen detection and analysis tool (NPIP) based on third generation nanopore sequencing technology. We also prove the practicability of nanopore sequencing and NPIP analysis tools in emergency and clinical pathogen detection by demonstrating its use in a practical case.</p> <p>Conclusions: This platform provides an effective, convenient, and fast analysis tool for clinicians and public health personnel to more successfully apply TGS in pathogen detection.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Sequencing summaries & Nanopolish eventalign outputs for downstream analysis of yeast RNA and synthetic oligonucleotides

<p>This dataset contains:</p> <ol> <li>A set of text files from running the tool&nbsp;Nanopolish eventalign on several nanopore direct RNA sequencing data sets produced by Jay&nbsp;Hesselberth&#39;s&nbsp;lab at the University of Colorado (BioProject accession number&nbsp;<a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA910992">PRJNA910992</a>), as well as external sequencing data sets from&nbsp;<a href="https://pubmed.ncbi.nlm.nih.gov/34893601/">PMID: 34893601</a>&nbsp;(synthetic oligonucleotides from Leger et al) and&nbsp;<a href="https://pubmed.ncbi.nlm.nih.gov/35252946/">PMID: 35252946</a>&nbsp;(yeast rRNA data from Stephenson et al).</li> <li>&quot;Sequencing summary&quot; files produced by MinKNOW, from nanopore sequencing of yeast mRNA and synthetic RNA oligos&nbsp;in the Hesselberth lab. We analyze these files in an associated manuscript to determine the &quot;end status&quot; of each read during the sequencing run.</li> </ol> <p>These files can be used as inputs to the R markdown documents at&nbsp;<a href="https://github.com/hesselberthlab/RNARePore">https://github.com/hesselberthlab/RNARePore</a>&nbsp;to reproduce the figures in the associated manuscript.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Read coverage information for analysis missing plasmid sequences

<p>This file contains two directories: read_coverage, and read_coverage_contigs. This directories should be decompressed and located within&nbsp;ecoli-binary-classifier/2021_11_missing_sequences_analysis/results/ in order to reproduce results described the in the mansucript.</p>

opencc-by-4.0Jul 2023View details →
dryad36/100

Dataset for: mRNA vaccine quality analysis using RNA sequencing

<p>The success of mRNA vaccines has been realised, in part, by advances in manufacturing that enabled billions of doses to be produced at sufficient quality and safety. However, mRNA vaccines must be rigorously analysed to measure their integrity and detect contaminants that reduce their effectiveness and induce side-effects. Currently, mRNA vaccines and therapies are analysed using a range of time-consuming and costly methods. Here we describe a streamlined method to analyse mRNA vaccines and therapies using long-read nanopore sequencing. Compared to other industry-standard techniques, VAX-seq can comprehensively measure key mRNA vaccine quality attributes, including sequence, length, integrity, and purity. We also show how direct RNA sequencing can analyse mRNA chemistry, including the detection of nucleoside modifications. To support this approach, we provide supporting software to automatically report on mRNA and plasmid template quality and integrity. Given these advantages, we anticipate that RNA sequencing methods, such as VAX-seq, will become central to the development and manufacture of mRNA drugs.</p>

opencc-zeroAug 2023View details →
dryad36/100

Data from: Adaptive radiation of the Callicarpa genus in the Bonin Islands revealed through double-digest restriction site–associated DNA sequencing analysis

Open the record for dataset details and reuse information.

publicAug 2024View details →
dryad36/100

Data from: Targeted genotyping-by-sequencing of potato and data analysis with R/polyBreedR

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad36/100

Annotation and evolutionary analysis of chemosensory gene sequence data in the Colorado potato beetle, Leptinotarsa decemlineata

Open the record for dataset details and reuse information.

publicOct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record