Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,106

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

3,106 results for “Sequence Analysis”

Learn how ShareScore rates datasets ↗
zenodo44/100

BAMBI ITS - Analysis of the fungal component (via ITS amplicon sequencing) of stool samples from preterm babies

<p>Amplicon analysis of ITS amplicons from preterm babies.</p> <p>Associated GitHub repository: <a href="https://github.com/quadram-institute-bioscience/bambi-its">https://github.com/quadram-institute-bioscience/bambi-its</a></p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

RNA sequencing dataset for prediction of liver hepatocellular carcinoma using SIMON analysis

<p>The LIHC dataset was used for data mining and for the generation of machine learning model for the detection of liver hepatocellular carcinoma cells (LIHC) using the SIMON platform as described in the &quot;SIMON: open-source knowledge discovery platform&quot; publication (<a href="https://doi.org/10.1101/2020.08.16.252767">https://doi.org/10.1101/2020.08.16.252767</a>). The LIHC dataset was obtained from the <em>GSEABenchmarkeR</em> package ( <a href="https://doi.org/10.1093/bib/bbz158">https://doi.org/10.1093/bib/bbz158</a>) and it contains RNA expression data from 374 liver hepatocellular carcinoma (LIHC) cells and 50 adjacent normal cells.</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Transcriptome analysis of the effect of over-expressing H2A.J mutants in proliferating WI38 fibroblasts for the paper entitled: The H2A.J histone variant contributes to Interferon-Stimulated Gene expression in senescence by its weak interaction with H1 and the derepression of repeated DNA sequences

<p>Abstract for overall study:</p> <p>The histone variant H2A.J was previously shown to accumulate in senescent human fibroblasts with persistent DNA damage to promote inflammatory gene expression, but its mechanism of action was unknown. We show that H2A.J accumulation contributes to weakening the association of histone H1 to chromatin and increasing its turnover. Decreased H1 in senescence is correlated with increased expression of some repeated DNA sequences, increased expression of STAT/IRF transcription factors, and transcriptional activation of Interferon-Stimulated Genes (ISGs). The H2A.J-specific Val-11 moderates the transcriptional activity of H2A.J, and H2A.J-specific Ser-123 can be phosphorylated in response to DNA damage with potentiation of its transcriptional activity by the phospho-mimetic S123E mutation. Our work demonstrates the functional importance of H2A.J-specific residues and potential mechanisms for its function in promoting inflammatory gene expression in senescence.</p> <p>Specific description for this dataset:</p> <p>H2A.J differs from canonical H2A only by a valine at position 11 instead of alanine, and the 7 C-terminal amino acids containing a potential minimal phosphorylation site SQ for DNA-damage response kinases. To test the functional importance of these H2A.J-specific sequences, we mutated Val-11 to Ala as is found in all canonical H2A sequences, and we mutated Ser-123 to either Glu to mimic a phospho-serine residue or to Ala to prevent phosphorylation. We also substituted the C-terminus of H2A.J with the C-terminus of H2A. These mutants, WT-H2A.J and canonical H2A-type1 were ectopically expressed in proliferating fibroblasts, and their microarray transcriptomes were compared to that of proliferating and senescent fibroblasts without ectopic histone expression. Genome-wide transcriptome analysis indicated that senescent fibroblasts clustered distinctly from proliferating fibroblasts, and proliferating fibroblasts expressing the H2A.J-V11A and H2A.J-S123E mutants clustered distinctly from fibroblasts expressing the other H2A.J mutants, WT-H2A.J, and H2A. Hallmark gene set enrichment analysis of the transcriptomes of fibroblasts expressing H2A.J-V11A or H2A.J-S123E versus control proliferating fibroblasts indicated that they showed the same highly significant enrichment for the Epithelial-Mesenchyme Transition, TNF-Alpha Signaling Via NF-kB, and Inflammatory Response gene sets. Notable inflammatory genes including IL1A, IL1B, IL6, CXCL8, and CCL2 are contained in these gene sets and are often induced in senescence as part of the senescence-associated secretory phenotype. Heat maps showed that the H2A.J-V11A and H2A.J-S123E mutants were particularly apt at activating the expression of these inflammatory genes in proliferating fibroblasts</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Flow diagram for analysis of high-throughput sequencing data

<p>Tex code and resulting pdf image, summarising the data processing pipeline of high-throughput sequencing data (fastq format files), through mapping the data to a reference genome, and then discovery and genotyping of sequence variants. The latter stage uses both 'GATK Haplotype Caller' for smaller variants, such as single-nucleotide polymorphisms and insertion-deletion polymorphisms, and Genomestrip for variants such as deletions and duplications greater than 1000 nucelotide bases in length. Note that the flow diagram is intended to represent what steps were take in the study, and does not necessarily represent the current optimum methods.</p> <p>The manuscript for which this image is a part of can be found open-access at F1000 Research "Whole genome resequencing of a laboratory-adapted <em>Drosophila melanogaster </em>population sample" https://f1000research.com/articles/5-2644/v1 doi: 10.12688/f1000research.9912.1</p> <p> </p>

opencc-by-4.0Nov 2016View details →
zenodo44/100

IBP-database and environmental IBP sequences for functional analysis of microalgae

<p>Database for analysis of ice binding protein (IBP) sequences (Uhlig et al. (2015)):</p> <p>(1) DUF3494_seqs_Uniprot.fasta:  full length sequences with DUF3494 domain used for the calculation of the backbone tree in the phylogenetic placement</p> <p>(2) env_IBPs.fasta: potential IBP sequences from one Arctic and five Antarctic sea ice metatranscriptomes (Sanger or 454)</p> <p>(3) DUF3494_substree_fig2a_UniprotIDs.txt: UniProtIDs for subtree in Fig 2a</p> <p>(4) DUF3494_confirmed_IBPactivity_UniprotIDs.txt: UniProtIDs for sequences with confirmed IBP function of the protein</p> <p>If using this dataset please cite the following publication: Uhlig, C., Kilpert, F., Frickenhaus, S., Kegel, J.U., Krell, A., Mock, T., Valentin, K., Beszteri, B., (2015) The significance of antifreeze proteins for eukaryotic microbial communities of Arctic and Antarctic sea ice, The ISME Journal, 9, 2537–2540, doi:10.1038/ismej.2015.43</p>

opencc-by-4.0Aug 2017View details →
zenodo44/100

Sample data for analysis of sequence variation in HIV

<p>These are downsampled interleaved paired fastq datasets from Jair et. 2019 (<a href="https://doi.org/10.1371/journal.pone.0214820">https://doi.org/10.1371/journal.pone.0214820</a>). The datasets were prepared by:</p> <ol> <li>Downloading original data from NCBI SRA (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA517147)</li> <li>Trimming contaminating Nextera adapters using trim-galore</li> <li>Mapping reads against nxb2 reference of HIV genome (K03455.1) with BWA MEM</li> <li>Restricting mapped reads to&nbsp;<em>pol</em>&nbsp;gene vicinity (K03455.1:2000-5100)</li> <li>Downsampling mapped data to ~10% of the original with Picard&#39;s DownsampleSam</li> <li>Converting BAM to Interleaved Fastq with Picard&#39;s SamToFastq</li> <li>Gzipping resultant interleaved paired fastq files</li> </ol>

opencc-by-4.0Oct 2019View details →
zenodo44/100

Mitochondrial genome sequencing and analysis of the invasive Microstegium vimineum: a resource for systematics, invasion history, and management

<p>Table S1: Accession data for Microstegium samples included in this study.</p> <p>File S1: Alignment of Mitochondrial CDS for Poales mitochondrial sequences.</p> <p>File S2: SNP data for Microstegium vimineum mitochondrial variants.</p> <p>Figure S1: Transposable element content in the Microstegium vimineum mitogenome.</p> <p>Figure S2: Summary of Kraken2 output.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Supplementary dataset to publication: "Genomic insight into Campylobacter jejuni isolated from commercial turkey flocks in Germany using whole-genome sequencing analysis"

<p><em>Campylobacter jejuni </em>is a zoonotic bacterium of public health significance. The present investigation was designed to assess the epidemiology and genetic heterogeneity of <em>Campylobacter jejuni</em> recovered from commercial turkey farms in Germany using whole-genome sequencing. The Illumina MiSeq<sup>&reg;</sup> technology was used to sequence 66 <em>Campylobacter jejuni </em>isolates obtained between 2010 and 2011 from commercial meat turkey flocks located in ten German federal states. Phenotypic antimicrobial resistance was determined. Phylogeny, resistome, plasmidome and virulome profiles were analyzed using whole-genome sequencing data. Genetic resistancemarkers were identified with bioinformatics tools (AMRFinder, ResFinder, NCBI and ABRicate) and compared with the phenotypic antimicrobial resistance.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Training data for 'Exome sequencing data analysis' tutorial (Galaxy Training Material)

<p>The data used in this tutorial are a subset of the data&nbsp;published previously in&nbsp;<a href="https://zenodo.org/record/3243160">Training material for the course &quot;Exome analysis with GALAXY&quot;</a>. Credit for uploading the original data goes to&nbsp;Paolo Uva and Gianmauro&nbsp;Cuccuru!</p> <p>Specifically, you may need the following datasets for following the tutorial:</p> <p><strong>Raw sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/father_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/father_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R2.fq.gz</a></li> </ul> <p><strong>Premapped sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_father.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_father.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_mother.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_mother.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_proband.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_proband.bam</a></li> </ul> <p><strong>Reference sequence (human chromosome 8)</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz?download=1">https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz</a></li> </ul> <p>&nbsp;</p> <p>If you would just like to play with GEMINI rather than work through the full tutorial, you&#39;ll find below a prebuilt GEMINI database (for GEMINI version 0.20.1) for the family trio. You can start exploring this database without having to run GEMINI load&nbsp;and, in fact, without having to install GEMINI&#39;s bundled annotation data.</p>

opencc-by-4.0May 2019View details →
zenodo40/100

Genbank accession numbers of sequences used in the phylogenetic analysis

<p><em>Phylogenetic analysis: </em>Sequences were assembled using Lasergene v15 (DNASTAR, Inc. Madison, USA), and combined with sequences obtained from Genebank</p>

opencc-by-4.0Sep 2020View details →
zenodo40/100

FIGURES 34 ­ 37. Lamyctes hellyeri n in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes ­ Henicops Group

FIGURES 34 ­ 37. Lamyctes hellyeri n. sp. QVMAG 23: 23048, female, pretarsus of leg 14, scales 10 m. 34 ­ 36, anterior, posterior, and ventral views; 37, detail of lateral pore and ornament on scutes of main claw.

opencc-zeroDec 2003View details →
zenodo40/100

FIGURES 11 ­ 17. Lamyctes hellyeri n in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes ­ Henicops Group

FIGURES 11 ­ 17. Lamyctes hellyeri n. sp. 11, 14 ­ 17, QVMAG 23: 23046, female. 11, anterior part of head shield and basal part of antennae, scale 100 m; 14, sensilla on dorsal side of antenna, scale 10 m; 15 ­ 16, antennal articles, dorsal side, scales 50 m; 17, cephalic pleurite with Tömösváry organ, scale 50 m. 12 ­ 13, QVMAG 23: 23047, female. 12, ventral view of clypeus and labrum, scale 100 m; 13, labral midpiece and inner parts of sidepieces, scale 30 m.

opencc-zeroDec 2003View details →
zenodo40/100

FIGURES 1 ­ 4 in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes ­ Henicops Group

FIGURES 1 ­ 4. Lamyctes coeculus (Brölemann). 1, 3, AM KS 57961, female, Mellong Range, NSW, Australia. 2, 4, MCZ DNA 100472, female, Cerro San Javier, Tucumán, Argentina. 1 ­ 2, ventral view of head, scales 100 m; 3 ­ 4, dental margin of maxillipede coxosternite, scales 50 m.

opencc-zeroDec 2003View details →
zenodo40/100

FIGURES 18 ­ 25. Lamyctes hellyeri n in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes ­ Henicops Group

FIGURES 18 ­ 25. Lamyctes hellyeri n. sp. QVMAG 23: 23046, female. 18, ventral view of maxillipede, scale 100 m; 19 ­ 20, dental margin of maxillipede coxosternite, scales 50 m, 10 m; 21, tarsus and claw of second maxilla, scale 50 m; 22, distal part of tarsus and claw of second maxilla, scale 10 m; 23, coxal projections and telopods of first maxillae, scale 50 m; 24, first maxillae, scale 100 m; 25, plumose setae on inner margins of telopods of first maxillae, scale 10 m.

opencc-zeroDec 2003View details →
zenodo40/100

Data Set for the Journal Article "Automated Preparation of Nanoscopic Structures: Graph-Based Sequence Analysis, Mismatch Detection, and pH-Consistent Protonation with Uncertainty Estimates"

<p>This repository containes the data generated by ASAP and discussed in the journal article [Csizi, K.-S. and Reiher, M., 2023, arXiv:2307.16344], including Cartesian coordinates of training and test set molecules, and MD trajectories.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

FIG. 1 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations

FIG. 1.—Map of the Hawaiian Islands with collection sitesfor Hawaiian hoary bat tissues used inthis study. Sites with n&gt; 1 are denoted with an asterisk.

opencc-by-4.0Aug 2020View details →
zenodo40/100

FIG. 2 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations

FIG. 2.—PCA result plot showing clustering of individual bats from four Hawaiian Islands using 21,808,031 SNPs. Sample information included in supplementary table S4, Supplementary Material online.

opencc-by-4.0Aug 2020View details →
zenodo40/100

FIG. 4 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations

FIG. 4.—SNAPP-based phylogenetic tree inference. (A) The maximum clade credibility or consensus tree, showing approximate divergence of hoary bats across the Hawaiian archipelago. The axis on the bottom of the figure corresponds to million years before present (Ma), using the emergence of Hawai'i (~0.43 Ma) as a calibration point (95% confidence intervals were given in square brackets). (B) The drawing of all sampled trees showing all ingroup nodes were supported by maximum posterior probabilities (1.00).

opencc-by-4.0Aug 2020View details →
zenodo40/100

Data from: "Rare earth elements sediment analysis tracing anthropogenic activities in the stratigraphic sequence of Alagankulam (India)"

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo40/100

FIG. 3 in Foraminiferal biostratigraphy, facies and sequence stratigraphy analysis across the K-Pg Boundary in Hazara, Lesser Himalayas (Dhudial Section)

FIG. 3. — Lithostratigraphic column showing the lithology, constituents and facies of the Dhudial Section.

opencc-zeroSep 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record