Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,109
datasets available to search
ShareScore release 0.9.0
Dataset results
3,109 results for “sequence analysis”
Fluorescent (C)LSM image sequences of Dictyostelium discoideum (Ax2 - LifeAct mRFP) for cell track and cell contour analysis
Open the record for dataset details and reuse information.
Supplementary material for: Phylogenetic analysis of allotetraploid species using polarized genomic sequences
Open the record for dataset details and reuse information.
DNA metabarcoding sequence data for diet analysis of caribou
Open the record for dataset details and reuse information.
Selection pressure analysis of dengue virus complete genome and E gene nucleotide sequences from Pakistan
Open the record for dataset details and reuse information.
Raw reads and metadata for 16S and 12S sequencing for microbiome and dietary analysis of Tasmanian devils
Open the record for dataset details and reuse information.
Dataset for: mRNA vaccine quality analysis using RNA sequencing
Open the record for dataset details and reuse information.
Bulk RNA sequencing analysis of Lin- leukemia BCR-ABL and BCR-ABL/MSI2-HOXA9 cells (post-transplantation)
Open the record for dataset details and reuse information.
Nanopore sequencing data analysis using Microsoft Azure cloud computing service
Open the record for dataset details and reuse information.
Phylogenetic and recombination analysis of adenovirus isolates reveals discordance between serotype and phylogeny: Multiple sequence alignments
Open the record for dataset details and reuse information.
Autism Detection Based on Eye Movement Sequences on the Web: A Scanpath Trend Analysis Approach
<p>We propose a novel approach to detect autism based on the eye-movement paths of users on the Web. This approach based on Scanpath Trend Analysis (STA) proposed by Eraslan et al. (2016, 2017). This dataset is created to provide supplementary data for our paper entitled "Autism Detection Based on Eye Movement Sequences on the Web: A Scanpath Trend Analysis Approach" presented at <a href="http://www.w4a.info/2020/">the 17th International Web for All Conference (W4A'20)</a>. This dataset includes all the individual paths used for the evaluation of our proposed approach. The dataset also includes the Python code to re-run the evaluation.</p> <p><strong>References:</strong></p> <ul> <li>Sukru Eraslan, Yeliz Yesilada, and Simon Harper. 2016. Scanpath Trend Analysis on Web Pages: Clustering Eye Tracking Scanpaths. ACM Transactions on the Web, (SCI-E), 10, 4, Article 20.</li> <li>Sukru Eraslan, Yeliz Yesilada, and Simon Harper. 2017. Engineering web-based interactive systems: trend analysis in eye tracking scanpaths with a tolerance. In Proceedings of the ACM SIGCHI Symposium on Engineering Interactive Computing Systems (EICS '17). ACM, New York, NY, USA, 3-8.</li> </ul>
data & analysis scripts of " Behavioral effects of rhythm, carrier frequency and temporal cueing on the perception of sound sequences"
<p>Analysis scripts and data accompanying the manuscript "Behavioral effects of rhythm, carrier frequency and temporal cueing on the perception of sound sequences"</p>
Database for mi-faser: Functional sequencing read annotation for high precision microbiome analysis
<p><strong>[Database for mi-faser]</strong></p> <p><strong>mi-faser: </strong><em>microbiome - functional annotation of sequencing reads</em></p> <p>A super-fast ( < 20min/10GB of reads ) and accurate ( > 90% precision ) method for annotation of molecular functionality encoded in sequencing read data without the need for assembly or gene finding.</p> <p>Web Service: http://services.bromberglab.org/mifaser/|<br> Repository: https://bitbucket.org/bromberglab/mifaser_base/</p>
NanoGalaxy: Nanopore long-read sequencing data analysis in Galaxy
<p>The data presented in "NanoGalaxy: A Galaxy tool kit with workflows for third-generation sequence analysis" to illustrate the functionality of the tools was obtained from: Wick, Ryan R., et al. "Completing bacterial genome assemblies with multiplex MinION sequencing." <em>Microbial genomics</em> 3.10 (2017).</p> <p>+</p> <p>Li, Ruichao, et al. "Efficient generation of complete sequences of MDR-encoding plasmids by rapid assembly of MinION barcoding sequencing data." <em>Gigascience</em> 7.3 (2018): gix132.</p>
standard with together) corner left bottom (analysis genetic the in included species among bold gene in shown oxidase-I are SE cytochrome and species the within at) % (divergence divergence sequence average The . pairwise) corner showing right upper (Matrix) %;. SE 4 ABLE (T error in Description of a new species of the Rhinolophus trifoliatus-group (Chiroptera: Rhinolophidae) from Southeast Asia
standard with together) corner left bottom (analysis genetic the in included species among bold gene in shown oxidase-I are SE cytochrome and species the within at) % (divergence divergence sequence average The . pairwise) corner showing right upper (Matrix) %;. SE 4 ABLE (T error
Sequenced-based paternity analysis to improve breeding and identify self-incompatibility loci in intermediate wheatgrass (Thinopyrum intermedium)
<p>In outcrossing species such as intermediate wheatgrass (IWG, Thinopyrum intermedium), polycrossing is often used to generate novel recombinants through each cycle of selection, but it cannot track pollen-parent pedigrees and it is unknown how self-incompatibility (SI) genes may limit the number of unique crosses obtained. This study investigated the potential of using next-generation sequencing to assign paternity and identify putative SI loci in IWG. Using a reference population of 380 individuals made from controlled crosses of 64 parents, paternity was assigned with 92% agreement using Cervus software. Using this approach, 80% of 4158 progeny (n = 3342) from a polycross of 89 parents were assigned paternity. Of the 89 pollen parents, 82 (92%) were represented with 1633 unique full-sib families representing 42% of all potential crosses. The number of progeny per successful pollen parent ranged from 1 to 123, with number of inflorescences per pollen parent significantly correlated to the number of progeny (r = 0.54, p < 0.001). Shannon's diversity index, assessing the total number and representation of families, was 7.33 compared to a theoretical maximum of 8.98. To test our hypothesis on the impact of SI genes, a genome-wide association study of the number of progeny observed from the 89 parents identified genetic effects related to non-random mating, including marker loci located near putative SI genes. Paternity testing of polycross progeny can impact future breeding gains by being incorporated in breeding programs to optimize polycross methodology, maintain genetic diversity, and reveal genetic architecture of mating patterns.</p>
Whole genome sequence analysis of porcine astroviruses reveals novel genetically diverse genotypes circulating in East African smallholder pig farms
<p>Supplementary materials for the porcine astrovirus study in East Africa.</p> <p><strong>Table S1</strong>: Pairwise comparison of nucleotide sequence identities of the complete (near complete, U460) genomes of the seven (7) astrovirus field strains (bold) and with sequences of other astroviruses available in GenBank </p> <p><strong>Table S2</strong>. Summary of nucleotide sequence identity matrix of the capsid protein (ORF2) among the seven (7) astroviruses field strains (bold) and the known reference strains in the GenBank using Clustal Omega</p> <p><strong>Table S3</strong>. Summary of amino acid sequence identity matrix of the capsid protein (ORF2) among the 7 astroviruses field strains (bold) and the known reference strains in the GenBank using Clustal Omega</p> <p><strong>Table S4</strong>: Estimates of evolutionary divergence between the East African PoAstVs and selected known AstV in the GenBank based on the amino acid sequences of complete ORF2 protein. The number of amino acid differences per site from between sequences is shown. Standard error estimate(s) are shown above the diagonal for our strains.</p> <p><strong>Table S5</strong>. Recommended potential linear antigenic epitopes predicted inside capsid protein (ORF2) of our field strains by SVMTriP web-based tool and corresponding antigenicity predicted by VaxiJen software</p>
Transcriptome analysis of WT versus H2A.J-KO MEFs for the paper entitled: The H2A.J histone variant contributes to Interferon-Stimulated Gene expression in senescence by its weak interaction with H1 and the derepression of repeated DNA sequences
<p>Abstract for overall study:</p> <p>The histone variant H2A.J was previously shown to accumulate in senescent human fibroblasts with persistent DNA damage to promote inflammatory gene expression, but its mechanism of action was unknown. We show that H2A.J accumulation contributes to weakening the association of histone H1 to chromatin and increasing its turnover. Decreased H1 in senescence is correlated with increased expression of some repeated DNA sequences, increased expression of STAT/IRF transcription factors, and transcriptional activation of Interferon-Stimulated Genes (ISGs). The H2A.J-specific Val-11 moderates the transcriptional activity of H2A.J, and H2A.J-specific Ser-123 can be phosphorylated in response to DNA damage with potentiation of its transcriptional activity by the phospho-mimetic S123E mutation. Our work demonstrates the functional importance of H2A.J-specific residues and potential mechanisms for its function in promoting inflammatory gene expression in senescence.</p> <p>Specific description for this dataset:</p> <p>We further tested a role for H2A.J in Interferon-Stimulated Gene expression by analyzing the transcriptome of WT and H2A.J MEFs induced into senescence by etoposide. TruSeq stranded DNA libraries were prepared from polyA-selected RNA and sequenced as 43 bp paired-end reads. The fastq sequences were mapped to Gencode.vM24.transcripts.fa.gz (GRCm38 transcriptome) with salmon. Read counts were then aggregated to the gene level with tximeta, and differential gene expression was analysed with DESeq2, edgeR, and limma-voom. Gene set enrichment analysis was performed with camera.</p> <p>The transciptomes of senescent WT and H2AFJ-KO showed strong separation from proliferating MEFs, and a weaker separation distinguished WT and H2A.J-KO MEFs. Strikingly, gene set enrichment analysis indicated highly significant defects in Interferon Response Gene Expression in the H2A.J-KO MEFs in senescence with significant down-regulation in senescent H2A.J-KO cells of a series of oligoadenylate synthase genes (Oas1g, Oas1a, Oasl1, Oas2, Oasl2) and several ISGs. Thus, H2A.J also contributes to ISG expression in the heterologous context of senescent MEFs.</p>
Kraken analysis from shotgun sequenced museum specimens via krona plot visualization
<p>This dataset contains html files with krona plots from Kraken analysis on shotgun sequenced museum specimens.</p>
Data from: Estimation of a killer whale (Orcinus orca) population's diet using sequencing analysis of DNA from feces
Estimating diet composition is important for understanding interactions between predators and prey and thus illuminating ecosystem function. The diet of many species, however, is difficult to observe directly. Genetic analysis of fecal material collected in the field is therefore a useful tool for gaining insight into wild animal diets. In this study, we used high-throughput DNA sequencing to quantitatively estimate the diet composition of an endangered population of wild killer whales (Orcinus orca) in their summer range in the Salish Sea. We combined 175 fecal samples collected between May and September from five years between 2006 and 2011 into 13 sample groups. Two known DNA composition control groups were also created. Each group was sequenced at a ~330bp segment of the 16s gene in the mitochondrial genome using an Illumina MiSeq sequencing system. After several quality controls steps, 4,987,107 individual sequences were aligned to a custom sequence database containing 19 potential fish prey species and the most likely species of each fecal-derived sequence was determined. Based on these alignments, salmonids made up >98.6% of the total sequences and thus of the inferred diet. Of the six salmonid species, Chinook salmon made up 79.5% of the sequences, followed by coho salmon (15%). Over all years, a clear pattern emerged with Chinook salmon dominating the estimated diet early in the summer, and coho salmon contributing an average of >40% of the diet in late summer. Sockeye salmon appeared to be occasionally important, at >18% in some sample groups. Non-salmonids were rarely observed. Our results are consistent with earlier results based on surface prey remains, and confirm the importance of Chinook salmon in this population's summer diet.
Data from: A pragmatic approach to the analysis of diets of generalist predators: the use of next-generation sequencing with no blocking probes
Predicting whether a predator is capable of affecting the dynamics of a prey species in the field implies the analysis of the complete diet of the predator, not simply rates of predation on a target taxon. Here, we employed the Ion Torrent next-generation sequencing technology to investigate the diet of a generalist arthropod predator. A complete dietary analysis requires the use of general primers, but these will also amplify the predator unless suppressed using a blocking probe. However, blocking probes can potentially block other species, particularly if they are phylogenetically close. Here, we aimed to demonstrate that enough prey sequence could be obtained without blocking probes. In communities with many predators, this approach obviates the need to design and test numerous blocking primers, thus making analysis of complex community food webs a viable proposition. We applied this approach to the analysis of predation by the linyphiid spider Oedothorax fuscus in an arable field. We obtained over two million raw reads. After discarding the low-quality and predator reads, the libraries still contained over 61 000 prey reads (3% of the raw reads; 6% of reads passing quality control). The libraries were rich in Collembola, Lepidoptera, Diptera and Nematoda. They also contained sequences derived from several spider species and from horticultural pests (aphids). Oedothorax fuscus is common in UK cereal fields, and the results showed that it is exploiting a wide range of prey. Next-generation sequencing using general primers but without blocking probes provided ample sequences for analysis of the prey range of this spider and proved to be a simple and inexpensive approach.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.