Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,109
datasets available to search
ShareScore release 0.9.0
Dataset results
3,109 results for “sequence analysis”
Comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity using Oxford Nanopore sequencing
<p><span>Metagenomics has become a prominent technology for studying the functional potential of all organisms in a microbial and eukaryotic community. The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. To achieve this, we have developed a universal PCR assay that targets the most conservative nuclear regions of the ribosomal gene for all cellular organisms, including plants, algae, fungi, protists, insects</span>,<span> and animals. The amplification product contains polymorphic regions of both ribosomal genes and the intergenic spacer. The size of the PCR products varies by class, kingdom</span>,<span> or domain, ranging from 2 kb for fungi to 7 kb for birds. This assay is also adapted for use with the Oxford Nanopore Rapid Barcoding Library Kit, which enables metagenomic biodiversity analysis. Our approach provides a rapid, sensitive</span>,<span> and equally efficient way to study the composition of eDNA from mixed species in the environment. This protocol reduces the time and cost of metagenomic biodiversity analysis using Oxford Nanopore sequencing. We can efficiently analyze the biodiversity of mixed species present in environmental samples.</span></span></p>
Engineering DszC Mutants from Transition State Macrodipole Considerations and Evolutionary Sequence Analysis
<p>Gaussian input and output files required to reproduce the results and analyses in the work.</p> <p>Code required to conduct analysis and representation.</p>
Oxford Nanopore sequencing for comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity
<p><span>The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here, we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. </span></span></p>
NASTRA: Accurate analysis of short tandem repeat markers by nanopore sequencing with repeat-structure-aware algorithm
<p><span>Forensic short-tandem repeats (STR) genetic markers are multi-allelic and widely utilized for individual identification, kinship testing, and cell-line authentication. Nanopore sequencing, known for its portability, is emerging as a promising approach for STR typing, facilitating real-time and in-field testing. However, its efficacy is often hampered by sequencing noise. Previous methods rely on alignment-based genotyping, necessitating known alleles, which limits their applicability to unknown alleles. Here, we introduced NASTRA, an innovative allele reference-free tool for precise germline analysis of STR genetic markers. NASTRA incorporates a recursive algorithm to infer repeat structures of allele sequences using only known repeat motifs. Our tests, conducted on 80 individual samples and 8 DNA standards, have demonstrated NASTRA's exceptional 100% accuracy in genotyping nearly all diploid STRs across various multiplex kits and flow cells. It surpasses alignment-based methods in accuracy and speed. In a paternity testing case study, NASTRA accurately identified three relationships among six individuals within an 18-minute sequencing duration. These results underscore NASTRA's ability to perform STR analysis on both NGS and nanopore sequencing platforms, significantly enhancing the utility of nanopore sequencing in relevant applications.</span></p>
MUFFIN : A suite of tools for the analysis of functional sequencing data - Example input data
<p>This repository contains the data required to run the example notebooks and to reproduce the figures from the paper : </p> <div> <div><strong>MUFFIN : A suite of tools for the analysis of functional sequencing data</strong></div> </div> <div><em>Pierre de Langen, Benoit Ballester</em></div> <div>bioRxiv 2023.12.11.570597; doi: <a href="https://doi.org/10.1101/2023.12.11.570597" target="_blank" rel="noopener">https://doi.org/10.1101/2023.12.11.570597</a></div> <div> </div> <div>Source code is located here :</div> <div><a href="https://github.com/pdelangen/Muffin" target="_blank" rel="noopener">https://github.com/pdelangen/Muffin</a></div> <div> </div> <ul> <li><strong>10k_pbmc_gene/ </strong>contains the data for 10k pbmc dataset in standard 10x sparse count table format.</li> <li><strong>genome_annot/</strong> contains gencode v38 and chromosomes for the human (used for gene set enrichment analyses)</li> <li><strong>GO_files/</strong> contains gene set information retrieved from the g:ProfileR website.</li> <li><strong>immune_chip/ </strong>contains the data required to re-run the ChIP-seq analyses, it will require to also launch the dl_data.smk to retrieve the data from ENCODE.</li> <li><strong>tcga_atac/</strong> contains the sample-genomic region ATAC tag count table, as well as the sample metadata and a gene set file of cancer hallmark genes.</li> <li><strong>scATAC/</strong> contains the cell barcode-genomic region ATAC tag count table, as well as the barcode metadata and the 10k pbmc dataset pre-analyzed in h5 AnnData format.</li> </ul> <div> </div>
Associated code and data for "A Practical Guideline for MicroRNA Sequencing Data Analysis in Chronic Lymphocytic Leukemia (doi: 10.1007/978-1-0716-4290-0_18)".
<p>This deposit contains the data, code, and analysis to recreate the results in the manuscript - Tuulikki Suomela, Liang Zhang, Julio Vera, Heiko Bruns, Xin Lai. A Practical Guideline for MicroRNA Sequencing Data Analysis in Chronic Lymphocytic Leukemia. Methods Mol. Biol., 2883, 403–426. <a href="https://www.researchgate.net/publication/387267721_A_Practical_Guideline_for_MicroRNA_Sequencing_Data_Analysis_in_Chronic_Lymphocytic_Leukemia">https://doi.org/10.1007/978-1-0716-4290-0_18</a>.</p> <p>The pipeline allows users to perform end-to-end analysis of bulk miRNA sequencing data, including quality control of FastQ files, mapping of read counts to miRNA genes using miRBase or Reference genome, quantification of miRNA read counts, differential gene expression analysis using DEseq2, gene set enrichment analysis using curated cancer hallmark gene sets, and identification of miRNA targets.</p> <p>If you have used the code for your research, please cite the original publication. Thank you very much.</p>
FIG. 2 in Foraminiferal biostratigraphy, facies and sequence stratigraphy analysis across the K-Pg Boundary in Hazara, Lesser Himalayas (Dhudial Section)
FIG. 2. — Generalized stratigraphy of the study area (Shah 2009).
GapA, RecA and DnaX sequences used for species delineation and phylogenetic analysis
<p>Survey of soft rot Pectobacteriaceae along a river stream highlights various ecological behaviour among species: GapA, RecA and DnaX sequences used for species delineation and phylogenetic analysis (fasta file format).</p>
Supplementary Materials from the article Characterization and molecular evolution analysis of Periploca forrestii inferred from its complete chloroplast genome sequence
<p>Table S1. Base composition of chloroplast genome in <em>P. forrestii</em>, Table S2. The lengths of introns and exons for the splitting genes, Table S3. The GC content of the codons from <em>P. forrestii </em>chloroplast genome, Table S4. Preferred codons in chloroplast genome of <em>P. forrestii</em>, Table S5. Long repeat sequences in the <em>P. forrestii </em>chloroplast genome, Figure S1. Codon bias analysis of P. forrestii chloroplast genome. (A) Neutrality plot analysis; (B) Analysis of PR2 bias plot; (C) Analysis on ENC and GC3 relationship.</p>
DNA metabarcoding sequence data for diet analysis of caribou
<p>Woodland caribou (<em>Rangifer tarandus caribou</em>) are threatened in Canada due to the drastic decline in population size caused primarily by human-induced landscape changes that decrease habitat and increase predation risk. Conservation efforts have largely focused on reducing predators and protecting critical habitat, whereas research on dietary niches and the role of potential food constraints in lichen-poor environments is limited. To improve our understanding of dietary niche variability, we used a next-generation sequencing approach with metabarcoding of DNA extracted from faecal pellets of woodland caribou located on Lake Superior in lichen-rich (mainland) and lichen-poor (island) environments. Amplicon sequencing of fungal ITS2 region revealed lichen-associated fungi as predominant in samples from both populations, but amplification at the chloroplast <em>trnL </em>region, which was only successful on island samples, revealed primary consumption of yew based on relative read abundance (<em>Taxus spp.</em>; 83.68%) with dogwood (<em>Cornus spp</em>.; 9.67%) and maple (<em>Acer spp.</em>; 4.10%) also prevalent. These results suggest that conservation efforts for caribou need to consider the availability of food resources beyond lichen to ensure successful outcomes. More broadly, we provide a reliable methodology for assessing ungulate diet from archived faecal pellets that could reveal important dietary shifts over time in response to climate change.</p>
Row fastq files generated by 16S rRNA sequencing for the metabarcoding analysis of Hidalgo-Villeda et al. paper
<p>Fastq files used for the microbiome profiling of the terminal ileal and caecum content.</p> <p>Sequencing of the terminal ileal content was performed at the Institute Hospitalo-Universitaire Méditerannée Infection, Marseille, France.</p> <p>Sequencing of the caecum content was performed at Genoscreen, Lille, France</p> <p>The sequencing methodology is described on the methods part of the paper "A physiological mouse model of severe acute malnutrition reveals prolonged dysbiosis and altered immunity under nutritional intervention, Hidalgo-Villeda et al.".</p> <p>The metadata.xlsx table present the description and correspondance of each fastq files.</p>
Nanopore sequencing data analysis using Microsoft Azure cloud computing service
<p>Genetic information provides insights into the exome, genome, epigenetics and structural organisation of the organism. Given the enormous amount of genetic information, scientists are able to perform mammoth tasks to improve the standard of health care such as determining genetic influences on outcome of allogeneic transplantation. Cloud-based computing has increasingly become a key choice for many scientists, engineers and institutions as it offers on-demand network access and users can conveniently rent rather than buy all required computing resources. With the positive advancements of cloud computing and nanopore sequencing data output, we were motivated to develop an automated and scalable analysis pipeline utilizing cloud infrastructure in Microsoft Azure to accelerate HLA genotyping service and improve the efficiency of the workflow at lower cost. In this study, we describe (i) the selection process for suitable virtual machine sizes for computing resources to balance between the best performance versus cost-effectiveness; (ii) the building of Docker containers to include all tools in the cloud computational environment; (iii) the comparison of HLA genotype concordance between the in-house manual method and the automated cloud-based pipeline to assess data accuracy. In conclusion, the Microsoft Azure cloud-based data analysis pipeline was shown to meet all the key imperatives for performance, cost, usability, simplicity and accuracy. Importantly, the pipeline allows for the ongoing maintenance and testing of version changes before implementation. This pipeline is suitable for data analysis from MinION sequencing platforms and could be adopted for other data analysis application processes.</p>
Exome sequence analysis identifies rare coding variants associated with a machine learning-based marker for coronary artery disease.
<p>*.sh and *.R are codes to test rare coding variants for association with ISCAD.</p> <p>Petrazzini_etal_2024_*_level_meta_analysis.txt.gz are summary statistics of variant- and gene-level associations of rare coding variants in the exome sequences of 604,914 individuals with an in-silico score for coronary artery disease (ISCAD).</p> <p>Chromosomal positions are mapped to the GRCh38 (hg38) human genome reference.</p> <p>Directions of effect correspond to associations in the UK Biobank, the All of Us Research Program, the BioMe Biobank sample 1 and the BioMe Biobank sample 2, in that order.</p>
Data from: Targeted genotyping-by-sequencing of potato and data analysis with R/polyBreedR
<p>"Mid-density" targeted genotyping-by-sequencing (GBS) combines trait-specific markers with thousands of genomic markers at an attractive price for linkage mapping and genomic selection. A 2.5K targeted GBS assay for potato was developed using the DArTag<sup>TM</sup> technology and later expanded to 4K targets. Genomic markers were selected from the potato Infinium<sup>TM</sup> SNP array to maximize genome coverage and polymorphism rates. The DArTag and SNP array platforms produced equivalent dendrograms in a test set of 298 tetraploid samples, and 83% of the common markers showed good quantitative agreement, with RMSE (root-mean-squared-error) less than 0.5. DArTag is suited for genomic selection candidates in the clonal evaluation trial, coupled with imputation to a higher-density platform for the training population. Using the software polyBreedR, an R package for the manipulation and analysis of polyploid marker data, the RMSE for imputation by linkage analysis was 0.15 in a small half-diallel population (N=85), which was significantly lower than the RMSE of 0.42 with the Random Forest method. Regarding high-value traits, the DArTag markers for resistance to potato virus Y, golden cyst nematode, and potato wart appeared to track their targets successfully, as did multi-allelic markers for maturity and tuber shape. In summary, the potato DArTag assay is a transformative and publicly available technology for potato breeding and genetics.</p>
Selection pressure analysis of dengue virus complete genome and E gene nucleotide sequences from Pakistan
<p>This dataset comprises 43 E gene and 44 complete genome nucleotide sequences of the dengue virus from serotypes DENV-1 to DENV-4, representing all documented sequences in Pakistan to date, sourced from the Virus Pathogen Resource (ViPR) database and NCBI. The E gene is critical as it is involved in serotype changes of the dengue virus, making it a pivotal target for understanding shifts in viral pathogenicity and immune escape mechanisms. The aim of compiling this dataset is to facilitate comprehensive genetic analysis and enhance understanding of the evolutionary dynamics of the dengue virus within the region. To assess the evolutionary pressures acting on these sequences, we conducted a selection pressure analysis utilizing computational methods. These methods include the Single Likelihood Ancestor Counting (SLAC), Fixed Effects Likelihood (FEL), adaptive Branch Site Random Effects Likelihood (aBSREL), Mixed Effects Model of Evolution (MEME), and the Genetic Algorithm for Recombination Detection (GARD), all implemented in the HyPhy software package. Our analysis focused on identifying genomic sites under both positive and negative selection pressures, providing insights into the adaptive evolutionary processes affecting the E gene of the dengue virus in Pakistan. Understanding the molecular evolution of this gene is crucial for predicting serotype evolution, potentially aiding in the development of effective vaccines and therapeutic strategies.</p>
Phylogenetic and recombination analysis of adenovirus isolates reveals discordance between serotype and phylogeny: Multiple sequence alignments
<p><strong>Background</strong></p> <p>Human adenovirus (HAdV) infections are caused by seven mastadenovirus species (A-G) and are the source for a variety of pathologies including gastrointestinal, respiratory, neurological, and ocular disease. While HAdV-D is the most common cause of adenovirus ocular infections, human adenoviruses B and E have also been isolated from the eye. </p> <p><strong>Results</strong></p> <p>In the course of classifying three new atypical ocular adenovirus samples, taken from the vitreous humor, we found that all three isolates were HAdV-B species, with isolate BP-AdV1 sorting with the B1 clade, and isolates BP-AdV2 and BP-AdV3 grouping into the B2 clade. The three Bascom Palmer HAdV-B genomes were then combined with over 300 HAdV-B genome sequences, including 9 ocular HAdV-B genome sequences. The whole genome phylogenetic analysis showed that 9 of the 11 ocular sequences grouped into the B1 clade, forming two clusters within B1. Attempts to categorize the penton, hexon and fiber serotypes using phylogeny of the three Bascom Palmer samples were inconclusive due to incongruence between serotype and phylogeny in the dataset. Recombination analysis using a subset of HAdV-B strains to generate a hybridization network detected recombination between non-human primate and human derived strains, recombination between one HAdV-B strain and the HAdV-E outgroup and limited recombination between the B1 and B2 clades. </p> <p><strong>Conclusions</strong></p> <p>The discordance between serotype and phylogeny detected in this study suggests that the current penton/hexon/fiber-based classification mechanism does not accurately describe the natural history and phylogenetic relationships amongst adenoviruses. A new adenovirus strain classification strategy may be beneficial to the field.</p>
Analysis of public single-cell sequencing database of COVID lung samples
<p>Lung endothelial cells from three published scRNA-seq datasets (GSE122960, GSE149878, GSE171668) of healthy subjects and COVID-19 patients were collected for further integrative analyses. The endothelial cells were classified into three sub-groups according to their distinguished expression of IL7R, DKK2, and EDNRB. For differential analysis of gene expression, counts per million of aggregated UMIs in each group were adopted in Wilcoxon rank-sum test.</p>
YfaL sequence analysis
<p><strong><span>Table S6A</span></strong><span>- Table showing parameters related to each yfaL sequence identified. The contig identifier and position of the sequence in it (column A), the strain name (column B),<span> </span>the phylogroup (column C), the size of each structural part of the sequence (column D to H), the number of mutations in these parts (column I to M), the related mutation frequency (column N to R). </span><strong><span>Table S6B</span></strong><span>- Table showing the number of yfaL sequences in each phylogroup as well as the related mutation frequences in the whole sequence<span> </span>and for each structural part of it. </span></p>
Original NGS dataset from publication "Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center".
<p>Original hepatitis C virus hypervariable region 1 NGS sequences in fastq format from patients analyzed in the study "Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center". </p> <p> </p>
Data from: Individual Movement - Sequence Analysis Method (IM-SAM): characterising spatio-temporal patterns of animal trajectories across scales and landscapes
<p>Dataset included in Zenodo supports the analyses performed in "<em>Individual Movement - Sequence Analysis Methods (IM-SAM) characterising spatio-temporal patterns of animal trajectories across scales and landscapes.</em>"</p> <p>The dataset includes one RDS file, that can be easily loaded into R using the readRDS function. The RDS file consists out of a list including two objects per animal:</p> <ul> <li>Object 1 contains a data frame with the real and simulated sequences for an animal. e.g., ls[[1]][[1]] </li> <li>Object 2 contains the home range in raster format of an animal. e.g., ls[[1]][[2]]</li> </ul> <p>The data frames in object 1 contain real habitat use sequences and corresponding simulated habitat use sequences generated in the home range of the specific individual (900 simulated sequences: 6 habitat selection rules x 3 selection coefficients x 50 repetitions). Open and closed habitats are respectively encoded by 0 and 1. The first 96 columns of each row in a data frame represent a 16-day habitat use sequence, with a fixed 4-hour relocation interval (0, 4, 8, 12, 16 and 20h). Column names are named as follows: Day_1_0h, Day_1_4h,..., Day_16_20h. In the next columns we provide the selection coefficients (columns 97-99), the habitat selection rules (or pattern, columns 100-102) and the number of missing values (mvs, columns, 103-104) for each of the real and simulated sequences. Note that simulated sequences have no missing values (i.e. values are always 0.00) and for real sequences there is no selection coefficient or habitat selection rule (i.e. values are always xxx).</p> <p>Rownames of simulated sequences are composed out of the habitat selection rule (c, o, a24, a33, a42 and u), the selection coefficient (5, 10, 50) and the replicate (1 to 50), separated by dashes. For example, the first simulated sequence in the first data frame (ls[[1]][[1]][1,]) is described as a24_10_1. The rownames of real sequences instead are composed out of the individuals' identifier, the biweekly period (1 to 23) and the year. For example, the first real sequence in the first data frame (ls[[1]][[1]][901,]) is described as 1_5_2006.</p> <p><br> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.