Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,574
datasets available to search
ShareScore release 0.9.0
Dataset results
1,574 results for “genome sequencing”
Whole genome sequence data of Lactiplantibacillus plantarum IMI507027
<p>The present data files are the annotation output from the whole genome sequencing of the strain Lactiplantibacillus plantarum IMI507027. </p>
Sequencing of individual barcoded cDNAs on Pacific Biosciences and Oxford Nanopore technologies reveals platform-specific error patterns (repository for Genome Research paper, 2022)
<p>Simulated ONT and PacBio RNA-Seq data for "Sequencing of individual barcoded cDNAs on Pacific Biosciences and Oxford Nanopore technologies reveals platform-specific error patterns" paper (Mikheenko et al., Genome Research, 2022). All details can be found in the Methods section of the paper.</p> <p><strong>PacBio.simulated_uniform_coverage.fasta.gz</strong> and <strong>ONT.simulated_uniform_coverage.fasta.gz files</strong> were used in Supplemental Note “Benchmarking of the read-to-isoform assignment algorithm”.</p> <p><strong>ONT.simulated_real_expression.fasta.gz</strong> file and all GTF files were used in the Section "Splice site correction improves transcript discovery precision". <strong>mouse.gencode.M26.spatial.15percent.reduced.gtf</strong> was used as the annotation file for all tools. <strong>mouse.gencode.M26.spatial.15percent.expressed.gtf </strong>contains the set of all expressed isoforms. <strong>mouse.gencode.M26.spatial.15percent.expressed_kept.gtf</strong> contains those of the isoforms that are in presented in the annotation file ("known" transcripts), <strong>mouse.gencode.M26.spatial.15percent.reduced.gtf</strong> contains expressed isoforms that were removed from the annotation ("novel" transcripts).</p>
Processed MinION genome sequencing data for strains ILHA G3AG5 and ILHA G3AA5
<p>Strain G3AA5 datasets include:</p> <p>Galaxy2750, Galaxy2754, Galaxy3522, Galaxy3523, Galaxy3524, Galaxy3541</p> <p> </p> <p>Strain G3AG4 datasets include:</p> <p>Galaxy2404, Galaxy2408, Galaxy3517, Galaxy3518, Galaxy3519, Galaxy3539</p>
Supplementary Materials from the article Characterization and molecular evolution analysis of Periploca forrestii inferred from its complete chloroplast genome sequence
<p>Table S1. Base composition of chloroplast genome in <em>P. forrestii</em>, Table S2. The lengths of introns and exons for the splitting genes, Table S3. The GC content of the codons from <em>P. forrestii </em>chloroplast genome, Table S4. Preferred codons in chloroplast genome of <em>P. forrestii</em>, Table S5. Long repeat sequences in the <em>P. forrestii </em>chloroplast genome, Figure S1. Codon bias analysis of P. forrestii chloroplast genome. (A) Neutrality plot analysis; (B) Analysis of PR2 bias plot; (C) Analysis on ENC and GC3 relationship.</p>
Genome-wide sequencing identifies a thermal tolerance related synonymous mutation in the mussel Mytilisepta virgata
<p><span>The roles of 'silent' synonymous mutations for organisms adapting to stressfully thermal environments are of fundamental biological and ecological interests but poorly understood. To study whether synonymous mutations influence the thermal adaptation of animals to heat stress at specific microhabitats, a genome-wide genotype-phenotype association analysis was carried out in the black mussels <em>Mytilisepta virgata</em> inhabiting different microhabitats. A synonymous mutation of Ubiquitin-specific Peptidase 15 (<em>MvUSP15</em>) was significantly associated with the physiological upper thermal limit of the mussel. The individuals carrying GG genotype (the G-type) at the mutant locus owned significantly lower heat tolerance compared to the individuals carrying GA and AA genotype (the A-type). Furthermore, when heated to sublethal temperature, the G-type exhibited higher inter-individual variations in the <em>MvUSP15 </em>expression, especially for the mussels on the sun-exposed microhabitats. Taken together, a synonymous mutation in <em>MvUSP15 </em>can affect the gene expression profile and interact with microhabitat heterogeneity to influence thermal resistance. This integrative study sheds light on the ecological importance of adaptive synonymous mutations as an underappreciated genetic buffer against heat stress and emphasizes the importance of integrative studies with the consideration of genetic, physiological, and environmental heterogeneity at a microhabitat scale for evaluating and predicting the impacts of climate change.</span></p>
First large-scale quantification study of DNA preservation in insects from natural history collections using genome-wide sequencing
<p>Insect declines are a global issue with significant ecological and economic ramifications. Yet we have a poor understanding of the genomic impact these losses can have. Genome-wide data from historical specimens has the potential to provide baselines of population genetic measures to study population change, with natural history collections representing large repositories of such specimens. However, an initial challenge in conducting historical DNA data analyses, is to understand how molecular preservation varies between specimens. Here, we highlight how Next Generation Sequencing methods developed for studying archaeological samples can be applied to determine DNA preservation from only a single leg taken from entomological museum specimens, some of which are more than a century old. An analysis of genome-wide data from a set of 113 red-tailed bumblebee (Bombus lapidarius) specimens, from five British museum collections, was used to quantify DNA preservation over time. Additionally, to improve our analysis and further enable future research we generated a novel assembly of the red-tailed bumblebee genome. Our approach shows that museum entomological specimens are comprised of short DNA fragments with mean lengths below 100 base pairs (BP), suggesting a rapid and large-scale post-mortem reduction in DNA fragment size. After this initial decline, however, we find a relatively consistent rate of DNA decay in our dataset, and estimate a mean reduction in fragment length of 1.9bp per decade. The proportion of quality filtered reads mapping our assembled reference genome was around 50 %, and decreased by 1.1 % per decade. We demonstrate that historical insects have significant potential to act as sources of DNA to create valuable genetic baselines. The relatively consistent rate of DNA degradation, both across collections and through time, mean that population level analyses - for example for conservation or evolutionary studies - are entirely feasible, as long as the degraded nature of DNA is accounted for. </p>
Data and processing scripts for PRISM barcode sequencing data used in "Massively parallel pooled screening reveals genomic determinants of nanoparticle-cell interactions"
<p>Sequencing data for the PRISM barcodes generated after nano-particle treatment is presented in this repository alongside the code to process the sequencing counts to generate the binning probabilities and weighted scores. <br> <br> For the details please see the original publication or the bioarxiv preprint: https://doi.org/10.1101/2021.04.05.438521<br> <br> The raw data is provided in PILOT_DATA_COUNTS.csv and EXPERIMENT_DATA_COUNTS.csv files, for the pilot and the actual experiment. <br> <br> For each of these files an R script is provided to process them, along with the output of the scripts (PILOT_DATA_PROBABILITIES.csv and EXPERIMENT_DATA_PROBABILITIES.csv)</p>
Long-read-based draft genome sequence of Indian black gram IPU-94-1 'Uttara': Insights into disease resistance and seed storage protein genes
<p>Black gram [Vigna mungo (L.) Hepper var. <em>mung<a>o</a></em>] [LAV1] is a warm-season legume highly prized for its protein content along with significant folate and iron proportions. To expedite the genetic enhancement of black gram, a high-quality draft genome from the center of origin of the crop is indispensable. Here, we established a draft genome sequence of an Indian black gram cultivar, 'Uttara' (IPU 94-1), known for its high resistance to mungbean yellow mosaic virus. Pacific Biosciences of California, Inc. (PacBio) single-molecule real-time (SMRT) and Illumina sequencing assembled a draft reference-guided assembly with a cumulative size of ~454.4 Mb, of which, 444.4 Mb was anchored on 11 pseudomolecules corresponding to 11 chromosomes. Uttara assembly denotes features of a high-quality draft genome illustrated through high N50 value (42.88 Mb), gene completeness (benchmarking universal single-copy ortholog [BUSCO] score 94.17%), and low levels of ambiguous nucleotides (N) percent (0.0005%). Gene discovery using transcript evidence predicted 28,881 protein-coding genes, from which, ~95% were functionally annotated. A global survey of genes associated with disease resistance revealed 119 nucleotide binding site–leucine rich repeat (NBS-LRR) proteins, while 23 genes encoding seed storage proteins (SSPs) were discovered in black gram. A large set of microsatellite loci were discovered for marker development in the crop. Our draft genome of an Indian black gram provides the foundational genomic resources for the improvement of important agronomic traits and ultimately will help in accelerating black gram breeding programs.</p>
Whole genome sequence variant discovery from Ethiopian Boran, N'Dama and Holstein cattle
<p>Whole genome sequence variants (SNPs) from forty samples of each Ethiopian Boran, N'Dama and Holstein cattle breeds which were utilised in the Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle study by Riggio et al., 2022, (in review). The NDama sequences included samples from Guinea (n =21), Nigeria (n =10) and Senegal (n = 9); and the Boran samples from Ethiopia (n = 30) and Kenya (n = 10). Variants discovery followed the protocol described in the Material and Method section of the paper.</p>
Draft Genome sequencing of Nocardia sp. strain WB46 Isolated from Salix purpurea Growing in a Site Chronically Contaminated by Petroleum Hydrocarbons
<p>Assembled contigs of the genome of <em>Nocardia</em> sp. strain WB46 isolated from <em>the rhizosphere of Salix purpurea </em>growing in a site chronically contaminated by petroleum hydrocarbons located at Varennes, Qubec, Canada.</p>
Data and scripts for the manuscript of svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data
<p>This upload include data and scripts supporting the results described in the manuscript of <em>svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data</em><em>. </em>Detailed description of the contents can be found in README.txt.</p>
Genome-resolved metagenomics using short-, long-read and metaHiC sequencing
<p>Reference-quality metagenome-assembled genomes (MAGs) are the key to exploring microbial compositions and microbe-phenotype associations. They can be recovered by different sequencing technologies and computational tools, which need an unbiased and comprehensive assessment to identify best practices. This work systematically evaluates 40 distinct strategies to recover high-quality MAGs generated by eight assemblers, eight metagenomics binners, and four sequencing technologies, including short-, long-read and metaHiC sequencing. We notice that the hybrid assemblies of short- and long-reads outperform either short- or long-read assemblies and generate more contigs with high contiguity. When the hybrid assemblies are combined with metaHiC-based binning (Hybrid-HiC), more high-quality MAGs with higher taxonomic diversity are recovered, and more tRNA and rRNA genes, phages, plasmids and antibiotic resistance genes are identified in the mock, simulated and real datasets.</p>
Single cell whole genome sequencing from Funnell, O'Flanagan, Williams et al
<p>This repository provides the processed data necessary to reproduce the results from: "Single cell genomic variation induced by mutational processes in cancer<strong> </strong><em>Funnell, O’Flanagan, Williams et al</em>"</p> <p>This includes the following:</p> <ul> <li>Single cell whole genome sequencing <ul> <li>Allele specific copy number profiles</li> <li>SNV counts per cell</li> <li>Structural variant counts per cell</li> <li>QC metrics</li> <li>clone assignments</li> <li>phylogenetic trees computed with sitka</li> <li>benchmarking results vs other methods</li> </ul> </li> <li>bulk whole genome sequencing <ul> <li>copy number profiles</li> <li>SNVs</li> </ul> </li> <li>10X single cell RNA sequencing <ul> <li>count matrices</li> <li>seurat Rdata objects</li> </ul> </li> <li>analysis tables <ul> <li>downstream processed results used to generate figures</li> </ul> </li> <li>oxford nanopore <ul> <li>phasing results</li> </ul> </li> </ul> <p> </p> <p>For further information please feel free to get in touch with Marc Williams (william1 [at] mskcc.org)</p> <p> </p>
New methods for the genotyping of Legionella pneumophila - Establishment, validation and implementation of a DNA-based microarray and a core genome multilocus sequence typing
<p>This data presented here are part a doctoral thesis with the focus on new genotyping methods for the human pathogen <em>Legionella pneumophila</em>. The data are partially published in articles. </p> <p>The thesis can be downloaded: update of the URL is coming soon</p>
Genotypes of Aedes aegypti mosquitoes derived from SNP chip and low-coverage whole genome sequencing for platform cross-validation
<p>The mosquito <em>Aedes aegypti </em>is the primary vector of many human arboviruses such as dengue, yellow fever, chikungunya, and Zika, which affect millions of people world-wide. Population genetics studies on this mosquito have been important in understanding its invasion pathways and success as a vector of human disease. The Axiom aegypti1 SNP chip was developed from a sample of geographically diverse <em>Ae. aegypti </em>populations to facilitate genomic studies on this species. Here we evaluate the utility of the Axiom aegypti1 SNP chip for population genetics and compare it with a low-depth shot-gun sequencing approach using mosquitoes from the species' native (Africa) and invasive range (outside Africa). These analyses indicate that the results from the SNP chip are highly reproducible and have a higher sensitivity to capture alternative alleles than a low-coverage whole-genome sequencing approach. Although the SNP chip suffers from ascertainment bias, results from population structure, ancestry, demographic, and phylogenetic analyses using the SNP chip were congruent with those derived from low coverage whole genome sequencing, and consistent with previous reports on Africa and outside Africa populations using microsatellites. More importantly, we identified a subset of SNPs that can be reliably used to generate merged databases, opening the door to combined analyses. We conclude that the Axiom aegypti1 SNP chip is a convenient, more accurate, low-cost alternative to low-depth whole genome sequencing for population genetic studies of <em>Ae. aegypti</em> that do not rely on full allelic frequency spectra. Whole genome sequencing and SNP chip data can be easily merged, extending the usefulness of both approaches. </p>
Genome sequence resources from three isolates of the apple canker pathogen Neonectria ditissima infecting forest trees
<p><em>Neonectria ditissima</em> is a generalist ascomycete plant pathogen causing canker diseases on a variety of hardwood tree species and can cross-infect many of them. The fungus enters the plants through wounds throughout the year. <em>N. ditissima</em> is considered a major threat to apple production responsible for the fruit tree canker disease which damages trees and causes rotting of fruits in storage. Nearby forests and shelter belts can serve as source of inoculum for well-managed apple orchards. Thus, knowledge about the <em>N. ditissima</em> isolates infecting different host species is essential for designing integrated pest management strategies. Here, we describe the genomes of three <em>N. ditissima</em> isolates, Nd_iso34, Nd_iso35, and Nd_iso36, infecting European beech, American tulip tree, and American beech, respectively. We obtained genome assemblies of ca. 45 megabases for all isolates, covering 94% of the <em>N. ditissima</em> reference annotation, and 97% of the universal single-copy orthologs (BUSCOs). We conclude that these genome assemblies are a highly relevant resource considering the scarcity of genomic data available for <em>N. ditissima</em>.</p>
Selection pressure analysis of dengue virus complete genome and E gene nucleotide sequences from Pakistan
<p>This dataset comprises 43 E gene and 44 complete genome nucleotide sequences of the dengue virus from serotypes DENV-1 to DENV-4, representing all documented sequences in Pakistan to date, sourced from the Virus Pathogen Resource (ViPR) database and NCBI. The E gene is critical as it is involved in serotype changes of the dengue virus, making it a pivotal target for understanding shifts in viral pathogenicity and immune escape mechanisms. The aim of compiling this dataset is to facilitate comprehensive genetic analysis and enhance understanding of the evolutionary dynamics of the dengue virus within the region. To assess the evolutionary pressures acting on these sequences, we conducted a selection pressure analysis utilizing computational methods. These methods include the Single Likelihood Ancestor Counting (SLAC), Fixed Effects Likelihood (FEL), adaptive Branch Site Random Effects Likelihood (aBSREL), Mixed Effects Model of Evolution (MEME), and the Genetic Algorithm for Recombination Detection (GARD), all implemented in the HyPhy software package. Our analysis focused on identifying genomic sites under both positive and negative selection pressures, providing insights into the adaptive evolutionary processes affecting the E gene of the dengue virus in Pakistan. Understanding the molecular evolution of this gene is crucial for predicting serotype evolution, potentially aiding in the development of effective vaccines and therapeutic strategies.</p>
Figure 1 in Genetic difference between two Schistosoma japonicum isolates with contrasting cercarial shedding patterns revealed by whole genome sequencing
Figure 1. Map of geographical locations of research sites.
Single Nucleotide Polymorphisms (SNPs) identified from the whole genome sequences of hilsa shad (Tenualosa ilisha) of the Bay of Bengal
<p>The data file contains 792,939 isolated SNPs identified by discoSnp++ v2.3.x (Uricaru et al., 2015) from the whole genome sequence of T. ilisha of the Bay of Bengal. The central sequence of length 2k-1 is seen in upper case, while the flanking sequences are seen in lower case. SNP_higher/lower: one of the two alleles. id: id of the SNP (each SNP has a unique id).</p> <p>FOR SNPs:</p> <p>P_i:pos_Alt1/Alt2: Information about a ith SNP (If more than a unique SNP is found, the following format is used: P_1:pos_Alt1/Alt2,P_2:pos_Alt1/Alt2,...</p> <p>pos: position of the SNP with respect to the starting position of the bubble, i.e. the starting of the upper case sequence.</p> <p>Alt1: One of the two alleles</p> <p>Alt2: the other</p> <p>FOR INDELs:</p> <p>P_1:pos_size_repeatSize</p> <p>pos: predicted position of the indel with respect to the starting position of the bubble, i.e. the starting of the upper case sequence.</p> <p>size: predicted size of the indel</p> <p>repeatSize: Size of the longest sequence both prefix of the indel and prefix of the sequence located just after the insertion.</p> <p>high/low: sequence complexity. If the sequence if of low complexity (e.g. ATATATATATATATAT) this variable would be low</p> <p>nb_pol: number of polymorphism.</p> <p>left_unitig_length: size of the full left extension.</p> <p>right_unitig_length: size of the right extension.</p> <p>left_contig_length: size of the full left extension.</p> <p>right_contig_length: size of the right extension.</p> <p>C1: number of reads mapping the central upper case sequence from the first read set.</p> <p>C2: number of reads mapping the central upper case sequence from the second read set.</p> <p>Q1 [if reads were given in fastq]: average phred quality of the central nucleotide from the mapped reads from the first read set.</p> <p>Q2 [if reads were given in fastq]: average phred quality of the central nucleotide from the mapped reads from the second read set.</p> <p>G1: Genotype of the variant in the first read set.</p> <p>G2: Genotype of the variant in the second read set.</p> <p>rank: ranks the predictions according to their read coverage in each condition favoring SNPs that are discriminant between conditions.</p>
Amino acid sequences of the proteins predicted from the whole genome of hilsa shad (Tenualosa ilisha) of the Bay of Bengal
<p>Gene prediction was performed by AUGUSTUS (Stanke et al., 2006) from the whole genome sequence of <em>T. ilisha</em> (<a href="https://www.ncbi.nlm.nih.gov/bioproject/400122">PRJNA400122</a>). The data contain amino acid sequences of 37,450 predicted protein coding genes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.