Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,293
datasets available to search
ShareScore release 0.9.0
Dataset results
1,293 results for “gene sequencing”
Data from: Recombination-dependent replication and gene conversion homogenize repeat sequences and diversify plastid genome structure
Open the record for dataset details and reuse information.
Data from: Allele phasing has minimal impact on phylogenetic reconstruction from targeted nuclear gene sequences in a case study of Artocarpus
Open the record for dataset details and reuse information.
Data from: A universal probe set for targeted sequencing of 353 nuclear genes from any flowering plant designed using k-medoids clustering
Open the record for dataset details and reuse information.
Data from: Dealing with the adaptive immune system during de novo evolution of genes from intergenic sequences
Open the record for dataset details and reuse information.
Mutant Harmonia axyridis raw sequencing data of CRISPR/Cas9 induced mutations in the genes laccase2 and scarlet
Open the record for dataset details and reuse information.
Data from: High-throughput amplicon sequencing of rRNA genes requires a copy number correction to accurately reflect the effects of management practices on soil nematode community structure
Open the record for dataset details and reuse information.
cDNA sequence of E2 gene family in Arabidopsis thaliana and data of statistical analysis
Open the record for dataset details and reuse information.
Data for: Longitudinal metatranscriptomic sequencing of Southern California wastewater representing 16 million people from August 2020-21 reveals widespread transcription of antibiotic resistance genes
Open the record for dataset details and reuse information.
Sequences of microbial eukaryotic genes obtained from the metagenome of the Mariana Trench
<p>nonredundant_eukaryotic_genes.faa contains the non-redundant amino acid sequences of predicted eukaryotic genes in all samples by MetaEuk.</p> <p>nonredundant_eukaryotic_genes.fna contains the non-redundant nucleotide sequences of predicted eukaryotic genes in all samples by MetaEuk.</p> <p>tax_for_per_gene.txt contains the taxonomic information of per non-redundant gene.</p>
Data from: Low-coverage, whole-genome sequencing of Artocarpus camansi (Moraceae) for phylogenetic marker development and gene discovery
Premise of the study: We used moderately low-coverage (17×) whole-genome sequencing of Artocarpus camansi (Moraceae) to develop genomic resources for Artocarpus and Moraceae. Methods and Results: A de novo assembly of Illumina short reads (251,378,536 pairs, 2 × 100 bp) accounted for 93% of the predicted genome size. Predicted coding regions were used in a three-way orthology search with published genomes of Morus notabilis and Cannabis sativa. Phylogenetic markers for Moraceae were developed from 333 inferred single-copy exons. Ninety-eight putative MADS-box genes were identified. Analysis of all predicted coding regions resulted in preliminary annotation of 49,089 genes. An analysis of synonymous substitutions for pairs of orthologs (Ks analysis) in M. notabilis and A. camansi strongly suggested a lineage-specific whole-genome duplication in Artocarpus. Conclusions: This study substantially increases the genomic resources available for Artocarpus and Moraceae and demonstrates the value of low-coverage de novo assemblies for nonmodel organisms with moderately large genomes.
Data from: Congruent deep relationships in the grape family (Vitaceae) based on sequences of chloroplast genomes and mitochondrial genes via genome skimming
Vitaceae is well-known for having one of the most economically important fruits, i.e., the grape (Vitis vinifera). The deep phylogeny of the grape family was not resolved until a recent phylogenomic analysis of 417 nuclear genes from transcriptome data. However, it has been reported extensively that topologies based on nuclear and organellar genes may be incongruent due to differences in their evolutionary histories. Therefore, it is important to reconstruct a backbone phylogeny of the grape family using plastomes and mitochondrial genes. In this study, next-generation sequencing data sets of 27 species were obtained using genome skimming with total DNAs from silica-gel preserved tissue samples on an Illumina HiSeq 2500 instrument. Plastomes were assembled using the combination of de novo and reference genome (of V. vinifera) methods. Sixteen mitochondrial genes were also obtained via genome skimming using the reference genome of V. vinifera. Extensive phylogenetic analyses were performed using maximum likelihood and Bayesian methods. The topology based on either plastome data or mitochondrial genes is congruent with the one using hundreds of nuclear genes, indicating that the grape family did not exhibit significant reticulation at the deep level. The results showcase the power of genome skimming in capturing extensive phylogenetic data: especially from chloroplast and mitochondrial DNAs.
Data from: Targeted sequencing of venom genes from cone snail genomes improves understanding of conotoxin molecular evolution
To expand our capacity to discover venom sequences from the genomes of venomous organisms, we applied targeted sequencing techniques to selectively recover venom gene superfamilies and non-toxin loci from the genomes of 32 cone snail species (family, Conidae), a diverse group of marine gastropods that capture their prey using a cocktail of neurotoxic peptides (conotoxins). We were able to successfully recover conotoxin gene superfamilies across all species with high confidence (> 100X coverage) and used these data to provide new insights into conotoxin evolution. First, we found that conotoxin gene superfamilies are composed of 1-6 exons and are typically short in length (mean = ~85bp). Second, we expanded our understanding of the following genetic features of conotoxin evolution: (a) positive selection, where exons coding the mature toxin region were often three times more divergent than their adjacent noncoding regions, (b) expression regulation, with comparisons to transcriptome data showing that cone snails only express a fraction of the genes available in their genome (24%-63%), and (c) extensive gene turnover, where Conidae species varied from 120-859 conotoxin gene copies. Finally, using comparative phylogenetic methods, we found that while diet specificity did not predict patterns of conotoxin evolution, dietary breadth was positively correlated with total conotoxin gene diversity. Overall, the targeted sequencing technique demonstrated here has the potential to radically increase the pace at which venom gene families are sequenced and studied, reshaping our ability to understand the impact of genetic changes on ecologically relevant phenotypes and subsequent diversification.
Data from: Plasticity of promoter-core sequences allows bacteria to compensate for the loss of a key global regulatory gene
Transcription regulatory networks (TRNs) are of central importance for both short-term phenotypic adaptation in response to environmental fluctuations and long-term evolutionary adaptation, with global regulatory genes often being targets of natural selection in laboratory experiments. Here, we combined evolution experiments, whole-genome resequencing, and molecular genetics to investigate the driving forces, genetic constraints, and molecular mechanisms that dictate how bacteria can cope with a drastic perturbation of their TRNs. The crp gene, encoding a major global regulator in Escherichia coli, was deleted in four different genetic backgrounds, all derived from the Long-Term Evolution Experiment (LTEE) but with different TRN architectures. We confirmed that crp deletion had a more deleterious effect on growth rate in the LTEE-adapted genotypes; and we showed that the ptsG gene, which encodes the major glucose-PTS transporter, gained CRP dependence over time in the LTEE. We then further evolved the four crp-deleted genotypes in glucose minimal medium, and we found that they all quickly recovered from their growth defects by increasing glucose uptake. We showed that this recovery was specific to the selective environment and consistently relied on mutations in the cis regulatory region of ptsG, regardless of the initial genotype. These mutations affected the interplay of transcription factors acting at the promoters, changed the intrinsic properties of the existing promoters, or produced new transcription initiation sites. Therefore, the plasticity of even a single promoter region can compensate by three different mechanisms for the loss of a key regulatory hub in the E. coli TRN.
Data from: Current methods for automated filtering of multiple sequence alignments frequently worsen single-gene phylogenetic inference
Phylogenetic inference is generally performed on the basis of multiple sequence alignments (MSA). Because errors in an alignment can lead to errors in tree estimation, there is a strong interest in identifying and removing unreliable parts of the alignment. In recent years several automated filtering approaches have been proposed, but despite their popularity, a systematic and comprehensive comparison of different alignment filtering methods on real data has been lacking. Here, we extend and apply recently introduced phylogenetic tests of alignment accuracy on a large number of gene families and contrast the performance of unfiltered versus filtered alignments in the context of single-gene phylogeny reconstruction. Based on multiple genome-wide empirical and simulated data sets, we show that the trees obtained from filtered MSAs are on average worse than those obtained from unfiltered MSAs. Furthermore, alignment filtering often leads to an increase in the proportion of well-supported branches that are actually wrong. We confirm that our findings hold for a wide range of parameters and methods. Although our results suggest that light filtering (up to 20% of alignment positions) has little impact on tree accuracy and may save some computation time, contrary to widespread practice, we do not generally recommend the use of current alignment filtering methods for phylogenetic inference. By providing a way to rigorously and systematically measure the impact of filtering on alignments, the methodology set forth here will guide the development of better filtering algorithms.
Data from: Coestimating reticulate phylogenies and gene trees from multilocus sequence data
The multispecies network coalescent (MSNC) is a stochastic process that captures how gene trees grow within the branches of a phylogenetic network. Coupling the MSNC with a stochastic mutational process that operates along the branches of the gene trees gives rise to a generative model of how multiple loci from within and across species evolve in the presence of both incomplete lineage sorting (ILS) and reticulation (e.g., hybridization). We report on a Bayesian method for sampling the parameters of this generative model, including the species phylogeny, gene trees, divergence times, and population sizes, from DNA sequences of multiple independent loci. We demonstrate the utility of our method by analyzing simulated data and reanalyzing an empirical data set. Our results demonstrate the significance of not only co-estimating species phylogenies and gene trees, but also accounting for reticulation and ILS simultaneously. In particular, we show that when gene flow occurs, our method accurately estimates the evolutionary histories, coalescence times, and divergence times. Tree inference methods, on the other hand, underestimate divergence times and overestimate coalescence times when the evolutionary history is reticulate. While the MSNC corresponds to an abstract model of ``intermixture," we study the performance of the model and method on simulated data generated under a gene flow model. We show that the method accurately infers the most recent time at which gene flow occurs. Finally, we demonstrate the application of the new method to a 106-locus yeast data set.
Data from: Bacterial characterization of Beijing drinking water by flow cytometry and MiSeq sequencing of the 16S rRNA gene
Flow cytometry (FCM) and 16S rRNA gene sequencing data are commonly used to monitor and characterize microbial differences in drinking water distribution systems. In this study, to assess microbial differences in drinking water distribution systems, 12 water samples from different sources water (groundwater, GW; surface water, SW) were analyzed by FCM, heterotrophic plate count (HPC), and 16S rRNA gene sequencing. FCM intact cell concentrations varied from 2.2 × 103 cells/mL to 1.6 × 104 cells/mL in the network. Characteristics of each water sample were also observed by FCM fluorescence fingerprint analysis. 16S rRNA gene sequencing showed that Proteobacteria (76.9–42.3%) or Cyanobacteria (42.0–3.1%) was most abundant among samples. Proteobacteria were abundant in samples containing chlorine, indicating resistance to disinfection. Interestingly, Mycobacterium, Corynebacterium, and Pseudomonas, were detected in drinking water distribution systems. There was no evidence that these microorganisms represented a health concern through water consumption by the general population. However, they provided a health risk for special crowd, such as the elderly or infants, patients with burns and immune-compromised people exposed by drinking. The combined use of FCM to detect total bacteria concentrations and sequencing to determine the relative abundance of pathogenic bacteria resulted in the quantitative evaluation of drinking water distribution systems. Knowledge regarding the concentration of opportunistic pathogenic bacteria will be particularly useful for epidemiological studies.
Data from: Identification and qualification of 500 nuclear, single-copy, orthologous genes for the Eupulmonata (Gastropoda) using transcriptome sequencing and exon capture
The qualification of orthology is a significant challenge when developing large, multiloci phylogenetic data sets from assembled transcripts. Transcriptome assemblies have various attributes, such as fragmentation, frameshifts and mis-indexing, which pose problems to automated methods of orthology assessment. Here, we identify a set of orthologous single-copy genes from transcriptome assemblies for the land snails and slugs (Eupulmonata) using a thorough approach to orthology determination involving manual alignment curation, gene tree assessment and sequencing from genomic DNA. We qualified the orthology of 500 nuclear, protein-coding genes from the transcriptome assemblies of 21 eupulmonate species to produce the most complete phylogenetic data matrix for a major molluscan lineage to date, both in terms of taxon and character completeness. Exon capture targeting 490 of the 500 genes (those with at least one exon >120 bp) from 22 species of Australian Camaenidae successfully captured sequences of 2825 exons (representing all targeted genes), with only a 3.7% reduction in the data matrix due to the presence of putative paralogs or pseudogenes. The automated pipeline Agalma retrieved the majority of the manually qualified 500 single-copy gene set and identified a further 375 putative single-copy genes, although it failed to account for fragmented transcripts resulting in lower data matrix completeness when considering the original 500 genes. This could potentially explain the minor inconsistencies we observed in the supported topologies for the 21 eupulmonate species between the manually curated and 'Agalma-equivalent' data set (sharing 458 genes). Overall, our study confirms the utility of the 500 gene set to resolve phylogenetic relationships at a range of evolutionary depths and highlights the importance of addressing fragmentation at the homolog alignment stage for probe design.
Data from: Successful recovery of nuclear protein-coding genes from small insects in museums using illumina sequencing
In this paper we explore high-throughput Illumina sequencing of nuclear protein-coding, ribosomal, and mitochondrial genes in small, dried insects stored in natural history collections. We sequenced one tenebrionid beetle and 12 carabid beetles ranging in size from 3.7 to 9.7 mm in length that have been stored in various museums for 4 to 84 years. Although we chose a number of old, small specimens for which we expected low sequence recovery, we successfully recovered at least some low-copy nuclear protein-coding genes from all specimens. For example, in one 56-year-old beetle, 4.4 mm in length, our de novo assembly recovered about 63% of approximately 41,900 nucleotides in a target suite of 67 nuclear protein-coding gene fragments, and 70% using a reference-based assembly. Even in the least successfully sequenced carabid specimen, reference-based assembly yielded fragments that were at least 50% of the target length for 34 of 67 nuclear protein-coding gene fragments. Exploration of alternative references for reference-based assembly revealed few signs of bias created by the reference. For all specimens we recovered almost complete copies of ribosomal and mitochondrial genes. We verified the general accuracy of the sequences through comparisons with sequences obtained from PCR and Sanger sequencing, including of conspecific, fresh specimens, and through phylogenetic analysis that tested the placement of sequences in predicted regions. A few possible inaccuracies in the sequences were detected, but these rarely affected the phylogenetic placement of the samples. Although our sample sizes are low, an exploratory regression study suggests that the dominant factor in predicting success at recovering nuclear protein-coding genes is a high number of Illumina reads, with success at PCR of COI and killing by immersion in ethanol being secondary factors; in analyses of only high-read samples, the primary significant explanatory variable was body length, with small beetles being more successfully sequenced.
A systematic evaluation of highly variable gene selection methods for single-cell RNA-sequencing
Open the record for dataset details and reuse information.
The invasive land flatworm Arthurdendyus triangulatus: repeated sequences in the mitogenome, extra-long cox2 gene and paralogous rRNA clusters
<p>Fasta and tbl files for the mitogenomes of various Rhynchodeminae</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.