Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,293
datasets available to search
ShareScore release 0.9.0
Dataset results
1,293 results for “gene sequencing”
Divergence in coding sequence and expression of different functional categories of immune genes between two wild rodent species
Open the record for dataset details and reuse information.
Cytb gene sequences of Fejervarya species from Lesser Sunda, Indonesia and other Asian countries
Open the record for dataset details and reuse information.
Alignments of ITS1 gene sequences from Harpacticella inopinata
Open the record for dataset details and reuse information.
Supplemental material for: Genome-wide association study and fine-mapping using imputed sequences to prioritize candidate genes for 30 complex traits in 50,309 Holstein bulls
Open the record for dataset details and reuse information.
High resolution diel transcriptomes of autotetraploid potato reveal expression and sequence conservation among rhythmic genes
Open the record for dataset details and reuse information.
16S rRNA gene sequencing data from: Breastmilk IgG engages the neonatal immune system to instruct immune responses to gut antigens
Open the record for dataset details and reuse information.
Sequences of the nuclear gene ADH for thirteen species of willows and two poplars: Investigating patterns of habitat specialization in fifteen co-occurring willow and poplar species.
Thirteen willow (Salix) species occur in southeastern Minnesota and often co-occur within the same wetlands. This high local diversity is challenging to explain since closely related species are often functionally similar and density-dependent interactions such as competition and susceptibility to pests and pathogens should limit their co-occurrence. However, if willow species are partitioning resources, or if they are phylogenetically structured so that closely related species rarely co-occur, then the impact of these density-dependent processes could be reduced. In this study, I examined the role of niche partitioning in maintaining local willow diversity by documenting species distributions in plots across a water availability gradient and comparing species physiology in the field and greenhouse. By taking a phylogenetic approach, I also investigated whether willow communities exhibit phylogenetic community structure and whether there is evidence for environmental filtering.
Soil Microbial Gene Sequences at the Kellogg Biological Station, Hickory Corners, MI (2004 to 2010)
Dataset Abstract Gene sequences extracted from soils. original data source http://lter.kbs.msu.edu/datasets/108
III Average nucleotide distances (%) based on the Kimura 2-parameter (K2P) model between Aselliscus spp., and associated outgroups based on complete mitochondrial Cytb (1,140 bp, below the diagonal) and COI (657 bp, above the diagonal) gene sequences in Description of a new species of the genus Aselliscus (Chiroptera, Hipposideridae) from Vietnam
III Average nucleotide distances (%) based on the Kimura 2-parameter (K2P) model between Aselliscus spp., and associated outgroups based on complete mitochondrial Cytb (1,140 bp, below the diagonal) and COI (657 bp, above the diagonal) gene sequences
standard with together) corner left bottom (analysis genetic the in included species among bold gene in shown oxidase-I are SE cytochrome and species the within at) % (divergence divergence sequence average The . pairwise) corner showing right upper (Matrix) %;. SE 4 ABLE (T error in Description of a new species of the Rhinolophus trifoliatus-group (Chiroptera: Rhinolophidae) from Southeast Asia
standard with together) corner left bottom (analysis genetic the in included species among bold gene in shown oxidase-I are SE cytochrome and species the within at) % (divergence divergence sequence average The . pairwise) corner showing right upper (Matrix) %;. SE 4 ABLE (T error
Genome sequencing of four culinary herbs reveals terpenoid genes underlying chemodiversity in the Nepetoideae
<p>Species within the mint family, Lamiaceae, are widely used for their culinary, cultural, and medicinal properties due to production of a wide variety of specialized metabolites, especially terpenoids. To further our understanding of genome diversity in the Lamiaceae and to provide a resource for mining biochemical pathways, we generated high-quality genome assemblies of four economically important culinary herbs, namely, sweet basil (<i>Ocimum basilicum </i>L<i>.</i>), sweet marjoram (<i>Origanum majorana </i>L.), oregano (<i>Origanum vulgare </i>L<i>.</i>), and rosemary (<i>Rosmarinus officinalis </i>L<i>.</i>), and characterized their terpenoid diversity through metabolite profiling and genomic analyses. A total 25 monoterpenes and 11 sesquiterpenes were identified in leaf tissue from the four species. Genes encoding enzymes responsible for the biosynthesis of precursors for mono- and sesqui-terpene synthases were identified in all four species. Across all four species, a total of 235 terpene synthases were identified, ranging from 27 in <i>O. majorana</i> to 137 in the tetraploid <i>O. basilicum</i>. This study provides valuable resources for further investigation of the genetic basis of chemodiversity in these important culinary herbs.</p>
Transcriptome analysis of WT versus H2A.J-KO MEFs for the paper entitled: The H2A.J histone variant contributes to Interferon-Stimulated Gene expression in senescence by its weak interaction with H1 and the derepression of repeated DNA sequences
<p>Abstract for overall study:</p> <p>The histone variant H2A.J was previously shown to accumulate in senescent human fibroblasts with persistent DNA damage to promote inflammatory gene expression, but its mechanism of action was unknown. We show that H2A.J accumulation contributes to weakening the association of histone H1 to chromatin and increasing its turnover. Decreased H1 in senescence is correlated with increased expression of some repeated DNA sequences, increased expression of STAT/IRF transcription factors, and transcriptional activation of Interferon-Stimulated Genes (ISGs). The H2A.J-specific Val-11 moderates the transcriptional activity of H2A.J, and H2A.J-specific Ser-123 can be phosphorylated in response to DNA damage with potentiation of its transcriptional activity by the phospho-mimetic S123E mutation. Our work demonstrates the functional importance of H2A.J-specific residues and potential mechanisms for its function in promoting inflammatory gene expression in senescence.</p> <p>Specific description for this dataset:</p> <p>We further tested a role for H2A.J in Interferon-Stimulated Gene expression by analyzing the transcriptome of WT and H2A.J MEFs induced into senescence by etoposide. TruSeq stranded DNA libraries were prepared from polyA-selected RNA and sequenced as 43 bp paired-end reads. The fastq sequences were mapped to Gencode.vM24.transcripts.fa.gz (GRCm38 transcriptome) with salmon. Read counts were then aggregated to the gene level with tximeta, and differential gene expression was analysed with DESeq2, edgeR, and limma-voom. Gene set enrichment analysis was performed with camera.</p> <p>The transciptomes of senescent WT and H2AFJ-KO showed strong separation from proliferating MEFs, and a weaker separation distinguished WT and H2A.J-KO MEFs. Strikingly, gene set enrichment analysis indicated highly significant defects in Interferon Response Gene Expression in the H2A.J-KO MEFs in senescence with significant down-regulation in senescent H2A.J-KO cells of a series of oligoadenylate synthase genes (Oas1g, Oas1a, Oasl1, Oas2, Oasl2) and several ISGs. Thus, H2A.J also contributes to ISG expression in the heterologous context of senescent MEFs.</p>
TA B L E 2 Estimates of pairwise sequence divergence (cyt-b gene) in pale-bellied Micronycteris, where M. minuta is divided in three clades. Below the diagonal: pairwise distance using the Kimura 2-parameter model (percentage). On the diagonal: within-clade distance using the Kimura 2-parameter model (percentage). Above the diagonal: pairwise p-distance values. Number of specimens sequenced in parenthesis. *Chimeric sequence obtained from two paratypes (Siles et al., 2013). in Revision of the pale-bellied Micronycteris Gray, 1866 (Chiroptera, Phyllostomidae) with descriptions of two new species
TA B L E 2 Estimates of pairwise sequence divergence (cyt-b gene) in pale-bellied Micronycteris, where M. minuta is divided in three clades. Below the diagonal: pairwise distance using the Kimura 2-parameter model (percentage). On the diagonal: within-clade distance using the Kimura 2-parameter model (percentage). Above the diagonal: pairwise p-distance values. Number of specimens sequenced in parenthesis. *Chimeric sequence obtained from two paratypes (Siles et al., 2013).
Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)
<p class="BodyA"><span><b>Background:</b> The great diversity in plant genome size and chromosome number is partly due to polyploidization (i.e., genome doubling events). The differences in genome size and chromosome number among diploid plant species can be a window into the intriguing phenomenon of past genome doubling that may be obscured through time by the process of diploidization. The genus <i>Hibiscus </i>L. (Malvaceae) has a wide diversity of chromosome numbers and a complex genomic history. <i>Hibiscus </i>is ideal for exploring past genomic events because although two ancient genome duplication events have been identified, more are likely to be found due to its diversity of chromosome numbers. To reappraise the history of whole genome duplication events, we tested three alternative scenarios describing different polyploidization events.</span></p> <p class="BodyA"><span><b>Results:</b> Using target sequence capture, we designed a new probe set for <i>Hibiscus </i>and generated 87 orthologous genes from four diploid species. We detected paralogues in >54% putative single-copy genes. 34 of these genes were selected for testing three different genome duplication scenarios using gene counting. All species of <i>Hibiscus</i> sampled shared one genome duplication with <i>H. syriacus</i> and one whole genome duplication occurred along the branch leading to <i>H. syriacus</i>.</span></p> <p class="BodyA"><span><b>Conclusions:</b> Here, we corroborated the independent genome doubling previously found in the lineage leading to <i>H. syriacus </i>and a shared genome doubling of this lineage and the remainder of <i>Hibiscus</i>. Additionally, we found a previously undiscovered genome duplication shared by the /Pavonia and /Malvaviscus clades (both nested within <i>Hibiscus</i>) with the occurrences of two copies in what were otherwise single-copy genes. Our results highlight the complexity of genomic diversity in some plant groups, which makes orthology assessment and accurate phylogenomic inference difficult.</span></p>
Data from: Phylogenetic relationships and timing of diversification in gonorynchiform fishes inferred using nuclear gene DNA sequences (Teleostei: Ostariophysi)
The Gonorynchiformes are the sister lineage of the species-rich Otophysi and provide important insights into the diversification of ostariophysan fishes. Phylogenies of gonorynchiforms inferred using morphological characters and mtDNA gene sequences provide differing resolutions with regard to the sister lineage of all other gonorynchiforms (Chanos vs. Gonorynchus) and support for monophyly of the two miniaturized lineages Cromeria and Grasseichthys. In this study the phylogeny and divergence times of gonorynchiforms are investigated with DNA sequences sampled from nine nuclear genes and a published morphological character matrix. Bayesian phylogenetic analyses reveal substantial congruence among individual gene trees with inferences from eight genes placing Gonorynchus as the sister lineage to all other gonorynchiforms. Seven gene trees resolve Cromeria and Grasseichthys as a clade, supporting previous inferences using morphological characters. Phylogenies resulting from either concatenating the nuclear genes, performing a multispecies coalescent species tree analysis, or combining the morphological and nuclear gene DNA sequences resolve Gonorynchus as the living sister lineage of all other gonorynchiforms, strongly support the monophyly of Cromeria and Grasseichthys, and resolve a clade containing Parakneria, Cromeria, and Grasseichthys. The morphological dataset, which includes 13 gonorynchiform fossil taxa that range in age from Early Cretaceous to Eocene, was analyzed in combination with DNA sequences from the nine nuclear genes and a relaxed molecular clock to estimate times of evolutionary divergence. This "tip dating" strategy accommodates uncertainty in the phylogenetic resolution of fossil taxa that provide calibration information in the relaxed molecular clock analysis. The estimated age of the most recent common ancestor (MRCA) of living gonorynchiforms is slightly older than estimates from previous node dating efforts, but the molecular tip dating estimated ages of Kneriinae (Kneria, Parakneria, Cromeria, and Grasseichthys) and the two paedomorphic lineages, Cromeria and Grasseichthys, are considerably younger.
Data from: Identification and characterization of sex-associated loci in sockeye salmon using genotyping-by-sequencing and comparison with a sex-determining assay based on the sdY gene
Loci that can be used to screen for sex in salmon can provide important information for study of both wild and cultured populations. Here, we tested for associations between sex and genotypes at thousands of loci available from a genotyping-by-sequencing (GBS) dataset to discover sex-associated loci in sockeye salmon (Oncorhynchus nerka). We discovered seven sex-associated loci, developed high-throughput assays for two loci, and tested the utility of these two assays in eight collections of sockeye salmon sampled throughout North America. We also screened an existing assay based on the master sex-determining gene in salmon (sdY) in these collections. The ability of GBS-derived loci to assign fish to their phenotypic sex varied substantially among collections suggesting that recombination between the loci that we discovered and the sex-determining gene has occurred. Assignment accuracy to phenotypic sex was much higher with the sdY assay but was still less than 100%. Alignment of sequences from GBS-derived loci to draft genomes for two salmonids provided strong evidence that many of these loci are found on the chromosome orthologous to the known sex chromosome in sockeye salmon. Our study is the first to describe the approximate location of the sex-determining region in sockeye salmon and indicates that sdY is also the master sex-determining gene in this species. However, discordances between sdY genotypes and phenotypic sex and the variable performance of GBS-derived loci warrant more research.
Data from: A NGS approach to the encrusting Mediterranean sponge Crella elegans (Porifera, Demospongiae, Poecilosclerida): transcriptome sequencing, characterization and overview of the gene expression along three life cycle stages
Sponges can be dominant organisms in many marine and freshwater habitats where they play essential ecological roles. They also represent a key group to address important questions in early metazoan evolution. Recent approaches for improving knowledge on sponge biological and ecological functions as well as on animal evolution have focused on the genetic toolkits involved in ecological responses to environmental changes (biotic and abiotic), development and reproduction. These approaches are possible thanks to newly available, massive sequencing technologies–such as the Illumina platform, which facilitate genome and transcriptome sequencing in a cost-effective manner. Here we present the first NGS (next-generation sequencing) approach to understanding the life cycle of an encrusting marine sponge. For this we sequenced libraries of three different life cycle stages of the Mediterranean sponge Crella elegans and generated de novo transcriptome assemblies. Three assemblies were based on sponge tissue of a particular life cycle stage, including non-reproductive tissue, tissue with sperm cysts and tissue with larvae. The fourth assembly pooled the data from all three stages. By aggregating data from all the different life cycle stages we obtained a higher total number of contigs, contigs with blast hit and annotated contigs than from one stage-based assemblies. In that multi-stage assembly we obtained a larger number of the developmental regulatory genes known for metazoans than in any other assembly. We also advance the differential expression of selected genes in the three life cycle stages to explore the potential of RNA-seq for improving knowledge on functional processes along the sponge life cycle.
Construction of genetic linkage map based on SNP markers, QTL mapping and detection of candidate genes of growth-related traits in Pacific abalone using genotyping-by-sequencing
<p><a name="_Hlk72585736"><span>Pacific abalone (<i>Haliotis discus hannai</i>) is a commercially important high valued molluscan species. Its wild population has decreased in recent years. Pacific abalone is widely cultured in Korea. Traditional breeding programs have been implemented for hatchery production of abalone seeds. To obtain more genetic information for the molecular breeding program, a high-density linkage map and quantitative trait locus (QTL) for three growth-related traits was constructed for Pacific abalone. F1 cross population with two parents were sampled to construct the linkage map using genotyping by sequencing (GBS). A total of 664,630,534 clean reads and 56,686 SNPs were generated. In sum, 3,345 segregating SNPs were used to construct a consensus linkage map. The map spanned 1,747.023 cM with 18 linkage groups and an average interval of 0.55 cM. QTL analysis revealed two significant QTL in LG10 on the consensus linkage map in each growth-related trait. Both the QTLs are located in the telomere region of the chromosome. Moreover, four potential candidate genes for growth-related traits were identified in the QTL region. Expression analysis revealed that identified genes are involved in growth regulation of abalone. The newly constructed genetic linkage map, growth-related QTLs and potential candidate genes identified in the present study can be used as valuable genetic resources and will be useful for marker-assisted selection (MAS) of Pacific abalone in molecular breeding program.</span></a></p>
Data from: The genetic architecture of reproductive isolation during speciation-with-gene-flow in lake whitefish species pairs assessed by RAD sequencing
During speciation-with-gene-flow, effective migration varies across the genome as a function of several factors, including proximity of selected loci, recombination rate, strength of selection, and number of selected loci. Genome scans may provide better empirical understanding of the genome-wide patterns of genetic differentiation, especially if the variance due to the previously mentioned factors is partitioned. In North American lake whitefish (Coregonus clupeaformis), glacial lineages that diverged in allopatry about 60,000 years ago and came into contact 12,000 years ago have independently evolved in several lakes into two sympatric species pairs (a normal benthic and a dwarf limnetic). Variable degrees of reproductive isolation between species pairs across lakes offer a continuum of genetic and phenotypic divergence associated with adaptation to distinct ecological niches. To disentangle the complex array of genetically based barriers that locally reduce the effective migration rate between whitefish species pairs, we compared genome-wide patterns of divergence across five lakes distributed along this divergence continuum. Using restriction site associated DNA (RAD) sequencing, we combined genetic mapping and population genetics approaches to identify genomic regions resistant to introgression and derive empirical measures of the barrier strength as a function of recombination distance. We found that the size of the genomic islands of differentiation was influenced by the joint effects of linkage disequilibrium maintained by selection on many loci, the strength of ecological niche divergence, as well as demographic characteristics unique to each lake. Partial parallelism in divergent genomic regions likely reflected the combined effects of polygenic adaptation from standing variation and independent changes in the genetic architecture of postzygotic isolation. This study illustrates how integrating genetic mapping and population genomics of multiple sympatric species pairs provide a window on the speciation-with-gene-flow mechanism.
Data from: Allele phasing has minimal impact on phylogenetic reconstruction from targeted nuclear gene sequences in a case study of Artocarpus
Premise of the study: Untapped information about allelic diversity within populations and individuals (i.e. heterozygosity) could improve phylogenetic resolution and accuracy. Many phylogenetic reconstructions ignore heterozygosity because it is difficult to assemble allele sequences and combine allelic data across unlinked loci and it is unclear how reconstruction methods accommodate variable sequences. We review the common methods of including heterozygosity in phylogenetic studies and present a novel method for assembling allele sequences from target enriched Illumina sequencing libraries. Methods: We perform supermatrix phylogeny reconstruction and species tree estimation of Artocarpus based on three methods of accounting for heterozygous sequences: a consensus method based on de novo sequence assembly, the use of ambiguity characters, and a novel method for phasing alleles. We characterize the extent to which highly heterozygous sequences impeded phylogeny reconstruction and determine whether the use of allele sequences improves resolution or decreases topological uncertainty. Key Results: We show that it is possible to infer phased alleles from target enriched Illumina libraries. We find that highly heterozygous sequences do not contribute disproportionately to poor phylogenetic resolution and that the use of allele sequences for phylogeny reconstruction does not have a clear effect on phylogenetic resolution or topological consistency. Conclusions: We provide a framework for inferring phased alleles from target enrichment data and for assessing the contribution of allelic diversity to phylogenetic reconstruction. In our dataset, the impact of allele phasing on phylogeny is minimal compared to the impact of using phylogenetic reconstruction methods that account for gene tree incongruence.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.