Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8
datasets available to search
ShareScore release 0.9.0
Dataset results
8 results for “single copy nuclear genes”
Set of 4 597 baits designed in collaboration with RapidGenomics (Gainesville, Florida, USA) to capture the identified low‐ to single‐copy nuclear genes (LSCN).
<p>This dataset presents a set of 4,597 baits designed in collaboration with RapidGenomics (Gainesville, Florida, USA) to capture the identified low- to single-copy nuclear genes (LSCN) in our study. These baits were instrumental in our research on the evolutionary relationships of the Neotropical magnolias based on plastome and nuclear phylogenomics. The data provided are crucial for understanding the methodology and results of our study.</p>
Capturing single-copy nuclear genes, organellar genomes, and nuclear ribosomal DNA from deep genome skimming data for plant phylogenetics: A case study in Vitaceae
<p>With the decreasing cost and availability of many newly developed bioinformatics pipelines, next-generation sequencing (NGS) has revolutionized plant systematics in recent years. Genome skimming has been widely used to obtain high-copy fractions of the genomes, including plastomes, mitochondrial DNA (mtDNA), and nuclear ribosomal DNA (nrDNA). In this study, through simulations, we evaluated the optimal (minimum) sequencing depth and performance for recovering single-copy nuclear genes (SCNs) from genome skimming data, by subsampling genome resequencing data and generating 10 datasets with different sequencing coverage <i>in silico</i>. We tested the performance of four datasets (plastome, nrDNA, mtDNA, and SCNs) obtained from genome skimming based on phylogenetic analyses of the <i>Vitis</i> clade at the genus level and Vitaceae at the family level, respectively. Our results showed that optimal minimum sequencing depth for high-quality SCNs assembly via genome skimming was about 10× coverage. Without the steps of synthesizing baits and enrichment experiments, coupled with incredibly low sequencing costs, we showcase that deep genome skimming (DGS) is as effective for capturing large datasets of SCNs as the widely used Hyb-Seq approach, in addition to capturing plastomes, mtDNA, and entire nrDNA repeats. DGS may serve as an efficient and economical alternative and may be superior to the popular target enrichment/Hyb-Seq approach.</p>
Phylogenomics of mulberries (Morus, Moraceae) inferred from plastomes and single copy nuclear genes
<p><span>Mulberry (genus <em>Morus</em>), belonging to the order Rosales, family Moraceae, is an important woody plant due to its economic value in sericulture as well as for its nutritional benefits and medicinal values. However, the taxonomy and phylogeny of <em>Morus</em> remain challenging due to its wide geographical distribution, morphological plasticity, and interspecific hybridization. To better understand the evolutionary history of <em>Morus</em>, we combined plastomes and a large-scale nuclear gene to investigate their phylogenetic relationships in the present study. We assembled the plastomes and screened 211 single-copy nuclear genes from 14 <em>Morus</em> species and related taxa. The plastomes of <em>Morus</em> species were relatively conserved in terms of genome size, gene content and order, IR boundary and codon usage. Using nuclear data, we yielded completely identical topologies based on coalescent and concatenation methods, and multiple individuals of the same species were intraspecific monophyletic. The genus <em>Morus</em> was supported as a monophyly, and <em>M. notabilis</em> was recovered as the first diverging, and the two North American <em>Morus</em> species, <em>M. celtidifolia</em> and <em>M. rubra</em>, were sister to the other Asian species. However, the relationships of <em>Morus</em> based on plastomes were strongly incongruent with those from nuclear genes, and intraspecific non-monophyly was retrieved in the plastid phylogeny. Comparisons of nuclear and plastid phylogenies, and combining with the result of network inference, hybridization/introgression was regarded as the main cause of the discordance between nuclear and plastid phylogenies in the genus <em>Morus</em>. Overall, the robust phylogenetic relationships of <em>Morus</em> described here will be useful for genetic resources development of this economically important genus and exploitation of sericulture industry.</span></p>
Capturing single-copy nuclear genes, organellar genomes, and nuclear ribosomal DNA from deep genome skimming data for plant phylogenetics: A case study in Vitaceae
Open the record for dataset details and reuse information.
Phylogenomics of mulberries (Morus, Moraceae) inferred from plastomes and single copy nuclear genes
Open the record for dataset details and reuse information.
Data from: Identification and qualification of 500 nuclear, single-copy, orthologous genes for the Eupulmonata (Gastropoda) using transcriptome sequencing and exon capture
The qualification of orthology is a significant challenge when developing large, multiloci phylogenetic data sets from assembled transcripts. Transcriptome assemblies have various attributes, such as fragmentation, frameshifts and mis-indexing, which pose problems to automated methods of orthology assessment. Here, we identify a set of orthologous single-copy genes from transcriptome assemblies for the land snails and slugs (Eupulmonata) using a thorough approach to orthology determination involving manual alignment curation, gene tree assessment and sequencing from genomic DNA. We qualified the orthology of 500 nuclear, protein-coding genes from the transcriptome assemblies of 21 eupulmonate species to produce the most complete phylogenetic data matrix for a major molluscan lineage to date, both in terms of taxon and character completeness. Exon capture targeting 490 of the 500 genes (those with at least one exon >120 bp) from 22 species of Australian Camaenidae successfully captured sequences of 2825 exons (representing all targeted genes), with only a 3.7% reduction in the data matrix due to the presence of putative paralogs or pseudogenes. The automated pipeline Agalma retrieved the majority of the manually qualified 500 single-copy gene set and identified a further 375 putative single-copy genes, although it failed to account for fragmented transcripts resulting in lower data matrix completeness when considering the original 500 genes. This could potentially explain the minor inconsistencies we observed in the supported topologies for the 21 eupulmonate species between the manually curated and 'Agalma-equivalent' data set (sharing 458 genes). Overall, our study confirms the utility of the 500 gene set to resolve phylogenetic relationships at a range of evolutionary depths and highlights the importance of addressing fragmentation at the homolog alignment stage for probe design.
Data from: Oligonucleotide primers for targeted amplification of single-copy nuclear genes in apocritan Hymenoptera
Open the record for dataset details and reuse information.
Data from: Identification and qualification of 500 nuclear, single-copy, orthologous genes for the Eupulmonata (Gastropoda) using transcriptome sequencing and exon capture
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.