Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
669
datasets available to search
ShareScore release 0.7.1
Dataset results
669 results for “Comparative genomics”
Aphidinae comparative genomics resource
<p>Here we provide early access to 18 new genome assemblies, including 8 assembled to chromosome-scale, for aphids from the subfamily Aphidinae. For consistency and to aid comparative analysis, all genomes have been annotated using the same repeat masking and RNA-seq-based gene prediction pipeline. Using this pipeline we also provide new annotations for three previously published genome assemblies.</p> <p>The genome assemblies and annotations are made freely available without restriction, we only request that this Zenodo resource is cited when using the data. Raw sequence data upload to NCBI is underway and full details of all accessions will be given in an updated version of this resource. Manuscripts are in preparation describing the individual genome assemblies in detail and larger comparative genome analyses and we will update this resource with additional citation information as papers are published.</p> <p>Full details of all genome assemblies and annotations included in this release are given in the attached "Data_Description.pdf" document. </p> <p><strong>Aphid species included in this release (bold type = chromosome-scale assembly):</strong></p> <p><em><strong>Aphis fabae</strong><br> Aphis glycines </em>(updated annotation)<br> <em><strong>Aphis gossypii</strong><br> Aphis thalictri<br> Aphis rumicis<br> Brachycaudus cardui<br> Brachycaudus helichrysi<br> Brachycaudus klugkisti<br> <strong>Brevicoryne brassicae</strong><br> Diuraphis noxia<br> <strong>Macrosiphum albifrons</strong><br> Metopolophium dirhodum<br> Myzus cerasi </em>(updated annotation)<br> <em>Myzus ligustri<br> Myzus lythri<br> Myzus varians<br> Pentalonia nigronervosa </em>(updated annotation)<br> <em><strong>Phorodon humuli</strong><br> <strong>Rhopalosiphum padi<br> Sitobion avenae<br> Sitobion miscanthi</strong></em></p>
Supplementary phylogenetic data for Rouïl et. al. 2020 "The protector within: Comparative genomics of APSE phages across aphids reveals rampant recombination and diverse toxin arsenals"
<p>Supplementary phylogenetic data for Rouïl <em>et. al.</em> 2020 "The protector within: Comparative genomics of APSE phages across aphids reveals rampant recombination and diverse toxin arsenals"</p> <p> </p> <p>The data set consists of the following sub-directories:</p> <p>1) "APSE_conserved_proteins_alns": Single-copy conserved genes codon sequences and alignments in FASTA format.</p> <p>2) "APSE_phylogeny": Files used for APSE phylogenetic and recombination analyses.</p> <p>3) "APSE_reannotations": GenBank-formatted files of the assemblies and re-annotations of APSE phages. Newly-sequenced phages deposited at the European nucleotide Archive are also included. ***New in this version***</p> <p>4) "APSE_toxin_lyzozyme": Files used for APSE toxin-cassette and lyzozyme-related gene phylogenies.</p> <p>5) "Arsenophonus_PHASTER": PHASTER phage annotation output files organised by organisim and contig/scaffold.</p> <p>6) "Hamiltonella_drafts": Newly-sequenced low-coverage draft <em>Hamiltonella</em> genomes in FASTA format.</p> <p>7) "Hamiltonella_phylogeny": files used for <em>Hamiltonella</em> phylogenetic analysis.</p> <p> </p> <p>See enclosed README.txt file for more details.</p> <p> </p> <p>* ver. 1.1.1: Updated annotations for APSE genomes including inteins missing in previous annotation files.</p>
Comparative genomics of Meloidogyne haplanaria
<p>Data related to the MSC thesis 'Comparative genomics of Meloidogyne haplanaria'.</p>
Data from: Full Bayesian comparative phylogeography from genomic data
A challenge to understanding biological diversification is accounting for community-scale processes that cause multiple, co-distributed lineages to co-speciate. Such processes predict non-independent, temporally clustered divergences across taxa. Approximate-likelihood Bayesian computation (ABC) approaches to inferring such patterns from comparative genetic data are very sensitive to prior assumptions and often biased toward estimating shared divergences. We introduce a full-likelihood Bayesian approach, ecoevolity, which takes full advantage of information in genomic data. By analytically integrating over gene trees, we are able to directly calculate the likelihood of the population history from genomic data, and efficiently sample the model-averaged posterior via Markov chain Monte Carlo algorithms. Using simulations, we find that the new method is much more accurate and precise at estimating the number and timing of divergence events across pairs of populations than existing approximate-likelihood methods. Our full Bayesian approach also requires several orders of magnitude less computational time than existing ABC approaches. We find that despite assuming unlinked characters (e.g., unlinked single-nucleotide polymorphisms), the new method performs better if this assumption is violated in order to retain the constant characters of whole linked loci. In fact, retaining constant characters allows the new method to robustly estimate the correct number of divergence events with high posterior probability in the face of character-acquisition biases, which commonly plague loci assembled from reduced-representation genomic libraries. We apply our method to genomic data from four pairs of insular populations of Gekko lizards from the Philippines that are not expected to have co-diverged. Despite all four pairs diverging very recently, our method strongly supports that they diverged independently, and these results are robust to very disparate prior assumptions.
Supplementary material Comparative Genomic Analysis of Antimicrobial-Resistant Escherichia coli from South American Camelids in Central Germany
<p>Supplementary material for publication González-Santamarina, B.; Weber, M.; Menge, C.; Berens, C. Comparative Genomic Analysis of Antimicrobial-Resistant <i>Escherichia coli</i> from South American Camelids in Central Germany. <i>Microorganisms</i> <strong>2022</strong>, <i>10</i>, 1697. https://doi.org/10.3390/microorganisms10091697 </p>
Comparative genomics of eight aphid subfamilies reveals variable relationships between host horizontally-transferred genes and symbiont peptidoglycan metabolism.
<p>Genome assemblies, annotations, and orthologs of aphids (<em>Geopemphigus sp.</em>, <em>Stegophylla sp.</em>, <em>Chaitophorus viminalis</em>, and<em> Pemphigus obesinymphae</em>) and their symbionts. </p> <p>step1_final_assemblies_and_annotations.tar.gz: Aphid genomes and annotations</p> <p>step2_protein_evidence_used_for_genome_annotation.tar.gz: Protein evidence used for aphid genome annotation</p> <p>step3_amino_acid_inputs_for_aphid_orthologs: amino acid inputs for aphid ortholog assignmentt</p> <p>step5_buchnera_genomes_and_annotations.tar.gz: Buchnera genomes and annotations</p>
The chloroplast genomes of Sanicula (Apiaceae): plastome structure, comparative analyses, and phylogenetic relationships
<p><em>Sanicula</em> (Apiaceae subfamily Saniculoideae) is a taxonomically difficult genus of medicinal value. Its distribution center is in China, where there are 18 species (11 of which are endemic). To provide plastid genome resources, whole chloroplast genomes of five <em>Sanicula</em> species (<em>S. flavovirens</em>, <em>S. giraldii</em>, <em>S. lamelligera</em>, <em>S. odorata</em>, and <em>S. rubriflora</em>) were sequenced and compared to the previously published <em>S. orthacantha</em> plastome. These genomes exhibit a typical quadripartite structure. All contain 129 different genes, including 84 protein-coding, 37 tRNA, and 8 rRNA genes. Loci <em>rpl2</em>, <em>matK</em>, <em>psbA</em>, and <em>ycf1</em> are the most variable. Results of maximum likelihood analysis of 90 whole plastome sequences from Apioideae and Saniculoideae and the outgroup <em>Hydrocotyle</em> (Araliaceae) reveal sectional relationships in <em>Sanicula</em> different from the traditional classification system, support the monophyly of Apioideae and its sister group relationship to Saniculoideae, and show concordant topologies to nrDNA ITS and other plastome-based phylogenies. <em>Sanicula orthacantha</em> and <em>S. chinensis</em> form a clade sister group to <em>S. lamelligera</em> and <em>S. odorata</em>, consecutively. These four species comprise a clade sister group to the clade of <em>S. rubriflora</em> and <em>S. flavovirens</em>, with this entire group sister to <em>S. giraldii</em>. The plastid genome resources provided herein will be important for future systematic, evolutionary, phylogenomic, and population-level studies of <em>Sanicula</em>.</p>
Comparative genomics of human distal lung Streptococci
<p><strong>Comparative genomics of human lung streptococcal isolates</strong></p> <p>1. All analysis pipelines and scripts are on the GitHub page of Slipa Kanungo: <a href="https://github.com/slipa17/Whole-genome-sequencing-and-comparative-genomics-of-human-lung-streptococcal-isolates">https://github.com/slipa17/Whole-genome-sequencing-and-comparative-genomics-of-human-lung-streptococcal-isolates</a></p> <p>2. Additional analysis pipelines (especially Dataset S12) are on the GitHub page of Garance Sarton-Lohéac: <a href="https://github.com/gsartonl/Publication_Sarton-Loheac_2022">https://github.com/gsartonl/Publication_Sarton-Loheac_2022</a></p> <p>3. All raw data were uploaded to NCBI SRA BioProject <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1001255">PRJNA1001255</a></p> <p><strong>Supplementary Table bundle: for peer review purposes</strong></p> <p><strong>Supplementary Datasets</strong></p> <ul> <li><strong>Dataset S1_Lung_Streptococcus_genomes_metaQUAST: </strong>MetaQUAST (Quality Assessment Tool for Metagenome Assemblies) output including HTML and PDF reports, summary statistics including total contigs, assembly size, and N50. Coverage analysis assesses how well reference genomes are represented, contig length distribution plots visualize contig length ranges, mis-assembly analysis detects potential errors and graphical representations to visualize assemblies.</li> <li><strong>Dataset S2_Lung_streptococcus_isolate_genomes: </strong>Nucleotide FASTA files of six lung streptococcal isolates obtained that were obtained via whole genome sequencing.</li> <li><strong>Dataset S3_Lung_isolates_genome_annotation_prokka: </strong>Output folders after annotation of six lung streptococcal isolates with PROKKA. This includes protein FASTA, GenBank files and GFF annotations.</li> <li> <p><strong>Dataset S4_TYGS_dDDH_analysis: </strong>Contains results of TYGS analysis from DSMZ including downloadable reports. Outputs including taxonomic identification with genus, species, and strain details, a TYGS index for tracking genomes, genome quality assessment metrics, GBDP whole genome and 16S rRNA phylogenetic tree files, comparisons with reference type strains in the TYGS database with table.</p> </li> <li> <p><strong>Dataset S5_Reference_type_strains_TYGS_genomes: </strong>Nucleotide FASTA files of 47 closely related reference <em>Streptococcus</em> genomes listed by TYGS and downloaded from NCBI.</p> </li> <li> <p><strong>Dataset S6_Reference_type_strains_TYGS_proteins: </strong>Protein FASTA files of 47 closely related reference <em>Streptococcus</em> genomes listed by TYGS and downloaded from NCBI. </p> </li> <li> <p><strong>Dataset S7_Lung_streptococcus_isolate_proteins: </strong>Protein FASTA files of 6 six lung streptococcal isolates. </p> </li> <li> <p><strong>Dataset S8_OrthoFinder_core_genome: </strong>OrthoFinder is a bioinformatics tool that offers comprehensive outputs for orthology inference across multiple genomes. The output includes overall statistics, gene duplication information, orthologous genes, orthologous gene tree, single copy orthologous genes and STAG evolutionary trees.</p> </li> <li> <p><strong>Dataset S9_Pan-Strep_BLAST_db: </strong>BLAST database using the <code>makeblastdb</code> command of NCBI datasets command line tool. This is constructed using protein FASTA files of 47 closely related reference <em>Streptococcus</em> genomes listed by TYGS and downloaded from NCBI.</p> </li> <li><strong>Dataset S10_OrthoVenn_cluster_files: </strong>OrthoVenn is a web-based tool for orthologous gene comparison. Downloadable results include Venn diagrams depicting shared and unique orthologous clusters amongst species, tabular results detailing genes within each cluster and their annotations. Functional enrichment analysis for Gene Ontology terms and KEGG pathways are also provided. enhances biological insights.</li> <li> <p><strong>Dataset S11_COG_analysis: </strong>Results of COG analysis of six lung streptococcal isolates individually using eggNOG (evolutionary genealogy of genes: Non-supervised Orthologous Groups) webtool. The output includes information on Clusters of Orthologous Groups (COGs) categorizing them into functional groups such as metabolism, information storage and processing, and cellular processes and signalling.</p> </li> <li> <p><strong>Dataset S12_CAZymes_lung_streptococci: </strong>Results of CAZyme analysis using a custom rule-based pipeline mostly based on dbCAN (Database for Carbohydrate-Active enZymes) provides information on the carbohydrate-active enzymes present in genomic datasets. The output includes the annotation of enzymes involved in the degradation, modification, or biosynthesis of carbohydrates: glycoside hydrolases (GH), glycosyltransferases (GT), carbohydrate-binding modules (CBM), Auxillary Activities (AA), Carbohydrates Esterases (CE) and Polysaccharide lyases (PL).</p> </li> <li> <p><strong>Dataset S13_pneumolysin_analysis: </strong>Results alignment and phylogeny of Pnuemolysin protein in <em>Streptococcus pneumoniae</em>, <em>Streptococcus pseudopneumoniae</em> and Streptococcus isolate P2E5 found by ABRIcate analysis. Visual plots by pyGenomeViz.</p> </li> <li> <p><strong>Dataset S14_capsule_analysis:</strong> Results from BLAST analysis of <em>Streptococcus pneumoniae </em>D39 capsular biosynthesis operon genes against the Pan-Strep (Dataset S10). Extracted of matching genes followed alignment and phylogeny. Visual plots by pyGenomeViz.</p> </li> <li> <p><strong>Dataset S15_Lung_isolate_HOMD_TYGS_comparison: </strong>Protein FASTA files of 47 closely related reference <em>Streptococcus</em> genomes listed by TYGS and downloaded from NCBI, 6 six lung streptococcal isolates and 47 streptococcal genomes from downloaded from human oral microbiome database (eHOMD).</p> </li> </ul>
Comparative genomics sheds new light on the convergent evolution of infrared vision in snakes
<p>Infrared vision is a highly specialized sensory system that evolved independently in three clades of snakes. Apparently, convergent evolution occurred in the transient receptor potential ankyrin 1 (<em>TRPA1</em>) proteins of infrared-sensing snakes. However, this gene can only explain how infrared signals are received, and not the transduction and processing of those signals. We sequenced the genome of <em>Xenopeltis unicolor</em>, a key outgroup species for pythons, and performed a genome-wide analysis of convergence between two clades of infrared-sensing snakes. Our results revealed pervasive molecular adaptation in pathways associated with neural development and other functions, with parallel selection on loci associated with trigeminal nerve structural organization. Additionally, we found evidence of convergent amino acid substitutions in a set of genes, including <em>TRPA1 </em>and<em> TRPM2</em>. Analysis also identified convergent accelerated evolution in non-coding elements near 12 genes involved in facial nerve structural organization and optic nerve development. Thus, convergent evolution occurred across multiple dimensions of infrared vision in vipers and pythons, as well as amino acid substitutions, non-coding elements, genes, and functions. These changes enabled independent groups of snakes to develop and utilize infrared vision.</p>
Fig. 2 in Comparative analyses of the fragmented mitochondrial genomes of wild pig louse Haematopinus apri from China and Japan
Fig. 2. The complete mitochondrial genome of wild pig louse Haematopinus apri form China. Each minichromosome has a coding region and a non-coding region (NCR, in black). The names and transcript orientation of genes are indicated in the coding region and the minichromosomes are placed in alphabetical order of protein-coding genes and rRNA genes. Abbreviations: atp6 and atp8, ATP synthase F0 subunits 6 and 8; cytb, cytochrome b; cox1-3, cytochrome c oxidase subunits 1–3; nad1-6 and nad4L, NADH dehydrogenase subunits 1–6 and 4L; rrnS and rrnL, small and large subunits of ribosomal RNA. tRNA genes are indicated with their single-letter abbreviations of the corresponding amino acids.
FIGURE 5 in Tracking the evolutionary pathways among Brazilian Lebiasina species (Teleostei: Lebiasinidae): a chromosomal and genomic comparative investigation
FIGURE 5 | First Row: Mitotic chromosome spreads of Lebiasina minuta males after CGH— interspecific comparisons (A–D). Male-derived genomic probe of L. minuta (A); L. melanoguttata (B); L. bimaculata (C) hybridized against male metaphase plates of L. minuta (D). Second Row: Mitotic chromosome spreads of Lebiasina minuta males after CGH— intraspecific comparisons (E–H). DAPI image (E); Male-derived genomic probe of L. minuta (F); Female-derived genomic probe of L. minuta (G) hybridized against male metaphase plates of L. minuta (H). The common genomic regions of both compared karyomorphs are depicted in yellow. Scale bar = 5 µm.
FIGURE 4 in Tracking the evolutionary pathways among Brazilian Lebiasina species (Teleostei: Lebiasinidae): a chromosomal and genomic comparative investigation
FIGURE 4 | Whole chromosome painting (WCP) highlighting the first chromosome pair of Lebiasina minuta completely hybridized with the probe from the first chromosome pair of L. bimaculata.
FIGURE 6 in Tracking the evolutionary pathways among Brazilian Lebiasina species (Teleostei: Lebiasinidae): a chromosomal and genomic comparative investigation
FIGURE 6 | Representative idiograms of L. bimaculata (A); L. melanoguttata (B) and L. minuta (C) highlighting the distribution of the 18S (green) and 5S (red) rDNA sequences; (CGG)n microsatellite (blue) and C-positive heterochromatin (black): Data for L. bimaculata and L. melanoguttata are from Sassi et al. (2019).
FIGURE 3 in Tracking the evolutionary pathways among Brazilian Lebiasina species (Teleostei: Lebiasinidae): a chromosomal and genomic comparative investigation
FIGURE 3 | Metaphase chromosomes of Lebiasina minuta hybridized with microsatellite probes (A, B and C) and telomeric probes (D), using red signals. Scale bar = 5 µm.
FIGURE 2 in Tracking the evolutionary pathways among Brazilian Lebiasina species (Teleostei: Lebiasinidae): a chromosomal and genomic comparative investigation
FIGURE 2 | Male and female karyotypes of Lebiasina minuta after A. Giemsa staining, B. C-banding, and C. "double-FISH" with 5S (red) and 18S (green) rDNA probes. Scale bar = 5 µm.
FIGURE 1 in Tracking the evolutionary pathways among Brazilian Lebiasina species (Teleostei: Lebiasinidae): a chromosomal and genomic comparative investigation
FIGURE 1 | Distribution of Lebiasina species with available cytogenetic data, highlighting the Brazilian state of Pará (orange) and Ecuadorian (purple) territories A. 1. L. bimaculata, 2. L. melanoguttata (Sassi et al., 2019), and 3. L. minuta (this study). B. Highlights the position of A in South America, and C. indicates that, although close, species 2 and 3 does not share an overlapped distribution.
Figure 3 in The complete mitochondrial genome of Lemyra melli (Daniel) (Lepidoptera: Erebidae) and a comparative analysis within the Noctuoidea
Figure 3. Putative secondary cloverleaf structures of the tRNA genes in the Lemyra melli mitogenome with mismatched bases. The blue dots, and red dots indicate Watson- Crick base pairing A-U and G-C, respectively, and the blanks indicate mismatched bases. Seven mismatches (five U-U, one A-A and one U-G) lie in five tRNA genes (three in the amino acid acceptor stems, three in the anticodon stems and one in pseudouridine (TΨC)).
Figure 1 in The complete mitochondrial genome of Lemyra melli (Daniel) (Lepidoptera: Erebidae) and a comparative analysis within the Noctuoidea
Figure 1. Map of the mitogenome of Lemyra melli. Genes lying outside and inside of the outer circle are transcribed in the counterclockwise and clockwise directions, respectively. The transfer RNA genes trnL1, trnL2, trnS1 and trnS2 are denoted trnL(UUR), trnL(CUN), trnS(AGN) and trnS(UCN), respectively. Area dashed darker gray in the inner circle denotes the GC content while the lighter gray denotes the AT content of the genome.
Figure 4 in The complete mitochondrial genome of Lemyra melli (Daniel) (Lepidoptera: Erebidae) and a comparative analysis within the Noctuoidea
Figure 4. The structure in the A+T-rich region of the Leymra melli mitogenome. The ATAGA + polyT, the duplicated 14-bp repeat element, the ATTTA + (AT)10 element, and the polyA structure are shown in the sequence.
Figure 2 in The complete mitochondrial genome of Lemyra melli (Daniel) (Lepidoptera: Erebidae) and a comparative analysis within the Noctuoidea
Figure 2. Relative Synonyous Codon Usage (RSCU) in L. melli, H. cunea, and A. formosae mitogenomes. Codons are provided on the x-axis.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.