Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
394
datasets available to search
ShareScore release 0.9.0
Dataset results
394 results for “Genomic Diversity”
Genome assemblies and respective cgMLST profiles of a diverse dataset comprising 1,874 Listeria monocytogenes isolates
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 1,748-loci core-genome (cg) Multiple Locus Sequence Type (MLST) profiles [Pasteur schema (<a href="https://pubmed.ncbi.nlm.nih.gov/27723724/">Moura et al. 2016</a>) available in <a href="https://chewbbaca.online/species/6/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)] of a final set of 1,874 <em>Listeria monocytogenes</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of Sequence Type [ST]). In total, 204 different STs are represented in this dataset, with ST121, ST6, ST9, ST1 and ST155 being in the top 5 and, together, corresponding to 37.9% of the dataset.</p> <p>File “Lm_metadata.xlsx” contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST.</p> <p>The directory “assemblies/” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file. </p> <p>The file “profiles/Lm_profile.tsv” corresponds to a tab separated file with the 1,748-loci cgMLST profile of each isolate presented in the metadata file. These profiles were determined as explained below.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>L. monocytogenes </em>genome assemblies, we collected information about the genetic diversity (STs) of the isolates available at <a href="https://bigsdb.pasteur.fr/listeria/">BIGSdb-Lm</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 1,957 samples associated with three previous studies (<a href="https://pubmed.ncbi.nlm.nih.gov/27723724/">Moura et al. 2016</a>; <a href="https://pubmed.ncbi.nlm.nih.gov/28827366/">Maury et al. 2017</a>; <a href="https://pubmed.ncbi.nlm.nih.gov/30775964/">Painset et al. 2019</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,874 isolates passed the dataset curation step and were included in the final dataset. cgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 1,748-loci Pasteur schema (<a href="https://pubmed.ncbi.nlm.nih.gov/27723724/">Moura et al. 2016</a>) available in <a href="https://chewbbaca.online/species/6/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on June 23<sup>rd</sup>, 2022.</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
Genome-wide single nucleotide polymorphisms reveal the genetic diversity and population structure of Creole goats from northern Peru
<p>Goat farming constitutes a significant source of income for farmers in northern Peru. There is currently an absence of information about the genetics of Peruvian Creole goats that would enable us to understand their origins and genetic spread. The objective of this study was to estimate the genetic diversity of Creole goats from northern Peru using SNP markers. This study involved the collection of 192 male Creole goats from three key goat production regions in northern Peru. These goat samples were genotyped using the GGPGoat70k SNP panel. To explore the genetic influence of other breeds on Peruvian Creole goats, our dataset was combined with previously published SNP genotypes. External data set includes multiple breeds genotypes sampled from Argentina, Brazil, Spain, and Alpine breed from Italy, France, and Switzerland. After quality control 52,832 autosomal SNPs were used to assess genetic diversity in the Peruvian goats. For the population structure analysis of the merged data 20,513 common SNPs were used. Estimations for expected heterozygosity (H<sub>e</sub>), observed heterozygosity (H<sub>o</sub>), and inbreeding coefficient (F<sub>IS</sub>) were computed for the Peruvian groups. AMOVA, principal component analysis and ADMIXTURE were conducted to evaluate the population structure in the two data sets, Peru and merged. The results revealed a considerable genetic diversity, with H<sub>o</sub> values ranging from 0.40 to 0.41 for the Peruvian sampling groups, and inbreeding coefficient was notably low for Peruvian goat. The population structure analysis demonstrated a distinction (p< 0.05) from other breeds. These findings suggest a level of genetic differentiation of the Peruvian goat population among other breeds, although further research is needed considering samples from other Peruvian areas. We expect this study will contribute to define genetic management strategies to prevent the loss of genetic diversity in Peruvian goat populations and for upcoming advancements in this field.</p>
Genealogical Forest Files for Simons Genome Diversity Project
<p>Genealogical forest files for the 1000 Genomes phase 3 autosomes. Converted using <a href="https://github.com/lukashuebner/gfkit">gfkit</a> from the <code>.trees</code> files available <a href="../records/3052359">here</a>.</p>
Data from: Responses of population structure and genomic diversity to climate change and fishing pressure in a pelagic fish
<p><span>The responses of marine species to environmental changes and anthropogenic pressures (e.g. fishing) interact with ecological and evolutionary processes that are not well understood. Knowledge of changes in the distribution range and genetic diversity of species and their populations into the future is essential for the conservation and sustainable management of resources.</span><span> Almaco jack (<em>Seriola rivoliana</em>) is<em> </em>a pelagic fish with high importance to fisheries and aquaculture in the Pacific Ocean. </span><span>In this study, we assessed contemporary genomic diversity and structure in loci that are putatively under selection (outlier loci) and determined their potential functions. Utilizing a combination of genotype-environment association, spatial distribution models, and demogenetic simulations, we modeled the effects of cl</span><span>imate change (under three different RCP scenarios) and fishing pressure on the species' geographic distribution and genomic diversity and structure to 2050 and 2100.</span><span> Our results show that most of the outlier loci identified were related to biological and metabolic processes that may be associated with temperature and salinity. Contemporary genomic structure showed three populations—two in the Eastern Pacific (</span><span>Cabo San Lucas </span><span>and Eastern Pacific) and one in the Central Pacific (</span><span>Hawaii</span><span>). Future projections suggest a loss of suitable habitat and potential range contractions for most scenarios, while fishing pressure decreased population connectivity. Our results suggest that future climate change scenarios and fishing pressure will affect the genomic structure and genotypic composition of <em>S. rivoliana</em> and lead to loss of genomic diversity in populations distributed in the eastern-central Pacific Ocean, which could have profound effects in fisheries that depend on this resource.</span></p>
Ontogenetic variation in the marine foraging of Atlantic salmon functionally links genomic diversity with a major life history polymorphism
<p>The ecological role of heritable phenotypic variation in free-living populations remains largely unknown. Knowledge of the genetic basis of functional ecological processes can link genomic and phenotypic diversity, providing insight into polymorphism evolution and how populations respond to environmental change. By quantifying the marine diet, of sub-adult Atlantic salmon, we assessed how foraging behavior changes along the ontogeny, and in relation to genetic variation in two loci with major effect on age-at-maturity (<em>six6</em> and <em>vgll3</em>). We used a two-component, zero-inflated negative binomial model to simultaneously quantify foraging frequency (zero-inflation components) and foraging outcome (count component), separately for fish and crustaceans in the diet. We found that older salmon forage for both prey types more actively (as evidenced by increased foraging frequency), but with a decreased efficiency (as evidenced by fewer prey items in the diet), suggesting an age-dependent shift in foraging dynamics. The <em>vgll3</em> locus was linked to age-dependent changes in foraging behavior: younger salmon with <em>vgll3<sup>LL</sup></em> (the genotype associated with late maturation) tend to forage crustaceans more often than those with <em>vgll3<sup>EE</sup></em> (the<em> </em>genotype associated with early maturation), while the pattern was reversed in older salmon. <em>Vgll3<sup> LL</sup></em> genotype was also linked to marginal increase in fish acquisition especially in younger salmon, while <em>six6</em> was not a factor explaining the diet variation. Our results suggest a functional role for variation in marine feeding behavior linking genomic diversity at <em>vgll3</em> with age-at-maturity among salmon, with potential age-dependent trade-offs maintaining the genetic variation. A shared genetic basis between dietary ecology and age-at-maturity likely subjects Atlantic salmon populations to evolution induced by bottom-up changes in marine productivity.</p>
Virulence and antibiotic resistance plasticity of Arcobacter butzleri: insights on the genomic diversity of an emerging human pathogen (genome assembly, annotation dataset, core- and pan-genome loci)
<p>This dataset refers to the analysis of 49 <em>Arcobacter butzleri</em> genomes and includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the predicted transcripts (CDS, rRNA, tRNA, tmRNA, misc_RNA) (.ffn files), the respective amino acid sequences of the translated CDS sequences (.faa files), the nucleotide alignments of all the 1165 core-genome loci, the nucleotide alignments of the genes <em>hecA</em>, <em>tetR </em>and <em>porA</em>, the categorized amino acid sequences of the six hypervariable regions of PorA, and the nucleotide sequences of the first allele of each of the 7474 pan-genome loci with the respective complete allelic profile matrix.</p> <p>All raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB34441).</p>
Comparative whole genome phylogeny of animal, environmental and human strains confirms the genogroups organization and the diversity of Stenotrophomonas maltophilia
<p>Reannotation of Smc genomes from Refseq (Prokka v1.13) and gene presence and absence spreadsheet from Roary.</p>
EGP Mitochondrial Genome Analysis on Human Genome Diversity Project Whole-Genome Sequencing Data
<p><strong>Summary: </strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the HGDP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Simons Genome Diversity Project</strong>: Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>2b388c1fa446ecec70e33ea0471e06f8</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>50b80ed32b1ae542c8967cc31986dd19</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>995f30b74c4bb094a674b1a994853246</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Copy Number</td> <td>e79e61efab4c491fa2825b7d1853df58</td> </tr> </tbody> </table> </div> <div>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</div>
EGP Mitochondrial Genome Analysis on Simons Genome Diversity Project Whole-Genome Sequencing Data
<p><strong>Summary: </strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the SGDP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Simons Genome Diversity Project</strong>: Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <table> <tbody> <tr> <th>Public Dataset</th> <th>EGP Result File Type</th> <th>MD5</th> </tr> </tbody> <tbody> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>86b09553f80926c1c29c57000ec1a88f</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>010026d77bee81e7b8daf5836bd12da3</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>f1ea3edf4a82b42f2028467fb3544dc4</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Copy Number</td> <td>4872eeb792c214ad49662e98e4b14620</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p>
Machine learning reveals the diversity of human 3D chromatin contact patterns (example predictions genome wide)
<p>Example data for the paper: Machine learning reveals the diversity of human 3D chromatin contact patterns</p> <p>GitHub: https://github.com/erin-n-gilbertson/3DGenome-diversity/tree/main</p> <p>biorXiv: https://www.biorxiv.org/content/10.1101/2023.12.22.573104v1.full</p> <p>Manuscript accepted at Molecular Biology and Evolution</p> <p>Of primary interest will be the example predictions genome wide for hg38 reference, human-archaic hominin ancestor and most divergent 1KG individual per genome along with the Jupyter notebook tutorial for making your own Akita predictions given any input 1MB sequence.</p> <div> <ul> <li>bin: contains python script for and qsub array shell script for generating example predictions. These scripts can be modified to take in any fasta files as input.</li> <li>akita_predictions: contains both Akita prediction output arrays and SVG files with predicted contact maps for the hg38 reference, human-archaic hominin ancestor and most divergent 1KG individual in each of 4,873 1MB windows</li> <li>anc_window_spearman.csv: spearman correlation between each 1KG individual and the ancestor for each 1MB window. To calculate 3D divergence subtract these values from 1.</li> <li>basenji: basenji dir from their github, necessary in the directory to run predictions - https://github.com/calico/basenji/tree/master</li> <li>genomes: fasta genomes for hg38 reference and human-archaic hominin ancestor used to make akita predictions</li> <li>divergent_windows: variants and expected divergence distributions for 392 more divergent than expected windows. Defined in the manuscript as windows where 3D divergence between 1KG indiivudals and the ancestor is greater than what would be expected based on sequence divergence. See manuscript Fig. S9 for more details. </li> <li>windows.txt: 4,873 1MB genomic windows with 100% coverage in hg38 used for Akita predictions</li> <li>making_examples.ipynb: jupyter notebook with tutorial instructions for making Akita predictions on any human genome sequence.</li> </ul> <br><br></div>
Data from: Aquatic adaptation and depleted diversity: a deep dive into the genomes of the sea otter and giant otter
Despite its recent invasion into the marine realm, the sea otter (Enhydra lutris) has evolved a suite of adaptations for life in cold coastal waters, including limb modifications and dense insulating fur. This uniquely dense coat led to the near-extinction of sea otters during the 18th-20th century fur trade and an extreme population bottleneck. We used the de novo genome of the southern sea otter (E. l. nereis) to reconstruct its evolutionary history, identify genes influencing aquatic adaptation, and detect signals of population bottlenecks. We compared the genome of the southern sea otter to the tropical freshwater-living giant otter (Pteronura brasiliensis) to assess common and divergent genomic trends between otter species, and to the closely related northern sea otter (E. l. kenyoni) to uncover population-level trends. We found signals of positive selection in genes related to aquatic adaptations, particularly limb development and polygenic selection on genes related to hair follicle development. We found extensive pseudogenization of olfactory receptor genes in both the sea otter and giant otter lineages, consistent with patterns of sensory gene loss in other aquatic mammals. At the population level, the southern sea otter and the northern sea otter showed extremely low genomic diversity, signals of recent inbreeding, and demographic histories marked by population declines. These declines pre-date the fur trade and appear to have resulted in an increase in putatively deleterious variants that could impact the future recovery of the sea otter.
Data from: A RAD-sequencing approach to genome-wide marker discovery, genotyping, and phylogenetic inference in a diverse radiation of primates
Until recently, most phylogenetic and population genetics studies of nonhuman primates have relied on mitochondrial DNA and/or a small number of nuclear DNA markers, which can limit our understanding of primate evolutionary and population history. Here, we describe a cost-effective reduced representation method (ddRAD-seq) for identifying and genotyping large numbers of SNP loci for taxa from across the New World monkeys, a diverse radiation of primates that shared a common ancestor ~20-26 mya. We also estimate, for the first time, the phylogenetic relationships among 15 of the 22 currently-recognized genera of New World monkeys using ddRAD-seq SNP data using both maximum likelihood and quartet-based coalescent methods. Our phylogenetic analyses robustly reconstructed three monophyletic clades corresponding to the three families of extant platyrrhines (Atelidae, Pitheciidae and Cebidae), with Pitheciidae as basal within the radiation. At the genus level, our results conformed well with previous phylogenetic studies and provide additional information relevant to the problematic position of the owl monkey (Aotus) within the family Cebidae, suggesting a need for further exploration of incomplete lineage sorting and other explanations for phylogenetic discordance, including introgression. Our study additionally provides one of the first applications of next-generation sequencing methods to the inference of phylogenetic history across an old, diverse radiation of mammals and highlights the broad promise and utility of ddRAD-seq data for molecular primatology.
Supplementary Data for: Whole genome sequencing elucidates the species-wide diversity and evolution of fungicide resistance in the early blight pathogen Alternaria solani
<p>Supplementary Data for: Whole genome sequencing elucidates the species-wide diversity and evolution of fungicide resistance in the early blight pathogen Alternaria solani</p> <p>This repository contains:</p> <p>SNP call data / VCF file</p> <p>Scripts for all processing steps from mapping up to PCA and phylogenetic analyses (script.ts)<br> Scripts for population genomic analyses with LEA and PopGenome (scripts.SE)<br> All script names are self explanatory.</p>
Methodological challenges in the genomic analysis of an endangered mammal population with low genetic diversity
<p><span>Recently, populations of various species with very low genetic diversity have been discovered. Some of these persist in the long term, but others could face extinction due to accelerated loss of fitness. In this work, we characterize 45 individuals of one of these populations, belonging to the Iberian desman (<em>Galemys</em> <em>pyrenaicus</em>). For this, we used the ddRADseq technique, which generated 1,421 SNPs. The heterozygosity values of the analyzed individuals were among the lowest recorded for mammals, ranging from 26 to 91 SNPs/Mb. Furthermore, the individuals from one of the localities, highly isolated due to strong barriers, presented extremely high inbreeding coefficients, with values above 0.7. Under this scenario of low genetic diversity and elevated inbreeding levels, some individuals appeared to be almost genetically identical. We used different methods and simulations to determine if genetic identification and parentage analysis were possible in this population. Only one of the methods, which does not assume population homogeneity, was able to identify all individuals correctly. Therefore, genetically impoverished populations pose a great methodological challenge for their genetic study. However, these populations are of primary scientific and conservation interest, so it is essential to characterize them genetically and improve genomic methodologies for their research.</span></p>
Genomic diversity gradients and functional differentiation put Northeast Pacific ribbon kelp lineages in the speciation grey zone
<p><span>The transition from reproductively isolated populations to species is not well understood. Genotyping entire genomes holds promise to enhance insights into the process of speciation and the evolutionary relationships among related taxa. Gulf of Alaska ribbon kelp was once recognized as four species before they were folded into <em>Alaria</em> <em>marginata</em> on the basis of DNA barcode markers, though several lineages have continued to be recognized. Here, we used whole genome sequencing datasets to test the hypothesis that these lineages represent incipient species. Whole genomes of 69 individuals from five genetically distinctive lineages in the Gulf of Alaska (USA) and Salish Sea (Canada) were analyzed, along with 63 genomes from three other species of <em>Alaria</em>. Our analysis of >3.4 million Single Nucleotide Polymorphisms reaffirms that organellar and nuclear phylogenetic signals are incongruent in <em>Alaria</em>, producing different topologies among five organellar and six nuclear <em>A</em>. <em>marginata</em> lineages. Lineages also display reproductive isolation, evidenced by a lack of recent admixture across genomes. Genetic distances between <em>A</em>. <em>marginata</em> lineages exceed levels expected of population-level divergence but fall short of distances between species of <em>Alaria</em>. Moreover, we provide evidence of functional genomic differences between the <em>A</em>. <em>marginata</em> lineages, exceeding differences expected between populations, but falling short of larger differences among species. Our results place <em>A</em>. <em>marginata</em> lineages in an evolutionary grey zone, where lineages display substantial differentiation, but not to the level expected of <em>Alaria</em> species. This information shifts taxonomic conversations towards a genome-scale framework that provides a more comprehensive picture of divergence, connectivity, and functional innovation for defining lineages.</span></p>
Genomic diversity and differentiation between island and mainland populations of White‐tailed Eagles (Haliaeetus albicilla)
<p>Using whole genome shotgun sequences from 92 white-tailed eagles (<em>Haliaeetus albicilla</em>) sampled from Greenland, Iceland, Norway, Denmark, Estonia, and Turkey between 1885–1950 and after 1990, we investigate the genomic variation within countries over time, and between countries. Clear signatures of ancient biogeographic substructure across Europe and the North‐East Atlantic are observed. The greatest genomic differentiation was observed between island (Greenland and Iceland) and mainland (Denmark, Norway and Estonia) populations. The two island populations share a common ancestry from a single mainland population, distinct from the other sampled mainland populations, and despite the potential for high connectivity between Iceland and Greenland they are well separated from each other and are characterized by inbreeding and little variation. Temporal differences also highlight a pattern of regional populations persisting despite the potential for admixture. All sampled populations generally showed a decline in effective population size over time, which may have been shaped by four historical events: I) isolation of refugia during the last glacial period 110‐115,000 years ago, II) population divergence following the colonization of the deglaciated areas ~10,000 years ago, III) human population expansion, which led to the settlement in Iceland ~1,100 years ago, and IV) human persecution and exposure to toxic pollutants during the last two centuries.</p>
Resolution of structural variation in diverse mouse genomes reveals chromatin remodeling due to transposable elements
<p>Structural variant calls, RepeatMasker annotations, and genome assemblies of diverse mouse genomes. </p>
Predictors of genomic diversity within North American squamates
<p>Comparisons of intraspecific genetic diversity across species can reveal the roles of geography, ecology, and life history in shaping biodiversity. The wide availability of mitochondrial DNA (mtDNA) sequences in open-access databases makes this marker practical for conducting analyses across several species in a common framework, but patterns may not be representative of overall species diversity. Here, we gather new and existing mtDNA sequences and genome-wide nuclear data (genotyping-by-sequencing; GBS) for 30 North American squamate species sampled in the Southeastern and Southwestern United States. We estimated mtDNA nucleotide diversity for two mtDNA genes, COI (22 species alignments; average 16 sequences) and cytb (22 species; average 58 sequences), as well as nuclear heterozygosity and nucleotide diversity from GBS data for 118 individuals (30 species; four individuals and 6,820–44,309 loci per species). We showed that nuclear genomic diversity estimates were highly consistent across individuals for some species, while other species showed large differences depending on the locality sampled. Range size was positively correlated with both cytb diversity (Phylogenetically Independent Contrasts: R<sup>2</sup> = 0.31, p = 0.007) and GBS diversity (R<sup>2</sup> = 0.21; p = 0.006), while other predictors differed across the top models for each dataset. Mitochondrial and nuclear diversity estimates were not correlated within species, although sampling differences in the data available made these datasets difficult to compare. Further study of mtDNA and nuclear diversity sampled across species' ranges is needed to evaluate the roles of geography and life history in structuring diversity across a variety of taxonomic groups.</p>
Supplemental data for: Endophyte genomes support greater metabolic gene cluster diversity compared with non-endophytes in Trichoderma
<p><em>Trichoderma</em> is a cosmopolitan genus with diverse lifestyles and nutritional modes, including mycotrophy, saprophytism, and endophytism. Previous research has reported greater metabolic gene repertoires in endophytic fungal species compared to closely-related non-endophytes. However, the extent of this ecological trend and its underlying mechanisms are unclear. Some endophytic fungi may also be mycotrophs and have one or more mycoparasitism mechanisms. Mycotrophic endophytes are prominent in certain genera like <em>Trichoderma</em>, therefore, the mechanisms that enable these fungi to colonize both living plants and fungi may be the result of expanded metabolic gene repertoires. Our objective was to determine what, if any, genomic features are overrepresented in endophytic fungi genomes in order to undercover the genomic underpinning of the fungal endophytic lifestyle. Here we compared metabolic gene cluster and mycoparasitism gene diversity across a dataset of thirty-eight <em>Trichoderma</em> genomes representing the full breadth of environmental <em>Trichoderma</em>'s diverse lifestyles and nutritional modes. We generated four new <em>Trichoderma endophyticum</em> genomes to improve the sampling of endophytic isolates from this genus. As predicted, endophytic <em>Trichoderma</em> genomes contained, on average, more total biosynthetic and degradative gene clusters than non-endophytic isolates, suggesting that the ability to create/modify a diversity of metabolites potential is beneficial or necessary to the endophytic fungi. Still, once the phylogenetic signal was taken into consideration, no particular class of metabolic gene cluster was independently associated with the <em>Trichoderma</em> endophytic lifestyle. Several mycoparasitism genes, but no chitinase genes, were associated with endophytic <em>Trichoderma</em> genomes. Most genomic differences between <em>Trichoderma</em> lifestyles and nutritional modes are difficult to disentangle from phylogenetic divergences among species, suggesting that <em>Trichoderma</em> genomes may be particularly well-equipped for lifestyle plasticity. We also consider the role of endophytism in diversifying secondary metabolism after identifying the horizontal transfer of the ergot alkaloid gene cluster to <em>Trichoderma</em>.</p>
Data from: Genome-wide epitope mapping reveals significant diversity in antibody responses to Coxiella burnetii vaccination and infection
<p><em>Coxiella burnetii</em> is an important zoonotic bacterial pathogen of global importance, causing the disease Q fever in a wide range of animal hosts. Ruminant livestock, in particular sheep and goats, are considered the main reservoir of infection. Vaccination is a key control measure and two commercial vaccines based on formalin-inactivated<em> C. burnetii </em>bacterins are currently available. However, their deployment is limited due to significant reactogenicity in individuals previously sensitized to <em>C. burnetii </em>antigens. Furthermore, these vaccines interfere with available serodiagnostic tests which are also based on <em>C. burnetii</em> bacterin preparations. Subunit vaccines based on recombinant proteins offer significant advantages, as they can be designed to reduce reactogenicity and can be co-designed with defined antigen serodiagnostic tests to allow discrimination between vaccinated and infected individuals. This study aimed to investigate the diversity of antibody responses to <em>C. burnetii </em>vaccination and/or infection in cattle, goats, humans, and sheep through genome-wide linear epitope mapping to identify candidate vaccine and diagnostic antigens within the predicted bacterial proteome. Using high-density peptide microarrays, we analyzed the seroreactivity in 156 serum samples from vaccinated and infected individuals to peptides derived from 2,092 ORFs in the <em>C. burnetii</em> genome. We found significant diversity in the antibody responses within and between species and across different <em>C. burnetii</em> exposure statuses. However, <em>C. burnetii</em> exposure did result in more uniform seroreactivity across species. Through the implementation of three different vaccine candidate methods, we identified 493 candidate protein antigens for protein subunit vaccine design or serodiagnostic, out of which 65 have been previously described. This is the first study to investigate seroreactivity against the entire <em>C. burnetii </em>genome presented as overlapping linear peptides and provides the basis for selection of antigen targets for next generation Q fever vaccines and diagnostic tests.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.