Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
183
datasets available to search
ShareScore release 0.9.0
Dataset results
183 results for “Genome conservation”
Conservation genomic analysis of the Croatian indigenous Black Slavonian and Turopolje pig breeds
<p>The majority of the nearly 400 existing local pig breeds are adapted to specific environments and human needs. The demand for large production quantities and the industrialized pig production have caused a rapid decline of many local pig breeds in recent decades. Black Slavonian pig and Turopolje pig, the latter highly threatened, are the two Croatian local indigenous breeds typically grown in extensive or semi-intensive systems. In order to guide a long-term breeding program to prevent the disappearance of these breeds, we analyzed their genetic diversity, inbreeding level and relationship with other local breeds across the world, as well as modern breeds and several wild populations, using high throughput genomic data obtained using the Illumina Infinium PorcineSNP60 v2 BeadChip. Multidimensional scaling analysis positioned Black Slavonian pigs close to the UK/North American breeds, while the Turopolje pig clustered within the Mediterranean breeds. Turopolje pig showed a very high inbreeding level (F<sub>ROH>4Mb</sub>=0.400 and F<sub>ROH>8Mb</sub>=0.332) that considerably exceeded the level of full-sib mating, while Black Slavonian pig showed much lower inbreeding (F<sub>ROH>4Mb</sub>=0.098 and F<sub>ROH>8Mb</sub>=0.074), indicating a planned mating strategy. In Croatian local breeds we identified several genome regions showing adaptive selection signals that were not present in commercial breeds. The results obtained in this study reflect the current genetic status and breeding management of the two Croatian indigenous local breeds. Given the small populations of both breeds, a controlled management activity has been implemented in Black Slavonian pigs since their commercial value has been recognized. In contrast, the extremely high inbreeding level observed in Turopolje pig argues for an urgent conservation plan with a long-term, diversity-oriented breeding program.</p>
Data From: Characterizing patterns of genomic variation in the threatened Utah prairie dog: implications for conservation and management
<p>Utah prairie dogs (<i>Cynomys parvidens</i>) are federally threatened due to eradication campaigns, habitat destruction, and outbreaks of plague. Today, Utah prairie dogs exist in small, isolated populations, making them less demographically stable and more susceptible to erosion of genetic variation by genetic drift. We characterized patterns of genetic structure at neutral and putatively adaptive loci in order to evaluate the relative effects of genetic drift and local adaptation on population divergence. We sampled individuals across the Utah prairie dog species range and generated 2,955 single nucleotide polymorphisms (SNPs) using double digest restriction site associated DNA sequencing (ddRAD). Genetic diversity was lower in low elevation sites compared to high elevation sites. Population divergence was high among sites and followed an isolation-by-distance (IBD) model. Our results indicate that genetic drift plays a substantial role in the population divergence of the Utah prairie dog, and colonies would likely benefit from translocation of individuals between recovery units, which are characterized by distinct elevations, despite the detection of environmental associations with outlier loci. By understanding the processes that shape genetic structure, better informed decisions can be made with respect to the management of threatened species to ensure that adaptation is not stymied.</p>
Genomic variation in the American pika: signatures of geographic isolation and implications for conservation
<p>Distributional responses by alpine taxa to repeated, glacial-interglacial cycles throughout the last two million years have significantly influenced the spatial genetic structure of populations. These effects have been exacerbated for the American pika (<i>Ochotona princeps</i>), a small alpine lagomorph constrained by thermal sensitivity and a limited dispersal capacity. As a species of conservation concern, long-term lack of gene flow has important consequences for landscape genetic structure and levels of diversity within populations. Here, we use reduced representation sequencing (ddRADseq) to provide a genome-wide perspective on patterns of genetic variation across pika populations representing distinct subspecies. To investigate how landscape and environmental features shape genetic variation, we collected genetic samples from distinct geographic regions as well as across finer spatial scales in two geographically proximate mountain ranges of eastern Nevada.</p> <p> </p> <p> </p>
Data from: Genome sequence and population declines in the critically endangered greater bamboo lemur (Prolemur simus) and implications for conservation
Background: The greater bamboo lemur (Prolemur simus) is a member of the Family Lemuridae that is unique in their dependency on bamboo as a primary food source. This Critically Endangered species lives in small forest patches in eastern Madagascar, occupying a fraction of its historical range. Here we sequence the genome of the greater bamboo lemur for the first time, and provide genome resources for future studies of this species that can be applied across its distribution. Results: Following whole genome sequencing of five individuals we identified over 152,000 polymorphic single nucleotide variants (SNVs), and evaluated geographic structuring across nearly 19k SNVs. We characterized a stronger signal associated with a north-south divide than across elevations for our limited samples. We also evaluated the demographic history of this species, and infer a dramatic population crash. This species had the largest effective population size (estimated between ~900,000 to one million individuals) between approximately 60,000-90,000 years before present (ybp), during a time in which global climate change affected terrestrial mammals worldwide. We also note the single sample from the northern portion of the extant range had the largest effective population size around 35,000 ybp. Conclusions: From our whole genome sequencing we recovered an average genomic heterozygosity of 0.0037%, comparable to other lemurs. Our demographic history reconstructions recovered a probable climate-related decline (60-90,000 ybp), followed by a second population decrease following human colonization, which has reduced the species to a census size of approximately 1,000 individuals. The historical distribution was likely a vast portion of Madagascar, minimally estimated at 44,259 km2, while the contemporary distribution is only ~1,700 km2. The decline in effective population size of 89-99.9% corresponded to a vast range retraction. Conservation management of this species is crucial to retain genetic diversity across the remaining isolated populations.
Data from: Genome-wide SNP discovery in the annual herb, Lasthenia fremontii (Asteraceae): genetic resources for the conservation and restoration of a California vernal pool endemic
California vernal pool (VP) ecosystems support a diverse community of endemic plants that are threatened by multiple anthropogenic pressures, generating a need for molecular tools to quantify the extent and distribution of genetic variation in native populations. Here, we used RADseq to discover single nucleotide polymorphisms (SNPs) for a widespread VP endemic plant species, Lasthenia fremontii. We discovered nuclear-based SNPs using a RAD-tag library of 12 individuals from different VP complexes using SbfI, a restriction enzyme that does not cleave the chloroplast genome in Lasthenia. A total of 316,106 catalog loci were obtained across all twelve individuals in the library. Of these, 713 loci were polymorphic, yielding 3918 candidate SNPs. Next, we genotyped an additional 285 additional plants to validate and characterize 71 of the candidate SNPs. Of these, 44 were polymorphic among VP complexes. A preliminary analysis of the distribution of genetic variation using these loci revealed significant isolation-by-distance across the species' geographic range. Weaker, but in some cases significant, genetic differentiation was detected among subpopulations from different pools within a single VP complex. Thus, in this study, RADseq allowed the discovery of SNP markers that can characterize patterns of genetic variation at multiple spatial scales in L. fremontii, which can be used to inform the conservation and mitigation of VP populations.
Data from: Sturgeon conservation genomics: SNP discovery and validation using RAD sequencing
Caviar-producing sturgeons belonging to the genus Acipenser are considered to be one of the most endangered species groups in the world. Continued overfishing in spite of increasing legislation, zero catch quotas and extensive aquaculture production have led to the collapse of wild stocks across Europe and Asia. The evolutionary relationships among Adriatic, Russian, Persian and Siberian sturgeons are complex because of past introgression events and remain poorly understood. Conservation management, traceability and enforcement suffer a lack of appropriate DNA markers for the genetic identification of sturgeon at the species, population and individual level. This study employed RAD sequencing to discover and characterize single nucleotide polymorphism (SNP) DNA markers for use in sturgeon conservation in these four tetraploid species over three biological levels, using a single sequencing lane. Four population meta-samples and eight individual samples from one family were barcoded separately before sequencing. Analysis of 14.4 Gb of paired-end RAD data focused on the identification of SNPs in the paired-end contig, with subsequent in silico and empirical validation of candidate markers. Thousands of putatively informative markers were identified including, for the first time, SNPs that show population-wide differentiation between Russian and Persian sturgeons, representing an important advance in our ability to manage these cryptic species. The results highlight the challenges of genotyping-by-sequencing in polyploid taxa, while establishing the potential genetic resources for developing a new range of caviar traceability and enforcement tools.
Data from: Conservation of the genome-wide recombination rate in white-footed mice
Despite being linked to the fundamental processes of chromosome segregation and offspring diversification, meiotic recombination rates vary within and between species. Recent years have seen progress in quantifying recombination rate evolution across multiple temporal and genomic scales. Nevertheless, the level of variation in recombination rate within wild populations – a key determinant of evolution in this trait – remains poorly documented on the genomic scale. To address this notable gap, we used immunofluorescent cytology to quantify genome-wide recombination rates in a wild population of male white-footed mouse, Peromyscus leucopus. For comparison, we measured recombination rates in a second population of male P. leucopus raised in the laboratory and in a lab strain of deer mouse, Peromyscus maniculatus bairdii. Although we found differences between individuals in the genome-wide recombination rate, levels of variation were low – within populations, between populations, and between species. Quantification of synaptonemal complex length and crossover positions along chromosome 1 using a novel automated approach also revealed conservation in broad-scale crossover patterning, including strong crossover interference. We propose stabilizing selection targeting recombination or correlated processes as the explanation for these patterns.
Dataset for: Conservation genomics of federally endangered Texella harvester species (Arachnida, Opiliones, Phalangodidae) from cave and karst habitats of central Texas
<p>Genomic-scale data for non-model taxa are providing new insights into landscape genomic structuring and species limits, leading to more informed conservation decisions, particularly in taxa with extremely restricted microhabitat preferences and small geographic distributions. This study applied sequence capture of ultraconserved elements (UCEs) to gather genomic-scale data for two federally endangered <i>Texella</i> harvester species distributed in Edwards Formation cave and karst habitats of central Texas, near Austin. We gathered UCE data for 51 <i>T. reyesi </i>specimens from 46 different caves, seven <i>T. reddelli</i> specimens from five caves, and from relevant outgroup species. For these UCE data we applied a combination of phylogenomic, multispecies coalescent phylogenetic, and single-nucleotide polymorphism (SNP) machine-learning analyses. We found that samples of <i>T. reddelli</i> and <i>T. reyesi</i> together form a single clade in phylogenetic analyses, but that <i>T. reddelli </i>samples are not recovered as monophyletic. Instead, <i>T. reddelli </i>samples from three northern caves are embedded within a larger <i>T. reyesi</i> genetic clade. Significantly, the genetic structuring of all samples closely follows geologic barriers defined for the region and formalized as karst fauna regions (KFRs). One exception is the Jollyville Plateau KFR, which includes two divergent, non-sister genetic lineages. Levels of troglomorphy, here assessed by a simple scoring of corneal and retinal development, also closely follows clade (and geographic) boundaries, implying that divergent genetic lineages might also have distinct ecologies. Overall, our study has important taxonomic implications, is the first to explore (and validate) regional KFR boundaries using intraspecific genetic data, and provides essential data for future management decisions involving these federally endangered species.</p>
Taxonomy based on limited genomic markers may underestimates species diversity of rockhopper penguins and threaten their conservation
<p><span><span><span><span><span><span><span><span><span><span><span>Delimiting recently diverged species is challenging. During speciation, genetic differentiation may be distributed unevenly across the genome, as different genomic regions can be subject to different selective pressures and evolutionary histories. Reliance on limited numbers of genetic markers that may be underpowered can make species delimitation even more challenging, potentially resulting in taxonomic inconsistencies. Rockhopper penguins of the genus <i>Eudyptes</i> comprise three broadly recognized taxa: northern (<i>E. moseleyi</i>), southern (<i>E. chrysocome</i>), and eastern rockhopper (<i>E. filholi</i>). Their taxonomic status has been controversial for decades, with researchers disagreeing about whether <i>E. chrysocome</i> and <i>E. filholi</i>are distinct species or conspecific. Our goal is to evaluate genome-wide patterns of divergence to evaluate genetic differentiation and species delimitation in<i> </i>rockhopper penguins<i>, </i>and to assess which mechanisms may underlie previous discordance among nuclear versus mitochondrial analyses. We generated reduced-representation genomic libraries using Double Digest Restriction-site Associated DNA (ddRAD) sequencing to evaluate genetic differentiation, contemporary migration rates and admixture among colonies of rockhopper penguins. The extent of genetic differentiation among the three taxa was consistently higher than population-level genetic differentiation found within these and other penguin species. There was no evidence of admixture among the three taxa, suggesting the absence of ongoing gene flow among them. Species delimitation analyses based on molecular data, along with other lines of evidence, provide strong support for the taxonomic distinction of three species of rockhopper penguins. Our results provide strong support for the existence of three distinct species of rockhopper penguins. The recognition of this taxonomic diversity is crucial for the management and conservation of this widely distributed species group. This study illustrates that widespread dispersive seabird lineages lacking obvious morphological differences may nevertheless have complex evolutionary histories and comprise cryptic species diversity. </span></span></span></span></span></span></span></span></span></span></span></p>
Comparative genomics of Nectriaceae, including freshwater fungi show environment adaptations and conservation strate-gies for fungi
Open the record for dataset details and reuse information.
Comparative genomics of Nectriaceae, including freshwater fungi show environment adaptations and conservation strategies for fungi - Supplementary Material
Open the record for dataset details and reuse information.
Data and scripts for "Seascape genomics as a new tool to empower coral reef conservation strategies: an example on north-western Pacific Acropora digitifera."
<p>The archive contains scripts and data used in the work "Seascape genomics as a new tool to empower coral reef conservation strategies: an example on north-western Pacific Acropora digitifera."</p> <p>See the readme_metadata.pdf for a description of the files in the archive.</p>
Comparative genomics of fungi in Nectriaceae reveals their environmental adaptation and conservation strategies
Open the record for dataset details and reuse information.
Datasets - Unveiling Host-Parasite Relationships through Conserved MITEs in Prokaryote and Viral Genomes
<p><em><span>Title:</span></em></p> <p><strong><span>Unveiling Host-Parasite Relationships through Conserved MITEs in Prokaryote and Viral Genomes<span> </span></span></strong></p> <p><em><span> </span></em></p> <p><em><span>Authors:</span></em></p> <p><span>Francisco Nadal-Molero<sup>(1)</sup>, Riccardo Roselli<sup>(1)</sup>, Silvia Garcia-Juan<sup>(1)</sup>, Alicia Campos-Lopez<sup>(1)</sup>, Ana-Belen Martin-Cuadrado<sup>(1*)</sup></span></p> <p> </p> <p><strong><span>SUPPLEMENTARY FILES</span></strong></p> <p><strong><span> </span></strong></p> <p><strong><span>Supplementary File S1.</span></strong><span> Sequences of cMITEs detected in Bacteria genomes (<em>fasta</em> format). The hosting microbial species and inferred NCBI-taxonomy are indicated in the name of each sequence. The structure of the MITE name is: “Accession|Genome|start|end|TSD|TIRlength|MITETracker_group|Lineage”.</span></p> <p><strong><span>Supplementary File S2.</span></strong><span> Sequences of cMITEs detected in the Archaea genomes (<em>fasta</em> format). The hosting microbial species and inferred NCBI-taxonomy are indicated in the name of each sequence. The structure of the MITE name is: “Accession|Genome|start|end|TSD|TIRlength|MITETracker_group|Lineage”.</span></p> <p><strong><span>Supplementary File S3.</span></strong><span> Sequences of vMITEs detected in the virus sequences from the NCBI and IMG/VR v.4.1 database (<em>fasta</em> format). Virus, microbial host (if known) and inferred NCBI-taxonomy is stated in the name of each sequence. The structure of the MITE name is: </span></p> <p><span>“Accession|Genome|start|end|TSD|TIRlength|MITETracker_group|Virus|Name|Host”.</span></p> <p><strong><span>Supplementary File S4.</span></strong><span> Sequences of <span>si-vMITEs</span> detected in the virus sequences from the NCBI and IMG/VR v.4.1 database (<em>fasta</em> format). Virus, microbial host (if known) and inferred NCBI-taxonomy are stated in the name of each sequence. The structure of the MITE name is: “Accession|Genome|start|end|Ident.Method.by.DB|Host”.</span></p> <p><strong><span>Supplementary Files S5. </span></strong><span>Cytoscape networks. (A) Figure 1A, (B) Figure 1B.</span></p> <p><strong><span>Supplementary File S6.</span></strong><span> Sequences of cMITEs obtained from <span>5837</span> genomes of Neisseriales. The structure of the MITE name is: </span></p> <p><span>“Accession|NucleotideID|start|end|TSD|TIRlength|MITETracker_group|Genome|Lineage”.</span></p> <p><strong><span>Supplementary File S7.</span></strong><span> Sequences of si-vMITEs obtained from <span>5837</span> genomes of Neisseriales. The structure of the MITE name is: “Accession|Genome|start|end|Host”.</span></p> <p><strong><span>Supplementary File S8.</span></strong><span> Sequences of cMITEs obtained from 46051 genomes of Bacteroidota. The structure of the MITE name is:</span></p> <p><span>“Accession|NucleotideID|start|end|TSD|TIRlength|MITETracker_group|Genome|Lineage”.</span></p> <p><strong><span>Supplementary File S9.</span></strong><span> Sequences of si-vMITEs obtained from 46051 genomes of Bacteroidota. The structure of the MITE name is: “Accession|Genome|start|end|Host”.</span></p>
Supporting data for: A comparative analysis of planarian genomes reveals regulatory conservation in the face of rapid structural divergence
<p>This upload contains genome assemblies and annotation files for four planarian species. </p> <h2>Genome assemblies</h2> <p>The genome assemblies are deposited as FASTA files with the ending '.fa.gz'.</p> <h2>Genome annotations</h2> <p>The genome annotations are deposited as gff3 files and contain the 'ENCODE' keyword.</p> <p>Integration into Wormbase are currently ongoing to provide userfriendly access.</p> <h2>Transposable element annotation</h2> <p>Transposable element annotation performed with the EDTA pipeline are deposited as GFF3 files with the file ending '.EDTA.TEanno.gff3.gz'.</p> <h2>Satellite DNA</h2> <p>Repetitive satellite regions annotated using RepeatExplorer and SRF are deposited as GFF3 files with the ending 'satDNA.gff3'.</p>
Conservation genomics in an endangered citrus relative
<p>Related paper: Conservation genomics insights into the spontaneous integration of multiple reproductive strategies in an endangered citrus relative</p> <p>Related source: https://github.com/wangnan9394/conservation_genomics</p> <p>#admixture.zip</p> <p>This file including population ancestry of <em>Fortunella</em>. The estimated admixture proportions ranged from K=2 to K=5.</p> <p>#apomixis_locus_200kb_vcf.zip</p> <p>This file including VCF file of the WILDAPO group. The proportion of introgressions from cultivars and the number of heterozygous introgressed variations in samples from WILDAPO was calculated based on this file.</p> <p>#deleterious_mutations.zip</p> <p>This file including the summary of deleterious mutations in <em>Fortunella.</em> The BED file annotated the types of each variation based on the ancestry state.</p> <p>#fd_and_D_statistics.zip</p> <p>There are population level D statistic (ABBA-BABA statistic) and genome-wide fd statistic, using the Dsuite, ANGSD and genomic_general programs.</p> <p>#genome-wide_haplotype_results.zip</p> <p>We calculated the genome-wide haplontic divergence in each sample from the WILDAPO group, and the summary was zipped in this file.</p> <p>#PCA_TajimaD_value.zip</p> <p>The PCA and Tajima’D value were calculated based on the variation map using PLINK. We have included the results of genome-wide Tajima’s D from five groups.</p> <p>#ROH_500kb.zip</p> <p>This file including ROH analysis of <em>Fortunella</em>.</p> <p>#S-locus_analysis.zip</p> <p>This file including VCF file of the S-locus (Chr. 1). The phylogeny of S-locus were constructed based on this VCF file.</p> <p>#SMC_files.zip</p> <p>This file including smc.gz files of five groups (CULAPO, CULSEX, WILDAPO, WILDSEX1, WILDSEX2). The approximate lookup table of each group in this zipped file.</p> <p>#SSMs.zip</p> <p>We calculated the introgression based on the genome-wide species-specific variations (this VCF format file).</p> <p>#Subtrees_TWISST.zip</p> <p>The combination of subtrees, which were calculated based on 25 kb non-overlapped windows. The input subtrees and the output files by TWISST in this file.</p> <p>#treemix.zip</p> <p>we estimated that the potential migration events ranged from 1 to 3 with five repeats for each test.</p> <p>#variation_map_kumquats.vcf.gz</p> <p>we filtered nuclear genomic variations based on depth of coverage and missing rates using VCFtools with the following criteria: variant quality (QD) > 2.0, quality score (QUAL) > 40.0, mapping quality (MQ) > 30.0, genotype calls with a depth > 2 or <100 and with < 20% of genotypes missing across all samples.</p>
Data from: Integrative approaches to guide conservation decisions: using genomics to define conservation units and functional corridors
Open the record for dataset details and reuse information.
Data from: Theory, practice, and conservation in the age of genomics: the Galápagos giant tortoise as a case study
Open the record for dataset details and reuse information.
Data from: Genome-wide SNP discovery in the annual herb, Lasthenia fremontii (Asteraceae): genetic resources for the conservation and restoration of a California vernal pool endemic
Open the record for dataset details and reuse information.
Data from: Sturgeon conservation genomics: SNP discovery and validation using RAD sequencing
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.