Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

448

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

448 results for “Genomic selection”

Learn how ShareScore rates datasets ↗
zenodo44/100

Highly parallel genomic selection response in replicated Drosophila melanogaster populations with reduced genetic variation

<p>Many adaptive traits are polygenic and frequently more loci contributing to the phenotype are segregating than needed to express the phenotypic optimum. Experimental evolution with replicated populations adapting to a new controlled environment provides a powerful approach to study polygenic adaptation. Since genetic redundancy often results in non-parallel selection responses among replicates, we propose a modified Evolve and Resequence (E&amp;R) design that maximizes the similarity among replicates. Rather than starting from many founders, we only use two inbred&nbsp;<em>Drosophila melanogaster</em>strains and expose them to a very extreme, hot temperature environment (29&deg;C). After 20 generations, we detect many genomic regions with a strong, highly parallel selection response in 10 evolved replicates. The X chromosome has a more pronounced selection response than the autosomes, which may be attributed to dominance effects. Furthermore, we find that the median selection coefficient for all chromosomes is higher in our two-genotype experiment than in classic E&amp;R studies. Since two random genomes harbor sufficient variation for adaptive responses, we propose that this approach is particularly well-suited for the analysis of polygenic adaptation.</p> <p>See the README.txt file to get&nbsp;a description of the uploaded files.&nbsp;Scripts.zip contains annotated command lines and scripts for the project&nbsp;(see internal README.txt file).</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Tracking Selection using Temporal Population Genomics Data

<p>This repository contains the implementation of a pipeline to run the simulations and to produce a reference table for the ABC-RF inference of demography and selection. In its new release, this repository contains the whole-genome polymorphism of contemporary and museum specimens of <em>Apis mellifera</em> feral populations analyzed&nbsp;by Cridland et al. (2018).</p>

opengpl-3.0Mar 2021View details →
zenodo44/100

Relate-estimated coalescence rates, allele ages, and selection p-values for the 1000 Genomes Project

<p><strong>Overview</strong></p> <p>Coalescence rates, allele ages, and p-values for evidence of positive selection calculated for 2478&nbsp;samples of the&nbsp;1000 Genomes Project&nbsp;using Relate.</p> <p>We estimated the joint genealogy of all 1000 GP populations and then extracted the embedded genealogy for each population.<br> For the genealogy of each population, we jointly estimated the population size history and branch lengths.&nbsp;<br> Variants segregating in more than one&nbsp;population&nbsp;therefore have&nbsp;correlated but different allele ages in each population.</p> <p>Please refer to&nbsp;<a href="https://www.nature.com/articles/s41588-019-0484-x">Speidel et al.&nbsp;Nature Genetics (2019)</a>&nbsp;for more details or email leo.speidel@outlook.com for any queries.</p> <p><strong>Coalescence rates</strong></p> <p>The zipped directory&nbsp;coalescence_rates.zip&nbsp;contains coalescence rates for 26 populations in the 1000 Genomes Project data set.</p> <ul> <li>The .coal files show the haploid coalescence rates, please refer to the&nbsp;<a href="https://myersgroup.github.io/relate/modules.html#PopulationSizeScript_FileFormats">Relate documentation</a>&nbsp;for the file format.</li> <li>The popsize.RData file is an R data frame storing the diploid population sizes (0.5/coalescence rate) calculated using the .coal files. The columns of this data frame, named &quot;pop_size&quot;,&nbsp;are <ul> <li>gens_ago: Time in generations at which epoch starts. (To get years from generations, we multiply by 28.)</li> <li>population_size: Diploid population size in this epoch.</li> <li>population: Name of population&nbsp;</li> <li>region: Name of region (AFR, AMR, EAS, EUR, SAS)</li> </ul> </li> </ul> <p><strong>Allele ages and selection p-values</strong></p> <p>The zipped directories&nbsp;allele_ages_*.zip&nbsp;contain&nbsp;R&nbsp;data frames for each 1000GP population storing allele ages and selection p-values.<br> Please note that only mutations that segregate in the population and map to a unique branch in the Relate-estimated marginal trees are included. Selection p-values are only provided for mutations of DAF &gt; 2 that pass quality filters (see Speidel et al., 2019).&nbsp;</p> <p>To get an age estimate for a neutral mutation, use&nbsp;0.5*(lower_age + upper_age). To get years from generations, we multiply by 28.</p> <p>The columns of these&nbsp;data frames, named &quot;allele_ages&quot;,&nbsp;are</p> <ul> <li>CHR: chromosome index</li> <li>BP: base-pair position (GRCh37)</li> <li>ID: id of SNP</li> <li>lower_age: Age in generations of coalescence event at the lower end of the branch onto which the mutation maps</li> <li>upper_age: Age in generations of coalescence event at the upper end of the branch onto which the mutation maps</li> <li>ancestral/derived: Ancestral/derived allele</li> <li>upstream: Upstream (5&#39;) allele</li> <li>downstream: Downstream (3&#39;) allele</li> <li>DAF: Derived-allele frequency</li> <li>pvalue: log10 p-value for selection evidence</li> </ul>

opencc-by-4.0May 2019View details →
dryad44/100

Data from: Genome-wide selection components analysis in a fish with male pregnancy

Open the record for dataset details and reuse information.

publicSep 2019View details →
zenodo40/100

skDER Representative Genomes for Select Bacterial Taxa

<p>Genomes belonging to a single genus or order were gathered using a loose search of taxonomic classifications in GTDB R214. By loose we required the string 'g__{GENUSNAME}' to be found in taxonomic info column by GTDB, thus allowing gathering of associated genera (which GTDB suggests are different, but literature/domain experts have yet to rename).</p><p>Genomes belonging to a taxa were dereplicated using skDER (v1.0.7) in "greedy" clustering mode with default values for parameters (99% ANI cutoff, 90% AF cutoff).</p><p>Overview of Files:</p><p>- The 'Genome_Dereplication_Overview.tsv' contains details of all the genomes considered as potential representatives for each taxonomic group and their GTDB R214 taxonomic classifications.</p><p>- 18 _Clustering_Information.txt files which contains the relationship information of non-representative genomes to their nearest representative genome. Generated using the `-n` argument in skder v.1.0.7. &nbsp;</p><p>-&nbsp;18 tar.gz compressed directories are provided. Each compressed directory features representative genomes in FASTA format determined for a particular taxon using skDER with greedy clustering and default cutoffs. Genome assemblies are renamed to feature both the GTDB taxonomic classification and the GCA identifier.<br>&nbsp; &nbsp; &nbsp; &nbsp; - Acinetobacter - 1,643 rep genomes (17.8% of 9,221 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Bacillales - 3,150 rep genomes (35.9% of 8,766 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Corynebacterium - 726 rep genomes (43.0% of 1,688 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Cutibacterium - 27 rep genomes (5.4% of 502 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Enterobacter - 878 rep genomes (19.9% of 4,408 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Enterococcus - 937 rep genomes (14.6% of 6,426 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Escherichia - 2,436 rep genomes (7.1% of 34,358 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Klebsiella - 1,022 rep genomes (5.6% of 18,145 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Lactobacillus - 541 rep genomes (30.9% of 1,747 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Listeria - 353 rep genomes (6.9% of 5,062 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Micromonospora - 211 rep genomes (73.3% of 288 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Mycobacterium - 744 rep genomes (6.9% of 10,657 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Neisseria - 414 rep genomes (12.8% of 3,235 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Pseudomonas - 2,666 rep genomes (18.9% of 14,066 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Salmonella - 308 rep genomes (2.2% of 14,109 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Staphylococcus - 496 rep genomes (2.5% of 19,627 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Streptococcus - 2,452 rep genomes (13.3% of 18,492 total genomes considered)<br>&nbsp; &nbsp; &nbsp; &nbsp; - Streptomyces - 1,555 rep genomes (57.7% of 2,697 total genomes considered)</p>

opencc-by-4.0Oct 2023View details →
dryad40/100

Efficient genomics based 'end-to-end' selective tree breeding framework

<p>Since their initiation in the 1950s, worldwide selective tree breeding programs followed the recurrent selection scheme of repeated cycles of selection, breeding (mating), and testing phases and essentially remained unchanged to accelerate this process or address environmental contingences and concerns. Here, we introduce an "end-to-end" selective tree breeding framework that: 1) leverages strategically preselected GWAS-based sequence data capturing trait architecture information, 2) generates unprecedented resolution of genealogical relationships among tested individuals, and 3) leads to the elimination of the breeding phase through the utilization of readily available wind-pollinated (OP) families. Individuals' breeding values generated from multi-trait multi-site analysis were also used in an optimum contribution selection protocol to effectively manage genetic gain/co-ancestry trade-offs and traits' correlated response to selection. The proof-of-concept study involved a 40-year-old spruce OP testing population growing on three sites in British Columbia, Canada, clearly demonstrating our method's superiority in capturing most of the available genetic gains in a substantially reduced timeline relative to the traditional approach. The proposed framework is expected to increase the efficiency of existing selective breeding programs, accelerate the start of new programs for ecologically and environmentally important tree species, and address climate-change caused biotic and abiotic stress concerns more effectively.</p>

opencc-zeroDec 2022View details →
dryad40/100

Selection shapes the genomic landscape of introgressed ancestry in a pair of sympatric sea urchin species

<p>A growing number of recent studies have demonstrated that introgression is common across the tree of life. However, we still have a limited understanding of the fate and fitness consequence of introgressed variation at the whole-genome scale across diverse taxonomic groups. Here, we implemented a phylogenetic hidden Markov model to identify and characterize introgressed genomic regions in a pair of well-diverged, non-sister sea urchin species: <em>Strongylocentrotus</em> <em>pallidus</em> and <em>S. droebachiensis</em>. Despite the old age of introgression, a sizable fraction of the genome (1% - 5%) exhibited introgressed ancestry, including numerous genes showing signals of historical positive selection that may represent cases of adaptive introgression. One striking result was the overrepresentation of hyalin genes in the identified introgressed regions despite observing considerable overall evidence of selection against introgression. There was a negative correlation between introgression and chromosome gene density, and two chromosomes were observed with considerably reduced introgression. Relative to the non-introgressed genome-wide background, introgressed regions had significantly reduced nucleotide divergence (<em>d</em><sub>XY</sub>) and overlapped fewer protein-coding genes, coding bases, and genes with a history of positive selection. Additionally, genes residing within introgressed regions showed slower rates of evolution (<em>d</em><sub>N</sub>, <em>d</em><sub>S</sub>, <em>d</em><sub>N</sub>/<em>d</em><sub>S</sub>) than random samples of genes without introgressed ancestry. Overall, our findings are consistent with widespread selection against introgressed ancestry across the genome and suggest that slowly evolving, low-divergence genomic regions are more likely to move between species and avoid negative selection following hybridization and introgression.</p>

opencc-zeroMar 2024View details →
zenodo40/100

Demographic history and natural selection shape patterns of deleterious mutation load and barriers to introgression across Populus genome

<p><br> Abbreviation of species names in each folder: Palb, P. alba; Pade, P. adenopoda; Pdav, P. davidiana; Ptra, P. tremula; Ptrs, P. tremuloides; Prot, P. rotundifolia; Pqio,P. qiongdaoensis.</p> <p>1. FST<br> Relative divergence (FST) for pairwise species comparisons was calculated for all sites with 100 Kbp non-overlapping windows.&nbsp;</p> <p>2. dxy<br> Absolute divergence (dxy) was calculated for all sites with 100 Kbp non-overlapping windows.&nbsp;</p> <p>3. Nucleotide diversity<br> Nucleotide diversity (&pi;) was calculated for all sites with 100 Kbp non-overlapping windows.&nbsp;</p> <p>4. Derived allele frequency<br> The derived frequencies of 4 different functional categories. Each folder contains seven Populus resluts</p> <p>5. Derived_allele_statistics<br> The statistics of homozygous and &nbsp;heterozygous derived alleles for loss of function, deleterious, tolerated and synonymous variants for each individual. The last two individuals in each file are outgroups&nbsp;</p> <p>6. dsuite-dinvestigate<br> The outputs of 10 trios using program Dinvestigate from Dsuite. The sliding window is 50 SNPs, and the step is 20 SNPs.</p> <p>7. Recombination rate<br> The result of population-scaled recombination rate was calculated by LDhat v2.2.</p> <p>8. Volcanofinder<br> Genome-wide scans of introgression sweeps within each species was implemented using VolcanFinder v.1.0 with the Model over 10 Kbp non-overlapping windows.</p> <p>9. ihh12<br> phased SNPs were used to computed ihh12 by selscan v1.3.0.&nbsp;</p> <p>10 populus162.phased.recode.vcf.gz<br> SNPs were phased with Beagle v.4.1 for the 162 non-hybrid individuals.</p> <p>11 populus227.snp.rm_indel.para_filter.biallelic.GQ30.max_miss20.bed.recode.vcf.gz&nbsp;<br> The vcf of 227 Populus samples.&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Genomic signatures of isolation, hybridization, and selection during speciation of island finches

<p><strong>Data associated</strong> to the study <em>Genomic signatures of isolation, hybridization, and selection during speciation of island finches</em></p> <p><strong>Contents</strong></p> <ul> <li>Table_S9_samples_accession_nos.xlsx: Editable Excel matrix with sample names and accession numbers.</li> <li>RAD_SNPs_stacks_42424_loci.vcf.tar.gz: VCF file (gzip-compressed tarball) containing SNPs in 42,424 loci, based on analyses of restriction site-associated DNA (RAD) sequencing using Stacks.</li> <li>RAD_SNPs_standard_variant_calling.vcf.tar.gz:&nbsp;VCF file (gzip-compressed tarball) containing 131,661 SNPs from standard variant calling pipelines.</li> <li>mitochondrial_markers_full_data.nex: Nexus file containing mitochondrial (mt) sequences used for mt-phylogeny. Partitioned for COX2, tRNA-Lys, ATP8, and ATP6.&nbsp;</li> <li>sequences_nuclear_genotype_with_zebra_finch_TG.tar.gz:&nbsp;Directory (gzip-compressed tarball)&nbsp;containing genotype sequence (heterozygous sites with IUPAC codes) alignments of nuclear markers in nexus files. In addition to the study species, the sequence for zebra finch <em>Taeniopygia guttata</em> is included with sample code TG.</li> <li>sequences_nuclear_phased_and_mitochondrial_haplotypes_matching.tar.gz:&nbsp;Directory (gzip-compressed tarball)&nbsp;containing phased sequence (haplotype) alignments of nuclear markers in nexus files. These include only those individuals that match&nbsp;individuals sequenced for mitochondrial markers (also included here). In case of recombining loci, both the full locus and the largest non-recombining block are represented.</li> <li>sequences_nuclear_phased_haplotypes_all.tar.gz:&nbsp;Directory (gzip-compressed tarball)&nbsp;containing phased sequence (haplotype) alignments of nuclear markers in nexus files. These include all individuals. In case of recombining loci, both the full locus and the largest non-recombining block are represented.</li> <li>microsatellite_dataset.xlsx: Microsatellite datasets for the study species and additional outgroups. <ul> </ul> <p>Sequences and short read datasets available from NCBI; accession numbers in Table S9 (Table_S9_samples_accession_nos.xlsx).</p> </li> </ul> <p>&nbsp;</p> <p><strong>Study summary</strong></p> <p>Sister species occurring sympatrically on islands are rare and offer unique opportunities to understand how speciation can proceed in the face of gene flow. The S&atilde;o Tom&eacute; grosbeak is a massive-billed, &lsquo;giant&rsquo; finch endemic to the island of S&atilde;o Tom&eacute; in the Gulf of Guinea, where it has diverged from its co-occurring sister species the Pr&iacute;ncipe seedeater, an average-sized finch that also inhabits two neighbouring islands. Here, we show that the grosbeak carries a large number of unique alleles different from all three Pr&iacute;ncipe seedeater populations, but also shares many alleles with the sympatric S&atilde;o Tom&eacute; population of the seedeater, a genomic signature signifying divergence in isolation as well as subsequent introgressive hybridization. Furthermore, genomic segments that remain unique to the grosbeak are situated close to genes, including genes that determine bill morphology, suggesting the preservation of adaptive variation through natural selection during divergence with gene flow. This study reveals a complex speciation process whereby genetic drift, introgression, and selection during periods of isolation and secondary contact all have shaped the diverging genomes of these sympatric island endemic finches.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Genome-wide selection signatures reveal widespread synergistic effects of two different stressors in Drosophila melanogaster: scripts and files

<p>Pipeline and analysis scripts as well as final data files</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

bin3C - GTDB metadata associated with reference genomes selected for the simulated community

<p>Supplementary data table&nbsp;S1 from the manuscript</p> <p>bin3C : Exploiting Hi-C sequencing data to accurately resolve metagenome-assembled genomes (MAGs)</p> <p>A simulated community was constructed for ground truth validation of bin3C results. This table lists the GTDB metadata associated with each of the 63 selected genomes.</p>

opencc-by-4.0Aug 2018View details →
zenodo40/100

Data release: Whole-genome sequencing of Schistosoma mansoni reveals extensive diversity with limited selection despite mass drug administration

<p>Source data used in the publication: Berger et al. (2021) - Provisional title: &#39;Whole-genome sequencing of <em>Schistosoma mansoni</em> reveals extensive diversity with limited selection despite mass drug administration&#39;. These data were used to generate all figures used in the publication and all files are organised and labelled specifically to run with the&nbsp;custom code that uses these data can be found at: http://doi.org/10.5281/zenodo.4975908.&nbsp;</p> <p><br> <strong>File descriptions:</strong></p> <p><strong>SOURCE DATA.zip - All source data for all figures.&nbsp;</strong></p> <p><strong>Figure 1b:</strong></p> <ul> <li>supplementary_data_9.txt - Metadata</li> </ul> <p><strong>Figure 2a&amp;b:</strong></p> <ul> <li>207_PCA.eigenvec -&nbsp;PCA eigenvectors</li> <li>207_PCA.eigenval&nbsp;- PCA eigenvalues</li> </ul> <p><strong>Figure 2c:</strong></p> <ul> <li>autosomes.mdist&nbsp;- PLINK distance&nbsp;matrix used to build the neighbour joining phylogeny</li> </ul> <p><strong>Figure 2d:</strong></p> <ul> <li>all.pi.pixy.schools.txt&nbsp;- Nucleotide diversity results for each school subpopulation.</li> </ul> <p><strong>Figure 2e:</strong></p> <ul> <li>autosomes.dxy.5kb.schools.txt&nbsp;- Autosomal D<sub>XY</sub>&nbsp;results between school subpopulations.&nbsp;</li> <li>autosomes.fst.5kb.schools.txt&nbsp;- Autosomal F<sub>ST</sub>&nbsp;results between school subpopulations.</li> </ul> <p><strong>Figure 2f:</strong></p> <ul> <li>admixture_all.txt&nbsp;- ADMIXTURE results for each sample and population sizes, column 1 represents number of populations (K), columns 3-8 represent admixture values for each population.&nbsp;</li> </ul> <p><strong>Figure 3a, Supplementary figure 10a:</strong></p> <ul> <li>sfs.csv&nbsp;- Site frequency spectra (allelic proportions at each frequency bin) for each school.&nbsp;</li> </ul> <p><strong>Figure 3b:</strong></p> <ul> <li>TD.all.txt&nbsp;- Tajima&#39;s D values calculated in 5 kb windows for each school subpopulation.&nbsp;</li> </ul> <p><strong>Figure 4a, Supplementary figures 13-18:&nbsp;</strong></p> <ul> <li>ALL.MAYUGE.IHS.ihs.out.100bins.norm.txt.zip&nbsp; - Normalised iHS scores for the Mayuge district parasite populations (Selscan output).</li> </ul> <p><strong>Figure 4b, Supplementary figures 13-18:&nbsp;</strong></p> <ul> <li>ALL.TORORO.IHS.ihs.out.100bins.norm.txt.zip&nbsp;-<strong> -&nbsp;</strong>Normalised iHS scores for the Tororo district parasite populations (Selscan output).</li> </ul> <p><strong>Figure 4c, Supplementary figures 13-18:&nbsp;</strong></p> <ul> <li>ALL.MAYUGEvsTORORO.xpehh.xpehh.out.norm.txt.zip&nbsp;- - Normalised XP-EHH scores between Mayuge and Tororo parasite populations.</li> </ul> <p><strong>Figure 4d, Supplementary figures 13-18:</strong></p> <ul> <li>MAYUGE_TORORO_2000.windowed.weir.txt.zip - F<sub>ST</sub> values calculated between Mayuge and Tororo populations in 2kb windows.&nbsp;&nbsp;</li> </ul> <p><strong>Figure 4e, Supplementary figures 12a&amp;c:</strong></p> <ul> <li>MAYUGE_PI.windowed.pi.zip&nbsp;- Nucleotide diversity values calculated in 2 kb windows for Mayuge populations.&nbsp;</li> <li>TORORO_PI.windowed.pi.zip&nbsp;- Nucleotide diversity values calculated in 2 kb windows for Kocoge populations (Tororo district).</li> </ul> <p><strong>Figure 5a:</strong></p> <ul> <li>all.pi.treat.fix.txt.zip&nbsp;- Nucleotide diversity results for each treatment subpopulation</li> </ul> <p><strong>Figure 5b</strong></p> <ul> <li>autosomes.dxy.5kb.treatment.txt&nbsp;- <strong>&nbsp;</strong>- Autosomal D<sub>XY</sub>&nbsp;results between clearance phenotype subpopulations.&nbsp;</li> <li>autosomes.fst.5kb.treatment.txt<strong>&nbsp;</strong>- Autosomal F<sub>ST</sub>&nbsp;results between clearance phenotype subpopulations.&nbsp;</li> </ul> <p><strong>Figure 5c:</strong></p> <ul> <li>fst.windows.2kb.treatment.txt.zip&nbsp;- F<sub>ST</sub> values for comparisons between different treatment groups (Pre-treatment, post-treatment (good clearers), post-treatment (poor clearers))</li> </ul> <p><strong>Figure 5d:&nbsp;</strong></p> <ul> <li>assoc_err_binary.txt.zip&nbsp;-&nbsp;Results of&nbsp;binary trait association between miracidia sampled from hosts with good clearance phenotypes (where treatment appeared to be highly effective) and miracidia isolated post-treatment from hosts with poor clearance phenotypes (where miracidia are potentially derived from parasites that survived treatment.</li> </ul> <p><strong>Figure 5e:</strong></p> <ul> <li>assoc_err_linear.txt.zip&nbsp;- - Results of&nbsp;linear regression genome-wide association study&nbsp;with the ERR estimates for all 198 samples, using the mean of the posterior ERR estimates from Crellen et al. (2016) as a quantitative trait.</li> </ul> <p><strong>Supplementary figure 1:</strong></p> <ul> <li>median.coverage.txt&nbsp;- Normalised depth of read coverage (column 4) calculated in 25 kb windows (columns 2&amp;3) across all samples for all chromosomes (column 1).</li> </ul> <p><strong>Supplementary figure 2a-f:&nbsp;</strong></p> <ul> <li>cohort.genotyped.txt.zip&nbsp;- <strong>&nbsp;</strong>- Variant quality site values (used to inform variant site retention or removal).&nbsp;</li> </ul> <p><strong>Supplementary figure 2g:</strong></p> <ul> <li>hard_filtered.imiss.txt&nbsp;- &nbsp;Per sample variant missingness (used to inform quality control).</li> </ul> <p><strong>Supplementary figure 2h:</strong></p> <ul> <li>hard_filtered_filtindv.lmiss.txt.zip&nbsp;- Per site missingness (used to inform quality control).</li> </ul> <p><strong>Supplementary figure 3a, 4a, 4b:</strong></p> <ul> <li>prunedData.eigenvec&nbsp;- PCA eigenvectors</li> <li>prunedData.eigenval&nbsp;- PCA eigenvalues</li> </ul> <p><strong>Supplementary figure 3b:</strong></p> <ul> <li>pruned_data.mdist.csv -&nbsp;Distance matrix used as the basis for the neighbour joining phylogeny.</li> </ul> <p><strong>Supplementary figure 5:</strong></p> <ul> <li>cv_scores.txt&nbsp;- ADMIXTURE coefficient of variation&nbsp;scores (column 2) for each population size (1).</li> </ul> <p><strong>Supplementary figure 6:</strong></p> <ul> <li>*_SMC_SE.csv&nbsp;- SMC++ results (from 25 subsampled replicates) for each school subpopulation and outgroup samples.&nbsp;</li> </ul> <p><strong>Supplementary Figure 7:</strong></p> <ul> <li>smcpp.csv&nbsp;-&nbsp;SMC++ results&nbsp;for each school subpopulation and outgroup samples.&nbsp;</li> </ul> <p><strong>Supplementary Figure 8a-d</strong></p> <ul> <li>pi.per_host.txt.zip&nbsp;- Nucleotide diversity values for each host infrapopulation.&nbsp;</li> </ul> <p><strong>Supplementary Figure 9:</strong></p> <ul> <li>sexing.csv&nbsp;- inferred sex (based on differential read coverage over pseudoautosomal and Z-specific regions of the Z chromosome).&nbsp;</li> </ul> <p><strong>Supplementary Figure 10b:</strong></p> <ul> <li>sfs_res.csv - residuals for the SFS analysis in 3a/10a.</li> </ul> <p><strong>Supplementary Figure 11:</strong></p> <ul> <li>MAYUGE_TAJIMA_D.Tajima.D.2kb.txt.zip&nbsp;- Tajima&#39;s D values calculated for the Mayuge population&nbsp;in 2kb windows.&nbsp;</li> <li>Tororo_TAJIMA_D.Tajima.D.2kb.txt.zip - Tajima&#39;s D values calculated for the Tororo population&nbsp;in 2kb windows.&nbsp;</li> </ul> <p><strong>Supplementary Figures 13-18:</strong></p> <ul> <li>genes.bed&nbsp;- Coordinates of gene models (<em>S. mansoni </em>v7 annotation).</li> <li>KOCOGE_SITE_PI.sites.pi.txt.zip - Per site nucleotide diversity values</li> <li>MAYUGE_TORORO_sites.weir.fst.txt.zip&nbsp;- Per site F<sub>ST</sub> values between Mayuge and Tororo populations.&nbsp;</li> <li>coverage_5kb.windows.txt.zip&nbsp;- Per sample depth of read coverage in 5 kb windows. Columns 4,5,6 represent the median, mean and sstev of coverage for each 5kb window (columns 2&amp;3) along each chromosome (column 1).&nbsp;</li> <li>median.sample.coverage.txt&nbsp;-&nbsp; Median chromosomal depth of read coverage for each sample.&nbsp;</li> </ul> <p><strong>Supplementary Figure 19:</strong></p> <ul> <li>kocoge_median.ld.txt.zip&nbsp;-&nbsp;<strong>&nbsp;</strong>- The decay of linkage disequilibrium with genomic distance between all sites within 50 kb for the Kocoge parasite samples. Chromosomes are shown in column 1, distance in column 2, median values in column 3.&nbsp;</li> <li>mayuge_median.ld.txt.zip&nbsp;-&nbsp;The decay of linkage disequilibrium with genomic distance between all sites within 50 kb for the Mayuge parasite samples. Chromosomes are shown in column 1, distance in column 2, median values in column 3.&nbsp;</li> </ul> <p><strong>Misc files:</strong></p> <p>schools.list - List of samples and schools where they were sampled.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2021View details →
dryad40/100

Data from: Selection on growth rate and local adaptation drive genomic adaptation during experimental range expansions in the protist Tetrahymena thermophila

<p>1. Populations that expand their range can undergo rapid evolutionary adaptation of life-history traits, dispersal behaviour, and adaptation to the local environment. Such adaptation may be aided or hindered by sexual reproduction, depending on the context.</p> <p>2. However, few empirical and experimental studies have investigated the genetic basis of adaptive evolution during range expansions. Even less attention has been given to the question how sexual reproduction may modulate such adaptive evolution during range expansions.</p> <p>3. We here studied genomic adaptation during experimental range expansions of the protist <em>Tetrahymena thermophila</em>in landscapes with a uniform environment or a pH-gradient. Specifically, we investigated two aspects of genomic adaptation during range expansion. Firstly, we investigated adaptive genetic change in terms of the underlying numbers of allele frequency changes from standing genetic variation and <em>de novo</em><span> variants. We focused on how sexual reproduction may alter this adaptive genetic change. Secondly, we identified genes subject to selection caused by the expanding range itself, and directional selection due to the presence or absence of the pH-gradient. We focused this analysis on alleles with large frequency changes that occurred in parallel in more than one population to identify the most likely candidate targets of selection. </span></p> <p><span>4. We found that sexual reproduction altered adaptive genetic change both in terms of <em>de novo</em></span><span> variants and standing genetic variation. However, sexual reproduction affected allele frequency changes in standing genetic variation only in the absence of long-distance gene flow. Adaptation to the range expansion affected genes involved in cell divisions and DNA repair, whereas adaptation to the pH-gradient additionally affected genes involved in ion balance, and oxidoreductase reactions. These genetic changes may result from selection on growth and adaptation to low pH. </span></p> <p><span>5. In the absence of gene flow, sexual reproduction may have aided genetic adaptation. Gene flow may have swamped expanding populations with maladapted alleles, thus reducing the extent of evolutionary adaptation during range expansion. Sexual reproduction also altered the genetic basis of adaptation in our evolving populations via <em>de novo </em>variants, possibly by purging deleterious mutations or by revealing fitness benefits of rare genetic variants. </span></p>

opencc-zeroOct 2021View details →
zenodo40/100

Balancing selection at a wing pattern locus is associated with major shifts in genome-wide patterns of diversity and gene flow

<p>Selection shapes genetic diversity around target mutations, yet little is known about how selection on specific loci affects the genetic trajectories of populations, including their genome-wide patterns of diversity and demographic responses. Here we study the patterns of genetic variation and geographic structure in a neotropical butterfly, <em>Heliconius numata</em>, and its closely related allies in the so-called melpomene-silvaniform clade. <em>H. numata</em> is known to have evolved an inversion supergene which controls variation in wing patterns involved in mimicry associations with distinct groups of co-mimics. Butterflies show disassortative mate preferences and heterozygote advantage at this locus. We contrasted patterns of genetic diversity and structure 1) among extant polymorphic and monomorphic populations of <em>H. numata</em>, 2) between <em>H. numata</em> and its close relatives, and 3) between ancestral lineages. We show that <em>H. numata</em> populations which carry the inversions as a balanced polymorphism show markedly distinct patterns of diversity compared to all other taxa. They show the highest genetic diversity and effective population size estimates in the entire clade, as well as a low level of geographic structure and isolation by distance across the entire Amazon basin. By contrast, monomorphic populations of <em>H. numata</em> as well as its sister species and their ancestral lineages all show lower effective population sizes and genetic diversity, and higher levels of geographical structure across the continent. One hypothesis is that the large effective population size of polymorphic populations could be caused by the shift to a regime of balancing selection due to the genetic load and disassortative preferences associated with inversions. Testing this hypothesis with forward simulations supported the observation of increased diversity in populations with the supergene. Our results are consistent with the hypothesis that the formation of a supergene triggered a change in gene flow, causing a general increase in genetic diversity and the homogenisation of genomes at the continental scale.</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Genomic data and common garden experiments reveal climate-driven selection on ecophysiological traits in two Mediterranean oaks

<p>This release includes the different genomic datasets used in the article entitled &quot;<em>Genomic data and common garden experiments reveal climate-driven selection on ecophysiological traits in two Mediterranean oaks</em> &quot; by Ram&iacute;rez-Valiente et al.,</p> <p>File description:</p> <p><strong>Samples.xlsx</strong>: Description of individual and population codes used in the different analyses and genomic datasets.</p> <p><strong>Quercus_faginea_p12r05m05minMAF001_all_loci.str</strong>: Input file used to perform genetic clustering analyses (STRUCTURE and DAPC) for <em>Quercus faginea</em> including all loci.</p> <p><strong>Quercus_faginea_p12r05m05minMAF001_neutral_loci.str</strong>: Input file used to perform genetic clustering analyses (STRUCTURE and DAPC) for <em>Quercus faginea</em> excluding outlier loci (i.e., putatively under selection) identified by either BAYESCAN or using the FDIST method in ARLEQUIN.</p> <p><strong>Quercus_lusitanica_p7r05m05minMAF001_all_loci.str</strong>: Input file used to perform genetic clustering analyses (STRUCTURE and DAPC) for <em>Quercus lusitanica </em>including all loci.</p> <p><strong>Quercus_ lusitanica_p7r05m05minMAF001_neutral_loci.str</strong>: Input file used to perform genetic clustering analyses (STRUCTURE and DAPC) for <em>Quercus lusitanica </em>excluding outlier loci (i.e., putatively under selection) identified by either BAYESCAN or using the FDIST method in ARLEQUIN.</p> <p><strong>Quercus_faginea_p12r05m05minMAF001_BAYESCAN.txt</strong>: Input file used to perform BAYESCAN analyses for <em>Quercus faginea</em>.</p> <p><strong>Quercus_lusitanica_p7r05m05minMAF001_BAYESCAN.txt</strong>: Input file used to perform BAYESCAN analyses for <em>Quercus lusitanica</em>.</p> <p><strong>Quercus_faginea_p12r05m05minMAF001_ARLEQUIN.arp</strong>: Input file used to perform ARLEQUIN analyses for <em>Quercus faginea</em>.</p> <p><strong>Quercus_lusitanica_p7r05m05minMAF001_ARLEQUIN.arp</strong>: Input file used to perform ARLEQUIN analyses for <em>Quercus lusitanica</em>.</p> <p><strong>Quercus_faginea_p12r05m05minMAF001_all_loci.vcf</strong>: Variant call format (VCF) file for <em>Quercus faginea</em> including all loci.</p> <p><strong>Quercus_faginea_p12r05m05minMAF001_neutral_loci.vcf</strong>: Variant call format (VCF) file for <em>Quercus faginea</em> excluding outlier loci (i.e., putatively under selection) identified by either BAYESCAN or using the FDIST method in ARLEQUIN.</p> <p><strong>Quercus_lusitanica_p7r05m05minMAF001_all_loci.vcf</strong>: Variant call format (VCF) file for <em>Quercus lusitanica </em>including all loci.</p> <p><strong>Quercus_ lusitanica_p7r05m05minMAF001_neutral_loci.vcf</strong>: Variant call format (VCF) file for <em>Quercus lusitanica </em>excluding outlier loci (i.e., putatively under selection) identified by either BAYESCAN or using the FDIST method in ARLEQUIN.</p> <p><strong>Quercus_faginea_Greenhouse_DRIFTSEL.txt</strong>: Input file used to run DRIFTSEL and evaluate selection on the different studied traits for <em>Quercus faginea </em>under common garden greenhouse experiments.</p> <p><strong>Quercus_faginea_Outdoor_DRIFTSEL.txt</strong>: Input file used to run DRIFTSEL and evaluate selection on the different studied traits for <em>Quercus faginea</em> under common garden outdoor experiments.</p> <p><strong>Quercus_lusitanica_Greenhouse_DRIFTSEL.txt</strong>: Input file used to run DRIFTSEL and evaluate selection on the different studied traits for <em>Quercus lusitanica </em>under common garden greenhouse experiments.</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Contrasting genome-wide signatures of selection in two closely related Epichloe plant pathogen species

<p>Deposited here composite plots for each species, each pairwise population combination and each of the seven chromosomes as shown and referred to in the manuscript.</p> <p>The filename contains [species abbrevation]_[chromosome number]_[population 1]_[population 2]. Chromosome-wide SNP data and sweeps identified for the population pair are shown. The top panel shows pairwise FST values, averaged across 5kb windows. Shaded rectangles represent the locations of AT-rich regions. The second panel shows the absolute values of the integrated haplotype score (iHS) calculated at each SNP locus for which the ancestral allele state was known. Scores for pop1 are shown at the top and scores for pop2 are negatively transformed and showed at the bottom. Horizontal dashed lines indicate the 99.9% percentile threshold which was used as a cutoff to identify outlier SNPs and inferred iHS sweeps are shown as shaded rectangles. The third panel shows the cross-population extended haplotype homozygosity (XP-EHH) scores calculated between the two populations. Dashed lines indicate 99.9% percentile threshold which was used as a cutoff to identify outlier SNPs and inferred divergent sweeps are shown as shaded rectangles. Positive and negative XP-EHH values refer to the direction of selection: positive values indicate selection in pop1 negative values indicate selection in pop2. In the bottom panel, composite likelihood ratio (CLR) scores are plotted for pop1 (black) and pop2 (blue), colored dashed lines indicate respective 99.9% threshold and colored rectangles highlight inferred CLR-sweeps.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

COVFlow: performing virus phylodynamics analyses from selected SARS-CoV-2 genome sequences

<p>This upload contains pipeline configuration files, output data, scripts and data identifiers (GISAID EPI_ISL_ID) required to reproduce the results of the article entitled &quot;COVFlow: performing virus phylodynamics analyses from selected SARS-CoV-2 genome sequences&quot;.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Integrative QTL mapping and selection signatures in Groningen White Headed cattle inferred from whole-genome sequences

<p>Here, we aimed to identify and characterize genomic regions that differ between Groningen White Headed (GWH) breed and other cattle, and in particular to identify candidate genes associated with coat color and/or eye-protective phenotypes. Firstly, whole genome sequences of 170 animals from eight breeds were used to evaluate the genetic structure of the GWH in relation to other cattle breeds by carrying out principal components and model-based clustering analyses. Secondly, the candidate genomic regions were identified by integrating the findings from: a) a genome-wide association study using GWH, other white headed breeds (Hereford and Simmental), and breeds with a non-white headed phenotype (Dutch Friesian, Deep Red, Meuse-Rhine-Yssel, Dutch Belted, and Holstein Friesian); b) scans for specific signatures of selection in GWH cattle by comparison with four other Dutch traditional breeds (Dutch Friesian, Deep Red, Meuse-Rhine-Yssel and Dutch Belted) and the commercial Holstein Friesian; and c) detection of candidate genes identified via these approaches. The alignment of the filtered reads to the reference genome (ARS-UCD1.2) resulted in a mean depth of coverage of 8.7X. After variant calling, the lowest number of breed-specific variants was detected in Holstein Friesian (148,213), and the largest in Deep Red (558,909). By integrating the results, we identified five genomic regions under selection on BTA4 (70.2&ndash;71.3 Mb), BTA5 (10.0&ndash;19.7 Mb), BTA20 (10.0&ndash;19.9 and 20.0&ndash;22.7 Mb), and BTA25 (0.5&ndash;9.2 Mb). These regions contain positional and functional candidate genes associated with retinal degeneration (e.g.,&nbsp;<em>CWC27</em>&nbsp;and&nbsp;<em>CLUAP1</em>), ultraviole<em>t</em>&nbsp;protection (e.g.,&nbsp;<em>ERCC8</em>), and pigmentation (e.g.&nbsp;<em>PDE4D</em>) which are probably associated with the GWH specific pigmentation and/or eye-protective phenotypes, e.g. Ambilateral Circumocular Pigmentation (ACOP). Our results will assist in characterizing the molecular basis of GWH phenotypes and the biological implications of its adaptation.</p>

opencc-by-4.0Oct 2022View details →
dryad40/100

Data From: Polygenic basis and the role of genome duplication in adaptation to similar selective environments

Open the record for dataset details and reuse information.

publicSep 2021View details →
dryad40/100

Data from: Whole-genome resequencing reveals polygenic signatures of directional and balancing selection on alternative migratory life histories

Open the record for dataset details and reuse information.

publicNov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record