Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,331
datasets available to search
ShareScore release 0.7.1
Dataset results
2,331 results for “polymorphism”
Dataset: Systematics of the color-polymorphic spider genus Cybaeolus, with comments on the phylogeny of the family Hahniidae (Araneae)
<p>Phylogenetic analysis of the spiders of the genus Cybaeulus, with outgroups in the marronoid clade. Data from six DNA markers, analyzed with maximum likelihood and parsimony.</p> <p><br>PHYLOGENETIC ANALYSIS</p> <p>We obtained sequences from 26 samples of the three known species of Cybaeolus, and of five additional species of Hahniidae. To these, we added legacy sequences of Cybaeolus and of other genera of Hahniidae, as well as representatives of the remaining families in the marronoid clade. For the new sequences, the extraction and amplification of DNA was made in the Laboratory of Molecular Tools at Museo Argentino de Ciencias Naturales (MACN), from tissues preserved in absolute alcohol at -18ºC. We targeted the markers histone H3 (H3), cytochrome oxidase subunit I (CO1), 28S ribosomal RNA (28S) and 16S ribosomal RNA (16S), previously used to estimate relationships of marronoid spiders (Wheeler et al., 2017). Details of extraction, primers and PCR protocols are the same as in Magalhaes & Ramírez (2022). Sequencing was outsourced to Macrogen Inc., South Korea. The resulting chromatograms were analyzed individually to detect contaminated sequences or ambiguous portions. In addition to these sequences obtained in the laboratory, we combined our data with additional sequences from previous work (Wheeler et al., 2017; Rivera-Quiroz et al., 2020), using the markers mentioned above plus 12S ribosomal RNA (12S) and 18S ribosomal RNA (18S). For the CO1 marker, additional sequences obtained by the Arachnology Division at MACN and deposited in the BOLDSYSTEMS platform (https://www.boldsystems.org/) were also used. Sequences were aligned with MAFFT Online v.7.463 (Katoh & Standley, 2013), using the L-INS-I algorithm. See Table 1 for list of vouchers and sequence identifiers.</p> <p>Maximum likelihood<br>For the maximum likelihood analyses we used the program IQ-TREE 2.2.0 (Minh et al., 2020), partitioning the data by marker, and selecting the best combination of partitions and evolution models by Bayesian information criterion (best fitting models were TPM2+I+G4 for H3, GTR+F+I+G4 for 18S, GTR+F+I+G4 for 16S and 12S together, GTR+F+I+G4 for CO1, and GTR+F+I+G4 for 28S). Since the relationships of outgroup taxa in the resulting trees were slightly different to that found in recent phylogenomic studies, we used the study of Gorneau et al. (2023) based on ultraconserved elements as a backbone topology to constrain our tree search, considering only the taxa in common with our analysis (see supplementary Fig. S1); this means that all the rest of the taxa are free to move anywhere during tree search. Support for groups (branches) was estimated by 1000 cycles of ultrafast bootstrapping. Ten independent runs were performed; of those, six converged into nearly identical log likelihood values (-57417.7725 to -57417.9604) and identical topologies; the tree with top-ranking log likelihood is presented in Results, after collapsing branches with bootstrap below 0.5. To estimate the support of an alternative topology with Cybaeolus as sister to the rest of the hahniids, we used TNT 1.6 (Goloboff & Morales, 2023) to modify the optimal tree placing Cybaeolus in such position, and asked for the frequency of the branch of interest (all hahniids except Cybaeolus) in the 1000 bootstrapped trees previously saved by IQTREE.<br>Ancestral character states for the arrangement of spinnerets (grouped; separated in a transversal line) were estimated by maximum likelihood on the optimal tree, using the R packages phytools and ape, under the models ER and ARD, and the best fitting model selected by the Akaike information criterion. </p> <p>Parsimony<br>For the parsimony analyses we used TNT 1.6. For the equal weights analysis, a heuristic search was made using a driven search with the default parameters of the “new technologies”, aiming for 10 independent hits to minimum length. The resulting trees were then submitted to an additional round of tree-bisection reconnection (TBR) branch swapping. These results were compared to a simpler search strategy of 300 random addition sequences, each followed by TBR, which produced 20 hits to minimal length. As both strategies reached the same trees with multiple independent hits, it is likely that the optimal trees were found. Finally, the strict consensus of all the optimal trees was obtained, and on this consensus the support values were calculated by means of 1000 bootstrap pseudoreplicates. </p>
LA1141 × OH8245 inbred backcross (IBC) single nucleotide polymorphism (SNP) markers for genetic studies
<p>The LA1141 × OH8245 157 polymorphic SNP markers from an optimized tomato panel Sim et al., 2012 were used for linkage map construction in the BC<sub>2</sub>S<sub>3</sub> IBC and composite interval mapping QTL analysis. Genetic map position and physical position corresponding to Sl4.0 (Hosmani et al., 2019), and flanking sequences are provided.</p>
Deciphering polymorphism in 61,157 Escherichia coli genomes via epistatic sequence landscapes
<p>We use computational models based on Direct Coupling Analysis - DCA - trained on PFAM domains of distant distant homologues to accurately predict the polymorphisms segregating in a panel of 61,157 <em>Escherichia coli </em>genomes.</p> <p>We show that the genetic context (<em>i.e. </em>the rest of the protein sequence) strongly constrains the tolerable amino acids in 30% to 50% of amino-acid sites. Our study also suggests the gradual build-up of genetic context over long evolutionary timescales by the accumulation of small epistatic contributions.</p> <p>Please refer to the README file for additional information on the structure of this dataset.</p> <p>Code to analyse this dataset is available at https://github.com/GiancarloCroce/DCA_polymorphism_Ecoli.</p> <p> </p>
Data from: Development of Single Nucleotide Polymorphism (SNP) Panel for determination of environmental influence on genome for wild Columbia River redband trout (Oncorhynchus mykiss gairdnerii) in Southwest Idaho streams
<p>DNA were derived from fin tissue samples taken from individual trout captured from Little Jacks Creek, Big Jacks Creek , and Duncan Creek of the Owyhee mountains and Keithly Creek and Upper Mann Creek in the Hitt mountains of Western Idaho, United States. Fin tissues were collected from individual trout from each stream during monthly sampling events in June through October 2020. </p> <p><em>DNA Extraction:</em> Extraction of DNA from caudal fin tissues were performed using Quick-DNA Miniprep Plus purification kits (Zymo Research Inc.©). Small sections of fin tissue (≤ 25 mg) were collected from each sample. This was mixed with a digesting solution comprised of ultra-pure water, solid tissue buffer (Zymo Research Inc.©) and proteinase K. All tissues were digested in sealed microcentrifuge tubes for at minimum 3 h at 55°C in a water bath. We then aliquoted 100 µL of digestion supernatant and combined with 200 µL of genomic binding buffer (Zymo Research Inc.©). DNA was eluted in 50, 75, and 100 µL of elution buffer to determine which volume provided sufficient DNA concentration for genotyping. After it was determined all quantities produced suitable concentrations, going forward, 50 µL of elution buffer used.</p> <p><em>Genotyping:</em> Following extraction, genotyping-in-thousands sequencing took place at the Hagerman National Fish Hatchery’s genetics research facility with the assistance of the Columbia River Intertribal Fish Commission (CRTFC). Genotyping protocols were as described in Campbell et al. (2015) and summarized below. First, samples were prepared for amplification via PCR by combining DNA extracts with a Qiagen Plus multiplex master mix and a species-specific pooled primer mix. This step added the Illumina sequencing primer sites to amplicons. Following the creation of the PCR cocktail, thermocycling was conducted for amplification. Amplified samples were then diluted 20-fold. Diluted samples were transferred to new 96-well PCR plates where two genetic indexes and barcodes provides a unique set of tagging primers to each well and plate. Tagged plates then underwent a second PCR step. After the second PCR, all DNA were transferred to Charm Biotech normalization plates where DNA was bound to wells, washed, and finally eluted. After normalization, all DNA was pooled together and a purification step using magnetized beads in two steps to selectively remove fragments of DNA that are both too large and too small for sequencing. Following purification, each plate was quantified via qPCR using Life Technologies QuantStudio 6 Flex Instrument (Life Technologies). Finally, sequencing was performed using an Illumina HiSeq 1500 instrument.</p> <p><strong>Ancillary peer-reviewed manuscripts:</strong><br> <em>Genotyping protocols</em><br> Campbell NR, Harmon SA, Narum SR. 2015. Genotyping-in-Thousands by sequencing (GT-seq): A cost effective SNP genotyping method based on custom amplicon sequencing. Mol Ecol Resour, 15: 855-867. https://doi.org/10.1111/1755-0998.12357<br> <em>SNP loci reference</em><br> Collins EE, Hargrove JS, Delomas TA, Narum SR. 2020. Distribution of genetic variation underlying adult migration timing in steelhead of the Columbia River basin. Ecology and Evolution, 10(17): 9486-9502. https://doi.org/10.1002/ece3.6641 </p> <p><strong>Data Use</strong>:<br> <em>License</em>: <a href="https://creativecommons.org/licenses/by/4.0/">CC-BY 4.0</a> <br> <em>Recommended Citation</em>: Wooding AP, Narum SR, Pradhan DS. 2022. Data from: Development of Single Nucleotide Polymorphism (SNP) Panel for determination of environmental influence on genome for wild Columbia River redband trout (Oncorhynchus mykiss gairdnerii) in Southwest Idaho streams (0.1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7055582</p> <p>Funding for this project is provided by US National Science Foundation and Idaho EPSCoR through award: OIA-1757324 </p>
Data from: Polymorphic tandem repeats shape single-cell gene expression across the immune landscape
<p>This dataset contains the association summary statistics (v0.1) for genome-wide tandem repeat (TR) expression quantitative trait (eQTL) analysis of TenK10K Phase 1 (https://doi.org/10.1101/2024.11.02.621562). </p> <p>Please access the README for a detailed description of file contents. </p> <p> </p>
gnomAD polymorphism and de novo mutation data for analysis of mutation rates in highly mutable gene classes
<p>We analyze the human mutation rate in three gene classes (IGK, RNU, and tRNA) which deviate from the expectations of a mutation rate model. We examine the distribution of allele frequencies for SNVs within these genes and we analyze the counts of de novo mutations stratified by whether the SNV was observed or not. </p> <p>{CHR}_IGK_SFS_v2_denovo.gz: allele frequencies, mutation rate estimates, and whether the de novo mutation was observed for IGK, RNU, and tRNA genes. Based on gnomAD v3. </p> <p>{CHR}_indiv_mu.csv: quality information for variants in these gene classes from the 1kg subset of gnomAD.</p> <p>"CHR", "POS", "REF", "ALT", "FILTER", "AC", "AN", "MQRankSum", "pab_max", "VQSLOD", "AB", "PN", "MR", "AR", "MG", "MC", "QUAL"</p> <p>all_variants_chr21_mu_h.csv.gz: all variants from chromosome 21 to use for comparing allele frequencies to those in our gene classes.</p> <p>21_indiv_mu_all.csv.gz: quality information from all variants on chromosome 21 from the 1kg subset of gnomAD to use for comparison with gene classes.</p> <p>"CHR", "POS", "REF", "ALT", "FILTER", "AC", "AN", "MQRankSum", "pab_max", "VQSLOD", "AB", "PN", "MR", "AR", "MG", "MC", "QUAL"</p> <p> </p>
Polymorphism and Perfection in Crystallization of Hard Sphere Polymers
<p>Data archive corresponding to the publications "Polymorphism and Perfection in Crystallization of Hard Sphere Polymers" by M. Herranz et al., Polymers 14, 4435 (2022); DOI: https://doi.org/10.3390/polym14204435</p> <p>Please see README.txt for instructions on how to access and read the files from the crystallographic analysis based on the CCE norm descriptor.</p> <p>All snapshots have been generated and successively analyzed by the Simu-D software.</p>
Static and Dynamic DFT Data Sets for Polymorphs of l-Cysteine - Stability and Terahertz Spectra
<p>Data sets associated with static and dynamic density functional calculations are reported for the four known polymorphs of l-cysteine and for models associated with the known disorder in Form I.. </p> <p>Static calculations are used to explore the relative free energies (within the harmonic approximation) of the polymorphs as a function of pressure. The energetics for dihedral angle rotation are explored and the barriers for rotation between the hydrogen bonding motifs have been calculated for each polymorph.</p> <p>Molecular dynamics calculations are reported for each polymorph and for models of hydrogen bond disorder which are known to exist at higher temperatures.</p> <p>Finally static and dynamic calculations of the infrared and terahertz spectra are performed.</p>
Simulation code and simulated data for: Transient polymorphisms in parental care strategies drive divergence of sex roles
<p>This repository contains C++ code, simulated datasets, an R-script for data analysis and a Mathematica notebook for mathetical analysis.</p><p>Datasets are organised into ZIP files named after the corresponding figure in the publication. All of the figures based on simulation data in the manuscript and supplementary materials can be created with the R-script. For further information see the article published in <i>Nature Communications (</i>doi:<i> </i>https://doi.org/10.1038/s41467-023-42607-6).</p><p> </p><p> </p><p> </p>
Data from: The genomics and evolution of inter-sexual mimicry and female-limited polymorphisms in damselflies
<p>The dataset contains intermediate output files required to reproduce the figures in the main text and Supporting Material of Willink <em>et al</em>. 2023. The genomics and evolution of inter-sexual mimicry and female-limited polymorphisms in damselflies.</p> <p>FILE OVERVIEW:</p> <p>1. Morph-specific assemblies<br> A. File names: Afem_1354_ragtag.fasta.gz, Ifem_1049_ragtag.fa.gz, Ofem_0081_ragtag.fa.gz, O054_Shasta_run2.PMDV.HAP1.purged.fasta.gz, A059_Shasta_run1.PMDV.HAP1.purged.fa.gz<br> B. Description: genome assemblies for different morphs of <em>Ischnura elegans</em> (Afem_1354, Ifem_1049, and Ofem_0081) and <em>Ischnura senegalensis</em> (A059 and O054), generated in this study from long-read Nanopore data using Shasta v 0.7.0 (https://github.com/paoloshasta/shasta).</p> <p>2. Assembly statistics<br> A. File names: Assembly_statistics.csv, Assembly_statistics_sen.csv<br> B. Description: Completeness and quality metrics for <em>de novo</em> genome assemblies of <em>I. elegans</em> and <em>I. senegalensis</em> female morphs. See Fig. S1-S2.</p> <p>3. Repetitive content annotation<br> A. File names: A1354_ragtag_RED.bed.repeats.bed.gz, Afem_Shasta1_polished_ragtag_UPPER.fa.out.gz, Ifem_Shasta2_polished_ragtag_UPPER.fa.out.gz, ioIscEleg1.1.primary_UPPER.fa.out.gz, ToL_RED.repeats.bed.gz<br> B. Description: Annotation of repetitive sequences in morph-specific assemblies. All morph assemblies (A, I and Darwin Tree of Life assemblies) were annotated using RepeatModeler v 2.0.1 and RepeatMasker v 1.0.93 (http://www.repeatmasker.org). The A morph and DToL assemblies were additionally annotated using Red v 0.0.1 (https://github.com/BioinformaticsToolsmith/Red). RepeatMasker annotations were then used to estimate TE coverage. See Extended Data Fig. 4 and Fig. S7.</p> <p>4. GWAS output<br> A. File names: A1354_ragtag_AvI.assoc_filtered.txt.gz, A1354_ragtag_AvO.assoc_filtered.txt.gz, A1354_ragtag_IvO.assoc_filtered.txt.gz, ToL_AvI.assoc_filtered.txt.gz, ToL_AvO.assoc_filtered.txt.gz, ToL_IvO.assoc_filtered.txt.gz<br> B. Description: filtered SNPs in pairwise association tests between morphs (n = 19 resequencing samples per morph) of<em> I. elegans</em>. Analyses were conducted in PLINK v 1.9 (http://pngu.mgh.harvard.edu/purcell/plink/), using either the A morph assembly (Fig. 2a-b), or the Darwin Tree of Life (DToL) reference assembly (Extended Data Figure 8a-b) as mapping reference.</p> <p>5. Population statistics<br> A. File names: Afem_pixy_30K_fst.txt.gz, A1354_30kb.Tajima.D.gz, Afem_pi_30K_pi.txt.gz, ToL_30K_fst.txt.gz, ToL_30kb.Tajima.D.gz, ToL_30K_onepop_pi.txt.gz<br> B. Description: Genetic differentiation (fst) between morphs, Tajima's D statistics, and nucleotide diversity across 30 kb windows of the<em> I. elegans</em> genome. Population statistics were computed using either the A morph assembly (Fig. 2c-e), or the DToL reference assembly (Extended Data Figure 8c-e) as mapping reference.</p> <p>6. k-mer based GWAS<br> A. File names: AvI_kmers.fa.gz, AvO_kmers.fa.gz, OvAI_kmers.fa.gz, AvI_kmers.fa_v_A1354_Shasta_run1_table.tsv.gz, AvO_kmers.fa_v_A1354_Shasta_run1_table.tsv.gz, OvAI_kmers.fa_v_A1354_Shasta_run1_table.tsv.gz, OvAI_kmers.fa_v_Ifem_1049_ragtag_table.tsv.gz<br> B. Description: List of significant k-mers (in fasta format) in three k-mer based association analyses (n = 19 resequencing samples per morph) between morphs of<em> I. elegans</em>. Significant k-mers were then mapped to morph-specific assemblies using Blast v 2.22.28 (https://blast.ncbi.nlm.nih.gov/Blast.cgi) for short sequences. We include mapping results shown in Fig. 3a-b.</p> <p>7. Read-depth coverage<br> A. File names: reseq_coverage_norepeat_500_window.bed.gz, nano_coverage_norepeat_500_window.bed.gz, Ifem_nano_coverage_norepeat_500_window.bed.gz, Ifem_reseq_coverage_norepeat_500_window_15Mb.bed.gz, poolseq_coverage_norepeat_500_window.bed.gz, morph_coverage_norepeat_diff_500.tsv.gz, SwD_popmap<br> B. Description: Read depth coverage of the morph locus and a 15 mb region used to estimate baseline read depths. 19 Illumina resequencing samples, and one long-read Nanopore sample of each morph of <em>I. elegans</em> were mapped to both the A and I assemblies to estimate read depth. Two poolseq samples (each pool consisting of 30 females of each morph) of<em> I. senegalensis</em> were mapped to the A assembly of<em> I. elegans</em> to estimate read depth. Read depth was estimated in mosdepth v 0.2.8 (https://github.com/brentp/mosdepth) across 500 bp windows after filtering windows with more than 10% repetitive content. For poolseq samples, the difference in coverage values between the A and O pools was computed across the entire genome. Sample information for resequencing samples is recorded in the file SwD_popmap. See Fig. 3c-d, 5b, and S8.</p> <p>8. Assembly alignment<br> A. File names: nucmer_aln_Ifem_1049_ragtag_Afem_1354_ragtag.qr1_filter.reformat.coords.gz, nucmer_aln_Ofem_0081_ragtag_Afem_1354_ragtag.qr1_filter.reformat.coords.gz, nucmer_aln_Afem_Isen_Afem_Iele.qr1_filter.reformat.coords.gz, nucmer_aln_Ofem_Isen_Afem_Iele.qr1_filter.reformat.coords.gz, karyotype_AI_RagTag.csv, karyotype_AO_RagTag.csv, karyotype_AIsen_AIele.cs, karyotype_OIsen_AIele.csv<br> B. Description: Assembly alignments using nucmer v 4.0.0 (https://github.com/mummer4/mummer) and contig synteny for plotting using RIdeogram v 0.2.2 (https://cran.r-project.org/web/packages/RIdeogram/vignettes/RIdeogram.html) in R v 4.2.2 (https://www.r-project.org/). The A morph assembly of <em>I. elegans</em> was aligned to the I and O morph assemblies of<em> I. elegans</em> and to the A and O-like assemblies of <em>I. senegalensis</em>. See Fig. 4a, 5c.</p> <p>9. Genotyping the Darwin Tree of Life assembly<br> A. File names: nucmer_aln_Afem_ragtag_ToL-haplotigs.qr1_filter.reformat.coords.gz, nucmer_aln_Afem_ragtag_ToL-primary.qr1_filter.reformat.coords.gz, ToL_500_norepeat.regions.bed.gz, karyotype_AToL_13_unloc_RagTag.csv, karyotype_AToL_RagTag_haplotigs.csv<br> B. Description: To genotype the DToL reference assembly of<em> I. elegans</em>, we estimated read-depth coverage of the DToL long-read Pacbio data mapped to the A morph assembly of <em>I. elegans</em> generated in this study, and aligned the A morph assembly to both the primary DToL assembly and to the purged haplotigs. Read depth was estimated in mosdepth v 0.2.8 (https://github.com/brentp/mosdepth) and assembly alignments were conducted using nucmer v 4.0.0 (https://github.com/mummer4/mummer). See Fig. S3.</p> <p>10. SV calling<br> A. File names: A_to_A.bam, A_to_A.bam.bai, A_to_I.bam, A_to_I.bam.bai, A_to_O.bam, A_to_O.bam.bai, A_to_ToL_2mb.bam, A_to_ToL_2mb.bam.bai, I_to_A.bam, I_to_A.bam.bai, I_to_I.bam, I_to_I.bam.bai, I_to_O.bam, I_to_O.bam.bai, I_to_ToL_2mb.bam, I_to_ToL_2mb.bam.bai, O_to_A.bam, O_to_A.bam.bai, O_to_I.bam, O_to_I.bam.bai, O_to_O.bam, O_to_O.bam.bai, O_to_ToL_2mb.bam, O_to_ToL_2mb.bam.bai<br> B. Description: mergede alignements of resequencing samples (n = 19 per morph) to alternative reference assemblies (A, I, O, and DToL) for<em> I. elegans</em>. The alignments have been filtered by quality and to contain only the unlocalized scaffold 2 of chromosome 13, which includes the morph locus. These files were used to call morph-specific structural variants using samplot v 1.3.0 (https://github.com/ryanlayer/samplot). See Extended Data Figs 2, 7, and Fig. S5-S6.</p> <p>11. Mapping of inversion breakpoint reads<br> A. File names: AvO_3K.tsv.gz, AvO_22K.tsv.gz, AvO_sen_3K.tsv.gz, AvO_sen_22K.tsv.gz, IvO_3K.tsv.gz<br> B. Description: Signatures of an inversion with breakpoints at ~ 3 kb and ~ 22 kb of the unlocalized scaffold 2 of chromosome 13 on the O assembly were found in A and I resequencing samples of <em>I. elegans</em> and in poolseq samples of A females of <em>I. senegalensis</em>. We queried the reads mapping to the inversion breakpoints and then tabulated their mapping locations of the A morph assembly of<em> I. elegans</em> (Fig. 6 and Extended Data Fig. 3, 7b-c). For the first inversion breakpoint, we also mapped reads on the I morph assembly of <em>I.</em> elegans (Fig. S12).</p> <p>12. Evidence of translocation in I<br> A. File names: Ifem_nano_SUPER_13_unloc_2.bam, Ifem_nano_SUPER_13_unloc_2.bam.bai<br> B. Description: Long-read Nanopore data of a I morph female of <em>I. elegans</em> mapped to the A morph of <em>I. elegans</em> and filtered to contain the entire unlocalized scaffold 2 of chromosome 13. Read mapping was conducted in minimap2 v 2.22-r1110 (https://github.com/lh3/minimap2) and used to identify a translocation signature in the I morph, relative to the A morph of <em>I. elegans</em>. See Extended Data Fig. 6.</p> <p>13. PCA output<br> A. File names: A1354_all.eigenval, A1354_all.eigenvec, I1049_all.eigenval, I1049_all.eigenvec<br> B. Description: Eigenvectors and eigenvalues of PCA analyses of population structure between morphs of <em>I. elegans</em>. PCA analysis were conducted on morph locus, using either the A morph or the I morph assembly as mapping reference in PLINK v 1.9 (http://pngu.mgh.harvard.edu/purcell/plink/). See Fig. S4.</p> <p>14. Linkage disequilibrium<br> A. File names: A1354_SUPER_1_allr.ld.gz, A1354_SUPER_2_allr.ld.gz, A1354_SUPER_3_allr.ld.gz, A1354_SUPER_4_allr.ld.gz, A1354_SUPER_5_allr.ld.gz, A1354_SUPER_6_allr.ld.gz, A1354_SUPER_7_allr.ld.gz, A1354_SUPER_8_allr.ld.gz, A1354_SUPER_9_allr.ld.gz, A1354_SUPER_10_allr.ld.gz, A1354_SUPER_11_allr.ld.gz, A1354_SUPER_12_allr.ld.gz, A1354_SUPER_13_allr.ld.gz, A1354_SUPER_13_unloc_1_allr.ld.gz, A1354_SUPER_13_unloc_2_allr.ld.gz, A1354_SUPER_13_unloc_3_allr.ld.gz, A1354_SUPER_13_unloc_4_allr.ld.gz, A1354_SUPER_X_allr.ld.gz<br> B. Description: Estimates of recombination rate (R2) between SNPs across the first 15 mb of each chromosome and unlocalized segments of chromosome 13 of <em>I. elegans</em>. Recombination rates were estimated based on 57 resequencing samples and using the A morph assembly as mapping reference in PLINK v 1.9 (http://pngu.mgh.harvard.edu/purcell/plink/). See Extended Data Fig. 5.</p> <p>15. Gene annotations<br> A. File names: Afem_all_ragtag.gtf.gz, Afem_all_transcripts.transdecoder.genome.gff3.gz, Isen.gtf.gz<br> B. Description: Annotation of the A morph assembly of <em>I. elegans</em> using RNAseq data to assemble transcripts <em>de novo </em>for <em>I. elengans</em> and <em>I. senegalensis</em> in Stringtie v 2.1.4 (https://ccb.jhu.edu/software/stringtie/). Peptide sequences for the <em>I. elegans</em> transcripts were then predicted using Transdecoder v 5.5.0 (https://github.com/TransDecoder/TransDecoder).</p> <p>16. Gene annotations in the morph locus<br> A. File names: gene_models_shared_trancripts_simple.csv, gene_models_shared_trancripts_simple_I.csv<br> B. Description: locations of exon features for genes in the morphs locus and expressed in at least one adult sample of both<em> I. elegans</em> and <em>I. senegalensis</em>. Locations are given for the A and I assemblies. See Fig. 6 and S12.</p> <p>17. Gene expression<br> A. File names: DToL_gene_count_matrix.csv.gz, DToL_transcript_count_matrix.csv.gz, gene_count_matrix.csv.gz, transcript_count_matrix.csv.gz, Isen_gene_count_matrix.csv.gz, Isen_transcript_count_matrix.csv.gz, Iele_phenodata.csv, Isen_phenodata.csv<br> B. Description: Sample information (phenodata), gene and transcript count matrices for gene expression analysis. For <em>I. elegans</em>, gene expression was quantified on thoracic tissue of six adult females of each morph and six adult males (three sexually mature and three sexually immature in each group). Reads were mapped to both the A morph assembly and the DToL reference assembly. For <em>I. senegalensis</em>, we used previously published data (NCBI BioProject PRJDB11387) from different tissues of adult females of each morph and males (one upon emergence and one two days after emergence for each group) mapped to the A morph assembly. Gene and transcript counts were generated using Stringtie v 2.1.4 (https://ccb.jhu.edu/software/stringtie/). See Fig. 6, S9-S11, S13.</p> <p>18. SNPs in the morph locus<br> A. File names: A1354-ragtag-allsites-candidate_gene_cds.vcf.gz, A1354-ragtag-allsites-candidate_gene_cds.vcf.gz.tbi, vcf_popmap<br> B. Description: SNPs in 57 resequencing samples across coding sequences of the morph locus of <em>I. elegans</em>. The A morph assembly was used as mapping reference. Sample information for resequencing samples is recorded in the file vcf_popmap. See Fig. S14a.</p> <p>19. Domains and orthologues of Gastrula zinc-finger transcription factor in the morph locus<br> A. File names: GZnf_domain_annot.csv, GZnF_orthologue.tre, GZnf_orthologue_annot.txt<br> B. Description: Functional domains were annotated using InterProScan (https://www.ebi.ac.uk/interpro/). The gene orthologue tree was inferred using OrthoFinder v 2.5.2 (https://github.com/davidemms/OrthoFinder). See Fig. S14.</p>
Geographical gradients of genetic diversity and differentiation among the southernmost marginal populations of Abies sachalinensis revealed by EST-SSR polymorphism
Research Highlights: We detected the longitudinal gradients of genetic diversity parameters, such as the number of alleles, effective number of alleles, heterozygosity, and inbreeding coefficient, and found that these might be attributable to climatic conditions, such as temperature and snow depth. Background and Objectives: Genetic diversity among local populations of a plant species at its distributional margin has long been of interest in ecological genetics. Populations at the distribution center grow well in favorable conditions, but those at the range margins are exposed to unfavorable environments, and the environmental conditions at establishment sites might reflect the genetic diversity of local populations. This is known as the central-marginal hypothesis in which marginal populations show lower genetic variation and higher differentiation than do central populations. In addition, genetic variation in a local population is influenced by phylogenetic constraints and the population history of selection under environmental constraints. In this study, we investigated this hypothesis in relation to Abies sachalinensis, a major conifer species in Hokkaido. Materials and methods: A total of 1,189 trees from 25 natural populations were analyzed using 19 EST-SSR loci. Results: The eastern populations; namely, those in the species distribution center, showed greater genetic diversity than did the western peripheral populations. Another important finding is that the southwestern marginal populations were highly differentiated from the other populations. Conclusions: These differences might be due to genetic drift in the small and isolated populations at the range margin. Therefore, our results indicated that the central-marginal hypothesis held true for the southernmost A. sachalinensis populations in Hokkaido.
Fig. 3. A. a in Untangling species identity in gastropods with polymorphic shells in the genus Bolma Risso, 1826 (Mollusca, Vetigastropoda)
Fig. 3. A. a, Bolma henica madagascarensis (Indian Ocean); b, Bo. henica abyssorum; c, Bo. henica henica, with type locality represented by a white star (Fiji Island, Southwest Pacific); d, Bo. cf. minutiradiosa. B. a–d, distinct shell morphs found in Bo. recens, with type locality represented by a white star (Kiwi seamount, Three Kings Ridge). C. a, Bo. mainbaza, with type locality (South Madagascar); b, Bo. pseudobathyraphis, with type locality (South New Caledonia); c, Bo. millegranosa; d, Bo. opaoana with type locality (South New Caledonia, Crypthélia Bank).
Fig. 4 in Untangling species identity in gastropods with polymorphic shells in the genus Bolma Risso, 1826 (Mollusca, Vetigastropoda)
Fig. 4. Shell diversity across the molecular phylogeny of the "deep-water" clade of the subfamily Turbininae (Williams 2007, i.e., the genera Astraea, Bellastraea , Bolma and Guildfordia). The phylogeny is based on Bayesian analyses of the concatenated sequences from cox1 and 28 S genes, incorporating an uncorrelated relaxed, log- normal clock produced using *BEAST. The tree is a maximum clade credibility tree with median node heights based in 9000 trees. Support values are posterior probabilities (PP); branches < 50% were collapsed. Species names are labelled on the right-hand side. Species hypotheses previously delineated by the integrative taxonomy approach are highlighted by the grey boxes.
Fig. 2 in Untangling species identity in gastropods with polymorphic shells in the genus Bolma Risso, 1826 (Mollusca, Vetigastropoda)
Fig. 2. [next page] Molecular based species delineation of the genus "Bolma". A. Ultrametric tree produced using BEAST based on cox1 sequences. B. PSHs derived from the GMYC model and labelled from 1 to 37. C. PSHs derived from the GMYC model using the lower limit of the equivalent of a 95% confidence interval, and labelled from A to ZD. D. SSHs drawn from congruency between cox1 and 28S. Boxes with a black outline indicate that the SSH was monophyletic in both cox1 and 28S trees. Boxes without a black outline highlight SSHs for which molecular data were either incomplete or non-informative. SSHs labelled from A to ZD (following step C) or with the species name when our sequences matched published data associated with the species names. E. PSHs derived from the Bayesian analysis based on 28S sequences. F. Bayesian, non-ultrametric tree produced using BEAST based on 28S sequences. G. Species names retained in the present study. For the SSH E-F-G-H, the name Bo. henica was retained; however, Bo. henica abyssorum, Bo. henica madagascarensis and Bo. henica henica are represented as sub-species separated by white dotted lines. For both trees, nodal support values are posterior probabilities (PP), shown only for PP> 50%. Branches with PP <50% were collapsed. Red and green branches correspond to monophyletic species hypotheses. Colour coded boxes: red corresponds to cox1 PSHs supported by PP> 95%; light red corresponds to cox1 PSH supported by PP <95%; light grey corresponds to cox1 and 28S singletons; a grey cross represents missing data; green corresponds to 28S species hypotheses supported by PP> 95%; light grey corresponds to groups of genotypes displaying diagnostic 28S sites. Specimen numbers are given in the Supplementary file.
Figure 7 in A new subterranean Maraenobiotus (Crustacea: Copepoda) from Slovenia challenges the concept of polymorphic and widely distributed harpacticoids
Figure 7. Maraenobiotus slovenicus sp. nov., line drawings, (A–C) holotype female; (D, E) allotype male: (A) genital segment with attached spermatophore, ventral; (B) last urosomite, anal somite and furcal rami, ventral; (C) last urosomite, anal somite and furcal rami, dorsal; (D) last urosomite, anal somite and furcal rami, ventral; (E) anal somite and furcal rami, dorsal.
Figure 4 in A new subterranean Maraenobiotus (Crustacea: Copepoda) from Slovenia challenges the concept of polymorphic and widely distributed harpacticoids
Figure 4. Maraenobiotus slovenicus sp. nov., SEM micrographs, (A–D) damaged paratype male 1; (E–H) paratype male 2: (A) mouth appendages, ventral; (B) maxillule and maxilla, ventral; (C) central part of left antennule, ventral; (D) antenna, ventral; (E) left antennule, lateral; (F) first three urosomites, lateral; (G) anal somite and caudal rami, lateral; (H) detail of first urosomite, with cuticular window, large pore, and sensillum, lateral.
Figure 3 in A new subterranean Maraenobiotus (Crustacea: Copepoda) from Slovenia challenges the concept of polymorphic and widely distributed harpacticoids
Figure 3. Maraenobiotus slovenicus sp. nov. (A, B) SEM micrographs, paratype female 3; (C–H) damaged paratype male 1: (A) first two urosomites and P5; (B) P1–P3, ventrolateral; (C) habitus with several large epibiotic ciliates, ventral; (D) P4–P6, ventral; (E) last two urosomites and caudal rami, ventral; (F) antennule, ventral; (G) distal part of right antennule, ventral; (H) central part of right antennule, ventral.
Figure 1 in A new subterranean Maraenobiotus (Crustacea: Copepoda) from Slovenia challenges the concept of polymorphic and widely distributed harpacticoids
Figure 1. Maraenobiotus slovenicus sp. nov., SEM micrographs, paratype female 1: (A) habitus, dorsal; (B) cephalothorax, dorsal; (C) anterior part of cephalothorax, dorsal; (D) right antennule, dorsal; (E) free pedigerous somites, dorsal; (F) first three urosomites, dorsal; (G) last three urosomites and caudal rami, dorsal; (H) right caudal ramus, dorsal.
Data from: Gene flow, ancient polymorphism, and ecological adaptation shape the genomic landscape of divergence among Darwin's finches
Genomic comparisons of closely related species have identified "islands" of locally elevated sequence divergence. Genomic islands may contain functional variants involved in local adaptation or reproductive isolation and may therefore play an important role in the speciation process. However, genomic islands can also arise through evolutionary processes unrelated to speciation, and examination of their properties can illuminate how new species evolve. Here, we performed scans for regions of high relative divergence (FST) in 12 species pairs of Darwin's finches at different genetic distances. In each pair, we identify genomic islands that are, on average, elevated in both relative divergence (FST) and absolute divergence (dXY). This signal indicates that haplotypes within these genomic regions became isolated from each other earlier than the rest of the genome. Interestingly, similar numbers of genomic islands of elevated dXY are observed in sympatric and allopatric species pairs, suggesting that recent gene flow is not a major factor in their formation. We find that two of the most pronounced genomic islands contain the ALX1 and HMGA2 loci, which are associated with variation in beak shape and size, respectively, suggesting that they are involved in ecological adaptation. A subset of genomic island regions, including these loci, appears to represent anciently diverged haplotypes that evolved early during the radiation of Darwin's finches. Comparative genomics data indicate that these loci, and genomic islands in general, have exceptionally low recombination rates, which may play a role in their establishment.
Differential associations between nucleotide polymorphisms and physiological traits in Norway spruce (Picea abies Karst.) provenances under contrasting water regimes
<p>Three datasets are provided here, yielded by a study on drought-stressed and control (well-watered) seedlings of Norway spruce (Picea abies Karst.), coming from 5 provenances distributed along a steep altitudinal gradient from 550 to 1,280 m a.s.l. in central Slovakia:</p> <p>1. physiological traits</p> <p>2. double-digest restriction-site associated sequencing data (ddRAD)</p> <p>3. nuclear microsatellite (nSSR) genotypes</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.