Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
637
datasets available to search
ShareScore release 0.7.1
Dataset results
637 results for “population analysis”
Van Dijk et al. (2021), A meta-analysis of projected global food demand and population at risk of hunger for the period 2010–2050, data and scripts
<p>This repository contains all data and R scripts to reproduce the figures in Van Dijk et al. (2021), A meta-analysis of global food demand and population at risk of hunger projections for the period 2010-2050, Nature Food. More specifically, it includes two databases: (1) A database with standardized information to describe the characteristics of 57 studies that were identified by the systematic literature review and (2) The Global Food Security Projections Database v1.0.1 with harmonized projections for three global food security indicators: food consumption in kcal per capita and total kcal, and population at risk of hunger. The database also includes projections for total global population that are required to derive the global food security indicators.</p> <p>The two scripts (nf_figures.r and nf_meta_regression.r) can be used to reproduce the figures and tables in the main paper and the supplementary information. Please start with the first script, which sources the second script. </p> <p>This is the first version of the Global Food Projections Database. We expect to update the data, including additional studies and variables in the future. For issues and suggestions, please contact michiel.vandijk@wur.nl.</p> <p> </p>
REPIN population analysis in 42 Pseudomonas chlororaphis genomes
<p>This dataset is the output of RAREFAN (http://rarefan.evolbio.mpg.de/) a webserver to identify REPIN populations across an entire bacterial species. The data was created using the following command "java -jar -Xmx10g rarefan.jar chlororaphis/in/ chlororaphis/out/ chlTAMOak81.fas 55 21 chlororaphis/in/yafM_SBW25.faa chlororaphis.nwk 1e-30 true 1"</p> <p>All input files are located in the folder chlororaphis/in/, all output data is located in chlororaphis/out/.</p> <p>The input files include the 42 fasta formatted <em>P. chlororaphis</em> genome files (*.fas) and a RAYT protein sequence called yafM_SBW25.faa.</p> <p>The output files include the following:</p> <p>A phylogenetic tree "chlororaphis.nwk" of all genomes generated with andi (<a href="http://github.com/evolbioinf/andi/">http://github.com/evolbioinf/andi/</a>) and clustDist (http://guanine.evolbio.mpg.de/problemsBook/node1.html).</p> <p>A file containing the frequencies of all 21bp long sequences found in the TAMOak81 genome: chlTAMOak81.wfr</p> <p>A file containing all 21bp long sequences that occur more frequently than 55 times in the TAMOak81 genome: chlTAMOak81.overrep</p> <p>A file containing information on the RAYTs and their cooccurrence with different REPIN populations: prox.stats</p> <p>A file containing the nucleotide sequences of all yafM_SBW25.faa relatives identified with BLAST+ in the <em>P. chlororaphis</em> species: yafM_relatives.fna</p> <p>maxREPIN_[0-5] Contains the most frequent REPIN identified for each sequence type in each <em>P. chlororaphis</em> strain.</p> <p> presAbs_[0-5].txt Contains for each strain information on the number of RAYTs, the number of REPINs, the master sequence, the number of master sequences, the entire REP/REPIN population size, the number of REPIN clusters that contain more than 10 sequences, all REPINs in the population as well as all all REPINs that differ to the master sequences in at most three nucleotides.</p> <p>rayt_[strain name].tab contains location information for each identified RAYT relative for each strain. The files can be viewed with artemis.</p> <p>results.txt contains for each strain the frequency of the six identified 21bp long seeds.</p> <p>There is one folder called groupSeedSequences, which includes the data for identifying the most common 21 bp long sequences in <em>P. chlororaphis</em> TAMOak81. All 21bp long sequences in the genome that occur more frequently than 55 times are sorted into 6 sequence groups. These sequence groups are stored in the files Group_chlTAMOak81_*.out and .out.fas. There is also a chlTAMOak81_words.tab file, which contains the locations of all overrepresented 21bp long sequences in the TAMOak81 genome. This file can be viewed in artemis (https://www.sanger.ac.uk/tool/artemis/) together with the TAMOak81 genome file. The most common sequence in each group is used as a seed sequence to determine REPIN populations across all 42 genomes.</p> <p> </p> <p>For each genome there are six output folders (ending in _0 to _5), for each sequence group one.</p> <p>Each folder contains the following files:</p> <p>*.dd: Degree distribution of the REPIN network, where each REPIN is a node. A REPIN is connected to another REPIN if they differ in exactly one position. The degree distribution is a histogram of the number of connections of all the nodes. </p> <p>*.hist For the largest sequence cluster determined by mcl that consists of REPINs (two REPs in inverted orientation) this file contains the number of REPINs in each sequence class. Sequence class 0 is the master sequence. By definition the most common REPIN in the sequence population. Sequence class 1 contains all REPINs differing in exactly one position to the master sequence. Sequence class 2 contains REPINs differing in 2 positions etc.</p> <p>*.mcl Contains the clustering output by mcl. Each line contains the member of a cluster. Lines are sorted by cluster size.</p> <p>*.mw Contains the most common 21bp long sequence and its frequency in the genome, which is the basis for identifying first all related REP sequences and from those the REPINs formed by these REP sequences.</p> <p>*.nodes The identity and frequency of all REPINs and REP sequences for either all sequences or only for the largest sequence cluster.</p> <p>*.ss Contains REPINs and REP sequences as well as their positions in fasta format. Position information starts with the location in genome fasta file (first sequence is 0...) followed by the start and end position of the entire REPIN/REP sequence. </p> <p>*.ss.REP REP sequence information in fasta format.</p> <p>*.tab Location in tab format. Can be used to display locations of REPs and REPINs in the genome via artemis.</p> <p>*_[0-9].ss Contains REPIN/REP sequence information for each subcluster separately.</p> <p>*_[0-9].tab Contains the location of REP/REPINs for each subcluster separately for viewing in artemis.</p> <p>*allSeed.nw Contains network connections between nodes of all sequences. Can be used to view network in for example R or cytoscape together with the nodes file.</p> <p>*largestCluster.nodes Information on nodes only from the largest REPIN cluster.</p> <p>*largestCluster.ss *.ss file for the largest REPIN cluster.</p> <p>*largestCluster.tab *.tab file for the largest REPIN cluster.</p> <p>*_rayt_repin_prox.txt shows which REPIN/REP cluster is in proximity to any of the RAYT genes identified in the genome (within 200bp).</p> <p>And a subfolder that contains the complete sequences (including the variable region) for all identified REPs and REPINs.</p> <p><strong>The dataset was generated using the following external tools:</strong></p> <p>andi for tree building:</p> <p>B Haubold, F Klötzl, and P Pfaffelhuber. <strong>andi: fast and accurate estimation of evolutionary distances between closely related genomes.</strong> Bioinformatics, 2015 vol. 31 (8) pp. 1169-1175.</p> <p>MCL for REPIN population clustering:</p> <p>A J Enright, S Van Dongen, and C A Ouzounis. <strong>An efficient algorithm for large-scale detection of protein families.</strong> Nucleic Acids Research, 2002 vol. 30 (7) pp. 1575-1584.</p> <p>BLAST+ for identifying RAYT relatives in the different genomes:</p> <p>C Camacho, G Coulouris, V Avagyan, N Ma, J Papadopoulos, K Bealer, and T L Madden. <strong>BLAST+: architecture and applications.</strong> BMC Bioinformatics, 2009 vol. 10 (1) pp. 421-9.</p>
REPIN population analysis in 130 Neisseria meningitidis and N. gonorrhoeae genomes
<p>This dataset is the output of RAREFAN (http://rarefan.evolbio.mpg.de/) a webserver to identify REPIN populations across an entire bacterial species. The data was created using the following command "java -jar -Xmx10g rarefan.jar neisseria/in/ neisseria/out/ Nmen_2594.fas 55 21 neisseria/in/NMAA_0235.faa neisseria.nwk 1e-80 true 1"</p> <p>All input files are located in the folder neisseria/in/, all output data is located in neisseria/out/.</p> <p>The input files include the 130 fasta formatted <em>Neisseria</em> genome files (*.fas) and a RAYT protein sequence called NMAA_0235.faa.</p> <p>The output files include the following:</p> <p>A phylogenetic tree "neisseria.nwk" of all genomes generated with andi (<a href="http://github.com/evolbioinf/andi/">http://github.com/evolbioinf/andi/</a>) and clustDist (http://guanine.evolbio.mpg.de/problemsBook/node1.html).</p> <p>A file containing the frequencies of all 21bp long sequences found in the Nmen_2594 genome: Nmen_2594.wfr</p> <p>A file containing all 21bp long sequences that occur more frequently than 55 times in the Nmen_2594 genome: Nmen_2594.overrep</p> <p>A file containing information on the RAYTs and their cooccurrence with different REPIN populations: prox.stats</p> <p>A file containing the nucleotide sequences of all NMAA_0235.faa relatives identified with BLAST+ in the <em>Neisseria </em>species: yafM_relatives.fna</p> <p>maxREPIN_[0-5] Contains the most frequent REPIN identified for each sequence type in each <em>Neisseria</em> strain.</p> <p> presAbs_[0-5].txt Contains for each strain information on the number of RAYTs, the number of REPINs, the master sequence, the number of master sequences, the entire REP/REPIN population size, the number of REPIN clusters that contain more than 10 sequences, all REPINs in the population as well as all REPINs that differ to the master sequences in at most three nucleotides.</p> <p>rayt_[strain name].tab contains location information for each identified RAYT relative for each strain. The files can be viewed with artemis.</p> <p>results.txt contains for each strain the frequency of the six identified 21bp long seeds.</p> <p>There is one folder called groupSeedSequences, which includes the data for identifying the most common 21 bp long sequences in <em>Neisseria meningitidis</em> WUE 2594. All 21bp long sequences in the genome that occur more frequently than 55 times are sorted into 6 sequence groups. These sequence groups are stored in the files Group_Nmen_2594_*.out and .out.fas. There is also a Nmen_2594_words.tab file, which contains the locations of all overrepresented 21bp long sequences in the Nmen_2594 genome. This file can be viewed in artemis (https://www.sanger.ac.uk/tool/artemis/) together with the Nmen_2594 genome file. The most common sequence in each group is used as a seed sequence to determine REPIN populations across all 130 genomes.</p> <p> </p> <p>For each genome there are six output folders (ending in _0 to _5), for each sequence group one.</p> <p>Each folder contains the following files:</p> <p>*.dd: Degree distribution of the REPIN network, where each REPIN is a node. A REPIN is connected to another REPIN if they differ in exactly one position. The degree distribution is a histogram of the number of connections of all the nodes. </p> <p>*.hist For the largest sequence cluster determined by mcl that consists of REPINs (two REPs in inverted orientation) this file contains the number of REPINs in each sequence class. Sequence class 0 is the master sequence. By definition the most common REPIN in the sequence population. Sequence class 1 contains all REPINs differing in exactly one position to the master sequence. Sequence class 2 contains REPINs differing in 2 positions etc.</p> <p>*.mcl Contains the clustering output by mcl. Each line contains the member of a cluster. Lines are sorted by cluster size.</p> <p>*.mw Contains the most common 21bp long sequence and its frequency in the genome, which is the basis for identifying first all related REP sequences and from those the REPINs formed by these REP sequences.</p> <p>*.nodes The identity and frequency of all REPINs and REP sequences for either all sequences or only for the largest sequence cluster.</p> <p>*.ss Contains REPINs and REP sequences as well as their positions in fasta format. Position information starts with the location in genome fasta file (first sequence is 0...) followed by the start and end position of the entire REPIN/REP sequence. </p> <p>*.ss.REP REP sequence information in fasta format.</p> <p>*.tab Location in tab format. Can be used to display locations of REPs and REPINs in the genome via artemis.</p> <p>*_[0-9].ss Contains REPIN/REP sequence information for each subcluster separately.</p> <p>*_[0-9].tab Contains the location of REP/REPINs for each subcluster separately for viewing in artemis.</p> <p>*allSeed.nw Contains network connections between nodes of all sequences. Can be used to view network in for example R or cytoscape together with the nodes file.</p> <p>*largestCluster.nodes Information on nodes only from the largest REPIN cluster.</p> <p>*largestCluster.ss *.ss file for the largest REPIN cluster.</p> <p>*largestCluster.tab *.tab file for the largest REPIN cluster.</p> <p>*_rayt_repin_prox.txt shows which REPIN/REP cluster is in proximity to any of the RAYT genes identified in the genome (within 200bp).</p> <p>And a subfolder that contains the complete sequences (including the variable region) for all identified REPs and REPINs.</p> <p>The dataset was generated using the following external tools:</p> <p>andi for tree building:</p> <p>B Haubold, F Klötzl, and P Pfaffelhuber. <strong>andi: fast and accurate estimation of evolutionary distances between closely related genomes.</strong> Bioinformatics, 2015 vol. 31 (8) pp. 1169-1175.</p> <p>MCL for REPIN population clustering:</p> <p>A J Enright, S Van Dongen, and C A Ouzounis. <strong>An efficient algorithm for large-scale detection of protein families.</strong> Nucleic Acids Research, 2002 vol. 30 (7) pp. 1575-1584.</p> <p>BLAST+ for identifying RAYT relatives in the different genomes:</p> <p>C Camacho, G Coulouris, V Avagyan, N Ma, J Papadopoulos, K Bealer, and T L Madden. <strong>BLAST+: architecture and applications.</strong> BMC Bioinformatics, 2009 vol. 10 (1) pp. 421-9.</p>
Supporting information for "Few-Shot Learning Enables Population-Scale Analysis of Leaf Traits in Populus trichocarpa"
<p><strong>Description</strong></p> <p>In this work, we use few-shot learning to segment the body and vein architecture of <em>P. trichocarpa</em> leaves from high-resolution scans obtained in the UC Davis common garden. Leaf and vein segmentation are formulated as separate tasks, in which convolutional neural networks (CNNs) are used to iteratively expand partial segmentations until reaching stopping criteria. Our leaf and vein segmentation approaches use just 50 and 8 manually traced images for training, respectively, and are applied to a set of 2,634 top and bottom leaf scans. We show that both methods achieve high segmentation accuracy and retain biologically realistic features. The leaf and vein segmentations are compared against a U-Net baseline model, and subsequently used to extract 68 morphological traits using traditional open-source image processing tools, which are validated using real-world physical measurements. For a biological perspective, we perform a genome-wide association study using the "vein density" trait to discover novel genetic architectures associated with multiple physiological processes relating to leaf development and function. In addition to sharing all of the few-shot learning code, we are releasing all images, manual segmentations, model predictions, 68 extracted leaf phenotypes, and a new set of SNPs called against the v4 <em>P. trichocarpa</em> genome for 1,419 genotypes.</p> <p><strong>Directories:</strong></p> <pre><code>Few-shot learning for p. trichocarpa leaf traits ├── data │ ├── genomes │ │ ├── Ptri_V4_Nisq1.[...].bed │ │ ├── Ptri_V4_Nisq1.[...].bim │ │ └── Ptri_V4_Nisq1.[...].fam │ ├── images │ │ └── *.jpeg │ ├── leaf_masks │ │ └── *.png │ ├── leaf_preds │ │ └── *.png │ ├── leaf_unet_preds │ │ └── *.png │ ├── results │ │ ├── digital_traits.tsv │ │ ├── gwas_results.csv │ │ ├── manual_traits.tsv │ │ ├── vein_density_blups.tsv │ │ └── vein_density_tps_adj.tsv │ ├── vein_bce_preds │ │ └── *.png │ ├── vein_bce_probs │ │ └── *.png │ ├── vein_fl_preds │ │ └── *.png │ ├── vein_fl_probs │ │ └── *.png │ ├── vein_masks │ │ └── *.png │ ├── vein_unet_bce_preds │ │ └── *.png │ ├── vein_unet_bce_probs │ │ └── *.png │ ├── vein_unet_fl_preds │ │ └── *.png │ ├── vein_unet_fl_probs │ │ └── *.png ├── figures │ └── *.png ├── logs │ ├── leaf_tracer_256.txt │ ├── leaf_unet_256.txt │ ├── vein_grower_bce_128.txt │ ├── vein_grower_fl_128.txt │ ├── vein_unet_bce_128.txt │ └── vein_unet_fl_128.txt ├── models │ ├── BuildCNN.py │ ├── BuildUNet.py │ ├── LeafTracer.py │ └── VeinGrower.py ├── notebooks │ ├── Figures.ipynb │ ├── GrowerInference.ipynb │ ├── GrowerTraining.ipynb │ ├── TracerInference.ipynb │ ├── TracerTraining.ipynb │ ├── UNetLeafSegmentation.ipynb │ └── UNetVeinSegmentation.ipynb ├── utils │ ├── GetLowestGPU.py │ ├── ImageLoader.py │ ├── LeafGenerator.py │ ├── ModelWrapperGenerator.py │ ├── TimeRemaining.py │ ├── TraceInitializer.py │ ├── UNetTileGenerator.py │ └── VeinGenerator.py └── weights ├── leaf_tracer_256_best_val_model.save ├── leaf_unet_256_best_val_model.save ├── vein_grower_bce_128_best_val_model.save ├── vein_grower_fl_128_best_val_model.save ├── vein_unet_bce_128_best_val_model.save └── vein_unet_fl_128_best_val_model.save </code></pre> <p><strong>Data:</strong></p> <p>The <code>data</code> folder includes all images, ground truth segmentations, predicted segmentations, and extracted leaf traits. All images encode the sample ID in the file name by indicating the treatment, block, row, position, and leaf side, respectively. For example, the file, <code>C_1_1_2_bot.jpeg</code>, indicates the control treatment, block 1, row 1, position 2, and the bottom side of the leaf. Tabulated results include position IDs as well as the corresponding genotype IDs.</p> <ul> <li>The <code>images</code> folder includes the 2,906 high-resolution leaf scans taken in the field.</li> <li>The <code>leaf_masks</code> folder includes 50 ground truth segmentations used for training the leaf tracing algorithm.</li> <li>The <code>leaf_preds</code> folder includes the 2,906 predicted segmentations from the leaf tracing algorithm.</li> <li>The <code>leaf_unet_preds</code> folder includes the 2,906 predicted segmentations from the U-Net model for leaf segmentation.</li> <li>The <code>vein_masks</code> folder includes 8 ground truth segmentations used for training the vein growing algorithm.</li> <li>The <code>vein_*_preds</code> folder includes the 1,453 predicted segmentations from the vein growing algorithm, where * specifies the loss function (bce: binary cross-entropy, fl: focal loss).</li> <li>The <code>vein_*_probs</code> folder includes the 1,453 predicted probability maps from the vein growing algorithm before thresholding, where * specifies the loss function (bce: binary cross-entropy, fl: focal loss).</li> <li>The <code>vein_unet_*_preds</code> folder includes the 1,453 predicted segmentations from the U-Net model for vein segmentation, where * specifies the loss function (bce: binary cross-entropy, fl: focal loss).</li> <li>The <code>vein_unet_*_probs</code> folder includes the 1,453 predicted probability maps from the U-Net model for vein segmentation before thresholding, where * specifies the loss function (bce: binary cross-entropy, fl: focal loss).</li> <li>The <code>genomes</code> folder includes the set of SNPs called against the v4 <em>P. trichocarpa</em> genome for 1,419 genotypes with a README file detailing the steps taken.</li> <li>The <code>results</code> folder includes <ul> <li>Raw values of the 68 predicted leaf traits in <code>digital_traits.tsv</code></li> <li>Manually measured values of petiole length and width in <code>manual_traits.tsv</code></li> <li>Thin plate spline (TPS) adjusted values of the vein density trait in <code>vein_density_tps_adj.tsv</code></li> <li>Best linear unbiased prediction (BLUP) adjusted values of the vein density trait in <code>vein_density_blups.tsv</code></li> <li>GWAS results for the vein density trait, including chromosome positions and corresponding P values, in <code>gwas_results.csv</code></li> </ul> </li> </ul> <p><strong>Figures:</strong></p> <p>The <code>figures</code> folder includes all figures and videos used in the manuscript. See <code>notebooks/Figures.ipynb</code> for the methods used to generate these figures.</p> <p><strong>Logs:</strong></p> <p>The <code>logs</code> folder includes logs of CNN convergence for the training and validation sets during model training for the leaf tracing CNN vein growing CNN, and U-Net models. The file names include the model, loss function (bce: binary cross-entropy, fl: focal loss), and size of the input window for each method (e.g., 128 for the vein growing CNN).</p> <p><strong>Models:</strong></p> <p>The <code>models</code> folder includes the CNN implementations in PyTorch as well as the leaf tracing and vein growing algorithms at inference time.</p> <ul> <li><code>BuildCNN.py</code> defines the CNN architecture for leaf tracing or vein growing, with user-specified input shape, output shape, layers, and output activation functions.</li> <li><code>BuildUNet.py</code> defines the U-Net architecture for leaf and vein segmentation, with user-specified input/output shape, layers, and output activation functions.</li> <li><code>LeafTracer.py</code> defines the leaf tracing algorithm at inference time.</li> <li><code>VeinGrower.py</code> defines the vein growing algorithm at inference time.</li> </ul> <p><strong>Notebooks:</strong></p> <p>The <code>notebooks</code> folder includes Jupyter notebooks used for model training, model inference, and figure generation.</p> <ul> <li><code>Figures.ipynb</code> is used to generate all of the manuscript figures.</li> <li><code>GrowerTraining.ipynb</code> is used to train the vein growing CNN.</li> <li><code>GrowerInference.ipynb</code> is used to apply the vein growing algorithm to the 1,453 leaf bottom images.</li> <li><code>TracerTraining.ipynb</code> is used to train the leaf tracing CNN.</li> <li><code>TracerInference.ipynb</code> is used to apply the leaf tracing algorithm to the 2,906 leaf top and bottom images.</li> <li><code>UNetLeafSegmentation.ipynb</code> is used to train and apply U-Net for leaf segmentation.</li> <li><code>UNetVeinSegmentation.ipynb</code> is used to train and apply U-Net for vein segmentation.</li> </ul> <p><strong>Utils:</strong></p> <p>The <code>utils</code> folder includes utility scripts implemented in Python that assist in model training and inference.</p> <ul> <li><code>ImageLoader.py</code> loads image/mask pairs for sampling training/validation tiles.</li> <li><code>LeafGenerator.py</code> generates inputs/outputs for the leaf tracing CNN.</li> <li><code>VeinGenerator.py</code> generates inputs/outputs for the vein growing CNN.</li> <li><code>UNetTileGenerator.py</code> generates inputs/outputs for the U-Net model.</li> <li><code>GetLowestGPU.py</code> identifies available GPUs using the <code>nvidia-smi</code> command and selects the one with lowest memory usage, if none available the device is set to CPU.</li> <li><code>ModelWrapperGenerator.py</code> wraps the PyTorch CNN and data loaders with similar functionality to the Keras Model class in TensorFlow (e.g., model.fit(...)).</li> <li><code>TimeRemaining.py</code> is used by the model wrapper to estimate remaining time left per epoch.</li> <li><code>TraceInitializer.py</code> is used by the tracing algorithm at inference time to initialize the leaf trace using automatic thresholding.</li> </ul> <p><strong>Weights:</strong></p> <p>The <code>weights</code> folder includes the CNN parameters from the epoch resulting in the best validation error. The file names include the model, loss function (bce: binary cross-entropy, fl: focal loss), and size of the input window for each method (e.g., 128 for the vein growing CNN). The weights are loaded into the CNN models for inference.</p> <p><strong>Citation:</strong></p> <pre><code>@article{ doi:10.34133/plantphenomics.0072, author = {John Lagergren and Mirko Pavicic and Hari B. Chhetri and Larry M. York and Doug Hyatt and David Kainer and Erica M. Rutter and Kevin Flores and Jack Bailey-Bale and Marie Klein and Gail Taylor and Daniel Jacobson and Jared Streich }, title = {Few-Shot Learning Enables Population-Scale Analysis of Leaf Traits in Populus trichocarpa}, journal = {Plant Phenomics}, volume = {0}, number = {ja}, pages = {}, year = {}, doi = {10.34133/plantphenomics.0072}, URL = {https://spj.science.org/doi/abs/10.34133/plantphenomics.0072}, eprint = {https://spj.science.org/doi/pdf/10.34133/plantphenomics.0072}, }</code></pre>
Datasets, reproducible codes, and results for evaluating differential expression analysis methods on population-level RNA-seq data
<p>This upload contains the necessary R codes and data to reproduce the FDR and Power results described in our correspondence "Neglecting normalization impact in semi-synthetic RNA-seq data simulation generates artificial false positives" to Li Y, Ge X, Peng F, Li W, Li JJ, Exaggerated false positives by popular differential expression methods when analyzing human population samples, <em>Genome Biology</em> 23, 79, 2022, DOI: 10.1186/s13059-022-02648-4.</p>
A scalable, accurate, and universal analysis framework using individual-level allele frequency for large-scale genetic association studies in an admixed population
<p>Inclusion of individuals with diverse or admixed genetic ancestries is crucial to discover novel findings that may be missed by genomics analyses rooted solely in Caucasian population. Here, we present an analysis framework, SPAmix, which is scalable to a large-scale biobank data analysis including hundreds of thousands of admixed individuals and is universally applicable to various types of complex traits including binary trait, quantitative trait, time-to-event trait, longitudinal traits, etc. For each genetic variant, SPAmix uses genotype data and genetic principal components (PCs) to estimate individual-level allele frequency, which is subsequently used to calibrate p values via a retrospective analysis. A hybrid strategy including saddlepoint approximation (SPA) can greatly increase the accuracy to analyze rare genetic variants, especially if the phenotypic distribution is unbalanced or extremely unbalanced. Compared to Tractor, SPAmix does not require local ancestry information and can be straightforwardly applicable to a multi-way admixed population. Meanwhile, SPAmix can also be extended to SPAmix<sub>local</sub> in which the local ancestry can be incorporated if available. In addition, we propose SPAmix<sub>CCT</sub> to combine the p values of SPAmix and SPAmix<sub>local</sub> via Cauchy combination (CCT). SPAmix<sub>local</sub> performs close to Tractor when analyzing quantitative traits and is more accurate when analyzing binary traits with an unbalanced case-control ratio. And SPAmix<sub>CCT </sub>is an optimal unified approach for various cross-ancestry genetic architectures. Extensive simulation studies and real data analyses of 369,314 UK Biobank individuals from multiple ancestries demonstrated that SPAmix is scalable and can discover novel hits while controlling type I error rates well.</p>
Fig. 1 in Population viability analysis of the population of Raffles' banded langurs Presbytis femoralis in Singapore
Fig. 1. Top left: baseline scenario shown with threat and management scenarios found to markedly impact upon population growth, run for 500 iterations over a 50-year period and showing the mean number of extant individuals for each. Top right: a retrospective analysis using parameters from the baseline scenario with estimates of 40 individuals in 2010 to predict the current population in 2021. Bottom right: a retrospective analysis using parameters from the baseline scenario with estimates of 10–14 individuals in 1990 to predict the current population in 2021. Bottom left: The Raffles' banded langur. Photos: Sabrina Jabbar.
FIG. 1 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations
FIG. 1.—Map of the Hawaiian Islands with collection sitesfor Hawaiian hoary bat tissues used inthis study. Sites with n> 1 are denoted with an asterisk.
FIG. 2 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations
FIG. 2.—PCA result plot showing clustering of individual bats from four Hawaiian Islands using 21,808,031 SNPs. Sample information included in supplementary table S4, Supplementary Material online.
FIG. 4 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations
FIG. 4.—SNAPP-based phylogenetic tree inference. (A) The maximum clade credibility or consensus tree, showing approximate divergence of hoary bats across the Hawaiian archipelago. The axis on the bottom of the figure corresponds to million years before present (Ma), using the emergence of Hawai'i (~0.43 Ma) as a calibration point (95% confidence intervals were given in square brackets). (B) The drawing of all sampled trees showing all ingroup nodes were supported by maximum posterior probabilities (1.00).
Global analysis of emperor penguin populations
<p>Like many polar animals, emperor penguin populations are challenging to monitor because of the species' life history and remoteness. Consequently, it has been difficult to establish its global status, a subject important to resolve as polar environments change. To advance our understanding of emperor penguins, we combined remote sensing, validation surveys, and using Bayesian modeling we estimated a comprehensive population trajectory over a recent 10-year period, encompassing the entirety of the species' range. Reported as indices of abundance, our study indicates with 81% probability that the global population of adult emperor penguins declined between 2009 and 2018, with a posterior median decrease of 9.6% (95% credible interval (CI) -26.4% to +9.4%). The global population trend was -1.3% per year over this period (95% CI = -3.3% to +1.0%) and declines likely occurred in four of eight fast ice regions, irrespective of habitat conditions. Thus far, explanations have yet to be identified regarding trends, especially as we observed an apparent population up-tick toward the end of time series. Our work potentially establishes a framework for monitoring other Antarctic coastal species detectable by satellite, while promoting a need for research to better understand factors driving biotic changes in the Southern Ocean ecosystem.</p>
Data and code for: Nonlinear life table response analysis: Decomposing nonlinear and nonadditive population growth responses to changes in environmental drivers
<p>Life table response experiments (LTREs) decompose differences in population growth rate between environments into separate contributions from each underlying demographic rate. However, most LTRE analyses make the unrealistic assumption that the relationships between demographic rates and environmental drivers are linear and independent, which may result in diminished accuracy when these assumptions are violated. In this study, we compare the relative efficacy of linear and second-order LTRE analyses in capturing changes in population growth rate caused by environmental driver changes. To explore this question, we analyze demographic data collected for three long-lived plant species: <em>Ardisia escallonioides</em> (Pascarella & Horvitz, 1998), <em>Silene acaulis</em>, and <em>Bistorta vivipara</em> (Doak & Morris, 2010). This repository includes data files containing vital rate (survival, growth, reproduction) observations or models for our three case studies, as well as an R script in which we use these demographic data to calculate linear and second-order LTRE approximations of changes in population growth rate for each system and generate the figures we present in our paper.</p>
REPIN population analysis in 49 Stenotrophomonas maltophilia genomes
<p>This dataset contains the genome sequences of 49 S. maltophilia strains (input.zip) and the processed data from four individual RAREFAN runs. RAREFAN identifies RAYTs REPINS for each supplied genome and a reference genome. Four different strains where used as a reference genome: AA1, AB550 FDARGOOS_649, ISMMS3, and Sm53. The data were processed with default RAREFAN job parameters. Links to the original RAREFAN jobs are given below.</p>
Whole-genome analysis of multiple wood ant population pairs supports similar speciation histories, but different degrees of gene flow, across their European ranges
<p>The application of demographic history modelling and inference to the study of divergence between species has become a cornerstone of speciation genomics. Speciation histories are usually reconstructed by analysing single populations from each species, assuming that the inferred population history represents the actual speciation history. However, this assumption may not be met when species diverge with gene flow, e.g., when secondary contact may be confined to specific geographic regions. Here, we tested whether divergence histories inferred from heterospecific populations may vary depending on their geographic locations, using the two wood ant species <em>Formica polyctena</em> and <em>F. aquilonia</em>. We performed whole-genome resequencing of 20 individuals sampled in multiple locations across the European ranges of both species. Then, we reconstructed the histories of distinct heterospecific population pairs using a coalescent-based approach. Our analyses always supported a scenario of divergence with gene flow, suggesting that divergence started in the Pleistocene (ca. 500 kya) and occurred with continuous asymmetrical gene flow from <em>F. aquilonia</em> to <em>F. polyctena</em> until a recent time, when migration became negligible (2-19 kya). However, we found support for contemporary gene flow in a sympatric pair from Finland, where the species hybridise, but no signature of recent bidirectional gene flow elsewhere. Overall, our results suggest that divergence histories reconstructed from a few individuals may be applicable at the species level. Nonetheless, the geographical context of populations chosen to represent their species should be taken into account, as it may affect estimates of migration rates between species when gene flow is spatially heterogeneous.</p>
Joint analysis of microsatellites and flanking sequences enlightens complex demographic history of interspecific gene flow and vicariance in rear-edge oak populations
<p><span>Inference of recent population divergence requires fast evolving markers and necessitates to differentiate shared genetic variation caused by ancestral polymorphism and gene flow. Theoretical research shows that the use of compound marker systems integrating linked polymorphisms with different mutational dynamics, such as a microsatellite and its flanking sequences, can improve estimation of population structure and inference of demographic history, especially in the case of complex population dynamics. However, empirical application in natural populations has so far been limited by lack of suitable methods for data collection. A solution comes from the development of sequence-based microsatellite genotyping which we used to study molecular variation at 36 sequenced nuclear microsatellites in seven <em>Quercus canariensis</em> and four <em>Q. faginea</em> rear-edge populations across Algeria. We aim to decipher their taxonomic relationship, past evolutionary history and recent demographic trajectory. First, we compare the estimation of population genetics parameters and simulation-based inference of demographic history from microsatellite sequence alone, flanking sequence alone or the combination of linked microsatellite and flanking sequence variation. Second, we apply random forest approximate Bayesian computation to identify which of these sequence types is most informative. Whereas analysing microsatellite variation alone indicates recent interspecific gene flow, additional information gained by integrating nucleotide variation in flanking sequences, by reducing homoplasy, suggests ancient interspecific gene flow followed by drift in isolation instead. The weight of each polymorphism in the inference also demonstrates the value of linked variations with contrasted mutation dynamic to improve estimation of both demographic and mutational parameters.</span></p>
DATASET: Environmental analysis of servicing centralised and decentralised wastewater treatment for population living in neighbourhoods
<p>DATASET:</p> <p>Article: Environmental analysis of servicing centralised and decentralised wastewater treatment for population living in neighbourhoods</p> <p>Journal of Water Process Engineering, Volume 37, October 2020, 101469</p> <p>https://doi.org/10.1016/j.jwpe.2020.101469</p>
Raw RADseq data for: Population genomics analysis with RAD, reprised: Stacks 2
<p>Restriction enzymes have been one of the primary tools in the population genetics toolkit for 50 years, being coupled with each new generation of technology to provide a more detailed view into the genetics of natural populations. Restriction site-Associated DNA protocols, which joined enzymes with short-read sequencing technology, have democratized the field of population genomics, providing a means to assay the underlying alleles in scores of populations. More than 10 years on, the technique has been widely applied across the tree of life and served as the basis for many different analysis techniques. Here, we provide a detailed protocol to conduct a RAD analysis from experimental design to de novo analysis—including parameter optimization—as well as reference-based analysis, all in Stacks version 2, which is designed to work with paired-end reads to assemble RAD loci up to 1000 nucleotides in length. The protocol focuses on major points of friction in the molecular approaches and downstream analysis, with special attention given to validating experimental analyses. Finally, the protocol provides several points of departure for further analysis.</p>
Data and analysis from: Body mass, temperature, and depth shape the maximum intrinsic rate of population increase in sharks and rays
<p>An important challenge in ecology is to understand variation in species' maximum intrinsic rate of population increase, 𝑟<sub>𝑚𝑎𝑥</sub>, not least because 𝑟<sub>𝑚𝑎𝑥</sub> underpins our understanding of the limits of fishing, recovery potential, and ultimately extinction risk. Across many vertebrate species, terrestrial and aquatic, body mass and environmental temperature are important correlates of 𝑟<sub>𝑚𝑎𝑥</sub>. In sharks and rays, specifically, 𝑟<sub>𝑚𝑎𝑥</sub> is known be lower in larger species, but also in deep-sea ones.</p> <p>We use an information-theoretic approach that accounts for phylogenetic relatedness to evaluate the relative importance of body mass, temperature and depth on 𝑟<sub>𝑚𝑎𝑥</sub>. We show that both temperature and depth have separate effects on shark and ray 𝑟<sub>𝑚𝑎𝑥</sub> estimates, such that species living in deeper waters have lower 𝑟<sub>𝑚𝑎𝑥</sub>. Furthermore, temperature also correlates with changes in the mass scaling coefficient, suggesting that as body size increases, decreases in 𝑟<sub>𝑚𝑎𝑥</sub> are much steeper for species in warmer waters.</p> <p>These findings suggest that there are (as-yet understood) depth-related processes that limit the maximum rate at which populations can grow in deep sea sharks and rays. While the deep ocean is associated with colder temperatures, other factors that are independent of temperature, such as food availability and physiological constraints, may influence the low 𝑟<sub>𝑚𝑎𝑥</sub> observed in deep sea sharks and rays. Our study lays the foundation for predicting the intrinsic limit of fishing, recovery potential, and extinction risk species based on easily accessible environmental information such as temperature and depth, particularly for data-poor species.</p> <p>This repository contains the data and a minimum working example of the model-fitting process used for the article "Body mass, temperature, and depth shape productivity in sharks and rays", which is currently in press at <em>Ecology and Evolution</em>.</p>
Construction of a SNP fingerprinting database and population genetic analysis of 329 cauliflower cultivars
<p>The VCF file contains the information of 1662 SNP sites of 820 cauliflower inbred lines that were filtered according to a series of stringent conditions.</p>
Data: Applying stochastic and Bayesian integral projection modeling to amphibian population viability analysis
<p>Integral projection models (IPMs) can estimate the population dynamics of species for which both discrete life stages and continuous variables influence demographic rates. Stochastic IPMs for imperiled species, in turn, can facilitate population viability analyses (PVAs) to guide conservation decision-making. Biphasic amphibians are globally distributed, often highly imperiled, and ecologically well-suited to the IPM approach. Herein, we present the first stochastic size- and stage-structured IPM for a biphasic amphibian, the U.S. federally threatened California tiger salamander (<em>Ambystoma</em> <em>californiense</em>; CTS). This Bayesian model reveals that CTS population dynamics show the greatest elasticity to changes in juvenile and metamorph growth and that populations are likely to experience rapid growth at low density. We integrated this IPM with climatic drivers of CTS demography to develop a PVA and examined CTS extinction risk under the primary threats of habitat loss and climate change. The PVA indicates that long-term viability is possible with surprisingly high (20–50%) terrestrial mortality, but simultaneously identified likely minimum terrestrial buffer requirements of 600–1000 m while accounting for numerous parameter uncertainties through the Bayesian framework. These analyses underscore the value of stochastic and Bayesian IPMs for understanding both climate-dependent taxa and those with cryptic life histories (e.g., biphasic amphibians) in service of ecological discovery and biodiversity conservation. In addition to providing guidance for CTS recovery, the contributed IPM and PVA supply a framework for applying these tools to investigations of ecologically-similar species.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.