Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
19,487
datasets available to search
ShareScore release 0.7.1
Dataset results
19,487 results for “Populations”
Variant, Metabolite and Source Data for: Population genomics uncover loci for trait improvement in the indigenous African cereal tef (Eragrostis tef)
<p>These files contain the variant and metabolome for a collection of 220 tef (<em>Eragrsotis tef)</em> accessions from an ethiopian diversity panel. The accessions were assembled and managed by the Ethiopian Institute of Agricultural Research (EIAR, Ethiopia). The variant data was produced at the John Innes Centre (UK). The metabolome data was produced at Aberystwyth University (UK). These dataset are described in Jones et al. (2024), <em>bioRxiv</em>, https://doi.org/10.1101/2024.09.30.615331. The source data for main figures in the publication are also included.</p> <p>The submission contains</p> <ol> <li>EIAR_filtered.vcf.gz: This is the variant data obtained from alignment of Illumina reads from all 220 teff accessions to the reference assembly of tef (Dabbi). Low quality variants were filtered out. This variant data was used for constructing the phylogenetic relationship between the accessions. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li>pooled_EIAR_filtered.vcf.gz: After the phylogentic analysis described above, reads from accessions that were found to be genetically redundant were pooled before variant calling. This file was used for the SNP GWAS analysis. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li> Metabolite_Profile.xlxs (source data for Figure 5): This file contains m/z feature intensities from untargeted metabolite fingerprinting using Flow Infusion Electrospray High-resolution Mass Spectrometry (FIE-HRMS). The sample names contains a combination of Location code and Plot number in Supplementary Table S10 e.g AT plot 1, CD plot 1, DZ plot 1, where AT, CD and DZ represent Alem Tena, Chefe Donsa and Debre Zeit, respectively. The data was used for the partial least squares discriminant analysis and differentially accumulated metabolites analysis presented in Figure 5.</li> <li>Source data: Numerical source data for graphs and charts in Figures 3 - 7.</li> <li>Tsedey TT2 Sequence from Improved Assembly: The 4A and 4B sequences around the TT2 orthologue in tef from the improved PacBio-based chromosome-scale assembly of tef. These sequences were used for plotting the LTR Copia alignments presented in Supplementary Figure 9. We thank Corteva for pre-publication access to this improved Tsedey genome assembly.</li> </ol>
Synthesized anthropometric data for the German working-age population
<p>The anthropometric datasets presented here are virtual datasets. The unweighted virtual dataset was generated using a synthesis and subsequent validation algorithm (Ackermann et al., 2023). The underlying original dataset used in the algorithm was collected within a regional epidemiological public health study in northeastern Germany (SHIP, see Völzke et al., 2022). Important details regarding the collection of the anthropometric dataset within SHIP (e.g. sampling strategy, measurement methodology & quality assurance process) are discussed extensively in the study by Bonin et al. (2022).</p><p>To approximate nationally representative values for the German working-age population, the virtual dataset was weighted with reference data from the first survey wave of the Study on health of adults in Germany (DEGS1, see Scheidt-Nave et al., 2012). Two different algorithms were used for the weighting procedure: (1) iterative proportional fitting (IPF), which is described in more detail in the publication by Bonin et al. (2022), and (2) a nearest neighbor approach (1NN), which is presented in the study by Kumar and Parkinson (2018). Weighting coefficients were calculated for both algorithms and it is left to the practitioner which coefficients are used in practice. Therefore, the weighted virtual dataset has two additional columns containing the calculated weighting coefficients with IPF ("WeightCoef_IPF") or 1NN ("WeightCoef_1NN"). Unfortunately, due to the sparse data basis at the distribution edges of SHIP compared to DEGS1, values underneath the 5th and above the 95th percentile should be considered with caution.</p><p>In addition, the following characteristics describe the weighted and unweighted virtual datasets: According to ISO 15535, values for "BMI" are in [kg/m2], values for "Body mass" are in [kg], and values for all other measures are in [mm]. Anthropometric measures correspond to measures defined in ISO 7250-1. Offset values were calculated for seven anthropometric measures because there were systematic differences in the measurement methodology between SHIP and ISO 7250-1 regarding the definition of two bony landmarks: the acromion and the olecranon. Since these seven measures rely on one of these bony landmarks, and it was not possible to modify the SHIP methodology regarding landmark definitions, offsets had to be calculated to obtain ISO-compliant values. In the presented datasets, two columns exist for these seven measures. One column contains the measured values with the landmarking definitions from SHIP, and the other column (marked with the suffix "_offs") contains the calculated ISO-compliant values (for more information concerning the offset values see Bonin et al., 2022). The sample size is N = 5000 for the male and female subsets. The original SHIP dataset has a sample size of N = 1152 (women) and N = 1161 (men). Due to this discrepancy between the original SHIP dataset and the virtual datasets, users may get a false sense of comfort when using the virtual data, which should be mentioned at this point. In order to get the best possible representation of the original dataset, a virtual sample size of N = 5000 is advantageous and has been confirmed in pre-tests with varying sample sizes, but it must be kept in mind that the statistical properties of the virtual data are based on an original dataset with a much smaller sample size.</p>
Selected properties of galaxy and SMBH populations (Spinoso et al. 2023)
<p>This record presents the catalogs of galaxy and Black Holes properties associated to the two runs of the modified version of the L-Galaxies Semi-Analytic Model (SAM) presented in Spinoso et al. 2023. These catalogs are aimed at providing the basic properties to study the population of Black Holes (BHs) and their host galaxies across cosmic times, obtained by running the L-Galaxies SAM over the whole Millennium-II box (see Boylan-Kolchin et al. 2009). The L-Galaxies SAM outputs summarized in these catalogs were obtained at several redshifts/snapshots, for two different runs which differ for the initial occupation fraction of BHs at the time of their formation. This initial occupation fraction is parametrized by the Gp parameter(see Spinoso et al. 2023 for details), with the two runs being characterized by Gp=1 and Gp=0.01. The catalogs are organized in two group of files, one group for each run. Each of these groups is composed by 18 different files, one per each availablle redshift, roughly corresponding to: z = 0, 0.5, 1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15. The two group of files can be easily distinguished by their names: indeed, the strings "Gp1" and "Gp001" are referred to the runs corresponding to the Gp=1 and Gp=0.01 values, respectively. In addition, every file includes a string of the form: "z[x.yz]" which specifies its redshift. </p> <p>The content of these catalogs is as follows: each redshift-file contains the same collection of arrays, each array being a galaxy or BH property. At each redshift, L-Galaxies outputs properties for the number 'NGAL' of galaxies identified in the Millennium-II box, at that specific redshift/snapshot. Therefore, most of the arrays have length equal to 'NGAL' (i.e. one value per each galaxy). Few of the arrays have a length of N * NGAL (i.e. N values per each galaxy). The content, units and data type of these arrays are as follows: </p> <ul> <li>"StellarMass" - Total stellar mass of each galaxy - [10^10 Msun / h] - array[NGAL]</li> <li>"Sfr" - Star formation rate of each galaxy - [Msun / yr] - array[NGAL]</li> <li>"SeedMass" - BH-seed mass. 7 values per galaxy; one value for each of the 7 possible BH-seed "flavors" modeled - [10^10 Msun / h] - array[NGAL, 7]</li> <li>"Rvir" - Virial radius of the DM halo hosting each galaxy - [Mpc / h] - array[NGAL]</li> <li>"Pos" - X, Y and Z position of each galaxy - [Mpc / h] - array[NGAL, 3]</li> <li>"Mvir" - Virial mass of the DM halo hosting each galaxy - [10^10 Msun / h] - array[NGAL]</li> <li>"Lbol" - Bolometric luminosity associated to the central AGN (==0 if the BH is not active) - [10^40 erg / s] - array[NGAL]</li> <li>"HotGas" - Mass of the hot-phase of each galaxy's gas component - [10^10 Msun / h] - array[NGAL]</li> <li>"fEDD" - Eddington ration (defined as Lbol/L_Edd, with L_Edd being the Eddington luminosity) for each AGN (==0 if the BH is not active) - [adim] - array[NGAL]</li> <li>"ColdGas" - Mass of the cold-phase of each galaxy's gas component - [10^10 Msun / h] - array[NGAL]</li> <li>"BlackHoleMass" - Mass of the central massive BH hosted by each galaxy (==0 if the galaxy does not host a central BH) - [10^10 Msun / h] - array[NGAL]</li> <li>"SeedType" - Identifier of the type of BH-seed which originated each BH (see below for details) - array[NGAL]</li> <li>"Redshift" - Redshift of each galaxy (within a single file, this is an array of identical values) - array[NGAL]</li> </ul> <p>NOTE:<br>The model presented in Spinoso et al. 2023 follows 7 different types of BH-seeds. The "SeedMass" array contains 7 mass values (one for each of these types of BH-seeds) for each galaxy in the Millennium-II box.This is the reason why the data type of "SeedMass" is [NGAL, 7]. Each of these 7 values is the sum, across the whole evolution of each galaxy, of the contributions to the total BH mass coming from each BH-seed who merged to form the final BH. In the vast majority of cases, BHs are associated to only one type of BH-seed. In those cases, 6 out of the 7 "SeedMass" values would be zero. Each element of "SeedMass" corresponds to one type of BH seed according the following scheme:<br>SeedMass[0] : total seed mass of light-seeds inherited from the GQd model (see Spinoso et al. 2023 for details)<br>SeedMass[1] : total seed mass of heavy-seeds inherited from the GQd model (see Spinoso et al. 2023 for details)<br>SeedMass[2] : un-resolved mass-growth driven by gas-accretion before the halo hosting the BH was resolved<br>SeedMass[3] : total seed mass formed as light-seeds in L-Galaxies<br>SeedMass[4] : total seed mass formed as Direct-Collapse BHs (DCBHs)<br>SeedMass[5] : total seed mass formed as intermediate-mass BH originated via Runaway Stellar Mergers (RSM)<br>SeedMass[6] : total seed mass formed as Merger-Induced Direct-Collapse BH (miDCBH)</p> <p>NOTE:<br>Similarly to "SeedMass", also the "Pos" array has more than one element per galaxy. These are the three cartesian positions of each galaxy.</p> <p>NOTE:<br>The possible values of the "SeedType" array are as follows (see Spinoso et al. 2023 for details):<br>-1 - No BH seed (the galaxy never hosted a BH)<br>1 - light seed (PopIII remnant)<br>6 - Direct-Collapse BH (DCBH)<br>7 - intermediate-mass BH originated via Runaway Stellar Mergers (RSM)<br>8 - Merger-Induced Direct-Collapse BH (miDCBH)<br>9 - mixed type: light+DCBH (the BH is the result of hierarchical mergers between light and DCBH seeds)<br>10 - mixed type: light+RSM (the BH is the result of hierarchical mergers between light and RSM seeds)</p>
Opinions and Views of the Population of Ukraine: May 2024 (KIIS Omnibus 2024/05) – Data from a nationwide public opinion poll conducted by KIIS in May 2024
"Opinions and Views of the Population of Ukraine" is a regular omnibus survey, conducted by Kyiv International Institute of Sociology (KIIS) among Ukraine's adult population and covering a wide range of topics. The data presented here is a subset of the survey conducted in May 2024 and include KIIS's own research questions. Questions included are: readiness for concessions for peace, views on Ukraine's relationship with Russia, perceptions of the war between Russia and Ukraine, views on security agreements, perceptions of Ukrainian society's unity, attitudes toward criticism of the government, attitudes toward the legalization of medical cannabis, and perceptions of Ukraine's statehood during the Soviet era. Data collection took place from May 16 to 22, 2024, with 1,067 respondents interviewed. The data is available in an SAV format (Ukrainian, English) and a converted CSV format (with a codebook). The Data Documentation (pdf file) also includes a short overview and discussion of survey results as well as the relevant parts of the original questionnaire.
Opinions and Views of the Population of Ukraine: February 2024 (KIIS Omnibus 2024/02) – Data from a nationwide public opinion poll conducted by KIIS in February 2024
"Opinions and Views of the Population of Ukraine" is a regular omnibus survey, conducted by Kyiv International Institute of Sociology (KIIS) among Ukraine's adult population and covering a wide range of topics. The data presented here is a subset of the survey conducted in February 2024 and include KIIS's own research questions. The questions cover the following topics: readiness for concessions for peace; perceptions of Russia, its people, and leadership; sources of information; perceptions of the war between Russia and Ukraine; views on Western support for Ukraine; factors contributing to Ukraine's success in the war; perceptions of recent investigations into large businesses and businessmen in Ukraine; state control over online information; state policy on the Russian language in Ukraine; the level of democracy in Ukraine; opportunities for personal success; and favorite national holidays. Data collection took place from February 17 to 28, 2024. Some of the survey questions were asked to all respondents (n=2,008), while others were directed to a sub-sample of 1,052 respondents. The data is available in an SAV format (Ukrainian, English) and a converted CSV format (with a codebook). The Data Documentation (pdf file) also includes a short overview and discussion of survey results as well as the relevant parts of the original questionnaire.
The effect of dynamical states on galaxy clusters populations. I. Classification of dynamical states
<p>This repository contains three figures mentioned in "The effect of dynamical states on galaxy clusters populations. I. Classification of dynamical states" <em>(DOI to follow on publication)</em>.</p> <p>We show the contours of the X-ray surface brightness distribution (solid green lines) and the distribution of galaxies belonging to the red sequence (solid gray lines). Black crosses symbolize the positions of the X-ray peaks, black "X" marks represent the positions of the X-ray centroids, and open red circles denote the positions of the BCGs. The blue circle corresponds to the R200 of each cluster.</p>
Modelled gridded population estimates for the Kasaï-Oriental Province in the Democratic Republic of Congo (2024) version 4.2
<h2><strong>Content</strong></h2> <p>This repository contains the input data and scripts used to create the modeled gridded population estimates for Kasaï-Oriental Province in the Democratic Republic of Congo. It also includes the grid-cell posterior distributions and scripts to aggregate them within user-defined geographic boundaries.</p> <p> In particular, this repository contains two compressed files (.zip):</p> <p><strong>1. <code>population_estimates.zip</code></strong></p> <ul> <li>Includes raster files (<code>.tif</code>) with summaries of population count posterior predictions at the grid-cell level, specifically the mean, median, lower credible interval, and upper credible interval.</li> <li>Includes spatial files (<code>.gpkg</code>) with summaries of population count posterior predictions at the health-area and health-zone levels, specifically the mean, median, lower credible interval, and upper credible interval.</li> </ul> <p><strong>2. <code>population_model.zip</code></strong></p> <p>This directory comprises five subdirectories with scripts, input data, and output data necessary to replicate the population model:</p> <ul> <li><code><strong>01_model_stan</strong></code>: Contains the Stan model, input data, and an R script (<code>01_model_stan.R</code>) with a function to run the model.</li> <li><code><strong>02_model_run</strong></code>: Includes an R script (<code>02_model_run.R</code>) for running the model, along with output data.</li> <li><code><strong>03_model_evaluate</strong></code>: Features a Quarto report template (<code>03_model_evaluate.qmd</code>) and model evaluation summary files(.pdf).</li> <li><code><strong>04_predict_posterior</strong></code>: Provides R scripts (<code>04_predict_posterior.R</code> and <code>04_predict_run.R</code>) for generating predictions, along with input and output data, namely the posterior predictions files (.rds).</li> <li><code><strong>05_aggregate_posterior</strong></code>: Contains R scripts (<code>05_aggregate_posterior.R</code> and <code>05_aggregate_run.R</code>) and associated input and output data, namely the population count posterior summaries as presented in the file <code>population_estimates.zip</code> .</li> </ul> <p>The work was carried out in <code>R</code> (version 4.4.0), with the packages <code>tidyverse</code> (version 2.0.0), <code>terra</code> (version 1.7-78), <code>sf</code> (version 1.0-16), <code>furrr</code> (version 0.3.1), <code>doParallel</code> (version 1.0.17), <code>foreach</code> (version 1.5.2), <code>rstudioapi</code> (version 0.16.0), and <code>rstan</code> (version 2.32.6), on macOS Sequoia (version 15.1.1). While the scripts are designed to be portable, minor adjustments may be required for compatibility with other operating systems.</p> <h2><strong>Important</strong></h2> <p>This version includes changes in the STAN model <code>10h_survey_survey_covariate_building_random_effect_hierarchy_building_covariate_density_fixed_effect_hierarchy_density.stan</code>. Consequentely, all the files generated in the previous versions are now changed.</p> <p> </p> <p>For inquiries regarding the model and the data, please contact Gianluca Boo at gianluca.boo@soton.ac.uk.</p>
The genetic population structure of Lake Tanganyika's Lates species flock, an endemic radiation of pelagic top predators
<p>Data associated with the manuscript "The genetic population structure of Lake Tanganyika’s Lates species flock, an endemic radiation of pelagic top predators," where we investigate the genetic population structure of the four endemic <em>Lates </em>species in Lake Tanganyika.</p> <p><strong>Abstract</strong>: Life history traits are important in shaping gene flow within species and can thus determine whether a species exhibits genetic homogeneity or population structure across its range. Understanding genetic connectivity plays a crucial role in species conservation decisions, and genetic connectivity is an important component of modern fisheries management in fishes exploited for human consumption. In this study, we investigated the population genetics of four endemic <em>Lates</em> species of Lake Tanganyika (<em>Lates stappersii</em>, <em>L. microlepis</em>, <em>L. mariae</em> and <em>L. angustifrons</em>), using reduced-representation genomic sequencing methods. We find the four species to be strongly differentiated from one another, with no evidence for contemporary admixture. We also find evidence for high levels of genetic structure within <em>L. mariae</em>, with the majority of individuals from the most southern sampling site forming a genetic group distinct from the individuals at other sampling sites<em>.</em> We find evidence for much weaker structure within the other three species, <em>L. stappersii,</em> <em>L. microlepis</em>, and <em>L. angustifrons</em>, although small and unbalanced sample sizes and imprecise geographic sampling locations may hinder our ability to detect weak population structure. We call for further research into the origins of the genetic differentiation that we observe in these four species, particularly that of <em>L. mariae</em>, which may be important for the conservation and management of this species.</p> <p>Code associated with the analysis of these data can be found on GitHub at <a href="https://github.com/jessicarick/lates-popgen">https://github.com/jessicarick/lates-popgen</a>.</p>
Night Population map
<p>Night Population map is a layer in support to Area Of Interest definition for an earthquake event. It provides an estimation of a population exposed to earthquake event in the examined area, based on the population occupancy during night hours.</p>
Orthophoto mosaic from UAV of a Pinna nobilis population within the Venice lagoon
<p>Orthophoto mosaic from a UAV survey on a tidal flat within the Venice lagoon (Italy) colonized by Pinna nobilis and Cymodocea nodosa. On June 23<sup>rd</sup>, 2020 at 06:05 a.m. (GMT) (07:05 solar local time) UAV images were collected at low tide with a DJI Zenmuse X4S camera (20 Mpixels, focal length 8.8 mm, 1-inch CMOS Sensor) mounted on a professional quadcopter DJI Matrice210v2. A total of 228 images were collected and the Agisoft Metashape Pro v 1.6.2 software was then used to produce an ortho-photo mosaic through Structure from Motion photogrammetric technique.</p>
diFUME Population Density V0.1
<p>Description:</p> <p>Annual statistics per city block on residential population (by age group) and workplace employees (<a href="https://www.basleratlas.ch/">https://www.basleratlas.ch/</a> ) are used to derive maps of annual night-time and daytime building-scale population density (inhabitants per m2) for weekdays and weekends. The spatial resampling of the population is based on the assumption of proportionality between building inhabitants and building volume (estimated as mean building height×building plan area, derived by land cover and DSM products). Considering the building type, building volume is separated to residential volume and workplace volume, so that population is redistributed between night-time, daytime, workdays and weekends.</p> <p> </p> <p>Data specifications:</p> <p>CRS: EPSG:32632 - WGS 84 / UTM zone 32N - Projected</p> <p>Spatial Extent: 392120.0,5266860.0 : 395160.0,5269840.0</p> <p>Temporal Extent: 2018 - 2020</p> <p>Units: meters</p> <p>Width: 608</p> <p>Height: 596</p> <p>Bands: 1</p> <p>Pixel Size: 5,-5</p> <p>Data type: Float32 - Thirty two bit floating point</p> <p>GDAL Driver Description: GTiff</p> <p>GDAL Driver Metadata: GeoTIFF</p>
The pan-genome of Aspergillus fumigatus provides a high-resolution view of its population structure revealing high-levels of lineage-specific diversity driven by recombination
<p><em>Aspergillus fumigatus </em>is a deadly agent of human fungal disease, where virulence heterogeneity is thought to be at least partially structured by genetic variation between strains. While population genomic analyses based on reference genome alignments offer valuable insights into how gene variants are distributed across populations, these approaches fail to capture intraspecific variation in genes absent from the reference genome. Pan-genomic analyses based on <em>de novo</em> assemblies offer a promising alternative to reference-based genomics, with the potential to address the full genetic repertoire of a species. Here, we use a combination of population genomics, phylogenomics, and pan-genomics to assess population structure and recombination frequency, phylogenetically structured gene presence-absence variation, evidence for metabolic specificity, and the distribution of putative antifungal resistance genes in <em>A. fumigatus</em>. We provide evidence for three distinct populations of <em>A. fumigatus</em>, structured by both gene variation (SNPs and indels) and distinct gene presence-absence variation with unique suites of accessory genes present exclusively in each clade. Accessory genes displayed functional enrichment for nitrogen and carbohydrate metabolism, hinting that populations may be stratified by environmental niche specialization. Similarly, the distribution of antifungal resistance genes and resistance alleles were often structured by phylogeny. Despite low levels of outcrossing, <em>A. fumigatus</em> demonstrated a large pan-genome including many genes unrepresented in the Af293 reference genome. These results highlight the inadequacy of relying on a single-reference based approach for evaluating intraspecific variation, and the power of combined genomic approaches to elucidate population structure, genetic diversity, and the putative ecological drivers of clinically relevant fungi.</p> <p>Accompanying manuscript is available as preprint at <a href="https://dx.doi.org/10.1101/2021.12.12.472145">https://dx.doi.org/10.1101/2021.12.12.472145</a> </p> <p>Lotus A. Lofgren, Brandon S. Ross, Robert A. Cramer, Jason E. Stajich. Combined Pan-, Population-, and Phylo-Genomic Analysis of <em>Aspergillus fumigatus</em> Reveals Population Structure and Lineage-Specific Diversity bioRxiv 2021.12.12.472145; doi: https://doi.org/10.1101/2021.12.12.472145</p>
X-rays across the galaxy population: The distribution of AGN accretion rates as a function of stellar mass and redshift
<p>We provide measurements of the probability distribution function of specific black hole accretion rates within a sample of galaxies of a given stellar mass and redshift, <span class="math-tex">\(p(\log \lambda_{sBHAR} | M_*,z)\)</span>. Measurements are provided for all galaxies, star-forming galaxies and quiescent galaxies. We also provide estimates of the AGN duty cycle, <span class="math-tex">\(f(\lambda_{sBHAR} >0.01)\)</span> i.e. the fraction of galaxies with an AGN above a given limit in specific accretion rate, based on the probability distribution functions. Full details are provided in Aird et al. (2018, MNRAS, 474, 1225); please cite this publication if you use these measurements. </p>
The population of merging compact binaries inferred using gravitational waves through GWTC-3 - Data release
<p>Data associated with Figures, Tables, and population parameter samples associated with <br><strong>The population of merging compact binaries inferred using gravitational waves through GWTC-3 , </strong><br><strong><a href="https://dcc.ligo.org/LIGO-P2100239/public">LIGO DCC</a>, <a href="https://arxiv.org/abs/2111.03634">arXiv</a>, <a href="https://journals.aps.org/prx/abstract/10.1103/PhysRevX.13.011048">PRX</a>. </strong><br>This is v3, superseding v2. Please see the README.md for more information.</p>
E4warning_2020_Population_Age_Sex
<p>Worldpop Human 2020 population by Age and Gender. </p> <p>Abstract: Human population estimates per pixel were extracted from MOOD partner Worldpop (www.worldpop.org) datasets for the MOOD extent. Gender age categories were summed to provide datasets for all males and all females as well as total populations. Filenames are are follows (MOWPGGGRRYY-OOCog.TIF where GGG =-gender (male = MAL, female = FEM), both = TOT); RR = Greater than (gt) or Less than (lt); YY = minimum age; OO= Maximum age</p>
Manipulating a host-native microbial strain compensates for low microbial diversity by increasing weight gain in a wild bird population
<h1>Manipulating a host-native microbial strain compensates for low microbial diversity by increasing weight gain in a wild bird population</h1> <h1> </h1> <p>These files contain data on bacteria present in the guts of wild great tit (Parus major) obtained from faecal samples and sequenced using Illumina MiSeq. These data resulted from an experiment which provided supplementary mealworms at the nest during the breeding season at number of woodland sites in Cork, Ireland. Approximately half of these nests were given mealworms covered in a freeze dried bacterial powder containing the bacteria Lactobacillus kimchicus, which had been isolated from great tit faeces from the previous season. This treatment aimed to disrupt the gut microbiota of the treatment birds in order to provide evidence for the gut microbiotas role in birds health and fitness. Included here are the 3 elements necessary to create a 'phyloseq object' containing the sample metadata, ASV (Amplicon Sequence Variant) count table and a taxonomy table. The metadata file includes the alpha diversity scores for each individual. The data include all negative control samples taken during sample collection and library preparation, which were removed before the main analyses. All analyses, except for the beta-diversity analyses, were conducted in R. All R code is available on GitHub (https://github.com/shan-e-s\). Raw Sequence data are available in the European Nucleotide Archive under access number PRJEB74941, and ERS18960426-ERS18960697.</p> <h2> </h2> <h2>## Description of the data and file structure </h2> <p>Taxonomy, ASV and metadata files required to create a phyloseq object in R. metadata.csv file contains data on individual birds (i.e. individual samples). The metadata includes descriptions of the bird itself and it's environment, namely:</p> <ul> <li>Rownames: unique sample ID for each sample, corresponds with asvTable.csv. </li> <li>Nest: unique identifier for the nest box associated with the bird being sampled. </li> <li>Sample.ID: unique identifier for the faecal sample or control sample.</li> <li>Bird.ID: Identity of the bird the sample came from, note some individuals sampled twice so some bird.ID's may reoccur in metadata with different Sample.ID.</li> <li>Date: Date the sample was taken dd/mm/yyyy.</li> <li>Day: Date the sample was taken, in days since 1st March.</li> <li>Ring.Mark: British Trust for Ornithology (BTO) metal ring ID where applicable. Birds only ringed at D15 so some young birds do not have IDRings.</li> <li>Site: ID of woodland site that bird was sampled at.</li> <li>Chick.LetterID: ID letter differentiates between different birds from the same nest. Either 'A'-'F' for nestlings, 'Fe' for females or 'M' for males.</li> <li>Age.code: BTO age code.</li> <li>Age.category: Age category that bird is in. D8 = 8 days post hatching, D15 = 15 days post hatching, adult = 1+ years post hatching.</li> <li>Sex: Bird's sex, only determined for adult birds. Fe = Female, M = Male.</li> <li>Wing_mm: Wing length in mm.</li> <li>Tarsus_mm: minimum tarsus length of bird in mm.</li> <li>Weight_g: bird's weight in grams.</li> <li>Faecal.Sample: bird's age at sampling.</li> <li>newRing: whether bird was fitted with a new BTO ring. Only relevant to adults.</li> <li>Treatment: the experimental treatment group that the bird was in. Either 'Treatment' when nest given L. kimchicus treated mealworms or 'Control' when nest given plain mealworms.</li> <li>Notes: field notes.</li> <li>Main.sample: indicates whether this sample was the main sample to be used for analysis, an alternative sample taken as a backup.</li> <li>Plate: the ID of the PCR plate which the sample was amplified on.</li> <li>Azenta_noPeriod: sample ID given to sequencing facility without special characters. Corresponds to fastq files and ASV table counts.</li> <li>Qubit_prePool: samples qubit score before pooling.</li> <li>Date_extracted: date the sample was extracted on dd/mm/yyyy.</li> <li>SampleType: whehther the sample was a 'main' sample intended for downstream analysis, a 'control' sample for detecting contamination during library preparation, a 'duplicate' for detecting PCR issues, a 'label_error' where sample was suspected of being mislabelled at some point, a 'repeat' sample intended to detect errors or issues, a 'contam' sample which was suspected of being contaminated, a 'common' sample used across different PCR plates to detect issues. Extraction_notes: notes regarding the DNA extraction of the sample. </li> <li>LibPrep_notes: notes regarding the library preparation of the sample.</li> <li>Ring.Mark.lab: the ring or sample ID written on the sample tube, recorded to help detect mislabelling.</li> <li>Post_lab_notes: notes regarding issues found post sequencing.</li> <li>NumberOfReads: number of sequence reads associated with the sample. </li> <li>DistanceToEdge: distance between nest and woodland edge in metres. </li> <li>BroodSize.D8: number of nestlings in the nest at day-8 post hatching. </li> <li>BroodSize.D15: number of nestlings in the nest at day-15 post hatching.</li> <li>firstEggLayDate: Date the first egg in the clutch was laid, in days since 1st March.</li> <li>lastEggLayDate: Date the last egg in the clutch was laid, in days since 1st March.</li> <li>Observed: number of unique ASV's (or taxa) detected in the sample.</li> <li>Chao1: Chao1 diversity of the sample.</li> <li>Shannon: Shannon diversity of the sample.</li> </ul> <p>The file 'taxonomy.csv' contains the taxonomic breakdown of each bacterial Amplicon Sequence Variant (ASV) found in the dataset from Phylum to Species. Obtained by using the Naive Bayes Classifier against the Silva (v138) taxonomic database.</p> <p>The file 'asvTable.csv' contains counts of each amplicon sequence variant's occurrence for each individual sample. Samples are rows and taxa are columns.</p> <p> </p> <h2>Sharing/Access information </h2> <p>All R code is available on GitHub (https://github.com/shan-e-s\). Raw Sequence data are available in the European Nucleotide Archive under access number PRJEB74941, and ERS18960426-ERS18960697.</p>
GLOBAL SNAPSHOT Physician Distribution and Density of Physicians per 1000 population - Worldwide 2021
<p>The chart presents the most up-to-date data (2021) available for 49 of the world’s 195 countries, focusing on the total number of physicians and the number of physicians per 1000 population(1). The countries are categorized into four income groups based on World Bank classifications, which are updated annually on July 1st each year(2).</p> <p>Only 25% of the countries present current data. This information is critical for decision-making for healthcare planning and policy development. Equally crucial, is for researchers to have comparable data to propose initiatives, to establish benchmarks and for crafting holistic strategies to gauge and advance progress in healthcare systems globally.</p> <p>Data sources: UnData <a href="https://data.un.org/">https://data.un.org/</a></p> <p>Visualization tools used: RAWGraphs <a href="https://www.rawgraphs.io/">https://www.rawgraphs.io/</a>, MS PowerPoint and Microsoft Excel</p> <p>Intended Audience: Academics and Researchers; Students and Educators; Healthcare Administrators and Policy Makers; Non-Governmental Organizations</p> <p>The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.</p> <p>The NNLM Data Visualization Challenge happens through work funded by the National Institutes of Health's National Library of Medicine, grant number U24LM013751</p> <p> </p> <p>References:</p> <p>1. United Nations, Department of Economic and Social Affairs. 10 Health Personnel. In: Statistical Yearbook. 66th issue (2023). New York: United Nations; 2023. (ST/ESA/STAT/SER.S/42). [Dataset available at UnData] <a href="https://data.un.org/_Docs/SYB/CSV/SYB66_154_202310_Health%20Personnel.csv">https://data.un.org/_Docs/SYB/CSV/SYB66_154_202310_Health%20Personnel.csv</a></p> <p>2 World Bank. World Bank Country and Lending Groups. World Bank Data Help Desk [Internet]. [cited 2024 Apr 5]. Available from:<a href="https://datahelpdesk.worldbank.org/knowledgebase/articles/906519-world-bank-country-and-lending-groups"> https://datahelpdesk.worldbank.org/knowledgebase/articles/906519-world-bank-country-and-lending-groups</a></p>
Data from Neutral genetic structuring of pathogen populations during rapid adaptation
<p><strong>Datasets and temporary dataframes relating to the article "Neutral genetic structuring of pathogen populations during rapid adaptation".</strong></p> <p>These datasets and temporary dataframes are necessary to run the scripts from the public GitLab repository: <a href="https://gitlab.com/saubin.meline/neutral-genetic-structuring-adaptation">https://gitlab.com/saubin.meline/neutral-genetic-structuring-adaptation</a>. Please refer to this public GitLab repository for the latest version of the codes and to perform all analyses presented in the article.</p> <p>Original datasets from the demogenetic model:</p> <ul> <li>Output_RandomDesign.txt</li> <li>Output_RegularDesign_With_host_alternation.txt</li> <li>Output_RegularDesign_Without_host_alternation.txt</li> <li>Output_RandomDesign_Mnull_Medoid_With_host_alternation.txt</li> <li>Output_RandomDesign_Mnull_Medoid_Without_host_alternation.txt</li> </ul> <p>All remaining files correspond to temporary dataframes generated by the scripts in the GitLab repository, provided here for reproducibility of the results and to save time at certain time-consuming scripts.</p>
Eco-evolutionary processes underlying early warning signals of population declines
<p>Datasets for the paper appearing in Journal of Animal ecology : "Eco-evolutionary processes underlying early warning signals of population declines". Also GitHub repository link :<a href="https://github.com/GauravKBaruah/ECO-EVO-EWS-DATA">https://github.com/GauravKBaruah/ECO-EVO-EWS-DATA</a></p>
Supplementary Table S27.1: Animal species native to South Africa that have invasive populations elsewhere.
<p>Animal species native to South Africa that have invasive populations elsewhere. Sorted by expected chronological appearance in the first place they were recorded as alien species. Notes are made on whether the introduction is known to be (Y) or not (N) from South Africa (or unknown U). Pathways are according to the CBD pathway classification scheme (Harrower et al. 2017), along with an indication of whether the introduction was intentional or accidental. Species that have multi-continental distributions, and which may in addition have some introduced populations are shown at the end of the table.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.