Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
22,445
datasets available to search
ShareScore release 0.7.1
Dataset results
22,445 results for “diversity”
Diverse Title Generation for Stack Overflow Posts with Multiple Sampling Enhanced Transformer
<p>Dataset for our paper "Diverse Title Generation for Stack Overflow Posts with Multiple Sampling Enhanced Transformer"</p> <p>(<a href="https://github.com/zfj1998/M3NSCT5">zfj1998/M3NSCT5: the code base for our paper "Diverse Title Generation for Stack Overflow Posts with Multiple Sampling Enhanced Transformer" (github.com)</a>)</p> <p>Including three files representing the train/val/test datasets. Each file contains all the collected data covering eight programming languages.</p>
Virome diversity of Hyalomma dromedarii ticks collected from camels in the United Arab Emirates
<p>Viruses are important components of the microbiome of ticks. Ticks are capable of transmitting several serious viral diseases to humans and animals. Hitherto, the composition of viral communities in <em>Hyalomma dromedarii</em> ticks associated with camels in the United Arab Emirates (UAE) remains unexplored. The purpose of this study was to characterize the RNA virome diversity in male and female <em>H. dromedarii</em> ticks collected from camels in Al Ain, UAE.<strong> </strong>We collected ticks, extracted and sequenced RNA, using Illumina (NovaSeq 6000) and Oxford Nanopore (MinION).<strong> </strong>From the total generated sequencing reads, 180,559 (~0.35 %) and 197,801 (~0.34 %) reads were identified as virus-related reads in male and female tick samples respectively. Taxonomic assignment of the viral sequencing reads was accomplished based on bioinformatic analyses. Further, viral reads were classified into 39 viral families. Poxiviridae, Phycodnaviridae, Phenuiviridae, Mimiviridae, and Polydnaviridae were the most abundant families in the tick viromes. Notably, we assembled the genomes of three RNA viruses, which were placed by phylogenetic analyses in clades that included the Bole tick virus.<strong> </strong>Overall, this study attempts to elucidate the RNA virome of ticks associated with camels in the UAE and the results obtained from this study improve the knowledge of the diversity of viruses in <em>H. dromedarii</em> ticks.</p>
Glycosylated models for: The diversity of the glycan shield of sarbecoviruses closely related to SARS-CoV-2
<p>Glycosylated models (as PDB files) of the sarbecovirus spike proteins used in the study: The diversity of the glycan shield of sarbecoviruses closely related to SARS-CoV-2.</p>
Multiple Sequence Alignment of a diverse dataset with 1788 Mycobacterium tuberculosis isolates
<p><strong>Multiple Sequence Alignment of a diverse dataset with 1788 <em>Mycobacterium tuberculosis</em> isolates used for <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> benchmarking</strong></p> <p>The dataset comprises whole-genome sequence data published by <a href="https://doi.org/10.1016/S1473-3099(15)00062-6">Walker et al. 2015</a>. For the multiple sequence analysis, we proceeded as follows:</p> <ol> <li>Reads were downloaded from ENA BioProject <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA282721">PRJNA282721</a> (accessed on March 16<sup>th</sup>, 2023) and trimmed using Trimmomatic (<a href="https://pubmed.ncbi.nlm.nih.gov/24695404/">Bolger et al., 2014</a>) with <a href="https://github.com/B-UMMI/INNUca">INNUca</a> default settings;</li> <li>Quality-processed reads were individually mapped against the H37Rv reference genome (Genbank accession: <a href="https://www.ncbi.nlm.nih.gov/nuccore/NC_000962.3/">NC_000962.3</a>) using <a href="https://github.com/tseemann/snippy">Snippy</a> v4.5.1 and SNP-calling was performed on variant sites with the following criteria: a minimum proportion of reads differing from the reference of 70%, a minimum mapping quality of 30 and a minimum coverage for SNP calling of 10;</li> <li>A full alignment was extracted using Snippy’s core module (snippy-core), with masking of SNPs falling within known <em>M. tuberculosis</em> genomic regions with high GC content, repetitive elements and resistance-associated positions (corresponding to ~8% of the genome), as previously described for surveillance purposes (<a href="https://pubmed.ncbi.nlm.nih.gov/30948181/">Macedo et al., 2019</a>);</li> <li><em>M. tuberculosis </em>lineages were determined using tb-profiler v4.4.1 (<a href="https://pubmed.ncbi.nlm.nih.gov/31234910/">Phelan et al., 2019</a>), with samples from the <em>M. tuberculosis</em> complex other than <em>M. tuberculosis</em>, representing a mix of multiple lineages, or with less than 95% of mapped positions in the reference, being excluded;</li> <li>A filtered alignment comprising the maximum number of informative sites (88,562 nucleotide sites with at least one mutation in a given sequence) was extracted from the full alignment using the alignment_processing.py v1.1.0 (default settings) of <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a>, and then used as input for the benchmarking.</li> </ol> <p>In this repository, we provide two alignment files:</p> <ul> <li>Core_MTB_1787_strs.full.aln: this corresponds to the full multiple sequence alignment comprising 1787 samples and the reference (corresponding to the point 4 of the methodology).</li> <li>MTb_original_align_profile.fasta: this corresponds to the multiple sequence alignment comprising 1787 samples and the reference and only presenting the alignment informative sites (corresponding to the point 5 of the methodology)</li> </ul>
Supplementary phylogenetic data for Rouïl et. al. 2020 "The protector within: Comparative genomics of APSE phages across aphids reveals rampant recombination and diverse toxin arsenals"
<p>Supplementary phylogenetic data for Rouïl <em>et. al.</em> 2020 "The protector within: Comparative genomics of APSE phages across aphids reveals rampant recombination and diverse toxin arsenals"</p> <p> </p> <p>The data set consists of the following sub-directories:</p> <p>1) "APSE_conserved_proteins_alns": Single-copy conserved genes codon sequences and alignments in FASTA format.</p> <p>2) "APSE_phylogeny": Files used for APSE phylogenetic and recombination analyses.</p> <p>3) "APSE_reannotations": GenBank-formatted files of the assemblies and re-annotations of APSE phages. Newly-sequenced phages deposited at the European nucleotide Archive are also included. ***New in this version***</p> <p>4) "APSE_toxin_lyzozyme": Files used for APSE toxin-cassette and lyzozyme-related gene phylogenies.</p> <p>5) "Arsenophonus_PHASTER": PHASTER phage annotation output files organised by organisim and contig/scaffold.</p> <p>6) "Hamiltonella_drafts": Newly-sequenced low-coverage draft <em>Hamiltonella</em> genomes in FASTA format.</p> <p>7) "Hamiltonella_phylogeny": files used for <em>Hamiltonella</em> phylogenetic analysis.</p> <p> </p> <p>See enclosed README.txt file for more details.</p> <p> </p> <p>* ver. 1.1.1: Updated annotations for APSE genomes including inteins missing in previous annotation files.</p>
Climate change threats to the global functional diversity of freshwater fish
<p>This dataset provides supplementary information for the paper entitled "Climate change threats to the global functional diversity of freshwater fish".</p> <p> </p> <p><strong>Fish trait data</strong></p> <p>fish_traits_removed.csv<br> - species with missing trait values were removed<br> - species coverage: 3,792</p> <p>fish_traits_imputed.csv<br> - missing trait values were imputed<br> - species coverage: 11,425</p> <p>Traits<br> - HLrel = relative head length<br> - BDrel = relative body depth<br> - Troph = trophic level<br> - K = relative growth rate</p> <p><br> <strong>Geospatial data</strong></p> <p>Files<br> Data under the assumption of no dispersal<br> - SR.tif: species richness<br> - FRic.tif: functional richness<br> - FEve.tif: functional evenness<br> - FDiv.tif: functional divergence<br> - FRic_loss.tif: functional richness loss<br> - FEve_loss.tif: functional evenness loss<br> - FDiv_loss.tif: functional divergence loss</p> <p>Data under the assumption of maximal dispersal<br> - SR_dispersal.tif: species richness<br> - FRic_dispersal.tif: functional richness<br> - FEve_dispersal.tif: functional evenness<br> - FDiv_dispersal.tif: functional divergence<br> - FRic_loss_dispersal.tif: functional richness loss<br> - FEve_loss_dispersal.tif: functional evenness loss<br> - FDiv_loss_dispersal.tif: functional divergence loss</p> <p>Layers<br> - imp_*: missing trait values were imputed<br> - rem_*: species with missing trait values were removed<br> - *_hist: historical reference scenario<br> - *_1p5: warming level of 1.5°C<br> - *_2p0: warming level of 2.0°C<br> - *_3p2: warming level of 3.2°C<br> - *_4p5: warming level of 4.5°C</p> <p>Spatial resolution: 0.08333333, 0.08333333 (x, y)<br> Spatial extent: -180, 180, -60, 85 (xmin, xmax, ymin, ymax)<br> Coordinate reference system: WGS84</p>
Spineless and overlooked: DNA metabarcoding of autonomous reef monitoring structures reveals intra- and interspecific genetic diversity in Mediterranean invertebrates
<p>Sequence data and stepwise pipeline outputs associated with the article "Spineless and overlooked: DNA metabarcoding of autonomous reef monitoring structures reveals intra- and interspecific genetic diversity in Mediterranean invertebrates".</p> <p>Preprint available here: <a href="https://doi.org/10.22541/au.167085544.47638352/v1">10.22541/au.167085544.47638352/v1</a></p> <p>Sequence data is deposited in fastq-format in folders by region (Palinuro.tar.gz, Livorno.tar.gz, and Rovinj.tar.gz) and a separate folder for controls (Controls.tar.gz). Each fastq-file contains sequences for a single PCR replicate named by sample and replicate number. Sample names are described in spineless_sample_names.csv. Positive control sequences are described in SM1_positive_controls.csv. Stepwise pipeline outputs are available in the folder Pipeline_outputs_stepwise.zip</p> <p>Scripts used to generate pipeline outputs as well as other aspects of the final article are available at <a href="https://github.com/thomasdotter/spineless-haplotypes">https://github.com/thomasdotter/spineless-haplotypes</a>.</p> <p> </p>
Raw Data for Publication: Fungal colonisation on wood surfaces weathered at diverse climatic conditions
<p>Colour_Changes_Izola.csv</p> <p>This file contains CIE Lab* color coordinates measured on the surface of Scots pine during the natural weathering test in Izola, Slovenia.</p> <p>Colour_Changes_Skelleftea.csv</p> <p>This file contains CIE Lab* color coordinates measured on the surface of Scots pine during the natural weathering test in Skelleftea, Sweden.</p> <p>Contact_Angles_Izola.csv</p> <p>This file contains dynamic contact angle with distilled water measured on the surface of Scots pine during the natural weathering test in Izola, Slovenia.</p> <p>Contact_Angles_Skelleftea.csv</p> <p>This file contains dynamic contact angle with distilled water measured on the surface of Scots pine during the natural weathering test in Skelleftea, Sweden.</p> <p>Gloss.csv</p> <p>This file contains the gloss value measured on the surface of Scots pine during the natural weathering test in Izola, Slovenia and Skelleftea, Sweden.</p> <p><strong>Note:</strong> The sample IDs are structured as follows: The first letter represents the treatment condition, with "R" indicating untreated wood. The second letter (A, B, C) represents the board's ID. The third letter represents the location, with "S" indicating Skelleftea and "I" representing Izola, Slovenia. The number indicates the exposure time in weeks.</p> <p>FUNGAL STRAIS_DNA sequence analysis.xlsx</p> <p>This file contains the Genomic DNA of the fungal strains detected on the surface of Scots pine during the natural weathering test in Izola, Slovenia and Skelleftea, Sweden.</p> <p>Weather_Data_Izola.csv</p> <p>This file contains hourly local weather conditions in Izola, Slovenia</p> <p>Weather_Data_Skelleftea.csv</p> <p>This file contains hourly local weather conditions in Skelleftea, Sweden</p> <p><strong>Note:</strong> The weather conditions including the following parameters:1. Air temperature (°C), 2. Dew point (°C), 3. Relative humidity (%), 4. One-hour precipitation total (mm), 5. Snow depth (mm), 6. Wind direction (°), 7. Average wind speed (km/h), 8.Sea-level air pressure (hPa)</p>
Diversity and evolution of cerebellar folding in mammals
<p>Coronal cerebellar mid-sections for 56 mammalian species, at the same scale.</p> <p> </p> <p>This figure is from our open access paper:</p> <p>Heuer, K., Traut, N., de Sousa, A. A., Valk, S., & Toro, R. (2022). Diversity and evolution of cerebellar folding in mammals. bioRxiv. <a href="https://doi.org/10.1101/2022.12.30.522292">https://doi.org/10.1101/2022.12.30.522292</a></p> <p> </p> <p>Abstract</p> <p>The process of brain folding is thought to play an important role in the development and organisation of the cerebrum and the cerebellum. The study of cerebellar folding is challenging due to the small size and abundance of its folia. In consequence, little is known about its anatomical diversity and evolution. We constituted an open collection of histological data from 56 mammalian species and manually segmented the cerebrum and the cerebellum. We developed methods to measure the geometry of cerebellar folia and to estimate the thickness of the molecular layer. We used phylogenetic comparative methods to study the diversity and evolution of cerebellar folding and its relationship with the anatomy of the cerebrum. Our results show that the evolution of cerebellar and cerebral anatomy follows a stabilising selection process. We observed 2 groups of phenotypes changing concertedly through evolution: a group of “diverse” phenotypes – varying over several orders of magnitude together with body size, and a group of “stable” phenotypes varying over less than 1 order of magnitude across species. Our analyses confirmed the strong correlation between cerebral and cerebellar volumes across species, and showed in addition that large cerebella are disproportionately more folded than smaller ones. Compared with the extreme variations in cerebellar surface area, folial anatomy and molecular layer thickness varied only slightly, showing a much smaller increase in the larger cerebella. We discuss how these findings could provide new insights into the diversity and evolution of cerebellar folding, the mechanisms of cerebellar and cerebral folding, and their potential influence on the organisation of the brain across species.</p>
Connecting the multiple dimensions of global soil fungal diversity
<p>How the multiple facets of soil fungal diversity vary worldwide remains virtually unknown, hindering the management of this essential species-rich group. By sequencing high-resolution DNA markers in over 4000 topsoil samples from natural and human-altered ecosystems across all continents, we illustrate the distributions and drivers of different levels of taxonomic and phylogenetic diversity of fungi and their ecological groups. We show the impact of precipitation and temperature interactions on fungal local species richness (alpha diversity) across different climates. Our findings reveal how temperature drives fungal compositional turnover (beta diversity) and phylogenetic diversity, linking them with regional species richness (gamma diversity). Our work integrates fungi into the principles of global biodiversity distribution and presents detailed maps for biodiversity conservation and modeling of global ecological processes.</p> <p><strong>### Data overview</strong></p> <p>These datasets contain comprehensive estimates of alpha, beta, and gamma diversity. The data are provided in two formats: TIFF (Tagged Image File Format) and GeoPackage formats, which are commonly used to store geospatially-referenced data.</p> <p><strong>Alpha Diversity</strong>:</p> <ul> <li>`<em>Alpha_S_</em>*` files: These files contain estimates of alpha diversity (local species diversity) for each grid cell of a raster file.</li> <li>`<em>Alpha_AOA_</em>*` files: These files outline the 'Area of Applicability' for the alpha diversity estimates.</li> <li>`<em>Alpha_Uncertainty_</em>*` files: These files contain data related to the uncertainty of the alpha diversity predictions. Uncertainty here represents the range or degree of error associated with the diversity estimates.</li> <li> `<em>Alpha_Hotspots_and_ProtectedAreas</em>` contains information on fungal diversity hotspots and their area under protection (based on IUCN classification). 'Hotspots' are areas with exceptionally high alpha diversity.</li> </ul> <p><strong>Beta Diversity</strong>:</p> <ul> <li>`<em>Beta_</em>*` files: These files include results of beta diversity analyses: maps of global compositional dissimilarity among soil fungal communities and maps of compositional turnover rate.</li> </ul> <p><strong>Other files</strong>:</p> <ul> <li>`<em>EcM_and_AM_GlobalDistribution</em>`: the global distribution of areas with high richness of ectomycorrhizal and arbuscular mycorrhizal fungi.</li> <li>`<em>Ecoregions_Alpha,Beta,Gamma_Diversities</em>`: estimates of alpha, beta, and gamma diversity at the level of ecoregion cf. Tedersoo et al., 2022 (DOI:10.1111/gcb.16398).</li> </ul> <p> </p> <p><strong>### Data description</strong></p> <p>Alpha diversity, which is a measure of local species richness (number of Operational Taxonomic Unit (OTU) representing distinct taxa, roughly corresponding to species level). Alpha diversity is represented by the residuals from a model adjusting for sequencing depth, with zero equating to the average OTU richness in the training data set.</p> <p><br> `<strong>Alpha_S_AllFungi_Consensus.tif</strong>`: This file provides consensus estimates for total fungal alpha diversity.<br> Within the file, there are two types of consensus estimates:</p> <ul> <li> <em>AvgW</em> - weighted consensus estimates for alpha diversity. The weighting takes into account both the area of applicability and the goodness-of-fit for the model used to generate the estimates.</li> <li> <em>Avg</em> - non-weighted consensus estimates for alpha diversity. Unlike <em>AvgW</em>, these estimates give equal weight to all models regardless of their goodness-of-fit or area of applicability.</li> </ul> <p><br> `<strong>Alpha_AOA_*</strong>`: Files containing Area of Applicability information:</p> <ul> <li> A raster value of '1' represents areas that are outside the Area of Applicability</li> <li> A raster value of '2' denotes areas that are inside the Area of Applicability</li> </ul> <p><br> In the files containing prediction uncertainties (`<strong>Alpha_Uncertainty_*</strong>`), two types of data are presented to quantify the amount of uncertainty in model predictions, each represented by a different band:</p> <ul> <li>The SD band represents the standard deviation of predictions based on different folds of cross-validation. A larger standard deviation indicates greater variability in the predictions.</li> <li>The IQR band represents the interquartile range (the difference between the upper and lower quartiles) of predictions. The wider the IQR, the greater variability in the predictions.</li> </ul> <p><br> `<strong>Alpha_Hotspots_and_ProtectedAreas.tif</strong>`: This file provides information on regions of exceptionally high species richness, referred to as 'hotspots', along with information about protected areas. Hotspots are identified as the top 2.5% quantiles of the richest grid cells on the map in terms of OTU richness.</p> <ul> <li><em>IUCN_1_4</em> - terrestrial protected areas that fall into categories I-IV, as classified by the International Union for Conservation of Nature (IUCN). These categories typically represent areas with high levels of protection, often prohibiting extractive and destructive activities to preserve biodiversity.</li> <li><em>IUCN_all</em> - all terrestrial protected areas as recorded in the World Database on Protected Areas (WDPA) database v.1.6. It includes a wider range of protected areas beyond the categories I-IV.</li> <li><em>All_Avg</em> - Hotspots of total fungal alpha diversity, based on the consensus map</li> <li><em>GSM_All</em> - Hotspots of total fungal alpha diversity, based on the GSMc dataset</li> <li><em>GSM_EcM</em> - Hotspots of ectomycorrhizal alpha diversity</li> <li><em>GSM_AM</em> - Hotspots of arbuscular mycorrhizal alpha diversity</li> <li><em>GSM_AgarNM</em> - Hotspots of non-EcM Agaricomycetes alpha diversity</li> <li><em>GSM_Mold</em> - Hotspots of mold alpha diversity</li> <li><em>GSM_Pathog</em> - Hotspots of opportunistic human parasitic fungal alpha diversity</li> <li><em>GSM_OHP</em> - Hotspots of putative pathogenic fungal alpha diversity</li> <li><em>GSM_Unicel</em> - Hotspots of unicellular, non-yeast fungal alpha diversity</li> <li><em>GSM_Yeast</em> - Hotspots of yeast alpha diversity</li> <li><em>GSMc_PD</em> - Hotspots of phylogenetic alpha diversity</li> <li><em>GSM_PDst</em> - Hotspots of phylogenetic dispersion</li> </ul> <p><br> `<strong>EcM_and_AM_GlobalDistribution.tif</strong>`: To illustrate the worldwide distribution of ectomycorrhizal (EcM) and arbuscular mycorrhizal (AM) fungi, we have categorized their richness into three distinct groups with low (1), medium (2), and high (3) alpha diversity. These categories have been encoded in the raster file using a bitcode system. Specifically, a value of '9' indicates that both EcM and AM fungal communities have low alpha diversity, while a value of '27' signifies that both groups of fungi are OTU-rich To assist with interpretation, a color legend has been provided in a separate QML style file (`<strong>EcM_and_AM_GlobalDistribution.qml</strong>`). This should be automatically recognized by geographic information system software, such as QGIS, to aid in visual analysis.</p> <p><br> `<strong>Beta_Taxonomic_AllFungi.tif</strong>` and `<strong>Beta_Phylogenetic_AllFungi.tif</strong>`: These files quantify the degree of difference in OTU composition of fungal communities. The measurements are based on the Generalized Dissimilarity Modelling (GDM) framework, as described by Mokany et al., 2022 (DOI:10.1111/geb.13459). Each file provides a different perspective on beta diversity: taxonomic (which is the change in species composition between different locations), and phylogenetic (the change in phylogenetic lineage composition between different locations). Each of these raster files contains three bands, with each band representing a scaled axis from a Principal Component Analysis (PCA) of the GDM-transformed environmental predictors.</p> <p><br> `<strong>Beta_LocalTurnover.tif</strong>`: This file contains estimates of local turnover in fungal communities composition estimated as the median expected compositional dissimilarity (taxonomic or phylogenetic) between each location and its closest neighbors within a 150 km radius. In addition, interquartile range (IQR) of dissimilarities is also provided.</p> <p> </p> <p>`<strong>Ecoregions_Alpha,Beta,Gamma_Diversities.gpkg</strong>`: Median alpha, beta, and gamma diversity estimates within ecoregions.</p> <ul> <li><em>Ecoregion</em> - Ecoregion name (cf. Tedersoo et al., 2022, DOI:10.1111/gcb.16398)</li> <li><em>area</em> - Ecoregion area, m<sup>2</sup></li> <li><em>Alpha_S_AllFungi_Consensus</em> - Richness of all fungi (S'<sub>tot</sub>), consensus map</li> <li><em>Alpha_S_AllFungi_GSMc</em> - Richness of all fungi (S'<sub>GSMc</sub>), based on GSMc dataset</li> <li><em>Alpha_S_EcM_GSMc</em> - Richness of ectomycorrhizal fungi (S'<sub>ecm</sub>)</li> <li><em>Alpha_S_AM_GSMc</em> - Richness of arbuscular mycorrhizal fungi (S'<sub>am</sub>)</li> <li><em>Alpha_S_NMA_GSMc</em> - Richness of non-EcM Agaricomycetes (S'<sub>nma</sub>)</li> <li><em>Alpha_S_Mold_GSMc</em> - Richness of molds (S'<sub>mold</sub>)</li> <li><em>Alpha_S_OHP_GSMc</em> - Richness of opportunistic human parasitic fungi (S'<sub>ohp</sub>)</li> <li><em>Alpha_S_Path_GSMc</em> - Richness of putative pathogenic fungi (S'<sub>path</sub>)</li> <li><em>Alpha_S_Ucel_GSMc</em> - Richness of unicellular, non-yeast fungi (S'<sub>ucel</sub>)</li> <li><em>Alpha_S_Yeast_GSMc</em> - Richness of yeasts (S'<sub>yeast</sub>)</li> <li><em>Alpha_SESPD_GSMc</em> - Phylogenetic dispersion of fungal communities (SES<sub>PD</sub>)</li> <li><em>Beta_Taxonomic_Median</em> - Median taxonomic dissimilarity of fungal communities (Simpson's index)</li> <li><em>Beta_Taxonomic_IQR</em> - Interquartile range of taxonomic dissimilarities of fungal communities</li> <li><em>Beta_Phylogenetic_Median</em> - Median phylogenetic dissimilarity of fungal communities</li> <li><em>Beta_Phylogenetic_IQR</em> - Interquartile range of phylogenetic dissimilarities of fungal communities</li> <li><em>Gamma_AllFungi</em> - Gamma diversity (regional species richness) for all fungi (G<sub>tot</sub>)</li> <li><em>Gamma_EcM</em> - Gamma diversity of ectomycorrhizal fungi (G<sub>ecm</sub>)</li> <li><em>Gamma_AM</em> - Gamma diversity of arbuscular mycorrhizal fungi (G<sub>am</sub>)</li> <li><em>Gamma_NMA</em> - Gamma diversity of non-EcM Agaricomycetes (G<sub>nma</sub>)</li> <li><em>Gamma_Mold</em> - Gamma diversity of molds (G<sub>mold</sub>)</li> <li><em>Gamma_Path</em> - Gamma diversity of opportunistic human parasitic fungi (G<sub>ohp</sub>)</li> <li><em>Gamma_OHP</em> - Gamma diversity of putative pathogenic fungi (G<sub>path</sub>)</li> <li><em>Gamma_Ucel</em> - Gamma diversity of unicellular, non-yeast fungi (G<sub>ucel</sub>)</li> <li><em>Gamma_Yeast</em> - Gamma diversity of yeasts (G<sub>yeast</sub>)</li> </ul> <p> </p> <p><strong>### Source code</strong></p> <p>The code used for data analysis and visualization of the main results of the study are available at GitHub:</p> <p><a href="https://github.com/Mycology-Microbiology-Center/Global_fungal_diversity">https://github.com/Mycology-Microbiology-Center/Global_fungal_diversity</a></p> <p> </p>
The effect of dietary bioactive on gut microbiome diversity (DIME) – a pilot study
<p>The DIME study consists of a randomised 2x2 cross-over human intervention where healthy participants (n = 20) are subjected to a diet high in bioactive-rich food for two weeks and a diet low in bioactive-rich food. There is a four-week washout between the two interventions. </p> <p>The continuous glucose monitoring was achieved using the Abbott freesylte libre flash glucose device. The baseline of the participants were determined 7 days before the start of the intervention, followed by the first arm and second arm. The period between the two arms (washout) was not recorded.</p> <p>We also included sleep data which consists of the amount of time spent in bed and during that time the amount of time spent in light, deep and rem in all 20 participants during the course of the dietary intervention, both the high and low bioactive diet. that was captured using Fitbit wearables during both stages of the dietary intervention,</p>
Supplementary file 1 from: Moliner Cachazo L, Makati K, Chadwick MA, Catford JA, Price BW, Mackay AW, Guiry MD, Murray-Hudson M, Murray-Hudson F (2023) A review of the freshwater diversity in the Okavango Delta and Lake Ngami (Botswana): taxonomic composition, ecology, comparison with similar systems and conservation status. Aquatic Sciences
<p>Dataset with 2,204 freshwater species from the Okavango Delta and Lake Ngami (Botswana), with additional 355 species found in other areas of Botswana that are likely to be present in the study region. The dataset covers the following groups: amphibians, birds, fishes, macroinvertebrates, macrophytes, mammals, reptiles, phytoplankton, and zooplankton. The following information is given for each species: status in the Okavango Delta and Lake Ngami (present/potentially present); conservation status globally, Phylum, Class, Order, Family, Genus, species name, cited synonyms, common name, habitat, presence in high water, presence in low water, ecology, distribution in continental Africa, confirmed locations in the Okavango Delta, site coordinates, references, notes.</p>
Voltage-based strategies for preventing battery degradation under diverse fast-charging conditions
<p>Here are the simulated datasets for the work 'Voltage-based strategies for preventing battery degradation under diverse fast-charging conditions', published in ACS Energy Letters, September 2023. The utilization of these datasets is demonstrated in the associated GitHub project folder https://github.com/zachkonz/Voltage-based-plating-prevention.</p>
Genetic diversity, population structure, and linkage disequilibrium among tropical quality protein maize (QPM) lines assessed with high-density SNP markers
<p>The study of genetic diversity (GD), population structure, and linkage disequilibrium (LD) provides a better understanding of the genetic relationships between individuals in a population which can be utilized in crop research and improvement. Genotyping-by-sequencing (GBS) was used to detect and genotype single nucleotide polymorphisms (SNPs) in a collection of 74 quality protein maize (QPM) lines and further to characterize their genetic diversity, population structure, and linkage disequilibrium. A total of 235,214 high-quality SNPs were used for different genetic analyses except for structure analysis where 11,950 SNPs were used. Analysis of molecular variance (AMOVA) based on these SNPs revealed high genetic heterozygosity among the five populations with 1% of the total genetic variation present among the subpopulations and 99% of the variation among individuals within the populations. Population structure analysis using Bayesian-based clustering revealed that the 74 lines could be clustered into four groups. However, neighbor-joining trees indicate the lines are grouped into three major clusters. Further analysis using principal component analyses (PCA) clustered the genotypes into five groups which are concordant with the groups based on pedigree information. Higher genetic diversity was detected in population 1 with a GD value of 0.484 and the lowest in population 5 (0.396) and overall, with a mean of 0.434. The LD pattern in the quality protein maize was investigated and we observed a relatively rapid LD decay of 3.53kb and 10.66kb at r<sup>2</sup> =0.2 and r<sup>2</sup>= 0.1, respectively. Our findings provide important information for future Linkage mapping studies, genome-wide association analyses, and marker-assisted selective breeding of maize as well as genomic prediction-based selection in tropical germplasm.</p>
Diversity loss from multiple interacting disturbances is regime-dependent
<p>Data and R code for 'Diversity loss from multiple interacting disturbances is regime-dependent'.</p> <p>Information about the files can be found in the ._README.txt file.</p>
CLDF dataset derived from Bowern et al.'s "Diversity in the Numeral Systems of Australian Languages" from 2012
<p>Cite the source of the dataset as:</p> <blockquote> <p>Bowern, Claire, and Jason Zentz (2012): Diversity in the Numeral Systems of Australian Languages. Anthropological Linguistics, 2012. http://www.jstor.org/stable/23621076</p> </blockquote>
Plumes and Blooms: Microbial eukaryote diversity and composition
These are amplicon sequencing data collected during Plumes and Blooms (PnB) cruises conducted from March, 2011, through September, 2014. The V9 hypervariable region of the 18S rRNA gene derived from microbial eukaryotic communities was amplified and sequenced from 345 discrete seawater samples. Sample collection and laboratory methods are described in Catlett et al. 2020 and Catlett et al. in review. Bioinformatic and data manipulation methods follow those employed in Catlett et al. in review. The data are provided in two tables: one includes amplicon sequence variant (ASV) sequences and relative sequence abundances for each sampling event, and the other includes ASV taxonomy predictions for each ASV sequence. References: Catlett, D., P. G. Matson, C. A. Carlson, E. G. Wilbanks, D. A. Siegel, and M. D. Iglesias‐Rodriguez. 2020. Evaluation of accuracy and precision in an amplicon sequencing workflow for marine protist communities. Limnol. Oceanogr.: Methods. 18(1): 20-40. https://doi.org/10.1002/lom3.10343. Catlett, D., D. A. Siegel, P. G. Matson, E. K. Wear, C. A. Carlson, T. S. Lankiewicz, and M. D. Iglesias‐Rodriguez. In review. Integrating phytoplankton pigment and DNA meta-barcoding observations to determine phytoplankton community composition in the coastal ocean. Limnol. Oceanogr.
Genome size influences plant growth and biodiversity responses to nutrient fertilization in diverse grassland communities
Experiments comparing diploids with polyploids and in single grassland sites show that nitrogen and/or phosphorus availability influences plant growth and community composition dependent on genome size; specifically plants with larger genomes grow faster under nutrient enrichments relative to those with smaller genomes. However, it is unknown if these effects are specific to particular site localities with speciifc plant assemblages, climates, and historical contingencies. To determine the generality of genome size dependent growth responses to nitrogen and phosphorus fertilisation, we combined genome size and species abundance data from 27 coordinated grassland nutrient addition experiments in the Nutrient Network that occur in the Northern Hemisphere across a range of climates and grassland communities. We found that after nitrogen treatment, species with larger genomes generally increased more in cover compared to those with smaller genomes, potentially due to a release from nutrient limitation. Responses were strongest for C3 grasses and in less seasonal, low precipitation environments, indicating that genome size effects on water-use-efficiency modulates genome size-nutrient interactions. Cumulatively the data suggest that genome size is informative and improves predictions of species’ success in grassland communities.
Data from publication: Castillioni, K., & Isbell, F. (2023). Early positive spatial selection effects of beta-diversity on ecosystem functioning. Landscape Ecology, 1-15.
Data from publication: Castillioni, K., & Isbell, F. (2023). Early positive spatial selection effects of beta-diversity on ecosystem functioning. Landscape Ecology, 1-15. Spatial beta-diversity may increase landscape productivity if there are positive spatial selection effects. Alternatively, dominant species in mixtures might not be the most productive species in monoculture leading to negative or neutral spatial selection effects. However, these hypotheses remain untested experimentally. Seedling survival can determine species establishment, influencing productivity later. To address this knowledge gap, we experimentally tested whether transplanted seedlings of dominant species optimally sort among habitat types (grassland dominated by Andropogon gerardii, savanna by Quercus macrocarpa, deciduous forest by Acer rubrum, coniferous forest by Pinus strobus, bog by Larix laricina), creating positive effects of landscape diversity on seedling survival and net biodiversity effects at Cedar Creek Ecosystem Science Reserve (CCESR) in Minnesota, USA. The study is named BetaDIV and consists of 100 plots (20 plots per habitat × 5 habitats). Each of the five habitats includes two true replicate monocultures for each of the five species and two true replicates for each of the five possible mixture compositions of four species (leaving each one out in turn to eventually explore the effect of species identity). Each plot is 1.5 by 1.5 m, with 12 seedlings planted 0.5 m apart in a 4 × 4 square grid, except in the plot corners. In the early June 2022, we tagged and planted all seedlings (i.e., bareroot seedlings for trees and plugs for the grass A. gerardii). Two weeks after the initial transplanting, we started tracking seedling survival (presented here) to investigate how seedlings responded to local habitat conditions. We conducted a seedling census for each of the 1200 tagged seedlings (12 seedlings per plot×100 plots), in early September 2022, which was two months at the end
NWFSC fish and invertebrate diversity derived from west coast groundfish trawl program
This dataset presents the community structure of groundfish and invertebrate in the West Coast since 1977. The community structure indices include the richness and Simpson’s evenness. The raw count data is from West Coast Groundfish Bottom Trawl survey conducted by the Northwest Fisheries Science Center. The spatial coverage of this dataset is between Pt Conception, California and north of U.S.-Mexico border.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.