Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

151

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

151 results for “genome size”

Learn how ShareScore rates datasets ↗
edi56/100

Examining genome size and nutrient influence on plant damage patterns

Data was collected to examine whether and how plant genome size (GS) interacts with environmental nutrient additions to influence the amount and patterns of damage plants sustain from invertebrate herbivores and fungal pathogens. Plants were selected based on visual abundance in treatment plots in which nitrogen (N), phosphorus (P), or NP combined had been annually added (Cont. is the abbreviation we used for the control plot with no nutrients added). Additionally, plant traits of percent foliar carbon (% C), percent foliar nitrogen (% N), and specific leaf area (SLA) were measured from all the same plants that damage values were observed from. Data was collected from 847 plants (626 forb individuals, 221 grass individuals) in eight grassland sites that are part of the Nutrient Network (https://nutnet.org), a globally distributed experiment in which plots have different nutrient amendment treatments that are administered identically to allow cross-site comparisons of the effects of nutrients on biodiversity patterning. The sites chosen varied along a north-south latitude, longitude, mean annual precipitation (MAP) and mean annual temperature (MAT) gradient in the United States. All field data was collected between May 2022 and August 2022. The sites included in this study are listed below with their respective Nutrient Network site codes. churn.us= Churning Rapids in Hancock, MI spin.us= Spindletop Farm in Lexington, KY temple.us= Temple in Temple, TX kbs.us= Kellogg Biological Station in Hickory Corners, MI konz.us= Konza Prairie Biological Station in Manhattan, KS cgbg.us= Chichaqua Bottoms Greenbelt in Maxwell, IA cdcr.us= Cedar Creek in East Bethel, MN msum.us= Minnesota State University at Moorhead in Moorhead, MN

openCC (other)Nov 2025View details →
edi52/100

Data in support of 'Mechanistic insights into plant community responses to environmental variables: genome size, cellular nutrient investments, and metabolic trade-offs.'

Data was collected to examine whether and how the plant genome size (GS) influences traits (stomata size, stomata density, cellular and tissue level carbon (C), nitrogen (N), and phosphorus (P) contents) and metabolic-tradeoffs (of photosynthesis, evapotranspiration, water-use, efficiency) of plants in treatment plots in which nothing, N, P, or NP had been annually added. Data was collected from ~500 plants from seven grassland sites that are all part of the Nutrient Network (https://nutnet.org), a globally distributed experiment in which plots have different nutrient amendment treatments that are administered identically to allow cross-site comparisons of the effects of nutrients on biodiversity patterning. The sites chosen varied along a North-South latitude, longitude, mean annual precipitation (MAP) and mean annual temperature (MAT) gradient.

openCC (other)Sep 2024View details →
zenodo48/100

PopDel identifies medium-size deletions jointly in tens of thousands of genomes - Variant call sets

<p>This data set contains the variant calls sets generated by different tools for the benchmarks in the paper <a href="https://www.nature.com/articles/s41467-020-20850-5">PopDel identifies medium-size deletions simultaneously in tens of thousands of genomes</a>. It includes the VCFs/BCFs for the following test cases:</p> <ul> <li>Random deletion simulation on up to 1000 chromosome 21 samples</li> <li>1000 Genomes Project deletions inserted into simulated chromosomes 17 to 22 of up to 500 samples</li> <li>HG001 (NA12878)</li> <li>Trio of <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG002_NA24385_son/NIST_HiSeq_HG002_Homogeneity-10953946/">HG002</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG003_NA24149_father/NIST_HiSeq_HG003_Homogeneity-12389378/">HG003</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG004_NA24143_mother/NIST_HiSeq_HG004_Homogeneity-14572558/">HG004</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Diversity-Cohort">Polaris Diversity cohort</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Kids-Cohort">Polaris Kids cohort</a></li> </ul> <p>Further, the long and short read reference call sets for HG001 are provided. For HG002 the reference call set and the high confidence regions by the Genome in a Bottle consortium are provided.</p> <p>For details on how the files have been created, please refer to the paper and the script repository on <a href="https://github.com/kehrlab/PopDel-scripts">GitHub</a>.</p>

opencc-by-4.0Aug 2020View details →
edi44/100

Genome size influences plant growth and biodiversity responses to nutrient fertilization in diverse grassland communities

Experiments comparing diploids with polyploids and in single grassland sites show that nitrogen and/or phosphorus availability influences plant growth and community composition dependent on genome size; specifically plants with larger genomes grow faster under nutrient enrichments relative to those with smaller genomes. However, it is unknown if these effects are specific to particular site localities with speciifc plant assemblages, climates, and historical contingencies. To determine the generality of genome size dependent growth responses to nitrogen and phosphorus fertilisation, we combined genome size and species abundance data from 27 coordinated grassland nutrient addition experiments in the Nutrient Network that occur in the Northern Hemisphere across a range of climates and grassland communities. We found that after nitrogen treatment, species with larger genomes generally increased more in cover compared to those with smaller genomes, potentially due to a release from nutrient limitation. Responses were strongest for C3 grasses and in less seasonal, low precipitation environments, indicating that genome size effects on water-use-efficiency modulates genome size-nutrient interactions. Cumulatively the data suggest that genome size is informative and improves predictions of species’ success in grassland communities.

openCC0Nov 2024View details →
zenodo40/100

Figure 2 in EVOLUTIONARY PATTERNS OF GENOME SIZE AND CHROMOSOME NUMBER VARIATION IN BEGONIACEAE

Figure 2. Genome sizes and chromosome numbers of 64 Begonia species and Hillebrandia sandwicensis. The sources of the data used to create this scatter plot are specified in Supplementary table 2. Species in the same sections are enclosed within an ellipse. Colours indicate the continent where each species is found. The Hillebrandia data point is labelled as 'Outgroup'.

opencc-by-4.0Aug 2022View details →
zenodo40/100

Figure 1 in EVOLUTIONARY PATTERNS OF GENOME SIZE AND CHROMOSOME NUMBER VARIATION IN BEGONIACEAE

Figure 1. Variation in haploid chromosome number across the Begonia sections recognised by Moonlight et al. (2018) and their chromosome data. Boxes in the box plot are grouped by clade. The colours indicate the continent where these sections are found. Bar charts indicate the proportion of the section with known chromosome counts. Dots indicate sections with polyploid species or with species with known interspecific chromosome number variation (including B chromosomes). AC-C, Asian clade C; AC-D, Asian clade D; EDAB, early diverging Asian Begonia; FFAB, fleshy-fruited African Begonia; MB, Malagasy Begonia; NC1, Neotropical clade 1; NC2-i, Neotropical clade 2-i; NC2-ii, Neotropical clade 2-ii; NC2-iii, Neotropical clade 2-iii; SB, Socotran Begonia; SDAAB1, seasonally dry adapted African Begonia 1; SDAAB2, seasonally dry adapted African Begonia 2; YFAB, yellow-flowered African Begonia. * Unresolved or polyphyletic in the phylogeny of Moonlight et al. (2018).

opencc-by-4.0Aug 2022View details →
zenodo40/100

Fig. 4 in Genome size of chrysophytes varies with cell size and nutritional mode

Fig. 4 Comparison of genome size within different taxonomic groups. Mixotrophic (blue) and heterotrophic (dark red) chrysophytes rank among the smallest eukaryotic genomes (values obtained from [1] Mohanta and Bae 2015; Egertová and Sochor 2017; [2] Gregory 2017; [3] Bennett 2012; [4] Courties et al. 1994)

opencc-by-4.0May 2018View details →
zenodo40/100

Fig. 1 in Genome size of chrysophytes varies with cell size and nutritional mode

Fig. 1 Cell volumes [μm 3] of different chrysophytes. Different colors represent the different nutritional modes present. Phototrophic chrysophytes (light green) do have highest cell volumes compared to

opencc-by-4.0May 2018View details →
zenodo40/100

Fig. 2 in Genome size of chrysophytes varies with cell size and nutritional mode

Fig. 2 Genome size [pg] of investigated chrysophytes. Different colors represent the different nutritional modes present. Heterotrophic chrysophytes (dark red) tend to have smaller genome sizes, compared to phototrophic chrysophytes (light green), while mixotrophic chrysophytes (blue) show intermediate genome sizes. *Dinobryon sociale var. americana cf. div. schauinslandii; HF = Heterotrophic flagellate

opencc-by-4.0May 2018View details →
zenodo40/100

Fig. 5 in Genome size of chrysophytes varies with cell size and nutritional mode

Fig. 5 Model of evolution of genome size, cell volume, and nutritional mode of chrysophytes: nutrient limitations may have driven genome size reduction in the ancestors of mixotrophic (and heterotrophic) chrysophytes, as well as the evolution of phagotrophic mechanisms to attain additional nutrients. Cell size reduction is supposedly a more gradual process, coming into play in taxa which were already able to obtain nutrients by phagotrophy, which optimized food uptake by the optimization of the predator-prey size ratio. This may have triggered the evolution of obligate heterotrophs in many chrysophyte lineages independently

opencc-by-4.0May 2018View details →
zenodo40/100

Genome Sizes of Bacterial Species Detected in Cell-Free DNA of Patients with Acute Leukemia and Sepsis, Including Those Undergoing Bone Marrow Transplantation

<p>Next Generation Sequencing (NGS) analysis of Cell-Free DNA provides valuable insights into a spectrum of pathogenic species (particularly bacterial) in blood. Patients with Sepsis often face problems like delays in treatment regimens (combination or cocktail of antibiotics) due to the long turnaround time (TAT) of classical and standard blood culture procedures. NGS gives results with lower TAT along with high-depth coverage. The use of NGS may be a possible solution to deciding treatment regimens for patients without losing precious time and more accurately possibly saving lives.</p> <p>Our curated dataset is of bacterial species or strains detected along with their genome size in 107 AML patients diagnosed with Sepsis clinically. Cell-free DNA profiles of patients were built and sequencing was done in Illumina (NovaSeq and NextSeq). Bioinformatic analysis was performed using two classification algorithms namely kraken2 and kaiju. For kraken2 &nbsp;based classification reference bacterial index developed by Carlo Ferravante et al (Zenodo 2020) &nbsp;(link: https://zenodo.org/records/4055180) was used, while for kaiju-based classification reference database named "nr_euk" dated "2023-05-10" (link: https://bioinformatics-centre.github.io/kaiju/downloads.html) was used.</p> <p>Genome size annotation is important in metagenomics since for the use of depth of coverage (abundance), genome size is required. In metagenomic classification algorithms like kraken/kraken2 and kaiju output computes reads assigned only and not abundance. In kaiju, the problem is more complicated since the reference database does not have a fasta file but only an index file from which alignment is done.&nbsp;</p> <p>To address the above challenges to compute "depth of coverage" or simply abundance, we build a Genome size annotator tool (https://github.com/patkarlab/Genome-Size-Annotation) which provides genome size for each species detected given its taxid is available. In this tool, the NCBI Datasets tool, NCBI Genome API check tool, and Data Mining from AI search engines like perplexity.ai are used.&nbsp;</p> <p>We have curated two datasets</p> <p>Kraken2 dataset named "FINAL METAGENOMIC DATA MASTERSHEET - kraken_genome_annotation"<br>Kaiju dataset named "FINAL METAGENOMIC DATA MASTERSHEET - kaiju_genome_annotation"</p> <p>*Please note that for kraken2 curated dataset, we used data mining from the AI search engine perplexity.ai while for kaiju we did not use perplexity, ai, and any species whose genome size was not found was labeled "NA"</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Genome streamlining: effect of mutation rate and population size on genome size reduction: simulated data

<p>Lineages data of populations simulated with Aevol (<a href="https://gitlab.inria.fr/aevol/aevol">https://gitlab.inria.fr/aevol/aevol</a>), and the Wild-Types sequences used for that.</p> <p>Conditions: change of mutation rate, population size, or both.<br>Mutational bias: none, insertion bias or deletion bias</p>

opencc-by-4.0Feb 2024View details →
dryad40/100

Whole genome demographic models indicate divergent effective population size histories shape contemporary genetic diversity gradients in a montane bumble bee

<p>Understanding historical range shifts and population size variation provides important context for interpreting contemporary genetic diversity. Methods to predict changes in species distributions and model changes in effective population size (N<sub>e</sub>) using whole genomes make it feasible to examine how temporal dynamics influence diversity across populations. We investigate N<sub>e</sub> variation and climate-associated range shifts to examine the origins of a previously observed latitudinal heterozygosity gradient in the bumble bee <em>Bombus</em> <em>vancouverensis</em> Cresson (Hymenoptera: Apidae: <em>Bombus</em> Latreille) in western North America. We analyze whole genomes from a latitude-elevation cline using sequentially Markovian coalescent models of N<sub>e</sub> through time to test whether relatively low diversity in southern high-elevation populations is a result of long-term differences in N<sub>e</sub>. We use Maxent models of the species range over the last 130,000 years to evaluate range shifts and stability. N<sub>e</sub> fluctuates with climate across populations, but more genetically diverse northern populations have maintained greater Ne over the late Pleistocene and experienced larger expansions with climatically favorable time periods. Northern populations also experienced larger bottlenecks during the last glacial period which matched the loss of range area near these sites, however, bottlenecks were not sufficient to erode diversity maintained during periods of large N<sub>e</sub>. A genome sampled from an island population indicated a severe postglacial bottleneck, indicating that large recent post-glacial declines are detectable if they have occurred. Genetic diversity was not related to niche stability or glacial-period bottleneck size. Instead, spatial expansions and increased connectivity during favorable climates likely maintain diversity in the north while restriction to high elevations maintains relatively low diversity despite greater stability in southern regions. Results suggest genetic diversity gradients reflect long-term differences in N<sub>e</sub> dynamics and also emphasize the unique effects of isolation on insular habitats for bumble bees. Patterns are discussed in the context of conservation under climate change.</p>

opencc-zeroJan 2023View details →
dryad40/100

Whole genome demographic models indicate divergent effective population size histories shape contemporary genetic diversity gradients in a montane bumble bee

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad40/100

Data from: Island size shapes genomic diversity in a great speciator (Aves: Zosterops)

Open the record for dataset details and reuse information.

publicMar 2025View details →
dryad36/100

Data from: Miniaturization, genome size, and biological size in a diverse clade of salamanders

Genome size (C-value) can affect organismal traits across levels of biological organization, from tissue complexity to metabolism. Neotropical salamanders show wide variation in genome and body sizes, including several clades with miniature species. Because miniaturization imposes strong constraints on morphology and development, and genome size is strongly correlated with cell size, we hypothesize that body size has played an important role in the evolution of genome size in bolitoglossine salamanders. If this hypothesis is correct, then genome size and body size should be correlated in this group. Using Feulgen Image Analysis Densitometry (FIAD), we estimated genome sizes for 60 species of neotropical salamanders. We also estimated the "biological size" of species by comparing genome size and physical body sizes in a phylogenetic context. We found a significant correlation between C-value and physical body size using optimal regression with an Ornstein-Uhlenbeck model, and report the smallest salamander genome found to date. Our index of biological size showed that some salamanders with large physical body size have smaller biological body size than some miniature species, and that several clades showed patterns of increased or decreased biological size compared to their physical size. Our results suggest a causal relationship between physical body size and genome size and show the importance of considering the impact of both on the biological size of organisms. Indeed, biological size may be a more appropriate measure than physical size when considering phenotypic consequences of genome size evolution in many groups.

opencc-zeroJun 2020View details →
dryad36/100

ModEst - Precise estimation of genome size from NGS data

<p>Accurate estimates of genome sizes are important parameters for both theoretical and practical biodiversity genomics. We present here a fast, easy-to-implement and precise method to estimate genome size from the number of bases sequenced and the mean sequencing depth. To estimate the latter, we take advantage of the fact that a precise estimation of the Poisson distribution parameter lambda is possible from truncated data, restricted to the part of the sequencing depth distribution representing the true underlying distribution. With simulations we could show that reasonable genome size estimates can be gained even from low-coverage (10X), highly discontinuous genome drafts. Comparison of estimates from a wide range of taxa and sequencing strategies with flow-cytometry estimates of the same individuals showed a very good fit and suggested that both methods yield comparable, interchangeable results.</p>

opencc-zeroJan 2022View details →
zenodo36/100

Nanopore MinION Run Metrics and genomic DNA fragment size analysis data from automated phenol-chloroform extractions (RBI LabDroid Maholo)

<p>Nanopore MinION run MinKNOW statistical metrics output, Agilent Femto Pulse and Tape Station gDNA fragment size analysis reports of genomic DNA isolated from automated&nbsp;RBI LabDroid&nbsp;Maholo organic extractions.</p>

opencc-by-4.0Jan 2022View details →
dryad36/100

Gigantic genomes of salamanders indicate body temperature, not genome size, is the driver of global methylation and 5-methylcytosine deamination in vertebrates

<p>Transposable elements (TEs) are sequences that replicate and move throughout genomes, and they can be silenced through methylation of cytosines at CpG dinucelotides. TE abundance contributes to genome size, but TE silencing variation across genomes of different sizes remains underexplored. Salamanders include most of the largest C-values -- 9 to 120 Gb. We measured CpG methylation levels in salamanders with genomes ranging from 2N = ~58 Gb to 4N = ~116 Gb. We compared these levels to results from endo- and ectothermic vertebrates with more typical genomes. Salamander methylation levels are ~90%, higher than all endotherms. However, salamander methylation does not differ from other ectotherms, despite a ~100-fold difference in nuclear DNA content. Because methylation affects the nucleotide compositional landscape through 5-methylcytosine deamination to thymine, we quantified salamander CpG dinucleotide levels and compared them to other vertebrates. Salamanders and other ectotherms have comparable CpG levels, and ectotherm levels are higher than endotherms. These data show no shift in global methylation at the base of salamanders, despite a dramatic increase in TE load and genome size. This result is reconcilable with previous studies by considering endothermy and ectothermy, which may be more important drivers of methylation in vertebrates than genome size.</p>

opencc-zeroFeb 2022View details →
dryad36/100

Does genome size increase with water depth in marine fishes?

<p><span>A growing body of research suggests that genome size in animals can be affected by ecological factors. Half a century ago, Ebeling et al. (1971; EEA71) proposed that genome size increases with depth in some teleost fish groups and discussed a number of biological mechanisms that may explain this pattern (e.g., passive accumulation, adaptive acclimation). Using phylogenetic comparative approaches, we revisit this hypothesis based on genome size and ecological data from up to 708 marine fish species in combination with a set of large-scale phylogenies, including a newly inferred tree. We also conduct modelling approaches of trait evolution and implement a variety of regression analyses to assess the relationship between genome size and depth. Our reanalysis of the EEA71 dataset shows a weak association between these variables, but the overall pattern in their data is driven by a single clade. While analyses based on the new dataset resulted in positive correlations, providing some evidence that genome size evolves adaptively as a function of depth, only a fraction of individual subclade analyses yielded statistically significant results. By contrast, negative correlations are rare and largely non-significant. All in all, we find modest evidence for an increase in genome size along the depth axis in marine fishes. We discuss some mechanistic explanations for the observed trends.</span></p>

opencc-zeroApr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record