Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
151
datasets available to search
ShareScore release 0.9.0
Dataset results
151 results for “genome size”
WorldClim, elevation and distribution data for all palms from: The ecology of palm genomes: Repeat-associated genome size expansion is constrained by aridity
<p>Genome size varies 2,400-fold across plants, influencing their evolution through changes in cell size and cell division rates which impact plants' environmental stress tolerance. Repetitive element expansion explains much genome size diversity, and the processes structuring repeat 'communities' are analogous to those structuring ecological communities. However, which environmental stressors influence repeat community dynamics has not yet been examined from an ecological perspective.</p> <p>We measured genome size and leveraged climatic data for 91% of genera within the ecologically diverse palm family (Arecaceae). We then generated genomic repeat profiles for 141 palm species, and analysed repeats using phylogenetically-informed linear models to explore relationships between repeat dynamics and environmental factors.</p> <p>We show that palm genome size and repeat 'community' composition are best explained by aridity. Specifically, <em>Ty3-gypsy</em> and <em>TIR </em>elements were more abundant in palm species from wetter environments, which generally had larger genomes, suggesting amplification. In contrast, <em>Ty1-copia</em> and <em>LINE </em>elements were more abundant in drier environments.</p> <p>Our results suggest that water stress inhibits repeat expansion through selection on upper genome size limits. However, elements which may associate with stress-response genes (e.g., <em>Ty1-copia</em>) have amplified in arid-adapted palm species. Overall, we provide novel evidence of climate influencing the assembly of repeat 'communities'. </p>
Data from: Unravelling hybridization in Phytophthora using phylogenomics and genome size estimation
<p>The genus <i>Phytophthora</i> comprises many economically and ecologically important plant pathogens. Hybrid species have previously been identified in at least six of the 12 phylogenetic clades. These hybrids can potentially infect a wider host range and display enhanced vigour compared to their progenitors. <i>Phytophthora</i> hybrids therefore pose a serious threat to agriculture as well as to natural ecosystems. Early and correct identification of hybrids is therefore essential for adequate plant protection but this is hampered by the limitations of morphological and traditional molecular methods. Identification of hybrids is also important in evolutionary studies as the positioning of hybrids in a phylogenetic tree can lead to suboptimal topologies. To improve the identification of hybrids we have combined genotyping-by-sequencing (GBS) and genome size estimation on a genus-wide collection of 614 <i>Phytophthora</i> isolates. Analyses based on locus- and allele counts and especially on the combination of species-specific loci and genome size estimations allowed us to confirm and characterize 27 previously described hybrid species and discover 16 new hybrid species. Our method was also valuable for species identification at an unprecedented resolution and further allowed correct naming of misidentified isolates. We used both a concatenation- and a coalescent-based phylogenomic method to construct a reliable phylogeny using the GBS data of 140 non-hybrid <i>Phytophthora</i> isolates. Hybrid species were subsequently connected to their progenitors in this phylogenetic tree. In this study we demonstrate the application of two validated techniques (GBS and flow cytometry) for relatively low cost but high resolution identification of hybrids and their phylogenetic relations.</p>
Data from: Genome size evolution and phenotypic correlates in the poison frog family Dendrobatidae
Open the record for dataset details and reuse information.
Data from: Is genomic diversity a useful proxy for census population size? Evidence from a species-rich community of desert lizards
Open the record for dataset details and reuse information.
WorldClim, elevation and distribution data for all palms from: The ecology of palm genomes: Repeat-associated genome size expansion is constrained by aridity
Open the record for dataset details and reuse information.
Simulated genotype data from: Effective population size estimation in large marine populations: Considering current challenges and opportunities when simulating large datasets with high-density genomic information
Open the record for dataset details and reuse information.
Data from: Scaling of thermal tolerance with body mass and genome size in ectotherms: a comparison between water-and air-breathers
Open the record for dataset details and reuse information.
Data from: Genomes of Galápagos mockingbirds reveal the impact of island size and past demography on inbreeding and genetic load in contemporary populations
Open the record for dataset details and reuse information.
ModEst - Precise estimation of genome size from NGS data
Open the record for dataset details and reuse information.
Gigantic genomes of salamanders indicate body temperature, not genome size, is the driver of global methylation and 5-methylcytosine deamination in vertebrates
Open the record for dataset details and reuse information.
Does genome size increase with water depth in marine fishes?
Open the record for dataset details and reuse information.
Data from: Miniaturization, genome size, and biological size in a diverse clade of salamanders
Open the record for dataset details and reuse information.
Selecting a window size for the analysis of whole genome alignments using AIC
Open the record for dataset details and reuse information.
Data from: Beyond population size: Whole-genome data reveal bottleneck legacies in the peninsular Italian wolf
Open the record for dataset details and reuse information.
Data from: Unravelling hybridization in Phytophthora using phylogenomics and genome size estimation
Open the record for dataset details and reuse information.
Data from: Competition among native and invasive Phragmites australis populations: an experimental test of the effects of invasion status, genome size, and ploidy level.
Open the record for dataset details and reuse information.
Austropuccinia psidii, causing myrtle rust, has a gigabase-sized genome shaped by transposable elements
<p><em>Austropuccinia psidii</em>, originating in South America, is a globally invasive fungal plant pathogen that causes rust disease on Myrtaceae. Several biotypes are recognized, with the most widely distributed pandemic biotype spreading throughout the Asia-Pacific and Oceania regions over the last decade. <em>Austropuccinia</em><em> psidii</em> has a broad host range with more than 480 myrtaceous species. Since first detected in Australia in 2010, the pathogen has caused the near extinction of at least three species and negatively affected commercial production of several Myrtaceae. To enable molecular and evolutionary studies into <em>A. psidii</em> pathogenicity, we assembled a highly contiguous genome for the pandemic biotype. With an estimated haploid genome size of just over 1 Gb (gigabases), it is the largest assembled fungal genome to date. The genome has undergone massive expansion via distinct transposable element (TE) bursts. Over 90% of the genome is covered by TEs predominantly belonging to the Gypsy superfamily. These TE bursts have likely been followed by deamination events of methylated cytosines to silence the repetitive elements. This in turn led to the depletion of CpG sites in transposable elements and a very low overall GC content of 33.8%. The overall gene content is highly conserved, when compared to other closely related Pucciniales, yet the intergenic distances are increased by an order of magnitude indicating a general insertion of TEs between genes. Overall, we show how transposable elements shaped the genome evolution of <em>A. psidii</em> and provide a greatly needed resource for strategic approaches to combat disease spread. Please cite the authors if using this data: https://academic.oup.com/g3journal/article/11/3/jkaa015/6007476</p>
Data from: Thoracic underreplication in Drosophila species estimates a minimum genome size and the dynamics of added DNA
Many cells in the thorax of <i>Drosophila </i>were found to stall during replication, a phenomenon known as underreplication. Unlike underreplication in nuclei of salivary and follicle cells, this stall occurs with less than one complete round of replication. This stall point allows precise estimations of early-replicating euchromatin and late-replicating heterochromatin regions, providing a powerful tool to investigate the dynamics of structural change across the genome. We measure underreplication in 132 species across the <i>Drosophila </i>genus and leverage this data to propose a model for estimating the rate at which additional DNA is accumulated as heterochromatin and euchromatin and also predict the minimum genome size for <i>Drosophila</i>. According to comparative phylogenetic approaches, the rates of change of heterochromatin differ strikingly between <i>Drosophila </i>subgenera. While these subgenera differ in karyotype, there were no differences by chromosome number, suggesting other structural changes may influence accumulation of heterochromatin. Measurements were taken for both sexes, allowing the visualization of genome size and heterochromatin changes for the hypothetical path of XY sex chromosome differentiation. Additionally, the model presented here estimates a minimum genome size in <i>Sophophora </i>remarkably close to the smallest insect genome measured to date, in a species over 200 million years diverged from <i>Drosophila</i>.
MicroCT data to 'Maximum CO2 diffusion inside leaves is limited by the scaling of cell size and genome size'
<p>This dataset is presented in the following publication. Please cite this publication if you use the dataset.</p> <p><em>Théroux-Rancourt Guillaume, Roddy Adam B., Earles J. Mason, Gilbert Matthew E., Zwieniecki Maciej A., Boyce C. Kevin, Tholen Danny, McElrone Andrew J., Simonin Kevin A. and Brodersen Craig R. 2021. <strong>Maximum CO<sub>2</sub> diffusion inside leaves is limited by the scaling of cell size and genome size.</strong> Proceedings of the Royal Society B. 288: 20203145. doi:<a href="https://doi.org/10.1098/rspb.2020.3145">10.1098/rspb.2020.3145</a></em></p> <p>We collected leaf samples from botanical gardens, greenhouses, and field sites, which were then transported to one of three synchrotron-based microCT beamlines for imaging. To- and three-dimensional data were then extracted from the microCT images to characterize the structural anatomy of the different study species. Those data were then analyzed and correlated with genome size data available from the Kew Plant DNA C-values database (https://cvalues.science.kew.org) or newly collected by us.</p> <p>From the botanical garden collections we selected representative species from across the vascular plant phylogeny, aiming to capture as much variation in anatomical/physiological/ecological traits as possible while also sampling species that represented important divergences in the phylogeny. Within clades, we also sampled species that spanned ecological breadth (e.g. xerophytic ferns). We focused solely on C3 terrestrial vascular plants, meaning that we did not sample C4 or CAM species, which have different photosynthetic biochemistry and associated anatomy.</p> <p>In most cases, single microCT scans were used per species. This constraint was primarily due to the extremely limited amount of time available at the microCT facility, as well as the labor intensive process of producing the final volume renderings and data analysis.</p> <p>MicroCT scans were collected by GTR, JME, ABR, CRB, AJM, CKB, MJZ, and DT at one of the three microCT beamlines from fresh leaf material. Leaves were cut at the base of the petiole or short stem segment, the cut end was wrapped in wet paper towels, and the entire shoot immediately put in a plastic bag before being transported to the synchrotron and scanned within 36 h of excision. Samples were prepared before each scan (less than 30 min) by excising a small sample that was then enclosed between two pieces of Kapton (polyimide) tape to prevent desiccation while allowing high X-ray transmittance.</p> <p>All microCT data were collected at the Lawrence Berkeley National Laboratory (LBNL) tomography beamline 8.3.2, the Swiss Light Source (SLS) TOMCAT Tomography beamline of the Paul Scherrer Institute, or the Advanced Photon Source tomography beamline 2-BM-A,B of Argonne National Laboratory (ANL). MicroCT datasets were reconstructed from the raw projection images obtained using TomoPy, an open-source Python-based framework for reconstructing tomographic data (LBNL), or using the in-house reconstruction platform of the beamline (SLS, ANL). LBNL and SLS data can be reconstruction using TomoPy, and this software is available at the following link: <a href="http://microct.lbl.gov/software">http://microct.lbl.gov/software</a></p> <p>Image stacks were segmented using the open-source software ImageJ either manually or using an automated machine learning algorithm (Théroux-Rancourt et al., 2020: doi:<a href="https://doi.org/10.1002/aps3.11380">10.1002/aps3.11380</a>). Traits were extracted from the segmented image stacks using ImageJ, the BoneJ plugin of ImageJ, or with an open-source Python-program available at: <a href="https://github.com/plant-microct-tools/leaf-traits-microct">https://github.com/plant-microct-tools/leaf-traits-microct</a>. Because of the multiple synchrotron and multiple sessions involved in acquiring this dataset, and because the analysis process lasted several years and was carried out by several persons, the size of each image stack and the details in each stack vary. What is present in each stack is the airspace and the whole mesophyll (i.e. the leaf without the epidermis), and in the majority of the cases, the vasculature is also segmented.</p>
Data from: Genomic dissection of variation in clutch size and egg mass in a wild great tit (Parus major) population
Clutch size and egg mass are life history traits that have been extensively studied in wild bird populations, as life history theory predicts a negative trade-off between them, either at the phenotypic or genetic level. Here, we analyse the genomic architecture of these heritable traits in a wild great tit (Parus major) population, using three marker-based approaches - chromosome partitioning, quantitative trait locus (QTL) mapping and a genome-wide association study (GWAS). The variance explained by each great tit chromosome scales with predicted chromosome size, no location in the genome contains genome-wide significant QTL, and no individual SNPs are associated with a large proportion of phenotypic variation, all of which may suggest that variation in both traits is due to many loci of small effect, located across the genome. There is no evidence that any regions of the genome contribute significantly to both traits, which combined with a small, non-significant, negative genetic covariance between the traits, suggests the absence of genetic constraints on the independent evolution of these traits. Our findings support the hypothesis that variation in life history traits in natural populations is likely to be determined by many loci of small effect spread throughout the genome, which are subject to continued input of variation by mutation and migration, although we cannot exclude the possibility of an additional input of major effect genes influencing either trait.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.