Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
27
datasets available to search
ShareScore release 0.9.0
Dataset results
27 results for “site heterogeneity”
Transcription start site analysis for heterogenous CD4+ T cells using 5′ scRNA-seq
<p>These datasets are generated by ReapTEC (read-level pre-filtering and transcribed enhancer call) using 5' single-cell RNA-seq data on human heterogenous CD4+ T cells. By taking advantage of a unique "cap signature" derived from the 5′-end of a transcript, ReapTEC simultaneously profiles gene expression and enhancer activity at nucleotide resolution using 5′-end single-cell RNA-sequencing (5′ scRNA-seq). The detail of ReapTEC pipeline is described in https://github.com/MurakawaLab/ReapTEC.</p>
Transcription start site analysis for heterogenous CD4+ T cells using 5′ scRNA-seq
Open the record for dataset details and reuse information.
Thermophysical properties and surface heterogeneity of landing sites on Mars from overlapping THEMIS observations-Data
<p>Files, scripts/functions, observed temperature data, and modeled results used in the Ahern et al. paper entitled: Thermophysical properties and surface heterogeneity of landing sites on Mars from overlapping THEMIS observations.</p>
Data sets for phylogenomic analyses in: Ant backbone phylogeny resolved by modelling compositional heterogeneity among sites in genomic data
<p>Ants are the most ubiquitous and ecologically dominant arthropods on Earth, and understanding their phylogeny is crucial for deciphering their character evolution, species diversification, and biogeography. Although recent genomic data have shown promise in clarifying intrafamilial relationships across the tree of ants, inconsistencies between molecular datasets have also emerged. Here I re-examine the most comprehensive published Sanger-sequencing and genome-scale datasets of ants using model comparison methods that model among-site compositional heterogeneity to understand the sources of conflict in phylogenetic studies. My results under the best-fitting model, selected on the basis of Bayesian cross-validation and posterior predictive model checking, identify contentious nodes in ant phylogeny whose resolution is <a>modelling-dependent. </a>I show that the Bayesian infinite mixture CAT model outperforms empirical finite mixture models (C20, C40 and C60) and that, under the best-fitting CAT-GTR+G4 model, the enigmatic <a><em>Martialis</em> </a><em>heureka</em> is sister to all ants except Leptanillinae, rejecting the more popular hypothesis supported under worse-fitting models, that place it as sister to Leptanillinae. These analyses resolve a lasting controversy in ant phylogeny and highlight the significance of model comparison and adequate modelling of among-site compositional heterogeneity in reconstructing the deep phylogeny of insects.</p>
DA Diffusion Simulation for: A fluorescent nanosensor paint reveals the heterogeneity of dopamine release from neurons at individual release sites
<p>Code developed for the revision of:</p> <p>Elizarova, S., Chouaib, A., Shaib, A., Mann, F., Brose, N., Kruss, S., & Daniel, J. A. (2021). A fluorescent nanosensor paint reveals the heterogeneity of dopamine release from neurons at individual release sites. BioRxiv, 2021.03.28.437019. https://doi.org/10.1101/2021.03.28.437019</p>
Data sets for phylogenomic analyses in: Ant backbone phylogeny resolved by modelling compositional heterogeneity among sites in genomic data
Open the record for dataset details and reuse information.
Data from: Occupancy models for data with false positive and false negative errors and heterogeneity across sites and surveys
False positive detections, such as species misidentifications, occur in ecological data, although many models do not account for them. Consequently, these models are expected to generate biased inference. The main challenge in an analysis of data with false positives is to distinguish false positive and false negative processes while modeling realistic levels of heterogeneity in occupancy and detection probabilities without restrictive assumptions about parameter spaces. Building on previous attempts to account for false positive and false negative detections in occupancy models, we present hierarchical Bayesian models that utilize a subset of data with either confirmed detections of a species' presence (CP model) or both confirmed presences and confirmed absences (CACP model). We demonstrate that our models overcome the challenges associated with false positive data by evaluating model performance in Monte Carlo simulations of a variety of scenarios. Our models also have the ability to improve inference by incorporating previous knowledge through informative priors. We describe an example application of the CP model to quantify the relationship between songbird occupancy and residential development, plus we provide instructions for ecologists to use the CACP and CP models in their own research. Monte Carlo simulation results indicated that, when data contained false positive detections, the CACP and CP models generated more accurate and precise posterior probability distributions than a model that assumed data did not have false positive errors. For the scenarios we expect to be most generally applicable, those with heterogeneity in occupancy and detection, the CACP and CP models generated essentially unbiased posterior occupancy probabilities. The CACP model with vague priors generated unbiased posterior distributions for covariate coefficients. The CP model generated unbiased posterior distributions for covariate coefficients with vague or informative priors, depending on the function relating covariates to occupancy probabilities. We conclude that the CACP and CP models generate accurate inference in situations with false positive data for which previous models were not suitable.
Data from: Mixture models of nucleotide sequence evolution that account for heterogeneity in the substitution process across sites and across lineages
Molecular phylogenetic studies of homologous sequences of nucleotides often assume that the underlying evolutionary process was globally stationary, reversible and homogeneous (SRH), and that a model of evolution with one or more site-specific and time-reversible rate matrices (e.g., the GTR rate matrix) is enough to accurately model the evolution of data over the whole tree. However, an increasing body of data suggests that evolution under these conditions is an exception, rather than the norm. To address this issue, several non-SRH models of molecular evolution have been proposed, but they either ignore heterogeneity in the substitution process across sites (HAS) or assume it can be modelled accurately using the Γ distribution. As an alternative to these models of evolution, we introduce a family of mixture models that approximate HAS without the assumption of an underlying predefined statistical distribution. This family of mixture models is combined with non-SRH models of evolution that account for heterogeneity in the substitution process across lineages (HAL). We also present two algorithms for searching model space and identifying an optimal model of evolution that is less likely to over- or under-parameterize the data. The performance of the two new algorithms was evaluated using alignments of nucleotides with 10,000 sites simulated under complex non-SRH conditions on a 25-tipped tree. The algorithms were found to be very successful, identifying the correct HAL model with a 75% success rate (the average success rate for assigning rate matrices to the tree's 48 edges was 99.25%) and, for the correct HAL model, identifying the correct HAS model with a 98% success rate. Finally, parameter estimates obtained under the correct HAL-HAS model were found to be accurate and precise. The merits of our new algorithms were illustrated with an analysis of 42,337 second codon sites extracted from a concatenation of 106 alignments of orthologous genes encoded by the nuclear genomes of Saccharomyces cerevisiae, S. paradoxus, S. mikatae, S. kudriavzevii, S. castellii, S. kluyveri, S. bayanus, and Candida albicans. Our results show that second codon sites in the ancestral genome of these species contained 49.1% invariable sites, 39.6% variable sites belonging to one rate category (V1), and 11.3% variable sites belonging to a second rate category (V2). The ancestral nucleotide content was found to differ markedly across these 3 sets of sites, and the evolutionary processes operating at the variable sites were found to be non-SRH and best modelled by a combination of 8 edge-specific rate matrices (4 for V1 and 4 for V2). The number of substitutions per site at the variable sites also differed markedly, with sites belonging to V1 evolving slower than those belonging to V2 along the lineages separating the 7 species of Saccharomyces. Finally, sites belonging to V1 appeared to have ceased evolving along the lineages separating S. cerevisiae, S. paradoxus, S. mikatae, S. kudriavzevii, and S. bayanus, implying that they might have become so selectively constrained that they could be considered invariable sites in these species.
Data for: Consistently heterogeneous structures observed at multiple spatial scales across fire-intact reference sites
<p>Geospatial polygons representing fire-suppressed control sites against which fire-intact reference sites were compared in Chamberlain et al. (2023). Control sites represent areas with 1) no record of fire history, 2) no record of late 20th century or early 21st century timber management, and 3) no "Fast Change" detected by the Landscape Change Monitoring System dataset. All sites are predominantly within the yellow pine and mixed-conifer zone of California's Sierra Nevada, USA. Polygon boundaries were defined using the NHDPlusV2 catchments, and were manually reshaped using aerial imagery to ensure that polygons were > 100 ha, represented primarily forested areas, and excluded major roads, infrastructure, and major rock outcrops.</p> <p>Detailed description of the methods used to produce this dataset provided in:<br> Chamberlain, C.P., Cova, G.R., Cansler, C.A., North, M.P., Meyer, M.D., Jeronimo, S.M.A., Kane, V.R., 2023. Consistently heterogeneous structures observed at multiple spatial scales across fire-intact reference sites. Forest Ecology and Management.</p>
Data from: Mixture models of nucleotide sequence evolution that account for heterogeneity in the substitution process across sites and across lineages
Open the record for dataset details and reuse information.
Data from: Occupancy models for data with false positive and false negative errors and heterogeneity across sites and surveys
Open the record for dataset details and reuse information.
data from "The chloroplast land plant phylogeny: analyses employing better-fitting tree- and site-heterogeneous composition models"
<p>The nucleotide, codon-degenerate and protein concatenated alignments of 83 chloroplast genes, with the respective character sets, in nexus format.</p>
Snow depth mapping with an Unmanned Aerial Vehicles in heterogeneous mountain site (Izas Catchment, Pyrenees)
<p>Recent developments in unmanned aircraft vehicles (UAV) and Structure for Motion (SfM photogrammetry or simply SfM) algorithms have proved their worth for determining snow depth distribution. This dataset presents UAV and Terrestrial Laser Scanner (TLS) snow depth obsrevations obtained in complex alpine terrain in the Pyrenees. UAV observations have been acqired with a fixed-wing UAV working in RTK mode with an RGB camera. During the 2018-19 season, seven field campaigns (13 UAV flights) were undertaken covering 0.48 km2 . Several UAV observations were obtained under different light conditions and flight block configurations (altitude and image overlaps) in the same day with the aim if evaluating UAV observations when compared ti a well-established close range remote sensing technique (TLS).</p>
Data from: Modeling site heterogeneity with posterior mean site frequency profiles accelerates accurate phylogenomic estimation
Proteins have distinct structural and functional constraints at different sites that lead to site-specific preferences for particular amino acid residues as the sequences evolve. Heterogeneity in the amino acid substitution process between sites is not modeled by commonly used empirical amino acid exchange matrices. Such model misspecification can lead to artefacts in phylogenetic estimation such as long-branch attraction. Although sophisticated site-heterogeneous mixture models have been developed to address this problem in both Bayesian and maximum likelihood (ML) frameworks, their formidable computational time and memory usage severely limits their use in large phylogenomic analyses. Here we propose a posterior mean site frequency (PMSF) method as a rapid and efficient approximation to full empirical profile mixture models for ML analysis. The PMSF approach assigns a conditional mean amino acid frequency profile to each site calculated based on a mixture model fitted to the data using a preliminary guide tree. These PMSF profiles can then be used for in-depth tree-searching in place of the full mixture model. Compared with widely used empirical mixture models with k classes, our implementation of PMSF in IQ-TREE (http://www.iqtree.org) speeds up the computation by approximately k /1.5-fold and requires a small fraction of the RAM. Furthermore, this speedup allows, for the first time, full nonparametric bootstrap analyses to be conducted under complex site-heterogeneous models on large concatenated data matrices. Our simulations and empirical data analyses demonstrate that PMSF can effectively ameliorate long-branch attraction artefacts. In some empirical and simulation settings PMSF provided more accurate estimates of phylogenies than the mixture models from which they derive.
Data from: Invariant versus classical quartet inference when evolution is heterogeneous across sites and lineages
One reason why classical phylogenetic reconstruction methods fail to correctly infer the underlying topology is because they assume oversimplified models. In this paper we propose a quartet reconstruction method consistent with the most general Markov model of nucleotide substitution, which can also deal with data coming from mixtures on the same topology. Our proposed method uses phylogenetic invariants and provides a system of weights that can be used as input for quartet-based methods. We study its performance on real data and on a wide range of simulated 4-taxon data (both time-homogeneous and nonhomogeneous, with or without among-site rate heterogeneity, and with different branch length settings). We compare it to the classical methods of neighbor-joining (with paralinear distance), maximum likelihood (with different underlying models), and maximum parsimony. Our results show that this method is accurate and robust, has a similar performance to ML when data satisfies the assumptions of both methods, and outperforms the other methods when these are based on inappropriate substitution models. If alignments are long enough, then it also outperforms other methods when some of its assumptions are violated.
Data from: The relative importance of modeling site pattern heterogeneity versus partition-wise heterotachy in phylogenomic inference
Large taxa-rich genome-scale data sets are often necessary for resolving ancient phylogenetic relationships. But accurate phylogenetic inference requires that they are analyzed with realistic models that account for the heterogeneity in substitution patterns amongst the sites, genes and lineages. Two kinds of adjustments are frequently used: models that account for heterogeneity in amino acid frequencies at sites in proteins, and partitioned models that accommodate the heterogeneity in rates (branch lengths) among different proteins in different lineages (protein-wise heterotachy). Although partitioned and site-heterogeneous models are both widely used in isolation, their relative importance to the inference of correct phylogenies has not been carefully evaluated. We conducted several empirical analyses and a large set of simulations to compare the relative performances of partitioned models, site-heterogeneous models and combined partitioned site heterogeneous models. In general, site-homogeneous models (partitioned or not) performed worse than site heterogeneous, except in simulations with extreme protein-wise heterotachy. Furthermore, simulations using empirically-derived realistic parameter settings showed a marked long-branch attraction (LBA) problem for analyses employing protein-wise partitioning even when the generating model included partitioning. This LBA problem results from a small sample bias compounded over many single protein alignments. In some cases, this problem was ameliorated by clustering similarly-evolving proteins together into larger partitions using the PartitionFinder method. Similar results were obtained under simulations with larger numbers of taxa or heterogeneity in simulating topologies over genes. For an empirical Microsporidia test data set, all but one tested site-heterogeneous models (with or without partitioning) obtain the correct Microsporidia+Fungi grouping, whereas site-homogenous models (with or without partitioning) did not. The single exception was the fully partitioned site-heterogeneous analysis that succumbed to the compounded small sample LBA bias. In general unless protein-wise heterotachy effects are extreme, it is more important to model site-heterogeneity than protein-wise heterotachy in phylogenomic analyses. Complete protein-wise partitioning should be avoided as it can lead to a serious LBA bias. In cases of extreme protein-wise heterotachy, approaches that cluster similarly-evolving proteins together and coupled with site-heterogeneous models work well for phylogenetic estimation.
Joint identification of groundwater contamination source and heterogeneous hydrogeological parameters in LNAPL contaminated site based on deep convolutional encoder-decoder neural networks
Open the record for dataset details and reuse information.
Data from: The relative importance of modeling site pattern heterogeneity versus partition-wise heterotachy in phylogenomic inference
Open the record for dataset details and reuse information.
Data from: Modeling site heterogeneity with posterior mean site frequency profiles accelerates accurate phylogenomic estimation
Open the record for dataset details and reuse information.
Data from: Invariant versus classical quartet inference when evolution is heterogeneous across sites and lineages
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.