Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
486
datasets available to search
ShareScore release 0.9.0
Dataset results
486 results for “hierarchic”
Hierarchical multi-grain models improve descriptions of species' environmental associations, distribution, and abundance
<p>The characterization of species' environmental niches and spatial distribution predictions based on them are now central to much of ecology and conservation, but implicitly requires decisions about the appropriate spatial scale (i.e. <i>grain</i>) of analysis. Ecological theory and empirical evidence suggest that range-resident species respond to their environment at two characteristic, hierarchical spatial grains: (i) <i>response grain</i>, the (relatively fine) grain at which an individual uses environmental resources, and (ii) <i>occupancy grain</i>,<i> </i>the (relatively coarse) grain equivalent to a typical home range. We use a multi-grain (MG) occupancy model, aided by fine-grain remotely sensed imagery, to simultaneously estimate species-environment associations at both grains, conduct grain optimization to measure response grain, and apply this analysis framework to an example species: a medium-sized bird (<i>Tockus deckeni</i>) in a heterogeneous East African landscape. Based on home range analysis of movement data, we calculate an occupancy grain of 1km for <i>T. deckeni</i>. Using a grain optimization procedure across 32 grains from 10m to 500m, we identify 60m as the most strongly supported response grain for a suite of environmental variables, slightly coarser than opportunistic behavioral observations would have suggested. Validation confirms that the accuracy of the optimized MG occupancy model substantially exceeds that of equivalent single-grain (SG) occupancy models. We further use a simulation approach to assess the potential impacts of accounting for the multi-scale structure of species' environmental requirements on estimates of population size. We find that the more strongly supported MG approach consistently predicts a minimum population sizes in the study landscape that is much lower than that provided by the SG model. This suggests that SG approaches commonly used in conservation applications could lead to overly optimistic abundance and population estimates and that the MG approach may be more appropriate for supporting species conservation goals. More generally, we conclude that multi-grain approaches of the sort presented, and increasingly enabled by growing high-resolution remotely sensed data, hold great promise for offering a more mechanistic framework for assessing the appropriate grain(s) for population monitoring and management and enable more reliable estimates of abundances and species' distributions.</p>
Data from: Disentangling elevational richness: a multi-scale hierarchical Bayesian occupancy model of Colorado ant communities
Understanding the forces that shape the distribution of biodiversity across spatial scales is central in ecology and critical to effective conservation. To assess effects of possible richness drivers, we sampled ant communities on four elevational transects across two mountain ranges in Colorado, USA, with seven or eight sites on each transect and twenty repeatedly sampled pitfall trap pairs at each site each for a total of 90 days. With a multi-scale hierarchical Bayesian community occupancy model, we simultaneously evaluated the effects of temperature, productivity, area, habitat diversity, vegetation structure, and temperature variability on ant richness at two spatial scales, quantifying detection error and genus-level phylogenetic effects. We fit the model with data from one mountain range and tested predictive ability with data from the other mountain range. In total, we detected 105 ant species, and richness peaked at intermediate elevations on each transect. Species-specific thermal preferences drove richness at each elevation with marginal effects of site-scale productivity. Trap-scale richness was primarily influenced by elevation-scale variables along with a negative impact of canopy cover. Soil diversity had a marginal negative effect while daily temperature variation had a marginal positive effect. We detected no impact of area, land cover diversity, trap-scale productivity, or tree density. While phylogenetic relationships among genera had little influence, congeners tended to respond similarly. The hierarchical model, trained on data from the first mountain range, predicted the trends on the second mountain range better than multiple regression, reducing root mean squared error up to 65%. Compared to a more standard approach, this modeling framework better predicts patterns on a novel mountain range and provides a nuanced, detailed evaluation of ant communities at two spatial scales.
River dams and the stability of bird communities: A hierarchical Bayesian analysis in a tropical hydroelectric power plant
<ol> <li>The effects of anthropogenic disturbance upon the stability of wildlife communities depend on the heterogeneity and connectivity of habitat remnants on multiple scales. The number of hydroelectric dams in biodiversity hotspots (Africa, South America and Asia) is growing rapidly. To establish their environmental impact, it is essential to understand the dynamics of wildlife communities before and following the establishment of dams.</li> <li>We evaluated the impacts of the filling of the Serra do Facão hydroelectric reservoir in the São Marcos river, central Brazil, upon the bird community. Using data from 1,145 surveys across 20 sampling sites over eight years, two years before and six years after the filling of the reservoir, we assessed the resistance, i.e., maintenance close to an equilibrium state during the disturbance, and resilience, i.e., ability to return to the original state following the disturbance, of the bird community. We used spatiotemporal hierarchical Bayesian models to assess the effects of reservoir filling on five community parameters: abundance, richness, phylogenetic diversity, functional diversity and species composition.</li> <li>In the period subsequent to reservoir filling, there was (i) a marked reduction in bird abundance, richness, phylogenetic diversity and functional diversity, and (ii) a reduction in the proportion of forest species, coupled with an increase in the proportion of savanna species. Except for bird abundance, none of the other community attributes returned to their original levels, even after six years. Our findings indicate that Cerrado bird communities have both low resistance and low resilience to habitat loss associated with the establishment of hydroelectric reservoirs.</li> <li> <i>Synthesis and applications.</i> The environmental costs of hydroelectric dams are still underestimated or neglected in Brazil. A new paradigm in the assessment of their environmental impacts is warranted, incorporating (i) models of spatiotemporal variations based on long-term monitoring with surveys initiated before disturbances and (ii) measures of functional and phylogenetic diversity, such that society can understand the costs and benefits of the establishment of new hydroelectric dams and make informed decisions. Biodiversity loss could be minimized by ensuring the preservation and connectivity of alluvial habitats, capable of maintaining the supply of resources and the functional and phylogenetic attributes of bird communities associated with such habitats.</li> </ol>
Experimental Data for: Hierarchical Software Landscape Visualization for System Comprehension: A Controlled Experiment
<p>In many enterprises the number of deployed applications is constantly increasing. Those applications - often several hundreds - form large software landscapes. The comprehension of such landscapes is frequently impeded due to, for instance, architectural erosion, personnel turnover, or changing requirements. Therefore, an efficient and effective way to comprehend such software landscapes is required. The current state of the art often visualizes software landscapes via flat graph-based representations of nodes, applications, and their communication.</p> <p>In our ExplorViz visualization, we introduce hierarchical abstractions aiming at solving typical system comprehension tasks fast and accurately for large software landscapes. To evaluate our hierarchical approach, we conduct a controlled experiment comparing our hierarchical landscape visualization to a flat, state-of-the-art visualization. In addition, we thoroughly analyze the strategies employed by the participants and provide a package containing all our experimental data to facilitate the verifiability, reproducibility, and further extensibility of our results.</p> <p>We observed a statistically significant increase of 14 % in task correctness of the hierarchical visualization group compared to the flat visualization group in our experiment. The time spent on the system comprehension tasks did not show any significant differences. The results backup our claim that our hierarchical concept enhances the current state of the art in landscape visualization.</p> <p>This package contains our experimental data.</p>
Data, code and supplementary plots for "A hierarchical spline model for correcting and hindcasting temperature data"
<p>Data, code and supplementary plots for the paper "A hierarchical spline model for correcting and hindcasting temperature data". Please see README.txt for detailed description of the files and relevant instructions.</p>
Hierarchical reject datasets
<p>Datasets used in the paper 'Uncertainty-aware single-cell annotation with a hierarchical reject option'. This paper uses 5 open-source datasets:</p><p>1. <strong>The Allen Mouse Brain (AMB) dataset</strong> [1]: Filtered_mouse_allen_brain_labels.csv and Filtered_mouse_allen_brain_data.csv</p><p>2. <strong>The COVID dataset</strong> [2]: CocidBALLabel.csv and CovidBALCounts.csv</p><p>3. <strong>The Azimuth PBMC dataset</strong> [3]: pbmc.multimodal.h5ad</p><p>4. from the Flyatlas [4] <strong>the Flyhead dataset</strong>: Flyatlas_Fbbt_head.csv, Flyatlas_head_10x.loom and Flyatlas_Labels_head.csv</p><p>5. from the Flyatlas [4]<i> <strong>the Flybody dataset</strong></i>: Flyatlas_Fbbt_body.csv, Flyatlas_body_10x.loom and Flyatlas_Labels_body.csv</p><p> </p><p>(All the credits of these datasets go to the original creators of the datasets.)</p><p> </p><p><strong>References</strong></p><p>[1] Tasic, B. et al. (2018). Shared and distinct transcriptomic cell types across neocortical areas. Nature, 563 (7729), 72–78. https://doi.org/10.1038/s41586-018-0654-5</p><p>[2] Chan Zuckerberg Initiative Single-Cell COVID-19 Consortia et al. (2020). Single cell profiling ofCOVID-19 patients: an international data resource from multiple tissues. Medrxiv preprint. https://doi.org/10.1101/2020.11.20.20227355</p><p>[3] Stuart, T. et al. (2019). Comprehensive Integration of Single-Cell Data. Cell, 177(7), 1888–1902.e21. https://doi.org/10.1016/j.cell.2019.05.031</p><p>[4] Li, H. et al. (2022). Fly Cell Atlas: A single-nucleus transcriptomic atlas of the adult fruit fly. Science, 375(6584), eabk2432. https://doi.org/10.1126/science.abk2432</p>
A hierarchical model for eDNA fate and transport dynamics accommodating low concentration samples
<p>Environmental DNA (eDNA) sampling is an increasingly important tool for answering ecological questions and informing aquatic species management . Challenges of using eDNA include determining species source location(s) and accurately and precisely measuring low concentration eDNA samples, especially considering inhibitory compounds and multiple sources of ecological and measurement variability. These challenges must be overcome to optimize our use of modeling frameworks like the eDNA Integrating Transport and Hydrology (eDITH) model. To better understand eDNA fate and transport dynamics, our ability to estimate parameters within the eDITH framework, and our ability to reliably quantify low concentration samples, we developed a hierarchical model and used it to evaluate a fate and transport experiment. Our model addresses several low concentration challenges by modeling the number of copies in each PCR replicate as latent variables with a count distribution and conditioning detection and quantification on replicate copy number. We provide evidence that the eDNA removal rate was not constant through time, estimating that over 80% of eDNA was removed over the first 10 m, traversed in 41 seconds. After this initial period of rapid decay, eDNA decayed slowly with consistent detection through our furthest site 1km from the release location, traversed in 250 seconds. We show that the eDITH model parameters can be difficult to estimate in this scenario. Our model further allowed us to detect extra-Poisson variation in the allocation of copies to replicates. Despite not observing evidence for inhibition as typically quantified using internal positive controls in conjunction with a binary decision rule (e.g., $\Delta$Cq>3), we hypothesized this overdispersion could be due to inhibitors. We extended our hierarchical model to accommodate a continuous effect of inhibitors, and used our model to provide evidence for the inhibitor hypothesis and explore the implications, if true. We show that inhibitors can cause substantial underestimation of eDNA site concentration, bias eDITH model parameter estimates, and attribute measurement variability erroneously to ecological variability. While our model is not a panacea for all challenges faced when quantifying low eDNA concentrations, it provides a framework for a more complete accounting of uncertainty that can be further tested and refined.</p>
Dataset of Paper "Optimization and parallelization of the Discrete Ordinate Method for radiation transport simulation in OpenFOAM: Hierarchical combination of shared and distributed memory approaches"
<p>Dataset of Paper "Optimization and parallelization of the Discrete Ordinate Method for radiation transport simulation in OpenFOAM: Hierarchical combination of shared and distributed memory approaches":</p> <ul> <li>Time profiling of the DOM model stages.</li> <li>Speed-up and parallel fraction of the “global results generation” stage in the DOM model.</li> <li>Comparison of computational time between original and modified DOM model stages.</li> <li>Scalability of the peer-to-peer communication (OpenFOAM) and master-slave communication (ANSYS Fluent) architectures.</li> <li>Results of the benchmarking of the model with ANSYS Fluent in three reactors.</li> <li>Mesh of the jerrycan.</li> <li>Mesh of the anular reactor</li> <li>Mesh of the tubular reactor couple to a compound parabolic collector.</li> </ul>
Fast Hierarchical Games for Image Explanations
<p>As modern complex neural networks keep breaking records and solving harder problems, their predictions also become less and less intelligible. The current lack of interpretability often undermines the deployment of accurate machine learning tools in sensitive settings. In this work, we present a model-agnostic explanation method for image classification based on a hierarchical extension of Shapley coefficients–Hierarchical Shap (h-Shap)–that resolves some of the limitations of current approaches. Unlike other Shapley-based explanation methods, h-Shap is scalable and can be computed without the need of approximation. Under certain distributional assumptions, such as those common in multiple instance learning, h-Shap retrieves the exact Shapley coefficients with an exponential improvement in computational complexity. We compare our hierarchical approach with popular Shapley-based and non-Shapley-based methods on a synthetic dataset, a medical imaging scenario, and a general computer vision problem, showing that h-Shap outperforms the state of the art in both accuracy and runtime. Code and experiments are made publicly available.</p>
Raw datasets and media accompanying the manuscript: Homogenous high enhancement surface-enhanced Raman scattering (SERS) substrates by simple hierarchical tuning of gold nanofoams
<p>Raw datasets and media accompanying the manuscript: Homogenous high enhancement surface-enhanced Raman scattering (SERS) substrates by simple hierarchical tuning of gold nanofoams</p>
Dataset and R code from: Positive and negative effects of land abandonment on butterfly communities revealed by a hierarchical sampling design across climatic regions
<p class="MsoNormal">Land abandonment may decrease biodiversity but also provides an opportunity for rewilding. It is therefore necessary to identify areas that may benefit from traditional land management practices and those that may benefit from a lack of human intervention. In this study, we conducted comparative field surveys of butterfly occurrence in abandoned and inhabited settlements in 18 regions of diverse climatic zones in Japan to test the hypotheses that species-specific responses to land abandonment correlate with climatic niches and habitat preferences. Hierarchical models that unified species occurrence and habitat preferences revealed that negative responses to land abandonment were associated with species that have cold climatic niches and utilize open habitats, suggesting that species negatively impacted by land abandonment will decline more due to future climate warming. Maps representing species gains and losses due to land abandonment, which were created from the model estimates, showed similar geographic patterns, but some areas exhibited high species losses relative to gains. Our hierarchical modelling approach was useful for scaling up local-scale effects of land abandonment to a macro-scale assessment, which is crucial to developing spatial conservation strategies in the era of depopulation.</p>
Data and code for the article "Hierarchical tensile structures with ultralow mechanical dissipation"
<p>Data and code for the article "Hierarchical tensile structures with ultralow mechanical dissipation", consisting of all the ringdown measurements, spectra and codes for data analysis and generation of the fabrication GDS masks.</p>
Data from: Hierarchical variation in phenotypic flexibility across timescales and associated survival selection shape the dynamics of partial seasonal migration
<p>Population responses to environmental variation ultimately depend on within-individual and among-individual variation in labile phenotypic traits that affect fitness, and resulting episodes of selection. Yet, complex patterns of individual phenotypic variation arising within and between time periods, and associated variation in selection, have not been fully conceptualised or quantified. We highlight how structured patterns of phenotypic variation in dichotomous threshold traits can theoretically arise and experience varying forms of selection, shaping overall phenotypic dynamics. We then fit novel multistate models to ten years of band-resighting data from European shags to quantify phenotypic variation and selection in a key threshold trait underlying spatio-seasonal population dynamics: seasonal migration versus residence. First, we demonstrate substantial among-individual variation alongside substantial between-year individual repeatability in within-year phenotypic variation ('flexibility'), with weak sexual dimorphism. Second, we demonstrate that between-year individual variation in within-year phenotypes ('supraflexibility') is structured and directional, consistent with the threshold trait model. Third, we demonstrate strong survival selection on within-year phenotypes, and hence on flexibility, that varies across years and sexes, including episodes of disruptive selection representing costs of flexibility. By quantitatively combining these results, we show how supraflexibility and survival selection on migratory flexibility jointly shape population-wide phenotypic dynamics of seasonal movement.</p>
Supplementary data for 'scROSHI - robust supervised hierarchical identification of single cells'
<p>This record contains some of the supplementary files for the manuscript "scROSHI - robust supervised hierarchical identification of single cells" by Prummer et al.</p> <p>The file "scROSHI_Fig03_counts.zip" contains SingleCellExperiment R objects as RDS files, "sce_A.RDS", "sce_B.RDS", "sce_C.RDS", corresponding to the three samples shown in Figure 3. The 'assay' slot of the SingleCellExperiment objects contains the gene x cell raw count matrix, the 'colData' slot contains the description for each cell: barcode, celltype_major, celltype_final, cnv_status.</p> <p>Please see https://github.com/ETH-NEXUS/scROSHI for code generating the cell type labels.</p>
Training and test data, plus saved models for the paper "Top-down effects in an early visual cortex inspired hierarchical Variational Autoencoder" submitted to the SVRHM 2022 Workshop @ NeurIPS
<p>Each .pkl file contains a training or test dataset in the form of a Python dictionary (generated with Python 3.8.5) with the following fields:</p><ul><li>'train_images': 640,000 float32 images used for model training. 20px images contain 400 pixel intensities, 40px images contain 1600 pixel intensities each.</li><li>'train_labels': float32 labels for each image in 'train_images'. All natural images are labeled with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0, according to their texture family.</li><li>'test_images': 64,000 float32 images used for model testing. 20px images contain 400 pixel intensities, 40px images contain 1600 pixel intensities each.</li><li>'test_labels': float32 labels for each image in 'test_images'. All natural images are labeled with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0, according to their texture family.</li></ul><p>Each .zip file contains a saved model. Details on these are coming soon.</p><p>For more details, see the paper "Top-down effects in an early visual cortex inspired hierarchical Variational Autoencoder" published at the SVRHM 2022 Workshop @ NeurIPS (<a href="https://openreview.net/forum?id=8dfboOQfYt3">link</a>).</p>
Predicting compound-protein interaction using hierarchical graph convolutional networks
<p>This repository contains the datasets which are used in the article "Predicting Compound-Protein Interaction using Hierarchical Graph Convolutional Networks".</p>
Mielke & Carvalho 2022 Chimpanzee play sequences are structured hierarchically as games - Data
<p>Data and scripts for the 2022 manuscript 'Chimpanzee play sequences are structured hierarchically as games' - preprint here: </p> <p>https://doi.org/10.1101/2022.06.14.496075</p> <p>Dataset and scripts generated on 20/09/2022. For potential changes and all information see:</p> <p>https://github.com/AlexMielke1988/Mielke-Carvalho_Chimpanzee-Play</p>
Figure 2. Hierarchical access to the image DB-Access Management in Medical Image Databases Based on New Format and Contents Protection with Inverse Pyramid Decomposition
<p>The structure of the hierarchical access to the image database contents is shown on Fig. 2.</p>
Replicate analysis from: Measuring complexity for hierarchical models using effective degrees of freedom
<p>Hierarchical models can express ecological dynamics using a combination of fixed and random effects, and measurement of their complexity (effective degrees of freedom, EDF) requires estimating how much random effects are shrunk towards a shared mean. Estimating EDF is helpful to (1) penalize complexity during model selection and (2) to improve understanding of model behavior. I apply the conditional Akaike Information Criterion (cAIC) to estimate EDF from the finite-difference approximation to the gradient of model predictions with respect to each datum. I confirm that this has similar behavior to widely used Bayesian criteria, and I illustrate ecological applications using three case studies. The first compares model parsimony with or without time-varying parameters when predicting density-dependent survival, where cAIC favors time-varying demographic parameters more than conventional AIC. The second estimates EDF in a phylogenetic structural equation model, and identifies a larger EDF when predicting longevity than mortality rates in fishes. The third compares EDF for a species distribution model (SDM) fitted for twenty bird species and identifies those species requiring more model complexity. These highlight the ecological and statistical insight from comparing EDF among experimental units, models, and data partitions, using an approach that can broadly adopted for nonlinear ecological models.</p>
Thalamocortical interactions shape hierarchical neural variability during stimulus perception dataset
<p>Dataset used in the Thalamocortical interactions shape hierarchical neural variability during stimulus perception article.</p> <p> </p> <p>Dataset contains neural activity recordings of a vibrotactile detection task recorded in four monkeys in the following areas: somatosensory thalamus (VPL), 3b and area 1 of the somatosensory cortex (S1)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.