Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25
datasets available to search
ShareScore release 0.9.0
Dataset results
25 results for “disaggregation”
Parameters for PITRI Precipitation Temporal Disaggregation over continental US, Mexico, and southern Canada, 1981-2013
<p>This dataset contains parameter values for the Precipitation Isosceles Triangle (PITRI) precipitation disaggregation method (Bohn et al., 2019) over the CONUS+Mexico domain (southern Canada, the continental US, and Mexico; 14.65 - 53° N latitude, 65-125° W longitude), at 1/16° (6 km) spatial resolution. There are two parameters: "dur" (mean event duration [minutes]) and "t_pk" (mean time of peak precipitation intensity [minutes from beginning of day]). In each land grid cell, each parameter has 12 climatological mean monthly values for the period 1981-2013.</p> <p>This dataset contains 2 NetCDF-format files:</p> <ul> <li>domain.CONUS_MX.L2015.nc - this contains parameters over the entire CONUS+Mexico domain, using the land mask of the Livneh et al. (2015) daily meteorology dataset.</li> <li>domain.USMX.L2015.nc - this contains the same parameters, but clipped to exclude Canada (to be consistent with datasets that cover only that part of the domain).</li> </ul> <p>These files are structured as input "domain" files for 2 applications:</p> <ul> <li>MetSim meteorology simulator (https://github.com/UW-Hydro/MetSim/releases/tag/2.0.0_alpha; Bennett et al., 2018). The PITRI algorithm has been implemented as an option in MetSim. To use this algorithm within MetSim, set the "prec_type" option to "triangle" or "mix" in the configuration file. The "mix" option is a blend of the "uniform" (previous) method and the "triangle" method that fixes biases in snow accumulation rates yielded by the "triangle" method in some climates. "mix" uses the "uniform" method on days for which minimum daily temperature falls below 0 C, and "triangle" method on all other days.</li> <li>Variable Infiltration Capacity (VIC) model, release 5.0 and later (Liang et al., 1994; Hamman et al., 2018; https://github.com/UW-Hydro/VIC). VIC does not use the PITRI parameters, but does use the other variables such as mask, elevation, area, etc. For VIC to use the output of MetSim (disaggregated meteorological fields) as input, VIC needs to use the same domain file as was used in MetSim.</li> </ul> <p>Algorithm details can be found in the following paper, which should be cited if you use this dataset:</p> <p>Bohn, T. J., K. M. Whitney, G. Mascaro, and E. R. Vivoni, 2019: A deterministic approach for approximating the diurnal cycle of precipitation for use in large-scale hydrological modeling, Journal of Hydrometeorology 20(2), 297-317, doi: 10.1175/JHM-D-18-0203.1.</p>
Social Accounting Matrix for Lithuania, 2017 (with disaggregated Rubber and Plastics activity)
<p>The dataset is based on doi: 10.5281/zenodo.5077893 but includes Rubber and plastics activity disaggregated to depict production of plastic bags with more details.</p>
Social Matrix for Lithuania, 2020 with disaggregated CPA_C22
<p>The dataset contains the Lithuanians social accounting matrix for 2020 with disaggregated CPA_C22. The FIGARO database’s 2022 edition (Eurostat (2022). ESA supply, use and input-output tables) is used to create product-by-product input-output table for the EU, while Eurostat’s data on non-financial transactions (a dataset called nasa_10_nf_tr , Eurostat (2022). Non-financial transactions - annual data) is used to cover the remaining parts of the social accounting matrix.</p> <p>The dataset includes a baseline (Reference) scenario without the simulation of sustainability practices and three scenarios simulating the substitution of plastic bags by paper bags (PaperBags), bioplastic bags (BioPlastics) and the reduction of the use of plastic bags (ConsReduction).</p> <p> </p> <p>This research was funded by the grant S-MIP-20-53 from the Research Council of Lithuania.</p>
Text-fig. 4. Lepidocarpon cone in the process of disaggregating as part of the dispersal strategy of the plants. When preserved isolated, the sporophylls are assigned to the fossil-genus Lepidostrobophyllum. Refigured from Thomas (1981). Grovesend Formation (upper Asrturian – lower Moscovian), Kilmersdon Tip, Radstock Coalfield, UK; Natural History Museum (London) specimen V.60431. in Naming Of Parts: The Use Of Fossil-Taxa In Palaeobotany
Text-fig. 4. Lepidocarpon cone in the process of disaggregating as part of the dispersal strategy of the plants. When preserved isolated, the sporophylls are assigned to the fossil-genus Lepidostrobophyllum. Refigured from Thomas (1981). Grovesend Formation (upper Asrturian – lower Moscovian), Kilmersdon Tip, Radstock Coalfield, UK; Natural History Museum (London) specimen V.60431.
FIGURE 3. Picked bryozoan fragments from the disaggregated, sieved samples from the Finis Shale. A – profile B15 in Stenolaemate bryozoans from the Graham Formation, Pennsylvanian (Virgilian) at Lost Creek Lake, Texas, USA
FIGURE 3. Picked bryozoan fragments from the disaggregated, sieved samples from the Finis Shale. A – profile B15; B – profile B16; C – profile C5; D – profile C13.
Data for: When Correlation Matters: On Uncertainty Propagation In The Case Of Data Disaggregation
<p>This is the data repository for our study on “When Correlation Matters: On Uncertainty Propagation In The Case Of Data Disaggregation” submitted to the Journal of Industrial Ecology (JIE).</p> <p>I contains:</p> <ul> <li>The data behind the all numeric plots</li> <li>The data needed to reproduce the case-study results in the Supplementary Information (check out V1)</li> </ul> <p> </p> <p>To reproduce our results you need to download the files from this repo, our code from Github (https://github.com/simschul/uncertainty_disaggregation) and put the data into the `./data` folder. </p> <p>The data is an intermediate output from an earlier study: https://essd.copernicus.org/articles/16/2669/2024/essd-16-2669-2024.html </p> <p>For more information how those intermediate results were created please refer to the paper and the code (https://github.com/simschul/uncertainty_GHG_accounts)</p> <p> </p> <p> </p>
Disaggregated soil map of the Designation of Origin Campo de Borja
<p>Disaggregated soil map of the Designation of Origin Campo de Borja in GeoPackage format. It contains a single-polygon vector layer and a table of the map legend.</p>
HydroWIRES B1: Monthly and Weekly Hydropower Constraints Based on Disaggregated EIA-923 Data
<p>This dataset provides both monthly and weekly constraints (maximum and minimum generation) and power targets for hundreds of hydropower plants across the United States. The data is intended for use in Production Cost Models (PCMs) and Capacity Expansion Models (CEMs). The hydropower data is based on disaggregated annual power data which is part of the EIA-923 dataset.</p> <p>The code to reproduce this data is available here: <a href="https://github.com/HydroWIRES-PNNL/B1-data">https://github.com/HydroWIRES-PNNL/B1-data</a></p> <p>The original disaggregation procedure is detailed here:</p> <p><a href="https://www.nature.com/articles/s41597-022-01748-x" target="_blank" rel="noopener">https://www.nature.com/articles/s41597-022-01748-x</a><br><a href="https://github.com/immm-sfa/turner_voisin_nelson_2022_scientific_data">https://github.com/immm-sfa/turner_voisin_nelson_2022_scientific_data</a><br><a href="https://github.com/pnnl/hydrofixr">https://github.com/pnnl/hydrofixr</a></p> <p>A paper describing the weekly data is in preperation.</p> <p>Corresponding author: cameron.bracken@pnnl.gov, nathalie.voisin@pnnl.gov</p> <p> </p> <p>Version 1.3.0 Uses RectiHyd 1.3 which includes several hundred new observed data locations from a variety of sources</p> <p>Version 1.2.0 Adds hydro plant data (forebay, inflow, outflow) for some Pacific Northwest plants and HUC4 flow data for most plants</p> <p>Version 1.1.2 Updates the data using final EIA 2022 data </p> <p>Version 1.1.1 Fixes the zip format </p> <p>Version 1.1.0 Extends the data to 2022 using updated versions of the underlying data</p> <p> </p>
Resources for "Disaggregating the carbon exchange of degrading permafrost peatlands using Bayesian deep learning"
<p>This dataset contains all predictors, fluxes, and footprint weights used and described in our manuscript.</p>
Sex-disaggregated Analysis of Risk Factors of COVID-19 Mortality Rates in India
<p>This Zenodo resource contains the data used to perform analysis in the article "Sex-disaggregated Analysis of Risk Factors of COVID-19 Mortality Rates in India".</p> <p>Data</p> <p>The data is organized in the form of tables.</p> <p>hypothesis-test-data</p> <p>This table contains data used to perform the two tailed hypothesis test on gender mortality in different regions.</p> <pre><code>* Region * Male_Deaths - Number of male COVID-19 deaths in region. * Female_Deaths - Number of female COVID-19 deaths in region. * Male_cases - Number of male COVID-19 positive in region. * Female_cases - Number of female COVID-19 positive in region. </code></pre> <p>lasso-covid19India</p> <p>This table contains data used for analysis on cases throughout India.</p> <p>Columns from COVID-19 India data</p> <pre><code>* State_Code * State * District * Confirmed * Active * Recovered * Deceased </code></pre> <p>Columns taken from NFHS data</p> <pre><code>* Sex_ratio_of_the_total_population_females_per_1000_males * Women_whose_Body_Mass_Index_BMI_is_below_normal_BMI__185_kgm214_ * Men_whose_Body_Mass_Index_BMI_is_below_normal_BMI__185_kgm2_ * Women_who_are_overweight_or_obese_BMI__250_kgm214_ * Men_who_are_overweight_or_obese_BMI__250_kgm2_ * All_women_age_1549_years_who_are_anaemic_ * Men_age_1549_years_who_are_anaemic_130_gdl_ * Women_Blood_sugar_level__high_140_mgdl_ * Men_Blood_sugar_level__high_140_mgdl_ * Women_Very_high_Systolic_180_mm_of_Hg_andor_Diastolic_110_mm_of_Hg_ * Men_Very_high_Systolic_180_mm_of_Hg_andor_Diastolic_110_mm_of_Hg_ </code></pre> <p>lasso-KA+TN-bulletin</p> <p>This table contains data used for analysis on the sub-cohort of Karnataka and Tamil Nadu.</p> <p>Data from Media Bulletin</p> <pre><code>* District * Total_Positives * total_deaths * male_deaths * female_deaths * Male_cases_in_data * Female_cases_in_data </code></pre> <p>Calculated Data</p> <pre><code>* Estimated_Male_cases - Estimated male cases using total positives column and existing case data * Estimated_Female_Cases - Estimated female cases using total positives column and existing case data * Male_Mortality - Estimated Male Cases / male_deaths * Female_Mortality - Estimated Female Cases / female_deaths </code></pre> <p>Columns taken from NFHS data</p> <pre><code>* Sex_Ratio_females_every_1000_males * State Women_whose_Body_Mass_Index_BMI_is_below_normal_BMI__185_kgm214_ * Men_whose_Body_Mass_Index_BMI_is_below_normal_BMI__185_kgm2_ * Women_who_are_overweight_or_obese_BMI__250_kgm214_ * Men_who_are_overweight_or_obese_BMI__250_kgm2_ * All_women_age_1549_years_who_are_anaemic_ * Men_age_1549_years_who_are_anaemic_130_gdl_ * Women_Blood_sugar_level__high_140_mgdl_ * Men_Blood_sugar_level__high_140_mgdl_ * Women_Very_high_Systolic_180_mm_of_Hg_andor_Diastolic_110_mm_of_Hg_ * Men_Very_high_Systolic_180_mm_of_Hg_andor_Diastolic_110_mm_of_Hg_ </code></pre> <p>Code</p> <p>The code is available at this <a href="https://github.com/harishpb26/Sex-disaggregated-Analysis-of-Risk-Factors-of-COVID-19-Mortality-Rates-in-India">Github Repository</a>.</p>
Neuron soma size and density measurements in male and female adult rat nucleus accumbens shell, nucleus accumbens core, and caudate-putamen disaggregated by sex and estrous cycle phase
Open the record for dataset details and reuse information.
Matsim Simulation Results (disaggregated)
<p>Matsim Simulation results at disaggregate level are based on the following 2 datasets files for each city i.e. Bologna, Hasselt and Vantaa. The details are as follows:</p> <p>1. The plan.xml file contains the population together with activity-travel details containing trip route information. All these results are for the base cases. </p> <p>2. Config.xml file contains parameters/details required to run Matsim simulation.</p>
The SmartNIALMeter Electrical Appliance Disaggregation Dataset
<p><strong>Intro</strong></p> <p>Electrical disaggregation, also known as non-intrusive load monitoring (NILM) or non-intrusive appliance load monitoring (NIALM), attempts to recognize the energy consumption of single electrical appliances from the aggregated signal. This capability unlocks several applications, such as giving feedback to users regarding their energy consumption patterns or helping distribution system operators (DSOs) to recognize loads which could be shifted to stabilize the electrical grid. The project SmartNIALMeter brought together universities, companies and DSOs and involved the collection of a large data corpus comprising 20 buildings with a total of 100 electrical appliances for a period of up to two years at a sampling interval of five seconds. The variability of the loads, including heat pumps and a charging station for electric vehicles, and the presence of single-phase and three-phase devices make this dataset suitable for several investigations. The total consumption was collected through smart meters and each appliance’s consumption was measured with a dedicated sensor, providing sub-metering for all loads. The dataset can be used to tackle several open research questions, for example to investigate new NILM algorithms able to learn with a limited amount of sub-metered data.</p> <p> </p> <p><strong>Data Description</strong></p> <p>For the residential data we chose the Hierarchical Data Format (HDF5), which has been developed for big datasets and fast access. We publish two versions of the SNM dataset - a raw version with minimal curation steps and a version with more extensive preprocessing applied. Both versions of the dataset are organized along the same structure: Each appliance is saved individually as HDF5 and grouped by the building they are measured in. Measurements from individual phases are denoted by the ending L1, L2 or L3 in the file header (e.g. active power L1). This leads to the following file structure: <type>/building_<x>/<appliance>.h5, where:<br>• <type> denotes the type of the dataset, i.e. raw or preprocessed.<br>• <x> is a unique integer assigned to the building.<br>• <appliance> is the name of the measured appliance. The naming follows the NILM metadata convention.</p> <p>A detailed description of the dataset, corresponding metadata and the measurement setup can be found in M. Vogel, M. Friedli, M. Camenzind, G. Kniesel, Ch. Klemenjak, G. Gugolz, P. Huber, A. Calatroni, L. Kaufmann, A. Rumsch, A. Paice, "The SmartNIALMeter Electrical Appliance Disaggregation Dataset".<br><strong>2024-05-02, Data-In-Brief (under review)</strong></p> <p> </p> <p><strong>Code</strong><br>The code to generate the preprocessed version of the dataset can be downloaded alongside the dataset. Check on <a href="https://github.com/ihomelab/snm-dataset">GitHub</a> if an updated versions is available.</p>
Interpretable semi-supervised population prediction and disaggregation using ancillary data: processed dataset
<p>Dataset ready for use for the tool "cross_validator.py", that runs the experiments.</p> <p> </p> <p>The data cannot be redistributed as-is as it includes non-redistributable work. Contact authors for access.</p>
Interpretable semi-supervised population prediction and disaggregation using ancillary data: experiment results
<p>Result of the experiments for the paper "Interpretable semi-supervised population prediction and disaggregation using ancillary data"</p>
Base data for "Interpretable semi-supervised population prediction and disaggregation using ancillary data"
<p>Dataset used for experiments for the paper "Interpretable semi-supervised population prediction and disaggregation using ancillary data".</p> <p>The data comes from various sources, some of them cannot be legally redistributed; in any case, the provider of the data is clearly indicated in each folder.</p>
Interpretable semi-supervised population prediction and disaggregation using ancillary data - Supplementary Material
<p>Interpretable semi-supervised population prediction and disaggregation using ancillary data - Supplementary Material</p> <p>(details about the datasets + access to datasets + access to source code)</p>
Identification of a HTT-specific binding motif in DNAJB1 essential for suppression and disaggregation of HTT
<p>Underlying data of the<em> in silico</em> work performed in this study. </p> <p>1. i-Tasser predictions for HTTExon1Q<sub>48 </sub></p> <p>2. HDOCK results of DNAJB1-HTTExon1Q<sub>48 </sub>docking</p> <p>3. TIGER2h structures and trajectories. The original data and filtered trajectories for the individual clusters</p> <p>4. MD Simulation structures and trajectories for DNAJB1 protein variants alone and as complexes with HTTExon1Q<sub>48.</sub></p>
A three-weight surface modeling approach for optimizing small-scale population disaggregation
<p><span>In recent decades, gridded population data at fine scales has become a popular data source for assessing and monitoring the Sustainable Development Goals (SDGs). However, current population disaggregation methods are facing challenges in generating high-precision population grids for small areas with limited data. To fill this gap, we proposed a lightweight population gridding method that combines basic dasymetric mapping and point-based surface modeling, named three-weight surface modeling. In this method, there are three weights designed to describe the population spatial heterogeneity from different perspectives. The first weight is building-volume weight, which is equivalent to the preliminary results of population assignment based on building volume information. The second weight, POI-center weight, incorporates POI categories and aggregation patterns to express the centers with high population density, which is calculated based on the neighborhood accumulation rule of Spearman's correlation coefficients between POIs and population size. The third weight called POI-distance weight, indicates different rates of population decay with distance from high-density centers. The three-weight surface model allows us to dynamically adjust the parameters so as to correct the building-volume weight according to the remaining two POI-related weights for a more accurate population surface. After analyzing the census population and the disaggregation results of 544 villages in three counties (Huishui, Luodian and Pingtang) in southern Guizhou Province, China, we found that the customized three-weight model constructed using the local parameter groups demonstrated better accuracy performance compared to separate dasymetric mapping or point-based surface modeling. Meanwhile, the 10-m population grid generated by the local parameter model (LPTW-POP) exhibited higher resolution and lower errors (RMSE, MAE and MRE) than widely used gridded population datasets like LandScan, WorldPop and GHS-POP.</span></p>
TM5-4DVAR Global Monthly Source-disaggregated Methane Emissions, 1999-2016
This dataset holds estimates of methane emissions derived from a dual tracer inversion of atmospheric measurements of CH4 mole fractions and d13CH4 isotopic values. The measurements were assimilated in the TM5 4-Dimensional Variational (TM5-4DVAR) source-sink inversion system to estimate methane emissions from fossil fuel, microbial, and pyrogenic sources. These estimates include monthly means of methane emissions from each source and all three sources combined at 1-degree longitude x 1-degree latitude spatial resolution globally and monthly totals across all global grid cells from each source and all three sources combined from 1999 to 2016. The data are provided in netCDF version 4 format.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.