Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
The INMCM-4.8 Earth system model data used in the paper by Guryanov V.V. et al. entitled ''The present-day and future lightning frequency as simulated by four CMIP6 models'
<p>The INMCM-4.8 Earth system model data used in the paper by Guryanov V.V. et al. entitled ''The present-day and future lightning frequency as simulated by four CMIP6 models'</p>
Support data for conference paper "Service-Oriented Model for Handling mMTC Subscribers' Traffic in a 5G Cluster"
<p>Support data for conference paper<br>V. Kovtun, and O. Kovtun, “Service-Oriented Model for Handling mMTC Subscribers’ Traffic in a 5G Cluster.” In Proc. 5th 5th International Workshop on Intelligent Information Technologies & Systems of Information Security, CEUR-WS, vol. 3675, 2024; pp. 236-246.<br>This research is part of the project No. 2022/45/P/ST7/03450 co-funded by the National Science Centre and the European Union Framework Programme for Research and Innovation Horizon 2020 under the Marie Skłodowska-Curie grant agreement No. 945339.</p>
Data repository for manuscript "A new approach to Health Benefits Package design: an application of the Thanzi La Onse model in Malawi"
<p>Dataset to accompany the publication <em>“A new approach to Health Benefits Package design: an application of the Thanzi La Onse model in Malawi”</em> by Margherita Molaro, Sakshi Mohan, Bingling She, Martin Chalkley, Tim Colbourn, Joseph H. Collins, Emilia Connolly, Matthew M. Graham, Eva Janoušková, Ines Li Lin, Gerald Manthalu, Emmanuel Mnjowe, Dominic Nkhoma, Pakwanja D. Twea, Andrew N. Phillips, Paul Revill, Asif U. Tamuri, Joseph Mfutso-Bengo, Tara Mangal, and Timothy B. Hallett.</p> <p>The Thanzi La Onse (TLO) model used to produce this data is open source and available for review and usage at<a href="https://github.com/UCL/TLOmodel"> https://github.com/UCL/TLOmodel</a>. In particular, the outputs analysed in this study can be reproduced from model tag "Molaro_et_al_2024_HBP_design" (accessible at https://github.com/UCL/TLOmodel/tags) using the scenario file src/scripts/healthsystem/impact_of_policy/scenario_impact_of_policy.py. All analysis scripts used to generate the plots in the manuscript are located in the same directory and have filenames beginning with "analysis_impact_of_policy_".</p> <p>This repository contains post-processed simulation outputs, which were generated using the script src/scripts/healthsystem/impact_of_policy/analysis_extract_data.py (available from the same tag). The data included have the following structure:</p> <p>"Draw": Represents a specific prioritisation-policy, identified by the acronyms listed in Table 1 of the publication.</p> <p>"Run": Represents a single simulation instance of a draw. Each draw was simulated 10 times, each with independent random sampling, resulting in 10 "runs" per draw.</p> <p>The data files included in this repository are:</p> <p><strong>DALYS_by_cause_with_time.csv</strong>: DALYs (as defined in the publication) incurred on a given year due to each of the causes of DALYs considered.</p> <p><strong>HSIs_requested_by_type_and_facility_level_with_time.csv</strong>: total number of requested HSIs on a given year, broken down by HSI type and the facility level at which they were requested.</p> <p><strong>HSIs_delivered_by_type_and_facility_level_with_time.csv</strong>:total number of HSIs delivered on a given year broken down by HSI type and the facility level at which they were delivered.</p> <p><strong>Population_with_time.csv</strong>:total population size on a given year. </p> <p> </p> <p> </p>
Data bundle for powerd-data: A transparent and reproducible data processing pipeline for energy system modeling based on egon-data
<div> <p><strong>powerd-data</strong> provides a transparent and reproducible open data based data processing pipeline for generating data models suitable for energy system modeling. Is is a fork from the open-source tool <strong>egon-data</strong>. </p> <p>powerd-data and egon-data retrieve and process data from several different external input sources. As not all data dependencies can be downloaded automatically from external sources, we provide a data bundle to be downloaded by egon-data.</p> <p>The following data sets are part of the available data bundle:</p> <ol> <li>district_heating_shares: <ul> <li>Assumed district heating share for all European countries in 2050</li> <li>Source: Own representation</li> <li>License: Attribution 4.0 International (CC BY 4.0)</li> </ul> </li> <li>egon_demandregio_cts_ind:<br> <ul> <li>Industrial and CTS demands per branch and NUTS3 region in Germany for the year 2050</li> <li>Source: egon-data, based on data from DemandRegio disaggregator tool</li> <li>License: Data license Germany – © FfE 2019, © Statistisches Bundesamt (Destatis), 2008-2017 – version 2.0</li> </ul> </li> <li>industrial_gas_demand: <ul> <li>This folder contains 5 files. The files CH4_for_industry_eGon100RE.json, CH4_for_industry_eGon2035.json, H2_for_industry_eGon100RE.json and H2_for_industry_eGon2035.json contain the industrial hourly demands for hydrogen and methane in NUTS3 resolution for the scenarios eGon100RE and eGon2035. The file region_corr.json provides information that make it possible to correlate each load to a geographical position.</li> <li>License: Attribution 4.0 International (CC BY 4.0) © FfE, eXtremOS Project</li> </ul> </li> </ol> <p> </p> </div>
Data for "Entanglement Dynamics in Monitored Kitaev Circuits: Loop Models, Symmetry Classification, and Quantum Lifshitz Scaling"
<p>We provide the data and scripts used to produce the figures shown in our publication "Entanglement Dynamics in Monitored Kitaev Circuits:<br>Loop Models, Symmetry Classification, and Quantum Lifshitz Scaling".</p>
Data table 5 from publication "Impaired interactions of ataxin-3 with protein complexes reveals their specific structure and functions in SCA3 Ki150 model" (doi.org/10.3389/fnmol.2023.1122308)
<div>An Excel table contains a list of proteins identified by MS from pull-down experiment using cerebellar cortex lysates with Dynabeads coated with anti-ataxin-3 1H9 mouse monoclonal antibodies. The false positive interactor proteins were excluded from this list by subtracting proteins found in "Supplementary_table 2_cortex isogenic mouse IgG control dynabeads.xlsx" </div> <div> </div>
Data table 4 from publication "Impaired interactions of ataxin-3 with protein complexes reveals their specific structure and functions in SCA3 Ki150 model" (doi.org/10.3389/fnmol.2023.1122308)
<div>An Excel table contains a list of proteins identified by MS from a pull-down experiment using cerebellum lysates with Dynabeads coated with anti-ataxin-3 1H9 mouse monoclonal antibodies. The false positive interactor proteins were excluded from this list by subtracting proteins found in "Supplementary_table 3_cerebellum isogenic mouse IgG control dynabeads.xlsx<span><br></span></div>
Data table 2 from publication "Impaired interactions of ataxin-3 with protein complexes reveals their specific structure and functions in SCA3 Ki150 model" (doi.org/10.3389/fnmol.2023.1122308)
<p>An excel table containing a list of proteins identified by MS from samples after a pull-down experiment using cerebral cortex lysates with Dyna beads coated with control isogenic mouse IgG. The proteins from the list were considered false positive interactors of ataxin-3 in the cortex.</p>
Data table 1 from publication "Impaired interactions of ataxin-3 with protein complexes reveals their specific structure and functions in SCA3 Ki150 model" (doi.org/10.3389/fnmol.2023.1122308)
<p>An excel table contains LFQ intensity and other raw MS data for proteins identified in fractions 3,4,5 (i), 11, 12, 13, (ii) 18, 19, and 20 (iii) from ki150 and ki21 model brains. These fractions showed enrichment in ATXN3 protein. The fractions are visualized in Figure 4 (<a href="http://doi.org/10.3389/fnmol.2023.1122308" target="_blank" rel="noopener">doi.org/10.3389/fnmol.2023.1122308</a>)</p>
FRAMEWORK FOR A VOLUME MODEL FOR MONOCLINIC AMPHIBOLE - Data and Code
<p>This upload includes the data and code I used to calibrate a volume model for monoclinic amphoboles.</p>
set of exploration data and parameters for h24/5ad agent based model
<p>Different set of data used for exploration, used for reproductibility, updated with HigherProp parameters</p>
Spreadsheet for analysis of illness-death model with aggregated data
<p>Spreadsheet for calculation of a recurrence equation and analysis of fixed points in the illness-death model.</p>
Demo-Dataset for publication "FAIR workflows in Earth system modelling: a use case with semantic data management"
<p>This demodataset is intended to be used to test the workflow described in the publication by Lennartz & Schlemmer "FAIR workflows in Earth System modelling: a use case with semantic data management". It contains example model output for an arbitrary biogeochemical model tracer (here: dissolved organic carbon, DOC) from an ocean model as a 4-dimensional dataset (latitude, longitude, depth, time), the corresponding grid point locations as well as a textfile specifying parameter inputs for the model. The file structure is adapted for seamless integration into the workflow described in Lennartz & Schlemmer, which builds on the open source semantic research data management system LinkAhead. The dataset contains the following structure: The folder DataAnalysis stores data required for data analysis, such as the grid point locations in the file TMM_grid_v2018a.mat. The folder SimulationData stores model output in the folder 2022_TMM, containing the parameter input file nl_in.txt and the model output TR_monthly.mat. Related instructions can be accessed here: https://gitlab.com/salexan/fairworkflows-demodataset .</p>
Southern Baltic Eddies from Model Data (2019-2023) with associated scripts in MATLAB
<p>The data regarding the eddies are located in the MATLAB file 'eddies.mat' and are structured into a single variable eddies{day,time}{eddy,1}.params, where day refers to consecutive days of the non-leap year Julian calendar (365 days in the model year), time denotes one of the 4 six-hour averages, and eddy numbers the subsequent eddies. params = {year, month, day, avgNo, is_ended_flag, xs, ys, sizes, omegas, duration}. The figures used in the article were generated using the script 'eddies_statistics.m'. The file 'constants.mat' contains the variables used to produce the figures. The 'calculate_g1.m' function was used to calculate the Gamma_1 function, 'eddy_size.m' for calculating the eddy sizes, 'cluster_points.m' for clustering points, and 'lat_lon_to_x_y.m' for converting geographic longitude and latitude into local x, y coordinates.</p>
Electrochemical data shown in A. Fasano, C. Baffert, C. Schumann, G Berggren, J. Birrell, V. Fourmond, C. Léger, "Kinetic modeling of the reversible or irreversible electrochemical responses of FeFe-hydrogenases", J. Am. Chem. Soc 146, 2, 1455–1466 (2024) doi: 10.1021/jacs.3c10693
Open the record for dataset details and reuse information.
Data and scripts used in: "Exploring Biological Neuronal Correlations with Quantum Generative Models"
<div>Data and script for the manuscript "Exploring Biological Neuronal Correlations with Quantum Generative Models", by Vinicius Hernandes and Eliska Greplova.</div> <h3>main scripts</h3> <div> <p><strong><em>generate_activity_dataset.py</em></strong></p> <p>reshape data in <em>neuronData.npy</em> to 50k samples of (neurons, timesteps) shape, saved in <em>activity_data.npy</em></p> <p><strong><em>create_target_distributions.py</em></strong></p> <p>based on the dataset, makes dicionary with the the target distribution for each (neurons, timesteps) pair, saved in <em>distribution_target_dictionary.pkl</em></p> <p><strong><em>create_hyperparameters_file.py</em></strong></p> <p>generates <em>hyperparameters.csv</em>, containing:</p> </div> <ul> <li>number of neurons</li> <li>number of timesteps</li> <li>number of auxiliary_qubits</li> <li>batch_size</li> <li>learning rate of generator</li> <li>learning rate of critic</li> <li>number parametrized layers</li> <li>number of training iterations</li> <li>loss type</li> </ul> <p>for each run</p> <p><strong><em>train_qgan.py</em></strong></p> <div> <p>trains models defined <em>models.py</em> using <em>activity_data.npy</em> dataset, and for the hyperparameters defined in <em>hyperparameters.csv</em></p> </div> <div>saves loss functions, and the trained models for each 10 iterations, in specific folders indexed by the run specified in the hyperparameters file</div> <div> </div> <div><strong><em>generate_fake_activity.py</em></strong></div> <div> </div> <div>uses trained models saved in <em>output/models/run{run}/i{training_step}.pth</em> for a specific <em>training_step</em> and <em>run</em> to generate fake data, and save them in <em>output/generated_data/run{run}/i{training_step}.npy</em> files</div> <div> </div> <div><strong><em>analyze_error.py</em></strong></div> <div> </div> <div>uses generated data saved in <em>output/generated_data/run{run}/i{training_step}.npy</em> to generate two statistical quantities (k-probs and firing rate), using the function in <em>metrics.py</em>, and compare the errors in those quantities between the models using k-loss and standard-loss</div> <div> </div> <div><strong><em>analyze_stats.py</em></strong></div> <div> </div> <div>uses generated data saved in <em>output/generated_data/run{run}/i{training_step}.npy</em> to generate:</div> <ul> <li>js diverge for each training step, and final distribution of generated states, stored in <em>distribution_target_dictionary.pkl</em></li> <li>other statistical quantities, using the function in <em>metrics.py</em> file</li> </ul> <h3>auxiliary scripts</h3> <div><strong><em>metrics.py</em></strong></div> <div> </div> <div>functions to calculate neuronal statistics</div> <div> </div> <div><strong><em>aux.py</em></strong></div> <div> </div> <div>auxiliary functions:</div> <ul> <li>to generate states distribution given a dataset</li> <li>custom js divergence</li> </ul> <h3>Data</h3> <p><strong><em>neuronData.npy</em></strong></p> <p>neuronal data from Marre et al., Multi-electrode array recording from salamander retinal ganglion cells (2017)</p> <p><strong><em>activity_data.npy</em></strong></p> <p>dataset obtained from <em>neuronData.npy</em>, taking 50 thousand samples of shape (neurons, timesteps)</p> <p><strong><em>output</em></strong></p> <p>results obtained from <em>train_qgan.py</em> and <em>generate_fake_activity.py</em> </p> <p>contains:</p> <ul> <li><strong><em>losses</em></strong></li> </ul> <p>generator and critic loss for all training runs and steps</p> <ul> <li><strong><em>models</em></strong></li> </ul> <p>saved torch models every 10 training steps, for all training runs</p> <ul> <li><strong><em>generated_data</em></strong></li> </ul> <p>generated data for all models saved in <em>models</em></p>
How to select predictive models for decision making or causal inference? Experiments data
<p>This is the full result data for the experiments of the paper : Doutreligne, M., & Varoquaux, G. (2023). How to select predictive models for decision making or causal inference?, https://hal.science/hal-03946902. <br><br>The code repository is : https://github.com/soda-inria/caussim/tree/main</p> <p>The files in this dataset are the one for the most computationnally costly experiments. There is one folder for each of the four datasets used in the paper. Then, one folder for each of the experimental setup. The files required for the main figure (Fig.3) of the paper are the one labelled #fig3 in the following descriptions.</p> <p>Details on the files : </p> <p>.<br>├── acic_2016_save<br>│ ├── acic_2016__nuisance_non_linear__candidates_hist_gradient_boosting__dgp_1-77__rs_1-5<br>│ │ └── run_logs.csv: results for the experiment with non linear models for both the nuisances and the candidates<br>│ ├── acic_2016__nuisance_non_linear__candidates_ridge__dgp_1-77__rs_1-10<br>│ │ └── run_logs.csv: results for the experiment with non linear models for the nuisances and linear models for the candidates<br>│ └── acic_2016__stacked_regressor__dgp_1-77__seed_1-10<br>│ └── run_logs.csv: results for the experiment with stacked models (linear and non linear) for the nuisances and non linear models for the candidates #fig3<br>├── acic_2018_save<br>│ └── acic_2018__nuisance_non_linear__candidates_hist_gradient_boosting__first_uid_432<br>│ └── run_logs.csv results for the experiment with stacked models (linear and non linear) for the nuisances models and non linear models for the candidates #fig3<br>├── caussim_save<br>│ ├── caussim__linear_regressor__test_size_5000__n_datasets_1000<br>│ │ ├── run_logs.csv: results for the experiment with stacked models for the nuisances models and linear models for the candidates <br>│ │ └── simu.yaml: configuration file of the experiment<br>│ ├── caussim__nuisance_non_linear__candidates_ridge__overlap_01-247_join_nuisance_train_set<br>│ │ └── run_logs.csv: results for the experiment with non linear models for the nuisances and linear models for the candidates, joined sets for the nuisances and the candidates<br>│ ├── caussim__nuisance_non_linear__candidates_ridge__overlap_01-247_separated_nuisance_train_set<br>│ │ └── run_logs.csv: results for the experiment with non linear models for the nuisances and linear models for the candidates, separated sets for the nuisances and the candidates<br>│ └── caussim__stacked_regressor__test_size_5000__n_datasets_1000<br>│ ├── run_logs.csv: results for the experiment with stacked models (linear and non linear) for the nuisances and linear models for the candidates #fig3<br>│ └── simu.yaml: configuration file of the experiment<br>└── twins_save<br> └── twins__stacked_regressor__rs_1-10__overlap_0.1-3<br> └── run_logs.csv: results for the experiment with stacked models (linear and non linear) for the nuisances and non linear models for the candidates #fig3</p>
Updated Smoke Exposure Estimate for Indonesian Peatland Fires using a Network of Low-cost PM2.5 sensors and a regional air quality model - Model Simulation Data
<p>WRF-Chem simulated daily mean PM2.5 concentrations for:</p> <p>1) with fires </p> <p>2) without fires</p> <p>simulations. </p>
Updated Smoke Exposure Estimate for Indonesian Peatland Fires using a Network of Low-cost PM2.5 sensors and a regional air quality model - Purple Air data
<p>Daily mean PM2.5 concentrations collected by Purple Air sensors between 2023-08-16 and 2023-12-01. Concentrations have been RH adjusted using the Nilson et al (2022) adjustment. </p>
Future decline of Antarctic Circumpolar Current model data and figures
<p>This upload contains all the post-processed files and code necessary to recreate the figures in the manuscript entitled "<em>Future decline of Antarctic Circumpolar Current due to polar ocean freshening</em>" by Sohail, Gayen and Klocker. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.