Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Raw data of the heterogeneous Hegselmann-Krause model on network ensembles
<p># Raw data of the heterogeneous Hegselmann-Krause model on network ensembles<br> This is the raw data underlying the results of the article *«On the effects of over-compromising: heterogeneity and network effects on a bounded confidence opinion dynamics model.»*.</p> <p>For each measured combination of the parameters, there is one gzipped file. The parameters are:</p> <p> - Lower and upper bounds of the confidence interval, [ε_l, ε_u].<br> - Topology: the different types of networks and the average degree with which the networks are generated.<br> - System size<br> - Number of realizations for the parameter combination<br> <br> The single files follow a naming scheme of `data_HK_uni[{eps_l},{eps_u}]_topo={topology}_N={N}_trajrecord=0_{m}real.dat.gz`, where:</p> <p> - `{eps_l},{eps_u}` are the values of the lower and upper bounds of the confidence interval.<br> - `{topology}` contains the type of network and the average degree. The possibilities are `BA_k=10`, `ER_c=10`, `sl1`, `sl2`, and `sl3`.<br> - `N` is the system size. The sizes are powers of two.<br> - `trajrecord=0` signals the fact that file contains only the final state.<br> - `{m}` is the number of realizations.</p> <p># Data format<br> Each file contains the final state of each realization back to back. Each final state is encoded as three lines:</p> <p> - The convergence time is a single integer with a line prefix '\# iterations:'<br> - The positions of all clusters in opinion space with a line prefix '\# ' (unsorted)<br> - The number of agents in each of the clusters without a line prefix</p> <p># Folders structure<br> The files are organized as follows:</p> <p> - **`phase_plots.tar`**: contains the data for the different phase plots (full exploration of the [ε_l, ε_u] space) with `N=16384` and `m=100` realizations.<br> - **`ER`** contains the data for Erdos Renyi with mean degree of 10 (c=10)<br> - **`BA`** contains the data for Barabasi Albert with a mean degree of 10 (k=10)<br> - **`SL`** contains the data for Square lattice with first, second and third nearest neighbors (k=4, 8, 12)<br> - **`swipes.tar`** contains the data for the finite size effects study at fixed ε_l with `m=1000` realizations.<br> - **`ER`** contains the data for Erdos Renyi with mean degree of 10 (c=10) with ε_l = 0.05<br> - **`BA`** contains the data for Barabasi Albert with a mean degree of 10 (k=10) with ε_l = 0.05<br> - **`SL`** contains the data for Square lattice with third nearest neighbors (k=12) with ε_l =0.03<br> - the different videos referenced in the main text and the SM follow various naming schemes:<br> - **`scatter3D_el_eu_Smax_uni_{topology}_N=16384.mp4`**: 360° rotation of the 3D visualisation of the data leading the average phase plots.<br> - **`scatter2D_el={eps_l}_eu_Smax_extremism_{topology}_SizeEffect.mp4`**: evolution of the scatter plot leading the finite size study as a function of N.<br> - **`scatter3D_el={eps_l}_eu_Smax_extremism_{topology}_SizeEffect.mp4`**: same as before, but in 3D where the Z-axis is the extremism.<br> - **`scatter_x0_xt_{topology}_N={N}_{realization_type}.mp4`**: time evolution of the scatter plot of the opinion at time `t` versus initial opinion, color-coded with the extremism. {realization_type} can be mild, skewed or U-turn.<br> - **`traj_2D_SL_k=12_N=16384_{realization_type}.mp4`**: because of the spatial embedding, the time evolution of those realizations on the Square Lattice can be visualized in 2D.</p> <p># Python example for reading the format<br> An example script, which visualizes <S\> vs ε_u graph for the largest size of the ER case, with a function to read this format is given in `example.py`.</p>
Bambu: a simple tool to generate QSAR models from bioassays data
<p>Quantitative Structure Activity Relationship (QSAR) is a computational method that allows the estimation of the properties of a molecule, including its biological activity, based on its structure. QSAR models have been widely employed in the search for potential drug candidates, but also for agrochemicals and other molecules with applications on different branches of the industry. Here we present Bambu, a simple command line tool to generate QSAR models from high-throughput screening bioassays datasets. </p> <p>The tool was developed using the Python programming language and relies mainly on RDKit for molecule data manipulation, FLAML for automated machine learning and the PubChem REST API for data retrieval. As a proof-of-concept we have employed the tool to generate QSAR models for melanoma cell growth inhibition based on HTS data and used them to screen libraries of FDA-approved drugs and natural compounds.</p> <p>Based on the developed tool we were able to produce QSAR models and identify a wide variety of molecules with potential melanoma cell growth inhibitors, many of which with anti-tumoral activity already described. The tool is available through the URL <a href="http://caramel.ufpel.edu.br">http://caramel.ufpel.edu.br</a>.</p> <p> Bambu is an free and open source tool which facilitates the creation of QSAR models and can be futurely applied in a wide variety of drug discovery projects.</p>
Transferability of data-driven models to predict urban pluvial flood water depth in Berlin, Germany
<p>The attached files include the predictive features and the water depth from 2D hydrodynamic simulations that were used to train data driven models to predict water depth in Berlin.</p>
Computational data on "A Model for the Rapid Assessment of Solution-Structures for 24-Atom Macrocycles: The Impact of β-Branched Amino Acids on Conformation"
<p>The archive contains 5 different folders:<br> <br> 1. mdp_files, that contains all the input files for the minimization, equilibration and molecular dynamics run (in gromacs format);<br> 2. topologies, that contains the equilibrated configurations and topology of the 4 systems studied (in gromacs format);<br> 3. trajectories, that contains the coordinates of the macrocycles during the metadynamics calculations and the relevant metadynamics output files (hills file and energy file);<br> 4. plumed_input, that contains the input file for the metadynamics calculations.<br> 5. gaussian, that contains input and output for the single-point charge calculations.</p>
Nonlinear sensitivity of glacier-mass balance to climate attested by temperature-index models; synthetic data
<p>Synthetic data and results of the PDD model used in the paper <a href="https://doi.org/10.5194/tc-2022-210">https://doi.org/10.5194/tc-2022-210</a></p>
Data and trained models for "Fourier Ring Correlation and anisotropic kernel density estimation improve deep learning based SMLM reconstruction of microtubules"
<p>Data and trained models for "Fourier Ring Correlation and anisotropic kernel density estimation improve deep learning based SMLM reconstruction of microtubules", https://github.com/CIA-CCTB/FRCnet</p>
Models and data in support of "EMFEM: a parallel 3D modeling code for frequency-domain electromagnetic method using goal-oriented adaptive finite element method"
<p>These directories contain the model and data files for "EMFEM: a parallel 3D modeling code for frequency-domain electromagnetic method using goal-oriented adaptive finite element method".<br> </p>
Incremental Model Transformations with Triple Graph Grammars for Multi-version Models Evaluation Data
<p>Java abstract syntax graphs for two software development projects in multi-version model and snapshot encoding.</p>
The code and data for "Deep-Learning Correction Methods for WRF Model Precipitation Forecasting from 2014 Through 2022 in Zhengzhou city, China"
<p>The code and data for “Deep-Learning Correction Methods for WRF Model Precipitation Forecasting from 2014 Through 2022 in Zhengzhou city, China”</p>
Research data for "Device-scale atomistic modelling of phase-change memory materials"
<p>This is a dataset related to the publication "Device-scale atomistic modelling of phase-change memory materials".</p> <p>Two folders have been provided for (1) the production data shown in this work, and (2) GAP models trained in this work:</p> <p>(1) The production data have been categorised according to the main text figures:</p> <ul> <li>"reference_database": Reference databases (i.e., training structures) of three GAP models discussed in this work. The structure data are provided in (extended) XYZ format, as labelled using either the PBEsol or the PBE functional. <ul> <li>"main_GST-GAP-22_PBEsol": the main GST-GAP-22 database, which was fitted using a two-step iterative training protocol. The resulting GAP model was used to obtain the results shown in the main text. The reference database is visualised in Fig. 1 of the main text. </li> <li>"refitted_GST-GAP-22_PBE": this dataset contains the same structures as the original GST-GAP-22 training data, with all structures having been relabelled using the PBE functional.</li> <li>"extended_GST-GAP-22_for_efield_PBEsol": an extension of the GST-GAP-22 database to a new, task-specific application, i.e., electromigration under an external electric field (cf. Extended Data Fig. 4).</li> </ul> </li> </ul> <ul> <li>"crystallization_simulations": three trajectories for the crystallization simulations shown in Fig. 2, of which all were obtained from GAP-MD. <ul> <li>"fig2b_growth_GAP-MD": Growth of Ge<sub>1</sub>Sb<sub>2</sub>Te<sub>4</sub> (1008 atoms).</li> <li>"fig2c_cumulative_set_cycles_GAP-MD": Cumulative set process of Ge<sub>1</sub>Sb<sub>2</sub>Te<sub>4</sub> (1008 atoms).</li> <li>"fig2d_crystallization_12096at_GAP-MD": Crystallization of Ge<sub>1</sub>Sb<sub>2</sub>Te<sub>4</sub> (12096 atoms), in which three crystalline seeds were used.</li> </ul> </li> </ul> <ul> <li>"RESET_mushroom_model": two non-isothermal simulations shown in Fig. 3, obtained from GAP-MD. <ul> <li>"Fig3b_small_pulse": The 70 ps NVE equilibrium process after a small heating pulse was imposed in the focal area, giving an excess kinetic energy of 1,650 eV for the atoms in the focal area.</li> <li>"Fig3d_large_pulse": The 70 ps NVE equilibrium process after a large heating pulse was imposed in the focal area, giving an excess kinetic energy of 3,900 eV for the atoms in the focal area.</li> </ul> </li> </ul> <ul> <li>"RESET_device_scale_simulations": the GAP-MD simulations of the melting and heat dissipation process of a device-scale structural model shown in Fig. 4. <ul> <li>"Fig4b_device_scale<strong>_</strong>heating_10ps": The melting process of the device-scale model over 10 ps.</li> <li>"Fig4c_device_scale<strong>_</strong>cooling_40ps": The heat dissipation process of the device-scale model over another 40 ps.</li> </ul> </li> </ul> <p>(2) The GAP models trained in this work:</p> <ul> <li>"main_GAP_potential": the main GAP model used for the production data of this work, which is fitted based on PBEsol data. The XML identifier of this GAP model is GAP_2022_4_7_480_18_6_12_970.</li> </ul> <ul> <li>"other_GAP_potentials": two derivatives of the original GAP model. <ul> <li>"refitted_GST-GAP-22_PBE": using the same reference structures as the original GAP model but re-labelled using the PBE functional. The XML identifier of this GAP model is GAP_2022_5_7_480_0_58_2_26.</li> <li>"extended_GST-GAP-22_for_efield_PBEsol": An extension of the original GAP model to a new, task-specific application, i.e., electromigration under an external electric field. The XML identifier of this GAP model is GAP_2023_3_19_480_18_14_10_174.</li> </ul> </li> </ul> <p> </p>
Data set for: Efficient and Stable Coupling of the SuperdropNet Deep Learning-based Cloud Microphysics (v0.1.0) to the ICON Climate and Weather Model (v2.6.5)
<p>This dataset provides the experiment results for the publication " Efficient and Stable Coupling of the SuperdropNet Deep<br> Learning-based Cloud Microphysics (v0.1.0) to the ICON Climate and Weather Model (v2.6.5) ", to be submitted.</p>
Evaluation and optimisation of the soil carbon turnover routine in the MONICA model (version 3.3.1) - MONICA model source code and data
<p>Zip file consisting of the data used in the manuscript "Evaluation and optimisation of the soil carbon turnover routine in the MONICA model".<br> The data consists of 11 German long term experimenting field sites, formatted in MONICA readable input files. Each consisting of one dataset for weather information (.met), and three .json files describing the management, site and soil properties of each treatment. The validation datasets consist of soil temperature and soil moisture measurement and are distinguishable by the .vals file type/format.</p> <p>For further information please refer to the manuscript.</p>
Data and Code for "An updated end-to-end ecosystem model of the Northern California Current"
<p>This upload contains all data and code for "An updated end-to-end ecosystem model of the Northern California Current reflecting ecosystem changes due to recent marine heat waves", submitted to PLOS ONE. Manuscript abstract is pasted below:</p> <p>The Northern California Current is a highly productive marine upwelling ecosystem that is economically and ecologically important. It is home to both commercially harvested species and those that are federally listed under the U.S. Endangered Species Act. Recently, there has been a global shift from single-species fisheries management to ecosystem-based fisheries management, which acknowledges that more complex dynamics can reverberate through a food web. Here, we have integrated new research into an end-to-end ecosystem model (i.e., physics to fisheries) using data from long-term ocean surveys, phytoplankton satellite imagery paired with a vertically generalized production model, a recently assembled diet database, fishery catch information, species distribution models, and existing literature. This spatially-explicit model includes 90 living and detrital functional groups ranging from phytoplankton, krill, and forage fish to salmon, seabirds, and marine mammals, and nine fisheries that occur off the coast of Washington, Oregon, and Northern California. This model was updated from previous regional models to account for more recent changes in the Northern California Current (e.g., increases in market squid and some gelatinous zooplankton such as pyrosomes and salps), to expand the previous domain to increase the spatial resolution, to include data from previously unincorporated surveys, and to add improved characterization of endangered species, such as Chinook salmon (<em>Oncorhynchus tshawytscha</em>) and southern resident killer whales (<em>Orcinus orca</em>). Our model is mass-balanced, ecologically plausible, without extinctions, and stable over 100-year simulations. Ammonium and nitrate availability, total primary production rates, and model-derived phytoplankton time series are within realistic ranges. As we move towards holistic ecosystem-based fisheries management, we must continue to openly and collaboratively integrate our disparate datasets and collective knowledge to solve the intricate problems we face. As a tool for future research, we provide the data and code to use our ecosystem model.</p>
Raw data for Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias
Open the record for dataset details and reuse information.
Construction of Perioperative Medical Data Platform and Its Typical Practice to Predict Postoperative Acute Moderate to Severe Pain With Machine Learning Models
ClinicalTrials.gov study NCT05569460. IPD Sharing: NO. Countries: 1. Publications: 0.
Data Acquisition to Model Glycemic Response
ClinicalTrials.gov study NCT03612999. IPD Sharing: Not stated. Countries: 1. Publications: 0.
An AI Platform Integrating Imaging Data and Models, Supporting Precision Care Through Prostate Cancer's Continuum
ClinicalTrials.gov study NCT05384002. IPD Sharing: YES. Countries: 1. Publications: 0.
Data Collection Protocol for the Development of CHLOE-OQ, an AI Model for Assessing Oocytes Quality.
ClinicalTrials.gov study NCT06928337. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Oral Health Policy for Patients With Type 2 Diabetes Mellitus in Chile: A Microsimulation Model Based on Real-world Data
ClinicalTrials.gov study NCT05809817. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Predicting Fall Risk in Stroke Patients Using a Machine Learning Model and Multi-Sensor Data
ClinicalTrials.gov study NCT06380049. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.