Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13,064
datasets available to search
ShareScore release 0.9.0
Dataset results
13,064 results for “Prediction”
Buoyancy and Brownian motion of plastics in aqueous media: Predictions and implications for density separation and aerosol internal mixing state (Data Underlying Figures)
<p>Data underlying figures in A. Bain 'Buoyancy and Brownian motion of plastics in aqueous media: Predictions and implications for density separation and aerosol internal mixing state' RSC Environmental Science: Nano, 2022. </p> <p>CA = citric acid<br> NaCl = sodium chloride<br> AS = ammonium sulfate</p> <p>rho = difference in density (g/cm^3)<br> Rh = % relative humidity<br> radius is in micrometers<br> Pe0 are the calculated dimensionless Peclet numbers<br> </p>
Data for the "Technical Note: assessing predicted cirrus ice properties between two deterministic ice formation parameterizations" manuscript
<p>This repository contains the post-processed ECHAM-HAM data files for comparing KM21_GCM and ML20 of "Technical Note: assessing predicted cirrus ice properties between two deterministic ice formation parameterizations" study.</p> <p>The files are all netCDF4.</p>
VEL-Ar trajectory prediction model linear, co-seismic, and post-seismic grids for interpolation
<p>VEL-Ar trajectory prediction model linear, co-seismic, and post-seismic interpolation grids in ASCII format. The generation of these grids is described in http://doi.org/10.1007/s00190-015-0871-8</p>
Data for: Presence-absence of marine macrozoobenthos does not generally predict abundance and biomass
<p>This repository contains data for the paper: Bijleveld, A. I. et al. (2018) Presence-absence of marine macrozoobenthos does not generally predict abundance and biomass. Scientific Reports 8, 3039, doi:10.1038/s41598-018-21285-1.</p>
Supplementary Data for NIPS Publication: Protein Interface Prediction using Graph Convolutional Networks.
<p>These data sets can be used to re-run the experiments from our paper, Protein Interface Prediction using Graph Convolutional Networks. The data are derived from protein complexes in the docking benchmark dataset v. 5.0. Each file is a python tuple that has been saved using cPickle and compressed using gzip.</p> <p>Links:</p> <p>Paper: https://papers.nips.cc/paper/7231-protein-interface-prediction-using-graph-convolutional-networks</p> <p>Poster: https://zenodo.org/record/1134154</p> <p>Code: https://github.com/fouticus/pipgcn</p> <p> </p> <p><strong>File Descriptions:</strong></p> <p>train.cpkl.gz and test.cpkl.gz have the data formatted for neighborhood based graph convolutions. The diffc_ files are the same data formatted for the diffusion convolutional neural networks that we compare against. </p> <p>train.cpkl.gz is a tuple of length 2:</p> <ul> <li>element 0 is a list of length 175 containing the PDB codes from the docking benchmark dataset</li> <li>element 1 is a list of length 175 containing features for each protein. Each element is a dictionary containing the following keys: <ul> <li>r_vertex: vertex (residue) features for the receptor. numpy array of shape (x, 70) where x is the number of residues in the receptor and 70 is the number of features.</li> <li>l_vertex: vertex (residue) features for the ligand. analogous to above, with shape (y, 70) where y is the number of residues in the ligand.</li> <li>complex_code: PDB code of the complex. matches the list of codes described above.</li> <li>l_edge: edge features for the neighborhood around each residue in the ligand. numpy array of shape (y, 20, 2) where y is defined as above. the second dimension is the edges to the 20 nearest neighboring residues, ordered by decreasing distance. The third dimension allows for two features per edge. </li> <li>r_edge: edge features for the neighborhood around each residue in the receptor. numpy array of shape (x, 20, 2) where x is as above. </li> <li>l_hood_indices: the index of the 20 closest residues to each residue, ordered by decreasing distance. numpy array of shape (y, 20, 1). "Index" means which row in l_vertex gives the vertex features for the closest neighbor, second closest neighbor, etc. </li> <li>r_hood_indices: analogous to above, shape (x, 20, 1).</li> <li>label: 1 or -1 label for each residue pair. numpy array of shape (x*y, 3). Each row looks like (i, j, k) where i is the index of the ligand residue, j is the index of the receptor residue, and k is either -1 (negative example) or 1 (positive example).</li> </ul> </li> </ul> <p>test.cpkl.gz matches the structure of train.cpkl.gz except it has the test set of 55 complexes. </p> <p>Descriptions of the vertex and edge features can be found in Appendix A of <a href="https://mountainscholar.org/handle/10217/185661">this.</a></p> <p>diffc_g2_p2_train.cpkl.gz is a tuple of length 2:</p> <ul> <li>element 0 is a list of the same 175 PDB codes as above. </li> <li>element 1 is a list of features for the 175 complexes. Each element is a dictionary of features with these keys: <ul> <li>r_vertex, l_vertex, complex_code, label: these are the same as described above. </li> <li>'r_power_series': Stacked diffusion matrices which are powers of the similarity matrix used in the DCNN method. numpy array of shape (x, 2, x) where x is the number of receptor residues. the middle dimension 2 indicates how many "hops" is used for that diffusion (1 vs. 2). In other words, element (i, 0, j) is the similarity after 1 hops between residues i and j. element (i, 1, j) is the similarity after 2 hops. See DCNN paper for details.</li> <li>'l_power_series': same as above but for the ligand. shape is (y, 2, y).</li> </ul> </li> </ul> <p>diffc_g2_p2_test.cpkl.gz is the same as diffc_g2_p2_train.cpkl.gz but for the 55 test complexes.</p> <p>diff_g2_p5_train.cpkl.gz and diff_g2_p5_test.cpkl.gz are the same as the p2 version above, except that the diffusion matrices have shape (x, 5, x) and (y, 5, y) because one of our comparisons against the DCNN model uses 5 hops instead of just 2. </p> <p> </p> <p>Note: these files were pickled with Python 2.7. If you're unpickling with Python 3.x you might have to specify encoding as 'latin1'. </p> <p> </p> <p>Please direct any questions to:</p> <ul> <li>Alex Fout (fout@colostate.edu)</li> <li>Jonathon Byrd (jonbyrd@colostate.edu)</li> <li>Basir Shariat (basir@cs.colostate.edu</li> <li>Asa Ben-Hur (asa@cs.colostate.edu)</li> </ul>
Ion Implantation Sensor and Process Target Data for Predicting Ion Beam Tuning in Semiconductor Manufacturing
<h2><strong>Dataset Description:</strong></h2> <p>This dataset is designed to predict ion beam tuning setup processes in semiconductor manufacturing, in terms of tuning success or failure, and tuning duration. It is split into <strong><code>X</code></strong> and <code><strong>y</strong></code> to allow for supervised learning approaches.</p> <ul> <li><code><strong>X</strong></code> represents the current equipment condition and the process targets of the currently processed and the upcoming lot, as defined within recipes.</li> <li><code><strong>y</strong></code> represents the ion beam tuning setup report, which informs about the tuning success ratio and tuning duration. These setups are necessary, when switching between recipes to prepare the equipment for processing the next lot. <strong><code>y</code></strong> contains three labels, enabling classification of (1) tuning success or fail, and (2) prolonged tuning, as well as (3) estimation of tuning duration as a regression task.</li> </ul> <p>About <strong><code>X</code></strong>:</p> <p>Each lot is processed with a specific recipe to achieve the process target. The tuning takes place before the first wafer of the to-be-tuned recipe is processed. Each row in <strong><code>X</code></strong> includes logistical information such as the equipment used for processing and parsed recipe / process target information for the current and upcoming lot. The majority of data consists out of aggregated metrics of equipment-internally tracked sensor traces, recording physical parameters such as gas flows, temperatures, voltages and currents. When analyzed in conjunction with the processed recipe, these sensors provide insights into the current equipment condition. </p> <p>About <code><strong>y</strong></code>:</p> <p>The <code>setup_result</code> column indicates the success or failure of tuning - with <code>setup_result=0</code> indicating tuning success, while <code>setup_result=1</code> signals tuning failure. If the first tuning attempt fails, there may be follow-up attempts, but these are not included in this dataset. The <code>duration</code> column represents the tuning duration in seconds, as used for regression analysis. The <code>duration_interval</code> column is a binary label for prolonged tunings, i.e. <code>duration_interval=1</code> for instances, which take more than 6 minutes to tune.</p> <p>For reproducibility of the corresponding paper's results:</p> <ol> <li>The dataset contains the same carefully curated subset of features.</li> <li>The train_test_split() has already been performed, thus we provide <code>x_train</code> and <code>x_valid</code> separately.</li> <li>To reduce the effect of outliers in the data, the sensor data has already been scaled, as derived from <code>x_train</code>.</li> </ol> <p>In summary, these datasets (<code><strong>X</strong></code>, <code><strong>y</strong></code>) provide comprehensive information for predicting ion beam tuning in semiconductor manufacturing, making it a valuable resource for researchers and practitioners in the field.</p> <h2><strong>Python Code for Reproducibility:</strong></h2> <p>Furthermore, we share a jupyter notebook <code>ionbeamtuning.ipynb</code> with Python code to train the best performing model on the provided data, as described in the paper. To execute the code, you may need to install any missing packages specified in the <code>requirements.txt</code>, as indicated within the notebook.</p>
West Nile Virus Predictions output data
<p><span>The testSubmission.csv was meant for validating the results in the Kaggle competition. It is a csv with 2 columns, one being a probability value from 0 to 1 and the second being an id to which that probability refers to. Those probabilities are the result of the classifications of a trained XGBoost classifier. </span></p>
Cladonema radiatum Alr and IgSF genes assembly and domain predictions
<p><span>This dataset is related to the </span><span>submitted paper</span><span> " A single gene determines allorecognition in hydrozoan jellyfish <em>Cladonema radiatum</em> inbred lines ".</span></p> <h1><span>Abstract:</span></h1> <p><strong><span> </span></strong></p> <p><span>Allorecognition—the ability of an organism to discriminate between self and non-self—is crucial to colonial marine animals to avoid invasion by other individuals in the same habitat. The cnidarian hydroid <em>Hydractinia</em> has long been a major research model in studying invertebrate allorecognition, establishing a rich knowledge foundation. In this study, we introduce a new cnidarian model <em>Cladonema radiatum</em> (<em>C. radiatum</em>). <em>C. radiatum</em> is a hydroid jellyfish which also forms polyp colonies interconnected with stolons. Allorecognition responses, fusion or regression of stolons, are observed when stolons encounter each other. By transmission electron microscopy, we observe rapid tissue remodelling contributing to gastrovascular system connection in fusion. Rejection responses are regulated by reconstruction of the chitinous exoskeleton perisarc, and induction of necrotic and autophagic cellular responses at cells in contact with the opponent. Genetic analysis identifies allorecognition genes: six<em> Alr </em>genes located on the putative Allorecognition Complex (ARC) and four immunoglobulin superfamily genes on a separate genome region. C. radiatum allorecognition genes show notable conservation with the <em>Hydractinia Alr</em> family. Remarkedly, stolon encounter assays of inbred lines reveal that genotypes of Alr1 solely determine allorecognition outcomes in <em>C. radiatum</em>.</span></p>
ΔvapHm-VOC: Standard Molar Vaporization Enthalpy Database for Machine Learning Prediction Models
<p>We present the full database of the article "Data-Driven, Explainable Machine Learning Model for Predicting Volatile Organic Compounds’ Standard Vaporization Enthalpy".</p> <p>This is the database used for data driven, explainable supervised ML model to predict Δ<sub>vap</sub><em>H</em><sub>m</sub>° of VOCs. The model was built on an established experimental database of 2410 unique molecules and 223 VOCs categorized by chemical groups. Using supervised ML regression algorithms, the Random Forest successfully predicted VOCs’ Δ<sub>vap</sub><em>H</em><sub>m</sub>° with a mean absolute error of 3.02 kJ mol<sup>-1</sup> and a 94% test score. The model was successfully validated through the prediction of Δ<sub>vap</sub><em>H</em><sub>m</sub>° for a known database of VOCs and through molecular group hold-out tests.</p> <div> <div> <div> <div> <p>The model's database was built with a variety of molecules from diverse chemical families with known experimental Δ<sub>vap</sub><em>H</em><sub>m</sub>° values. Entries were collected from <a href="https://doi.org/10.1063/1.3309507" target="_blank" rel="noopener">Acree and Chickos’ 2010 compilation</a>, curated by <a href="https://doi.org/10.1016/j.fluid.2013.09.021" target="_blank" rel="noopener">Gharagheizi (2013)</a>, with experimental vaporization enthalpy at the standard temperature of 298.15 K. This database was selected as it is an open-access repository, generally presenting experimental values with low uncertainties and corrected for the real-to-ideal behavior of the gas phase. We introduced a routine to convert and present each chemical entry into a SMILES string, along with chemical family categorization. For VOCs, we built a specific database of compounds documented in a VOC regulatory environmental guideline (<a href="https://www.s-t-a.org/Files%20Public%20Area/Documents/The%20Categorisation%20of%20Volatile%20Organic%20Compounds%20HMIP%20(1996).pdf">Marlowe <em>et al</em>., 1995</a>), and we used our web-scrapping routine to gather experimental Δ<sub>vap</sub><em>H</em><sub>m</sub>° values. The external dataset for validation studies was also collected from <a href="https://doi.org/10.1016/j.fluid.2013.09.021" target="_blank" rel="noopener">Gharagheizi (2013)</a>.</p> <p>Along with Δ<sub>vap</sub><em>H</em><sub>m</sub>° experimental values, each molecule is represented by its CAS number, SMILES string and InChlKey. We generated 106 chemical descriptors for every molecule in the database, using <a href="http://http//www.rdkit.org/">RDKit</a> software version 2022.09.4, running on top of Python 3.9. Descriptors were calculated from the “MolFromSmiles” function in “RDKIT.Chem” as descriptors with non-numerical values were removed. The descriptors encode significant chemical information and are used to present physicochemical characteristics of compounds, building a relationship between structure and Δ<sub>vap</sub><em>H</em><sub>m</sub>°.</p> </div> </div> </div> </div> <p>Through chemical feature importance analysis, the explainable model revealed that VOC polarizability, connectivity indexes and electrotopological state are key for the model’s prediction accuracy. We thus present a replicable and explainable model, which can be further expanded towards the prediction of other thermodynamic properties of VOCs.</p>
1 million cMSSM parameter space points with low-energy predictions from SPheno and MicrOMEGAs
<p>This dataset was produced and used in the paper <a href="https://arxiv.org/abs/2405.18471">Symbolically Regressing Beyond the Standard Model Physics</a>. The code used to generate and to analyse these data can be found <a href="https://gitlab.com/miguel.romao/symbolic-regression-bsm">here</a>.</p> <p>The dataset specifications:</p> <ul> <li>Randomly sampled 1 million points of the cMSSM parameter space and respective low-energy observables.</li> <li>Low-energy observables computed using using `SPheno` and `MicrOMEGAs`. <ul> <li>Only points that produced `SPheno` output and neutral LSP are processed by `MicrOMEGAs`.</li> <li>The dataset includes all points, even if they are "unphysical", i.e. points without `SPheno` output or neutral LSP. In the paper, this was used to train a classifier to filter out "unphysical" points.</li> </ul> </li> <li>The columns are <ul> <li>'m0', 'm12', 'A0', 'tanb': the four physical parameters of the theory sampled in the priori <ul> <li>'m0': [0, 10] TeV</li> <li>'m12': [0, 10] TeV</li> <li>'A0': [-60,60] TeV</li> <li>'tanb': [1.5,50]</li> <li>The sign of the 'mu' parameter was fixed to positive (+1)</li> </ul> </li> <li>'idx': an utility identifier used during generation, can/should be ignored</li> <li>Flattened `SPheno` outputs. These are obtained by reading the resulting slha spectrum file outputted by SPheno and flatten the blocks. For example from the 'MINPAR' block, the key-value pairs are given by the columns 'MINPAR_1', 'MINPAR_2', 'MINPAR_3', 'MINPAR_4', 'MINPAR_5', and likewise for all blocks in the slha file.</li> <li>`MicrOMEGAs` outputs. These inlcude: 'dm_Omega', 'dm_spin', 'dm_candidate`, `mo_output`, `dm_c_{bino,wino,higgsino1,higgsino2}`, which are, respectively: dark matter relic density value, dark matter candidate spin, dark matter candidate, the whole `MicrOMEGAs` output, and the coefficient of `{bino,wino,higgsino1,higgsino2}` components of the dark matter state.</li> </ul> </li> </ul> <p>Versions:</p> <ul> <li>SPheno 4.0.5, with a patch to output a warning when the LSP is charged. This version can be found <a href="https://gitlab.com/lip_ml/blackboxbsm">here</a>.</li> <li>MicrOMEGAs 5.3.41, with the MSSM model adapted for low-scale slha inputs.</li> </ul> <p>The datasets are provided in <a href="https://parquet.apache.org/">Apache `parquet`</a> format. In order to read them using `pandas`, an installation with the optional flag `[parquet]` should be used. Alternatively, one can use <a href="https://arrow.apache.org/docs/python/index.html">`pyarrow`</a>.</p> <p> </p>
Weather-related Disease Prediction Dataset
<p>This dataset integrates medical symptoms and weather conditions to facilitate research into the prediction of diseases influenced by meteorological factors. The data spans a range of weather parameters alongside reported medical symptoms from a sizable number of anonymized individuals.</p> <p> </p> <p><strong>Variables:</strong></p> <p>Age: Age of the patient.</p> <p>Gender: Gender of the patient (encoded numerically).</p> <p>Temperature (C): Daily average temperature in Celsius.</p> <p>Humidity: Daily average humidity percentage.</p> <p>Wind Speed (km/h): Daily average wind speed in kilometers per hour.</p> <p>Symptoms: Various symptoms such as nausea, joint pain, abdominal pain, high fever, chills, fatigue, runny nose, pain behind the eyes, etc., encoded as binary values (1 for present, 0 for absent).</p> <p>Pre-existing Conditions: Conditions like asthma history, high cholesterol, diabetes, obesity, HIV/AIDS, nasal polyps, high blood pressure, encoded as binary values.</p> <p><strong>Data Collection Method:</strong></p> <p>The data was collected from anonymous medical records and corresponding local weather stations. All personal identifiers have been removed to ensure patient confidentiality and data privacy, adhering to ethical standards for medical data handling.</p> <p><strong>Usage Notes:</strong></p> <p>This dataset is intended for use in academic and research settings, especially in studies focused on the impact of weather on human health. It could be particularly useful for developing machine learning models to predict the likelihood of disease outbreaks based on weather patterns.</p>
Binding Affinity Prediction Workflow - Simulation Input Files and Absolute Binding Free Energies
<p>The Binding Affinity Prediction (BAP) workflow calculates absolute binding free energies for protein-ligand complexes by taking their crystal structures, converting them into input files for molecular dynamics (MD) simulations with GROMACS after they have passed extensive quality checks, and analysing the resulting trajectories with the Generalised Born model of implicit solvation as implemented in gmx_MMPBSA to obtain the free-energy estimates. The workflow was designed for soluble proteins without post-translational modifications, co-factors and non-standard amino acids, and it has limited support for coordinated ions.</p> <p>For the dataset published here, the BAP workflow was run on the PDBbind 2020 (http://www.pdbbind.org.cn/index.php) refined set. This entry contains the MD simulation input files (BAPSimulationInputFiles.tar.gz) and the ABFE estimates (BAPBindingFreeEnergyEstimates.csv) obtained from four 250 ns trajectories for each complex. The MD simulations for more than 4000 complexes were run on the Leonardo supercomputer while the implicit-solvent calculations were carried out on Galileo, both operated by Cineca (Italy). The MD trajectories will be stored at Cineca for approx. 1 year after publication of this entry; contact Cineca's user support if you are interested in the trajectories.</p> <p>The README file describes how to reproduce the MD trajectories and the subsequent implicit-solvent calculations yielding the free-energy estimates. The workflow scripts can be downloaded from GitHub (https://github.com/LigateProject/Binding-Affinity-Prediction-workflow). The MD simulations were run with GROMACS 2023.2 (https://manual.gromacs.org/2023.2/index.html), and the implicit-solvent calculations were carried out with gmx_MMPBSA 1.6.1 (https://valdes-tresanco-ms.github.io/gmx_MMPBSA/v1.6.1/).</p>
P2PXML Dataset: Deep Geometric Framework to Predict Antibody-Antigen Binding Affinity
<p>In drug development, the efficacy of an antibody depends on how the antibody interacts with the target antigen. The strength of these interactions indicates how successful an antibody is in neutralizing an antigen. Therefore, the strength, measured by “binding affinity”, is a critical aspect of antibody engineering. In theory, the higher the binding affinity, the higher the chances are that the antibody is successful against the target antigen. Currently, techniques such as molecular docking and molecular dynamics are utilized in quantifying the binding affinity. However, owing to the computational complexity of the aforementioned techniques, running simulations for large antibodies/antigens remains a daunting task. Despite the commendable improvements in deep learning-based binding affinity prediction, such approaches are highly dependent on the quality of the antibody-antigen structures and they tend to overlook the importance of capturing the evolutionary details of proteins upon mutation. Further, most of the existing datasets for the task only include antibody-antigen pairs related to one antigen variant and, thus, are not suitable for developing comprehensive data-driven approaches. To circumvent the said complexities, we first curate the largest and most generalized datasets for antibody-antigen binding affinity prediction, consisting of both protein sequences and structures. Subsequently, we propose a deep geometric neural network comprising a structure-based model and a sequence-based model that considers both atomistic and evolutionary details when predicting the binding affinity. The proposed framework exhibited a 10% improvement in mean absolute error compared to the state-of-the-art models while showing a strong correlation between the predictions and target values. We release the datasets and code publicly https://drug-discovery-entc.github.io/p2pxml/ to support the development of antibody-antigen binding affinity prediction frameworks for the benefit of science and society. </p>
Data and code from: Climate-based prediction of carbon fluxes from deadwood in Australia
This repository contains the code for the publication 'Climate-based prediction of carbon fluxes from deadwood in Australia'.
AdsMT: Multi-modal Transformer for Predicting Global Minimum Adsorption Energy
<p>We built three Global Minimum Adsorption Energy (GMAE) benchmark datasets named OCD-GMAE, Alloy-GMAE and FG-GMAE from OC20-Dense, Catalysis Hub, and `functional groups' (FG)-dataset datasets through strict data cleaning, and each data point represents a unique combination of catalyst surface and adsorbate. These new benchmark datasets can be beneficial for future ML study on GMAE prediction.</p> <p>In addition, a similar data cleaning procedure was employed on the OC20 dataset to create a new dataset named OC20-LMAE, which comprises surface/adsorbate pairings along with their local minimum adsorption energies (LMAE). The OC20-LMAE dataset contains 363,937 data points and serves as an effective resource for model pretraining.</p>
Datasets for "Advancing Drug-Target Interactions Prediction: Leveraging a Large-Scale Dataset with a Rapid and Robust Chemogenomic Algorithm"
<p>All datasets required to reproduce the results of publication "Drug-Target Interactions Prediction at Scale: the Komet Algorithm with the LCIdb Dataset"</p>
First-principles prediction of the Co-Al phase diagram including configurational, vibrational and magnetic contributions
<p>Documentation for the Dataset used in the publication entitled "First-principles prediction of the Co–Al phase diagram including configurational, vibrational and magnetic contributions" <br>** These datasets comprise all configurations used in Co-Al system and their formation enthalpies at different temperatures, where configurational, vibrational and magnetic contributions were considered. Hcp Co and fcc Al were used as reference states. **<br>** More details about the methodology can be found in the paper "First-principles prediction of the Co-Al phase diagram including configurational, vibrational and magnetic contributions, Journal of Materials Research and Technology, 2024" **</p> <p>1. bcc-Co-Al.zip<br>- Description: bcc-Co-Al.zip is a compressed folder. It contains Al1-xCox configurations with bcc lattice used to fit the cluster expansion (CE). Each folder contains a POSCAR file that correspons to a configuration. The POSCAR can be opened with Notepad and visualized with VESTA software.</p> <p>2. fcc-Co-Al.zip<br>- Description: fcc-Co-Al.zip is a compressed folder. It contains Al1-xCox configurations with fcc lattice used to fit the CE. Each folder contains a POSCAR file that correspons to a configuration. The POSCAR can be opened with Notepad and visualized with VESTA software.</p> <p>3. hcp-Co-Al.zip<br>- Description: hcp-Co-Al.zip is a compressed folder. It contains Al1-xCox configurations with hcp lattice used to fit the CE. Each folder contains a POSCAR file that correspons to a configuration. The POSCAR can be opened with Notepad and visualized with VESTA software.</p> <p><br>4. Formation enthalpies of bcc-Co-Al.xlsx<br>- Description: Formation enthalpies of bcc lattice in Co-Al system at different temperatures, which includes the effect of lattice vibration and magnetic excitation. Fcc Al and hcp Co were used as reference states.</p> <p>- Variable description by columns:<br> 1-(Folder name) - type: numerical (integer)<br> Description: Each folder name in the bcc-Co-Al.zip corresponds to a configuration.<br> 2- (at. fraction of Co (%)) - type: numerical (float)<br> Description: The atomic fraction of Co in each configuration.<br> 3- (H_f^(conf)(DFT) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 0 K calculated by density functional theory (DFT) following eq.(18) in the paper.<br> 4- (H_f^(conf)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 0 K fitted by CE. <br> 6- (at. fraction of Co (%)) - type: numerical (float)<br> Description: The atomic fraction of Co in each configuration.<br> 7- (H_f^(conf+vib+mag)(Cal.) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 400 K calculated by DFT, the bond length vs. bond stiffness relationship and Monte Carlo simulation of the Heisenberg Hamiltonian following eq.(20) in the paper.<br> 8- (H_f^(conf+vib+mag)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 400 K fitted by CE. <br> 10- (at. fraction of Co (%)) - type: numerical (float)<br> Description: The atomic fraction of Co in each configuration.<br> 11- (H_f^(conf+vib+mag)(Cal.) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 800 K calculated by DFT, the bond length vs. bond stiffness relationship and Monte Carlo simulation of the Heisenberg Hamiltonian following eq.(20) in the paper.<br> 12- (H_f^(conf+vib+mag)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 800 K fitted by CE. <br> 14- (at. fraction of Co (%)) - type: numerical (float)<br> Description: The atomic fraction of Co in each configuration.<br> 15- (H_f^(conf+vib+mag)(Cal.) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 1200 K calculated by DFT, the bond length vs.bond stiffness relationship and Monte Carlo simulation of the Heisenberg Hamiltonian following eq.(20) in the paper.<br> 16- (H_f^(conf+vib+mag)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 1200 K fitted by CE.<br> 18- (at. fraction of Co (%)) - type: numerical (float)<br> Description: The atomic fraction of Co in each configuration.<br> 19- (H_f^(conf+vib+mag)(Cal.) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 1600 K calculated by DFT, the bond length vs.bond stiffness relationship and Monte Carlo simulation of the Heisenberg Hamiltonian following eq.(20) in the paper.<br> 20- (H_f^(conf+vib+mag)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 1600 K fitted by CE.</p> <p><br>5. Formation enthalpies of fcc Co-Al.xlsx<br>- Description: Formation enthalpies of fcc lattice in Co-Al system at different temperatures, which includes the effect of lattice vibration and magnetic excitation. Fcc Al and hcp Co were used as reference states.</p> <p>- Variable descriptions by columns are the same as those of Formation enthalpies of bcc-Co-Al.xlsx.</p> <p><br>6. Formation enthalpies of hcp-Co-Al.xlsx<br>- Description: Formation enthalpies of hcp lattice in Co-Al system at different temperatures, which includes the effect of lattice vibration and magnetic excitation. Fcc Al and hcp Co were used as reference states.</p> <p>- Variable descriptions by columns are the same as those of Formation energies of bcc-Co-Al.xlsx.</p> <p><br>7. ECIs of bcc-Co-Al at different temperatures.txt<br>- Description: ECIs of bcc lattice in Co-Al system from 0 to 2000 K with increment step of 10 K. The ECIs at different temperatures are separated by blank lines. ECIs at 0 K means that only configurational contribution was considered. ECIs at finite temperature means that configurational, vibrational and magnetic contributions were considered.</p> <p><br>8. ECIs of fcc-Co-Al at different temperatures.txt<br>- Description: ECIs of fcc lattice in Co-Al system from 0 to 2000 K with increment step of 10 K. The ECIs at different temperatures are separated by blank lines. ECIs at 0 K means that only configurational contribution was considered. ECIs at finite temperature means that configurational, vibrational and magnetic contributions were considered.</p> <p><br>9. ECIs of hcp-Co-Al at different temperatures.txt<br>- Description: ECIs of hcp lattice in Co-Al system from 0 to 2000 K with increment step of 10 K. The ECIs at different temperatures are separated by blank lines. ECIs at 0 K means that only configurational contribution was considered. ECIs at finite temperature means that configurational, vibrational and magnetic contributions were considered.</p> <p><br>10. Clusters of bcc-Co-Al.txt<br>- Description: Cluster information of bcc lattice in Co-Al system. Each cluster is separated by a blank line. Each cluster contains: multiplicity; Length of the longest pair within the cluster; number of points in cluster; coordinates of point. They are arranged in a row.</p> <p><br>11. Clusters of fcc-Co-Al.txt<br>- Description: Cluster information of fcc lattice in Co-Al system. Each cluster is separated by a blank line. Each cluster contains: multiplicity; Length of the longest pair within the cluster; number of points in cluster; coordinates of point. They are arranged in a row.</p> <p><br>12. Clusters of hcp-Co-Al.txt<br>- Description: Cluster information of hcp lattice in Co-Al system. Each cluster is separated by a blank line. Each cluster contains: multiplicity; Length of the longest pair within the cluster; number of points in cluster; coordinates of point. They are arranged in a row.</p>
First principles prediction of the Al-Li phase diagram including configurational and vibrational entropic contributions
<p>Documentation for the Dataset used in the publication entitled "First principles prediction of the Al-Li phase diagram including configurational and vibrational entropic contributions" <br>** These datasets comprise all configurations used in Al-Li system and their formation enthalpies at different temperatures, where both configurational and vibrational contribution were considered. Bcc Li and fcc Al were used as reference state. **<br>** More details about the methodology can be found in the paper "Wei Shao, Sha Liu, Javier LLorca, First principles prediction of the Al-Li phase diagram including configurational and vibrational entropic contributions, Computational Materials Science, 2023"**</p> <p>1. bcc-Al-Li.zip<br>- Description: bcc-Al-Li.zip is a compressed folder. It contains Al1-xLix configurations with bcc lattice used to fit the cluster expansion (CE). Each folder contains a POSCAR file that corresponds to a configuration. The POSCAR can be opened with Notepad and visualized with VESTA software.</p> <p><br>2. fcc-Al-Li.zip<br>- Description: fcc-Al-Li.zip is a compressed folder. It contains Al1-xLix configurations with fcc lattice used to fit the CE. Each folder contains a POSCAR file that corresponds to a configuration. The POSCAR can be opened with Notepad and visualized with VESTA software.</p> <p>3. Formation enthalpies of bcc-Al-Li.xlsx<br>- Description: Formation enthalpy of each configuration in bcc Al-Li system at different temperatures, which includes the effect of lattice vibration. Bcc Li and fcc Al were used as reference state.</p> <p>- Variable descriptions by columns:<br> 1-(Folder nam) - type: numerical (integer)<br> Description: Each folder name in the bcc-Al-Li.zip corresponds to a configuration.<br> 2- (at. fraction of Li (%)) - type: numerical (float)<br> Description: The atomic fraction of Li in each configuration.<br> 3- (H_f^(conf)(DFT) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 0 K calculated by density functional theory (DFT).<br> 4- (H_f^(conf)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration fitted by CE at 0 K. <br> 6- (at. fraction of Li (%)) - type: numerical (float)<br> Description: The atomic fraction of Li in each configuration.<br> 7- (H_f^(conf+vib)(DFT+L-S) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 100 K calculated by DFT and bond length vs. bond stiffness relationship (L-S).<br> 8- (H_f^(conf+vib)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 100 K fitted by CE. <br> 10- (at. fraction of Li (%)) - type: numerical (float)<br> Description: The atomic fraction of Li in each configuration.<br> 11- (H_f^(conf+vib)(DFT+L-S) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 200 K calculated by DFT and L-S.<br> 12- (H_f^(conf+vib)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 200 K fitted by CE. <br> 14- (at. fraction of Li (%)) - type: numerical (float)<br> Description: The atomic fraction of Li in each configuration.<br> 15- (H_f^(conf+vib)(DFT+L-S) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 300 K calculated by DFT and L-S.<br> 16- (H_f^(conf+vib)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 300 K fitted by CE.<br> 18- (at. fraction of Li (%)) - type: numerical (float)<br> Description: The atomic fraction of Li in each configuration.<br> 19- (H_f^(conf+vib)(DFT+L-S) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 400 K calculated by DFT and L-S.<br> 20- (H_f^(conf+vib)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 400 K fitted by CE.<br> 22- (at. fraction of Li (%)) - type: numerical (float)<br> Description: The atomic fraction of Li in each configuration.<br> 23- (H_f^(conf+vib)(DFT+L-S) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 500 K calculated by DFT and L-S.<br> 24- (H_f^(conf+vib)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 500 K fitted by CE. <br> 26- (at. fraction of Li (%)) - type: numerical (float)<br> Description: The atomic fraction of Li in each configuration.<br> 27- (H_f^(conf+vib)(DFT+L-S) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 600 K calculated by DFT and L-S.<br> 28- (H_f^(conf+vib)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 600 K fitted by CE.<br> 30- (at. fraction of Li (%)) - type: numerical (float)<br> Description: The atomic fraction of Li in each configuration.<br> 31- (H_f^(conf+vib)(DFT+L-S) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 700 K calculated by DFT and L-S.<br> 32- (H_f^(conf+vib)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 700 K fitted by CE.<br> 34- (at. fraction of Li (%)) - type: numerical (float)<br> Description: The atomic fraction of Li in each configuration.<br> 35- (H_f^(conf+vib)(DFT+L-S) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 800 K calculated by DFT and L-S.<br> 36- (H_f^(conf+vib)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 800 K fitted by CE. <br> 38- (at. fraction of Li (%)) - type: numerical (float)<br> Description: The atomic fraction of Li in each configuration.<br> 39- (H_f^(conf+vib)(DFT+L-S) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 900 K calculated by DFT and L-S.<br> 40- (H_f^(conf+vib)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 900 K fitted by CE.<br> 42- (at. fraction of Li (%)) - type: numerical (float)<br> Description: The atomic fraction of Li in each configuration.<br> 43- (H_f^(conf+vib)(DFT+L-S) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 1000 K calculated by DFT and L-S.<br> 44- (H_f^(conf+vib)(CE) (eV/atom)) - type: numerical (float)<br> Description: Formation enthalpy of each configuration at 1000 K fitted by CE.</p> <p>4. Formation enthalpies fcc-Al-Li.xlsx<br>- Description: Formation enthalpy of each configuration in fcc Al-Li system at different temperatures, which includes the effect of lattice vibration. Bcc Li and fcc Al were used as reference state.</p> <p>- Variable descriptions by columns are the same as those of Formation enthalpies of bcc-Al-Li.xlsx.</p> <p><br>5. ECIs of bcc-Al-Li at different temperatures.txt<br>- Description: ECIs of bcc lattice in Al-Li system from 0 to 2000 K with increment step of 10 K. The ECIs at different temperatures are separated by blank lines. ECIs at 0 K means that only configurational contribution was considered. ECIs at finite temperature means that both configurational and vibrational contributions were considered.</p> <p><br>6. ECIs of fcc-Al-Li at different temperatures.txt<br>- Description: ECIs of fcc lattice in Al-Li system from 0 to 2000 K with increment step of 10 K. The ECIs at different temperatures are separated by blank lines. ECIs at 0 K means that only configurational contribution was considered. ECIs at finite temperature means that both configurational and vibrational contributions were considered.</p> <p><br>7. Clusters of bcc-Al-Li.txt<br>- Description: Cluster information of bcc lattice in Al-Li system. Each cluster is separated by a blank line. Each cluster contains: multiplicity; Length of the longest pair within the cluster; number of points in cluster; coordinates of point. They are arranged in a row.</p> <p><br>8. Clusters of fcc-Al-Li.txt<br>- Description: Cluster information of fcc lattice in Al-Li system. Each cluster is separated by a blank line. Each cluster contains: multiplicity; Length of the longest pair within the cluster; number of points in cluster; coordinates of point. They are arranged in a row.</p>
Metadata for RadPhysBio: A Radiobiological Database for the Prediction of Cell Survival upon Exposure to Ionizing Radiation
<p>This is the Metadata of our recent publication https://doi.org/10.3390/ijms25094729 </p> <p>Based on the need for radiobiological databases, in this work, we mined experimental ionizing radiation data of human cells treated with X-rays, γ-rays, carbon ions, protons and α-particles, by manually searching the relevant literature in PubMed from 1980 until 2024. In order to calculate normal and tumor cell survival α and β coefficients of the linear quadratic (LQ) established model, as well as the initial values of the double-strand breaks (DSBs) in DNA, we used WebPlotDigitizer and Python programming language. We also produced complex DNA damage results through the fast Monte Carlo code MCDS in order to complete any missing data. In the attached files you will find</p> <ol> <li>Current database for photons</li> <li>Current database for particle radiation</li> <li>Helping supplementary information</li> <li>Tips for help with our Database</li> </ol>
Data archive and code for "Predicting September Arctic Sea Ice: A Multi-Model Seasonal Skill Comparison"
<p>This upload contains data and code related to the paper "Predicting September Arctic Sea Ice: A Multi-Model Seasonal Skill Comparison" by M. Bushuk, S. Ali, D. Bailey, Q. Bao, L. Batte, U. S. Bhatt, E. Blanchard-Wrigglesworth, E. Blockley, G. Cawley, J. Chi, F. Counillon, P. Goulet Coulombe, R. Cullather, F. X. Diebold, A. Dirkson, E. Exarchou, M. Gobel, W. Gregory, V. Guemas, L. Hamilton, B. He, S. Horvath, M. Ionita, J. E. Kay, E. Kim, N. Kimura, D. Kondrashov, Z. M. Labe, W. Lee, Y. J. Lee, C. Li, X. Li, Y. Lin, Y. Liu, W. Maslowski, F. Massonnet, W. N. Meier, W. J. Merryfield, H. Myint, J. C. Acosta Navarro, A. Petty, F. Qiao, D. Schroder, A. Schweiger, Q. Shu, M. Sigmond, M. Steele, J. Stroeve, N. Sun, S. Tietsche, M. Tsamados, K. Wang, J. Wang, W. Wang, Y. Wang, Y. Wang, J. Williams, Q. Yang, X. Yuan, J. Zhang, and Y. Zhang, published in the Bulletin of the American Meteorological Society, DOI: https://doi.org/10.1175/BAMS-D-23-0163.1.</p> <p>See README.txt for a description of the datasets and code.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.