Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,782
datasets available to search
ShareScore release 0.7.1
Dataset results
1,782 results for “algorithms”
Data and scripts for "Simulating AMOC tipping driven by internal climate variability with a rare event algorithm."
<p>This dataset contains supplementary material for the paper Simulating AMOC tipping driven by internal climate variability with a rare event algorithm." (M.Cini, G. Zappa, F. Ragone, S. Corti, 2023) submitted to <i>npj Climate and Atmospheric Science. </i> Preprint is available at https://www.researchsquare.com/article/rs-3215995/latest.<br>Here we uploaded most relevant data and scripts concerning this study. Feel free to contact us for other resources.<br><br>This datasets contains:</p><p>1) The output of one 125-years ensemble simulation performed with the Algorithm.</p><p>2) Time series of the AMOC index for all the simulations performed.<br>3) All data-analysis related scripts. These scripts have been used to plot all the figures in the paper.</p><p>The script for the rare event algorithm is already available in the Supplementary material for "Rare event algorithm study of extreme warm summers and heatwaves over Europe" zenodo repository, available at https://zenodo.org/records/4763283.<br><br><strong>Simulation Architecture and model setup</strong></p><p>All the simulations have been performed with an intermediate complexity coupled climate model, composed by the Planet Simulator (PlaSim) and the Large Scale Geostrophic Ocean (LSG). All simulations are performed at stationary greenhouse gases forcing. More information about model setup and scope of the simulations can be found in the paper.<br><br>We performed 10 100-member ensemble simulations. First 125 years of the simulation are performed with the algorithm on (k=3). Then, simulations have been restarted at year 120 with the algorithm off (k=0) up to year 400. This last 380 years of simulation have been performed with only 20 members.<br><br><strong>y2480_k3_ntraj100</strong></p><p>y2480_k3_ntraj100.tar.gz contains the output of one ensemble simulation (y2480, i.e. the one that starts at year 2480 of the control run simulation) performed with the algorithm, i.e. contains data of 125 years simulation of 100 members.<br>Data is organized in blocks for each years. Every block contains the 4 NetCDF light file for each member and 4 files with full-size output that represent the mean state as the average of the 100 members. The 4 different file name accounts for the 4 different module output of the model: "data" for the atmosphere, "ice" for sea ice, "ocean" for the slab ocean layer in PlaSIM, lsg for LSG dynamical ocean.</p><p> </p><p><strong>Time Series</strong></p><p>Time series of the AMOC index are contained in 10 files representing the 10 different simulations performed. In each file are present 125 .txt files, one for each year of the simulation with the algorithm on, with the annual average AMOC indices of the 100 members, and 380 .txt files, one for each year of the simulation with the algorithm off, with the annual average AMOC indices of the 20 members.</p><p> </p><p><strong>Response Analysis REA</strong></p><p>Response Analysis REA contains the script that generates Fig.2, Fig. S3, Fig. S5, Fig. S6 and Fig. S7 of the paper. In general it provides tools for data analysis of the climate response to an AMOC slowdown. Be aware that data of the lsg module needs different processing. </p><p> </p><p><strong>AMOC Evolution REA</strong></p><p>AMOC Evolution REA contains the script that generates Fig.1, Fig. 4, Fig. 5, Fig. 6 and Fig. S2 of the paper. In general it provides tools for analysis of time series and scatter plots of the AMOC evolution. <br><br> </p><p><strong>Causes REA</strong></p><p>Causes REA contains the script that generates Fig.3, Fig. S4 of the paper. . In general it provides tools for analysis of driving elements of the AMOC decline. More information about these methods can be found in the "Triggering mechanisms" section of the paper. Be aware that data of the lsg module needs different processing. </p>
Temporal Patterns and Trends in Corporate Donations Using PageRank and Node Similarity Graph Algorithm
<p>Corporate donations wield considerable influence within political arenas, shaping policies and influencing decision-making processes. This study uses Neo4j, an advanced graph database tool, to explore a comprehensive company dataset, focusing on unraveling temporal patterns and evolving trends in corporate contributions. Visual representations, such as bar charts, reveal significant fluctuations in donations, indicating potential cyclic patterns occurring every six years. The study explores intricate relationships between donor entities and recipients, highlighting diverse donation patterns—both focused and widespread. The study's derived PageRank scores offer a comprehensive portrayal of the varying degrees of influence among diverse entities receiving donations within the network. Notably, the Conservative and Unionist Party emerges as the most prominent entity, boasting a striking score of 1.86, indicating a substantial influx of financial support likely to significantly shape its political endeavors. Despite a lower score of 0.62, the Labor Party still signifies a noteworthy level of financial backing, albeit less extensive than its counterpart. In contrast, the Liberal Democrats, The In Campaign Ltd, and Network for Animals Ltd exhibit comparatively restrained financial backing, warranting deeper investigation into the factors affecting their funding. Moreover, undisclosed findings regarding 170 similarity scores using Node Similarity algorithm disclose a prevalent similarity trend among entities, notably observed between Company 1 and Company 2, implying potential synergistic partnerships in donation-related endeavors. This high similarity often indicates shared values, highlighting prospects for collaborative initiatives or partnerships to augment positive impacts. Utilizing these insights supports the formulation of targeted donation strategies, circumventing donation redundancies, and ensuring optimal resource allocation for maximal societal benefit within specified sectors.</p> <p>Keywords—Company Dataset, Corporate Donations, Neo4j, Node Similarity, PageRank, Political Influence </p> <p> </p>
Datasets: Enhancing Anger Management via Reinforcement Learning: A Comparative Analysis of the PPO Algorithm with Optimised Hyperparameters
Open the record for dataset details and reuse information.
Improving Algorithm-Selectors and Performance-Predictors via Learning Discriminating Training Samples - Code and Data
<p>This repository contains the code and data for reproducibility of the paper 'Improving Algorithm-Selectors and Performance-Predictors via Learning Discriminating Training Samples'. </p> <p>The following files are included:</p> <ul> <li>Plots: Additional plots not in the paper;</li> <li>Code: Python scripts to generate trajectories and perform classification/regression;</li> <li>best_algo.csv : Labels for the classification;</li> <li>performances.csv : Performances used for the regression;</li> <li>SA_parameters.csv : SA parameters for all machine learning tasks;</li> <li>irace_scenario.txt : scenario used for the tuning.</li> </ul>
Identifying Easy Instances to Improve Efficiency of ML Pipelines for Algorithm-Selection - Code and Data
<p>This repository contains the code and data for reproducibility of the paper 'Identifying Easy Instances to Improve Efficiency of ML Pipelines for Algorithm-Selection'. </p> <p>The following files are included:</p> <ul> <li>best_algo.csv : labels for the classification and median performance of algorithms;</li> <li>ML_models.ipyb : jupyter notebook with the definition of the neural networks for both classifiers;</li> <li>pickle.zip : pickled models for the hardness classification and the algorithm selector;</li> <li>trajectories.zip : raw data files containing parts of the trajectories of each algorithm;</li> <li>Results.zip : results obtained using the approach on the stream of instances.</li> </ul>
Sea-surface pCO2 maps for the Bay of Bengal based on machine learning algorithms
<p>The dataset contains two products, the first being sea-surface <em>p</em>CO<sub>2</sub> and the second being air-sea CO<sub>2</sub> flux for the Bay of Bengal region. The data is climatological data, with 12 months of the climatological year. Each of these data has a spatial resolution of 1/12<sup>o</sup>. The positive value of CO<sub>2</sub> flux indicates the outgassing of CO<sub>2</sub>, and the negative value shows the uptake of atmospheric CO<sub>2</sub>. This version contains an elaborate description in the attribute section of the NC files, which is missing from the previous versions.</p>
Accompanying data for the paper "Robustness of the Data-Driven Identification algorithm with incomplete input data"
<h2>Links</h2> <ul> <li>isSupplementTo <em>publication-article</em> <a href="https://doi.org/10.46298/jtcam.12590">https://doi.org/10.46298/jtcam.12590</a></li> <li>isNewVersionOf <em>dataset</em> <a href="../records/10090469">https://zenodo.org/records/10090469</a></li> </ul> <h2>Authors</h2> <ul> <li><strong>Leygue, Adrien</strong>, Ecole Centrale de Nantes, ORCID: <a href="https://orcid.org/0000-0003-0714-822X">0000-0003-0714-822X</a></li> </ul> <h2>Language</h2> <ul> <li>English</li> </ul> <h2>License</h2> <ul> <li>Creative Commons Attribution 4.0</li> </ul> <h2>Funding sources</h2> <ul> <li>This work was performed by using HPC resources of Centrale Nantes Supercomputing Center on the cluster Liger, granted and identified D1705030 by the High Performance Computing Institute(ICI).</li> </ul> <h2>Data structure and information</h2> <p>Synthetic data used in the case study (section 3) of the paper.</p> <p>The data in XDMF (Milou.xdmf ) + hdf5 (Milou.hdf5) format comprises:</p> <ol> <li>The 2D computational mesh with triangular linear elements</li> <li>The nodal Forces for all loading steps (nodal quantity)</li> <li>The displacement for all loading steps (nodal quantity)</li> <li>Cauchy stress fields for all loading steps (cell quantity)</li> </ol>
Data from: Quantifying the impact of internal variability on the CESM2 control algorithm for stratospheric aerosol injection dataset
<p>Earth system models are a powerful tool to simulate the response to hypothetical climate intervention strategies, such as stratospheric aerosol injection (SAI). Recent simulations of SAI implement tools from control theory, called "controllers", to determine the quantity of aerosol to inject into the stratosphere to reach or maintain specified global temperature targets, such as limiting global warming to 1.5C above pre-industrial temperatures. This work explores how internal (unforced) climate variability can impact controller-determined injection amounts using the Assessing Responses and Impacts of Solar climate intervention on the Earth system with Stratospheric Aerosol Injection (ARISE-SAI) simulations. Since the ARISE-SAI controller determines injection amounts by comparing global annual-mean surface temperature to predetermined temperature targets, internal variability that impacts temperature can impact the total injection amount as well. Using an offline version of the ARISE-SAI controller and data from CESM2 earth system model simulations, we quantify how internal climate variability and volcanic eruptions impact injection amounts. While idealized, this approach allows for the investigation of a large variety of climate states without additional simulations and can be used to attribute controller sensitivities to specific modes of internal variability.</p>
Auxiliary data release for "Fast marginalization algorithm for optimizing gravitational wave detection, parameter estimation and sky localization"
<p>This release contains parameter estimation runs on synthetic injections, described in https://arxiv.org/abs/2404.02435 .</p> <p>Important note: The posterior samples provided are weighted, the weights are stored in a column named 'weights' . </p>
Source code and simulation results: Efficient rational approximation of optical response functions with the AAA algorithm
<p>This publication provides data published in the article "Efficient rational approximation of optical response functions with the AAA algorithm" [1] in tabulated form along with the Matlab scripts that have been used to produce them. These scripts interface the finite element method solver JCMsuite [2,3]. The article presents rational approximations of optical response functions based on an extended version of the AAA algorithm [4] that allows to efficiently reconstruct sensitivty spectra and gives access to sensitivities of poles, residues, and zeros. Furthermore, the rational approximation of a scalar observalbe is used to construct solutions of the source free Maxwell's equation, i.e., a nonlinear eigenvalue problem. </p> <p><strong>The physical Structure</strong></p> <p>The example is based on the chiral metasurface introduced in [5]. For the sake of simplicity we added infinite layers of SiO\(_2\) to the top and the bottom of the structure. The original structure has a SiO\(_2\) substrate and a layer of PMMA polymethyl methacrylate (PMMA) deposited on top. PMMA can be modelled with the same refractive index of 1.45 as SiO\(_2\). Furthermore, our simulations include the 13 nm indium tin oxide (ITO) coating which drastically reduces the Q-factor as it is slightly absorbing. The accuracy of the discrete model is verified by assessing reflection, transmission, and absorption at 241 evenly spaced points within the specified range. Energy conservation requires that the discrepancy between their sum and the energy entering the system is zero. The numerical discretization is chosen such that the maximum relative error is less than \(3\times10^{−5}\).</p> <p><strong>Dispersion</strong></p> <p>Tabulated data for ITO has been taken from the <a href="https://refractiveindex.info/?shelf=other&book=In2O3-SnO2&page=Konig">refractiveindex.info</a> database (T. A. F. König et al., 2014, https://doi.org/10.1021/nn501601e) and the data for TiO2 was kindly provided the authors of [5]. The permittivity \(\varepsilon = (n+ik)^2\) is locally approximated as a rational function, i.e., only data in a vicinity of the frequency range of interest is considered. As we aim for a function with the symmetry \(f^\ast(\omega) = f(-\omega^\ast)\) we add the complex conjugated data at negative frequencies and enforce the symmetry in a second step. The partial fraction decomposition of the required function is of the form: \(\varepsilon(\omega) = \varepsilon_\infty + \sum_{j=1}^{4}a_j/(\omega-\omega_j) - a_j^\ast/(\omega+\omega_j^\ast)\) with the residues \(a_j\) and the poles \(\omega_j\). We expect 4 pairs of poles to sufficiently approximate the data within the range of interest (4 with positive and 4 with negative real parts).</p> <h4><strong>Requirements</strong></h4> <ul> <li>JCMsuite (at least 6.2.0)</li> <li>MATLAB (tested with version R2023b)</li> </ul> <p>In order to run the simulations with JCMsuite you must replace corresponding place holders with a path to your installation of JCMsuite. Free trial licenses are available, please refer to the homepage of <a href="https://jcmwave.com/">JCMwave</a>.</p> <p><strong>Usage</strong></p> <p>With the content of 'spectra.zip' you can reproduce results presented in the paper. Running the script 'plots.m' will not start any expensive simulation but use the provided data. With 'dispersion.m' the fits to the material data can be reproduced. Additionally, tabulated data is contained in 'data/ascii'. The archive 'eigenmodes.zip' must be extracted in the same directory as 'spectra.zip'.</p> <p><strong>References</strong></p> <p>[1] Fridtjof Betz, Martin Hammerschmidt, Lin Zschiedrich, Sven Burger, Felix Binkowski: Efficient rational approximation of optical response functions<br>with the AAA algorithm, https://doi.org/10.48550/arXiv.2403.19404.</p> <p>[2] Jan Pomplun, Sven Burger, Lin Zschiedrich, Frank Schmidt, Adaptive finite element method for simulation of optical nano structures, Physica Status Solidi B <strong>244</strong>, 3419 (2007), http://dx.doi.org/10.1002/pssb.200743192.</p> <p>[3] Fridtjof Betz, Felix Binkowski, Sven Burger, RPExpand: Software for Riesz projection expansion of resonance phenomena, SoftwareX <strong>15</strong>, 100763 (2021), https://doi.org/10.1016/j.softx.2021.100763.</p> <p>[4] Y. Nakatsukasa, O. Sète, and L. N. Trefethen, The AAA Algorithm for Rational Approximation, SIAM Journal on Scientific Computing <strong>40</strong>, A1494 (2018), http://dx.doi.org/10.1137/16M1106122.</p> <p>[5] X. Zhang, Y. Liu, J. Han, Y. Kivshar, and Q. Song, Chiral emission from resonant metasurfaces, Science <strong>377</strong>, 1215 (2022), http://dx.doi.org/%2010.1126/science.abq7870.</p>
Automatic message sequence chart creation from simulation run of the parametric colored Petri net model of the Chandy-Lamport algorithm with four processes
<p><span>The video shows the creation of the message sequence chart from a simulation run of the proposed parametric colored Petri net model of the Chandy-Lamport algorithm using the CPN tool with four constituting processes. The picture shows the resulting message sequence chart. </span></p> <p><strong><span>Message Sequence Chart of Parametric Model With 4 Processes via Automatic Simulation Run_SuppInfo.mp4</span></strong><span>: This video shows the automatic generation of a message sequence chart of the proposed parametric colored Petri net model of the Chandy-Lamport distributed global snapshot algorithm using the CPN tool version 4.0.0. The model's number of constituting processes is parametric and was set to four. The video was generated using the authors' updated CPN tool extension server. The automatic simulation run of the model has been used to create this video. The CPN tool randomly selects the enabled transition at each step in an automatic simulation run.</span></p> <p><strong><span>Picture of Message Sequence Chart of Parametric Model With 4 Processes_SuppInfo.png:</span></strong><span> This picture shows the automatically generated message sequence chart of the proposed parametric colored Petri net model of the Chandy-Lamport algorithm that is visible in the above video clip. The number of constituting processes was set to four. </span></p>
PERFORMANCE OF MACHINE LEARNING ALGORITHMS FOR LUNG CANCER PREDICTION: A COMPARATIVE STUDY
<p>This study compares the performance of five machine learning algorithms—logistic regression, support vector machines, random forests, gradient boosting, and neural networks—for lung cancer prediction using demographic, lifestyle, and medical data from the UCI Machine Learning Repository. Gradient boosting and random forests achieved the highest accuracy (89% and 87%, respectively) and AUC-ROC scores (0.93 and 0.92), while neural networks reached 90% accuracy but presented interpretability limitations. Key predictors included smoking history, chronic disease, and respiratory symptoms, aligning with established risk factors. Ensemble methods, particularly gradient boosting and random forests, provided an optimal balance of accuracy and interpretability, highlighting their potential for clinical applications in early lung cancer detection.</p>
The Multi-Temporal Dual Channel Algorithm (MT-DCA)
<p>I) SUMMARY</p> <p>This soil moisture and vegetation optical depth product is called the Multi-Temporal Dual Channel Algorithm (MT-DCA). It retrieves surface soil moisture and vegetation optical depth (directly related to total water volume in the vegetation canopy) from <a href="https://nsidc.org/data/SPL1CTB_E">SMAP level 1C brightness temperature</a> observations using a robust estimation technique. It is an in-house MIT algorithm and is not an official SMAP product. The data are freely available on 9km and 36km grids from April 2015 to July 2021 in daily time steps.</p> <p>No co-authorship is required for use of this data in publications. However, to properly acknowledge the dataset when publishing any research using the MT-DCA, we ask data users to (1) cite the DOI as an in-text citation and/or in the data acknowledgements in any publication and (2) reference <a href="https://www.sciencedirect.com/science/article/pii/S0034425717302961">Konings et al. (2017</a>) when referring to the MT-DCA in the text. Feel free to send us an email at <a href="mailto:afeld24@mit.edu">afeld24@mit.edu</a> to let us know how you are using the data. </p> <p>The version 5 update is a re-implementation of the MT-DCA using the updated SMAP L1C brightness temperatures. It extends the data through July 2021.</p> <p>II) CONTACT</p> <p>For questions, please email Andrew Feldman at <a href="mailto:afeld24@mit.edu">afeld24@mit.edu</a>.</p> <p>III) ALGORITHM DESCRIPTION</p> <p>The algorithmic approach uses both horizontally and vertically polarized brightness temperatures to retrieve soil moisture and VOD simultaneously. The key innovation of the MT-DCA is that it recognizes that classical dual-channel algorithms are under-determined: brightness temperature observations are correlated and cannot retrieve two unknowns (soil moisture and VOD) (as illustrated in Konings et al, RSE 2016). This creates amplifying errors in retrievals from snapshot dual-channel algorithms. The MT-DCA uses a viable assumption that VOD changes more slowly than soil moisture between overpasses, and uses information from multiple SMAP overpasses to stabilize the retrieval. It is considered a regularization approach similar to the Sobolev Norm regularization. Specifically, this approach is applied to each temporally adjacent pair of overpasses (for SMAP, two overpasses approximately 2-3-days apart), which includes four brightness temperature measurements. For each overpass pair, the soil moisture at both overpasses is retrieved, along with a constant VOD for both overpasses. This leads to two retrievals of each of soil moisture and VOD at any given overpass time: one where the parameters are retrieved using additional TB information from the overpass before and one from the overpass after. Both retrievals of VOD and soil moisture values at each overpass are averaged. Ultimately, VOD is not held constant, but rather is slowed in time between overpasses. A second key innovation of the MT-DCA is that, because the retrievals are no longer under determined, it is also possible to retrieve a constant single scattering albedo for each pixel. The single scattering albedo is estimated through model selection of the value of the parameter that minimizes the sum of all overpass cost functions. The retrieved albedo is also included in the files here. VOD is reported at nadir.</p> <p>The single scattering albedo is assumed constant over the full record of SMAP data, as is currently accepted practice across approaches with SMAP, SMOS, and AMSR. There is a high amount of computational power required to retrieve an albedo over more than three years of SMAP data. Therefore, an adjustment was made: the single scattering albedo was retrieved over the third year of SMAP data (April 1st, 2017 to March 31st, 2018). This constant value was then applied to the other years without requiring the albedo optimization loop. Tests across many individual pixels revealed that albedo in the third year does not differ greatly from albedo over all years and the other individual years. </p> <p>The algorithm is described in more detail in Konings et al. (2017). The algorithm is based on principles explained in more detail in Konings et al. (2016), which describes the original algorithm development using Aquarius observations. See also the related Konings et al. (2015) publication for quantitative justification for the approach. While the dataset has not been officially validated, the MT-DCA soil moisture retrievals show in-situ comparison statistics similarly to the official baseline SMAP soil moisture product (SMAP soil moisture retrieval in-situ assessment can be found in Chan et al. (2016)). Finally, the MT-DCA vegetation optical depth retrievals are not validated due to only sparsely available ground information related to vegetation water content. Nevertheless, information about error propagation into the MT-DCA soil moisture and VOD retrievals as well as VOD error reductions using the MT-DCA regularization technique can be found in Feldman et al. (2021).</p> <p>Chan, S.K., Bindlish, R., O’Neill, P.E., Njoku, E., Jackson, T., Colliander, A., Chen, F., Burgin, M., Dunbar, S., Piepmeier, J., Yueh, S., Entekhabi, D., Cosh, M.H., Caldwell, T., Walker, J., Wu, X., Berg, A., Rowlandson, T., Pacheco, A., McNairn, H., Thibeault, M., Martinez-Fernandez, J., Gonzalez-Zamora, A., Seyfried, M., Bosch, D., Starks, P., Goodrich, D., Prueger, J., Palecki, M., Small, E.E., Zreda, M., Calvet, J.C., Crow, W.T., Kerr, Y., 2016. Assessment of the SMAP Passive Soil Moisture Product. IEEE Trans. Geosci. Remote Sens. 54, 4994–5007. <a href="https://doi.org/10.1109/TGRS.2016.2561938">https://doi.org/10.1109/TGRS.2016.2561938</a></p> <p>Feldman, A.F., D. Chaparro, and D. Entekhabi (2021). Error propagation in microwave soil moisture and vegetation optical depth retrievals. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. In Press.</p> <p>Konings, A.G., M. Piles, N. Das, and D. Entekhabi (2017). L-band vegetation optical depth and effective scattering albedo estimation from SMAP. Remote Sensing of Environment, 198:460-470. <a href="https://doi.org/10.1016/j.rse.2017.06.037">https://doi.org/10.1016/j.rse.2017.06.037</a></p> <p>Konings, A.G., M. Piles, K. Rötzer, K.A. McColl, S. Chan, and D. Entekhabi (2016). Vegetation optical depth and scattering albedo retrieval using time-series of dual-polarized L-band radiometer observations. Remote Sensing of Environment. 172, 178-189. https://doi.org/10.1016/j.rse.2015.11.009</p> <p>Konings, A.G., K.A. McColl, M. Piles and D. Entekhabi (2015): How many parameters can be maximally estimated from a set of measurements? IEEE Geoscience and Remote Sensing Letters, 12(5), 1081-1085. https://doi.org/10.1109/LGRS.2014.2381641</p> <p>IV) QUALITY CONTROL</p> <p>Several conditions can create uncertainty in the MT-DCA retrievals including surface water bodies (lakes, rivers, coastal areas, etc.), radio frequency interference (RFI), highly sloped surfaces (mountainous regions), dense vegetation, frozen ground, and others. The MT-DCA removes time periods of frozen ground and removes pixels with water body fractions of greater than 0.5. SMAP L1C brightness temperatures are adjusted considering RFI and surface water body information. Nevertheless, the MT-DCA retrievals are purposefully not substantially quality controlled to increase the range of science applications of the data. Therefore, the retrievals are subject to uncertainty in regions where and times when these aforementioned issues occur. We suggest the data user familiarize themselves with quality flags in the SMAP algorithm theoretical basis document in <a href="https://nsidc.org/data/SPL3SMP_E">https://nsidc.org/data/SPL3SMP_E</a>. Conservative quality control can be applied using SMAP quality flag information directly applicable to the dataset here. These quality flags can be downloaded from the SMAP official product files at <a href="https://nsidc.org/data/SPL3SMP_E">https://nsidc.org/data/SPL3SMP_E</a>.</p> <p>V) DATA FORMATTING AND FILE NAMES </p> <p>Data are provided in zipped folders in both netcdf4 (.nc) and matfile (.mat) formats. Each zipped folder contains soil moisture, vegetation optical depth, single scattering albedo, latitude, longitude, and time vector information. Note that as an update in Version 5, the zipped folders for 9km .mat files are separated into soil moisture and vegetation optical depth to reduce zip folder size. The other zipped folders still have all variables within them. These variables are provided at a 9km resolution as well as upscaled to 36km. For both .nc and .mat files, the 9km data are provided in 3-month periods with a naming convention of ‘YYYYMM_YYYYMM’ where YYYY is the 4-digit year, and MM is the 2-digit month. The first YYYYMM string represents the first month and the second YYYYMM string is the final month of the period. The 36km data are provided in 12-month periods with the same naming conventions in the file names.</p> <p>Retrievals are obtained from enhanced-resolution brightness temperatures from SMAP that are gridded at 9km. As such, they are on a 9km EASE2-grid. These retrievals are upscaled to 36km and gridded on a 36km EASE2 grid. Additional information and geolocation tools are available at <a href="https://nsidc.org/data/ease/ease_grid2.html">https://nsidc.org/data/ease/ease_grid2.html</a>. </p> <p>Information specific to folders with .nc and .mat formats is given below:</p> <p>a) NETCDF Files (.nc): The folders with netcdf files contain files with the convention MTDCA_YYYYMM_YYYYMM_Xkm_VX.nc where VX is the version number, Xkm is the grid scale, and YYYYMM strings are the first and last months of the range of data saved in the file. Soil moisture, vegetation optical depth, latitude, longitude, and time index information are provided in these files. A map of single scattering albedo for the full time series is saved in a separate file as MTDCA_OMEGA_Xkm_VX.nc along with latitude and longitude information.</p> <p>b) MATFILES (.mat): The folders with matfiles contain individual files for:</p> <ol> <li>Soil moisture: MTDCA_VX_SM_YYYYXX_YYYYXX_Xkm.mat</li> <li>Vegetation Optical Depth: MTDCA_VX_TAU_YYYYXX_YYYYXX_Xkm.mat</li> <li>Single Scattering Albedo: MTDCA_VX_OMEGA_Xkm.mat</li> <li>Latitude/Longitude: SMAPCenterCoordinatesXKM.mat</li> </ol> <p>A datevector variable in each soil moisture and vegetation optical depth file contains information on the year, month, and day corresponding to the timestep of each variable.</p>
Explaining human mobility predictions through a pattern matching algorithm
<p>The name of the file indicate information:<br> {type of sequence}_{type of measure}_{sequence properites}_{additional information}.csv</p> <p>{type of sequence} - 'synth' for synthetic or 'london' for real mobility data from London, UK.<br> {type of measure} - 'r2' for R-squared measure or 'corr' for Spearman's correlation<br> {sequence properties} - for synthetic data there are three types of sequences, described in the research article (random, markovian, nonstationary). For real mobility data this part includes information about data processing parameters: (...)_london_{type of mobility sequence}_{DBSCAN epsilon value}_{DBSCAN min_pts value}. {type of mobility sequence} is 'seq' for next-place sequences and '30min' or '1H' for the next time-bin sequences and indicate the size of the time-bin.<br> Files with 'predictability' at the end of the file contain R-squared and Spearman's correlation of measures calculated in relation to the predictability measure.</p> <p>R2 files include values of R-squared for all types of modelled regression functions.<br> 'line' indicates {y = ax + b} for single variable and {y = ax + by + c} for two variables.<br> 'expo' indicates {y = a*x^b + c} for single variable and {y = a*x^b + c*y^d + e} for two variables<br> 'log' indicates {y = a*log(x*b) + c} for single variable and {y = a * x + c * log(y) + e + d*x * log(y)} for two variables.<br> 'logf' indicates {y = a*log(x) + c * log(y) + e + b*log(x) * log(y)} for two variables</p>
Benchmarks for ApproxCov and ApproxMaxCov algorithms evaluation
<p>This deposit contains the benchmarks used for the evaluation of ApproxCov and ApproxMaxCov algorithms and their extensions. Implementations of the algorithms can be found at https://github.com/meelgroup/approxcov.<br> The folder two_values/cnf contains constraints of configurable systems. It is a subset of benchmarks from https://zenodo.org/record/4022395 and https://zenodo.org/record/3793090. In this set of benchmarks all features can have two values.<br> The folder two_values/samples contains sets of configurations computed with baital, quicksampler, and waps tools.<br> The folder mult_values contains samples and constraints of configurable systems where features can have finite number of values. These benchmarks are taken from the evaluation materials of the paper [1] and are converted to the input format of ApproxCov and ApproxMaxCov algorithms.</p> <p>[1] Brady J Garvin, Myra B Cohen, and Matthew B Dwyer. 2009. An improved meta-heuristic search for constrained interaction testing. In 2009 1st International Symposium on Search Based Software Engineering. IEEE, 13–22.</p>
An event-based precipitation dataset with life cycle evolution using resilient algorithms
<p>The dataset covers eastern Asia at a temporal range of April to June 2016-2020. We identified initial rain clusters (RCs) from the Global Precipitation Measurement 2ADPR dataset and Mesoscale Convective Systems (MCSs) from the Himawari-8 Advanced Himawari Image gridded product. Based on the contours of the initial RCs and MCSs, we then carried out a series of resilient processes, including filtration, segmentation, and consolidation, to obtain the final RCs. The final RCs had a one-to-one correspondence with the relevant MCS. We extracted the RC area, central location, average radar reflectivity profile, average droplet size distribution profile and other precipitation information from the final RCs and retrieved the life cycle evolution of the MCS area, location, and cloud-top brightness temperature from the corresponding MCSs and tracking algorithms. This dataset facilitates studies of the life cycle evolution of precipitation and provides a good foundation for convection parameterizations in precipitation simulations.</p>
Monte-Carlo-simulated MR spectra of the rat brain and their quantification results obtained with QUEST and QUEST-MM jMRUI algorithms
<p>MC-simulated spectra of the rat brain along with the corresponding basis set (metabolite signals simulated using NMRScopeB from jMRUI) and the QUEST-MM, QUEST, QUEST(Met+Back) and QUEST(Met+MM) quantification results are stored in the MC_results folder.</p> <p>Results of an in-vivo rat brain MRS signal (SPECIAL, dead time t0 = 0.134 ms, TE = 2.8 ms, at 9.4 T) quantification performed with the QUEST-MM jMRUI algorithm (origin for MC-simulation) are stored in the folder rat_results.</p> <p>All files can be loaded to jMRUI software version 5 and later.</p>
Data for the paper: Water-Food-Energy nexus: Learning from global cities using machine learning algorithms
<p>This data respository contains the datasets used to produce the "<strong>Water-Food-Energy nexus: Learning from global cities using machine learning algorithms" </strong>article. </p>
Path-finding algorithm as a dispersal assessment method for invasive species with human-vectored long-distance dispersal event
<p><strong>Aim</strong>: An assessment method that can precisely represent human-vectored long-distance dispersals (HVLDD) is currently in need for effective management of invasive species. Here, we focused on HVLDD happening along roads and proposed a path-finding algorithm as a more precise dispersal assessment tool than the most widely used Euclidean distance method by using pine wilt disease (PWD) as a case study.</p> <p><strong>Location</strong>: Busan Metropolitan City, Republic of Korea</p> <p><strong>Methods</strong>: A path-finding algorithm, which calculates distances by considering spatial distribution of road networks, was tested for its effectiveness in estimating dispersal distances of HVLDD events. To this end, annual HVLDD cases were classified from entire PWD occurrence data from 2016 to 2019 and their dispersal distances were calculated using the path-finding algorithm and the Euclidean distance method. We constructed potential dispersal ranges based on the occurrence points in 2016, 2017, and 2018 using the respective year's mean dispersal distance for both methods, and their performances in accounting for each subsequent year's HVLDD cases were compared to determine which method calculated more precise distances. The information on which road class contributed more to dispersal occurrences and distances was analysed as well using the proposed algorithm.</p> <p><strong>Results</strong>: The potential dispersal ranges of the path-finding algorithm accounted for more future anthropogenic infection cases than the ones that used the Euclidean distance method, validating its higher functionality. It also revealed that most HVLDDs started and ended on small roads, and large roads constituted the majority of the total dispersal length.</p> <p><strong>Main Conclusions</strong>: The path-finding algorithm has proven to be a more effective dispersal assessment method for HVLDD events. It can help design effective control strategies. Thus, we encourage using the path-finding algorithm for dispersal assessment of invasive species that move along road networks, as well as for the development of more powerful HVLDD prediction models.a</p>
Machine learning algorithm evaluation
<p>This is the data management plan for the purpose of this report, to compare three different classifiers through supervised machine learning on two diverse datasets. The whole machine learning process was applied and conducted in different experiments. The exploration of the data sets as well as the preprocessing strategies are outlined in the following. Furthermore, the modelling processes and the performance measures on which their results are evaluated will be explained. Finally, different parameter adjustments and settings are compared and discussed which leads to a conclusion.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.