Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,782
datasets available to search
ShareScore release 0.7.1
Dataset results
1,782 results for “algorithms”
Data from: Development of a sustainability assessment algorithm and its validation using case studies on cryogenic machining
This work presents a comprehensive structure for evaluating the sustainability of machining processes. Industries can contribute towards developing a sustainable future by using algorithms that evaluate the sustainability of their processes. Inspired by the literature, the proposed model involves a set of metrics that are critical in evaluating the impact of a process on society, environment, and economy. The flexibility of this model allows decision-makers to use the available responses to identify the most favorable process. The entropy weight method was suggested for objectively calculating the weights of each indicator. A multi-criteria decision-making method i.e., Technique for Order Preference based on Similarity to Ideal Solution (TOPSIS), was used to rank processes in the decreasing order of their sustainability. The proposed algorithm was successfully validated with case studies from the published literature. A MATLAB code was also created so that industries may expeditiously apply this method to evaluate the sustainability of machining processes.
Data from: Algorithm for post-clustering curation of DNA amplicon data yields reliable biodiversity estimates
DNA metabarcoding is promising for cost-effective biodiversity monitoring, but reliable diversity estimates are difficult to achieve and validate. Here we present and validate a method, called LULU, for removing erroneous molecular operational taxonomic units (OTUs) from community data derived by high-throughput sequencing of amplified marker genes. LULU identifies errors by combining sequence similarity and co-occurrence patterns. To validate the LULU method, we use a unique data set of high quality survey data of vascular plants paired with plant ITS2 metabarcoding data of DNA extracted from soil from 130 sites in Denmark spanning major environmental gradients. OTU tables are produced with several different OTU definition algorithms and subsequently curated with LULU, and validated against field survey data. LULU curation consistently improves α-diversity estimates and other biodiversity metrics, and does not require a sequence reference database; thus, it represents a promising method for reliable biodiversity estimation.
Supporting datasets PubFig05 for: "Heterogeneous Ensemble Combination Search using Genetic Algorithm for Class Imbalanced Data Classification"
<p><strong>Faces Dataset: PubFig05</strong></p> <p>This is a subset of the ''PubFig83'' dataset [1] which provides 100 images each of 5 most difficult celebrities to recognise (referred as class in the classification problem). For each celebrity persons, we took 100 images and separated them into training and testing sets of 90 and 10 images, respectively:</p> <p><strong>Person: </strong>Jenifer Lopez; Katherine Heigl; Scarlett Johansson; Mariah Carey; Jessica Alba</p> <p> </p> <p><strong>Feature Extraction</strong></p> <p>To extract features from images, we have applied the HT-L3-model as described in [2] and obtained 25600 features.</p> <p><strong>Feature Selection</strong></p> <p>Details about feature selection followed in brief as follows:</p> <ol> <li> <p><strong>Entropy Filtering:</strong> First we apply an implementation of Fayyad and Irani's [3] entropy base heuristic to discretise the dataset and discarded features using the minimum description length (MDL) principle and only 4878 passed this entropy based filtering method.</p> </li> <li> <p><strong>Class-Distribution Balancing:</strong> Next, we have converted the dataset to binary-class problem by separating into 5 binary-class datasets using one-vs-all setup. Hence, these datasets became <em>imbalanced</em> at a ratio of 1:4. Then we converted them into <em>balanced binary-class</em> datasets using random sub-sampled method. Further processing of the dataset has been described in the paper.</p> </li> <li> <p><strong>(alpha,beta)-k Feature selection:</strong> To get a good feature set for training the classifier, we select the features using the approach based on the (alpha,beta)-k feature selection [4] problem. It selects a minimum subset of features that maximise both within class similarity and dissimilarity in different classes. We applied the entropy filtering and (alpha,beta)-k feature subset selection methods in three ways and obtained different numbers of features (in the Table below) after consolidating them into binary class dataset.</p> </li> </ol> <ul> <li> <p><strong>UAB:</strong> We applied (alpha,beta)-k feature set method on each of the balanced binary-class datasets and we took the <em>union</em> of selected features for each binary-class datasets. Finally, we applied the (alpha,beta)-k feature set selection method on each of the binary-class datasets and get a set of features.</p> </li> <li> <p><strong>IAB:</strong> We applied (alpha,beta)-k feature set method on each of the balanced binary-class datasets and we took the <em>intersection</em> of selected features for each binary-class datasets. Finally, we applied the (alpha,beta)-k feature set selection method on each of the binary-class datasets and get a set of features.</p> </li> <li> <p><strong>UEAB:</strong> We applied (alpha,beta)-k feature set method on each of the balanced binary-class datasets. Then, we applied the entropy filtering and (alpha,beta)-k feature set selection method on each of the balanced binary-class datasets. Finally, we took the <em>union</em> of selected features for each <em>balanced binary-class</em> datasets and get a set of features.</p> </li> </ul> <p>All of these datasets are inside the compressed folder. It also contains the document describing the process detail.</p> <p> </p> <p><strong>References</strong></p> <p>[1] Pinto, N., Stone, Z., Zickler, T., & Cox, D. (2011). Scaling up biologically-inspired computer vision: A case study in unconstrained face recognition on facebook. In Computer Vision and Pattern Recognition Workshops (CVPRW), 2011 IEEE Computer Society Conference on (pp. 35–42).</p> <p>[2] Cox, D., & Pinto, N. (2011). Beyond simple features: A large-scale feature search approach to unconstrained face recognition. In Automatic Face Gesture Recognition and Workshops (FG 2011), 2011 IEEE International Conference on (pp. 8–15).</p> <p>[3] Fayyad, U. M., & Irani, K. B. (1993). Multi-Interval Discretization of Continuous-Valued Attributes for Classification Learning. In International Joint Conference on Artificial Intelligence (pp. 1022–1029).</p> <p>[4] Berretta, R., Mendes, A., & Moscato, P. (2005). Integer programming models and algorithms for molecular classification of cancer from microarray data. In Proceedings of the Twenty-eighth Australasian conference on Computer Science - Volume 38 (pp. 361–370). 1082201: Australian Computer Society, Inc.</p> <p> </p>
THE STATUS OF THE SERBIAN TERMINOLOGY DEFINED BY THE SERBIAN LANGUAGE POLICY THROUGHOUT ITS CONTEMPORARY AND FUTURE PLANS. AN OUTLINE OF ONE TERMINOLOGICAL ALGORITHM
<p>This lecture aims to point out the necessity of creating a digital terminology database of the Serbian professional terminology which would provide systematically collecting, inventory making, precise identifying, defining, linguistic analyzing and interpreting terminologies of different professions for making a stable ground for the codification and standardization of professional terminology of the Serbian language. It also implicitly points to a need to organize terminology workshops through which the experts that “build” vocabularies of their profession would become familiar with the basic principles of creating or designing terminology expressed in the work with the database via the network interface and thus become able to work on different professional terminologies. In addition, the paper suggests the most common mistakes in contemporary terminographic work as a result of lack of understanding of the meaning of inter- and multi-disciplinary approach to the study of terminology as a special branch of linguistic research.</p>
Raw data sets for: A simple calculation algorithm to seperate high-resolution CH4 flux measurements into ebullition- and diffusion derived components (AMT)
<p>Raw data sets for the research article "A simple calculation algorithm to seperate high-resolution CH4 flux measurements into ebullition- and diffusion derived components", published in "Atmospheric Measurment Techniques" (AMT). Data sets include raw data sets for the field and laboratory study, as well as calculated CH4 fluxes (field).</p>
A Benchmark Set for Multilevel Hypergraph Partitioning Algorithms
<p>DESCRIPTION<br> -------------------------------------------------------------------------------------------------------<br> This archive contains a large benchmark set for hypergraph partitioning algorithms.<br> All hypergraphs are unweighted (i.e., have unit edge and vertex weights) and use<br> the hMetis hypergraph input file format [1].</p> <p>BENCHMARK SETS<br> -------------------------------------------------------------------------------------------------------<br> Hypergraphs are derived from the following benchmark sets:<br> - The ISPD98 Circuit Benchmark Suite [2]<br> - The DAC 2012 Routability-Driven Placement Contest [3]<br> - The international SAT Competition 2014 [4]<br> - The University of Florida Sparse Matrix Collection (UF-SPM) [5]</p> <p>The benchmark set contains all ISPD98 and DAC2012 instances. Furthermore,<br> it contains 92 randomly selected instances from the application track of the SAT Competition 2014.<br> The Sparse Matrix Collection is organized into 172 groups and each group contains<br> matrices of different application areas. From each group, we chose one matrix for each application <br> area that has between 10 000 and 10.000.000 columns. In case multiple matrices fulfill<br> our criteria, we randomly selected one. In total, we include 192 matrices.</p> <p><br> HYPERGRAPH REPRESENTATION<br> -------------------------------------------------------------------------------------------------------<br> VLSI instances [2,3] are transformed into hypergraphs by converting the netlist into a<br> set of hyperedges. Sparse Matrices are translated into hypergraphs using the row-net model [6],<br> i.e. each row is treated as a net and each column as a vertex. SAT instances are converted into<br> three different hypergraph representations: In the literal model, each boolean literal is mapped to one<br> vertex and each clause constitutes a net [7]. In the primal model each variable is represented by a vertex<br> and each clause is represented by a net, whereas in the dual model the opposite is the case [8].</p> <p>FILE NAMES<br> -------------------------------------------------------------------------------------------------------<br> The origin of each hypergraph (and for SAT instances the hypergraph model) is encoded<br> into the file names as follows:<br> - Sparse Matrices : *.mtx.hgr<br> - DAC2012 : dac2012_superblue*.hgr<br> - ISPD98 : ISPD98_ibm*.hgr<br> - SAT-14 primal : sat14_*.cnf.primal.hgr<br> - SAT-14 dual : sat14_*.cnf.dual.hgr<br> - SAT-14 literal : sat14_*.cnf.hgr</p> <p>REFERENCES<br> -------------------------------------------------------------------------------------------------------<br> [1] http://glaros.dtc.umn.edu/gkhome/fetch/sw/hmetis/manual.pdf<br> [2] C. J. Alpert. The ISPD98 Circuit Benchmark Suite. In Proc. of the 1998 Int. Symp. on Physical Design, pages 80–85, New York, 1998. ACM.<br> [3] N. Viswanathan, C. Alpert, C. Sze, Z. Li, and Y/ Wei. The dac 2012 routability-driven placement contest and benchmark suite. In Proceedings of the 49th Annual Design Automation Conference, DAC ’12, pages 774–782<br> [4] A. Belov, D. Diepold, M. Heule, and M. Järvisalo. The SAT Competition 2014. http://www.satcompetition.org/2014/, 2014.<br> [5] T. A. Davis and Y. Hu. The University of Florida Sparse Matrix Collection. ACM Trans. Math. Softw.,38(1):1:1–1:25, 2011.<br> [6] Ü. V. Catalyürek and C. Aykanat. Hypergraph-partitioning-based decomposition for parallel sparse-matrix vector multiplication. IEEE Transactions on Parallel and Distributed Systems, 10(7):673–693, Jul 1999.<br> [7] D. A. Papa and I. L. Markov. Hypergraph Partitioning and Clustering. In T. F. Gonzalez, editor, Handbook of Approximation Algorithms and Metaheuristics. Chapman and Hall/CRC, 2007.<br> [8] Zoltan Mann and Pal Papp. Formula partitioning revisited. In Daniel Le Berre, editor, POS-14. Fifth Pragmatics of SAT workshop, volume 27 of EPiC Series in Computing, pages 41–56. EasyChair, 2014.</p>
Data Release: A Domain Specific Language for Performance Portable Molecular Dynamics Algorithms
<p>The archive contains the supporting data for the results described in the paper titled "A Domain Specific Language for Performance Portable Molecular Dynamics Algorithms".</p> <p>For more information see either the individual README files or consult the project git repository:</p> <p>https://bitbucket.org/wrs20/ppmd</p> <p> </p> <p>Copyright W.R.Saunders 2017</p>
Electrical impedance tomography - Depency of cardiac related impedance change amplitudes on body position and reconstruction algorithm
<p>Dataset of a publication investigating the influence of body position and reconstruction algorithm on the amplitudes of cardiac related impedance changes in healthy adult volunteers.</p>
Example data to run the STABLE algorithm
<p>Datasets from ERA5 (1 degree resolution) and NCAR (2.5 degree resolution) from both hemispheres to run the SubTropical Atmospheric ridge and BLocking Events (STABLE) algorithm. See the STABLE GitHub repository (https://github.com/mikaslima/STABLE) for details.</p>
Net primary production from the Eppley-VGPM, Behrenfeld-VGPM, Behrenfeld-CbPM, Westberry-CbPM and Silsbe-CAFE algorithms
<p>Net primary production (mg C m<sup>-2</sup> d<sup>-1</sup>) calculated from the Eppley-VGPM, Behrenfeld-VGPM, Behrenfeld-CbPM, Westberry-CbPM and Silsbe-CAFE algorithms. MLD data taken from HADLEY EN4.2.2 using the density criterion of 0.03 kg m<sup>-3</sup>. </p> <p>Data on a regular 25km grid at a 8 day resolution.</p> <p><strong>Version 1.1</strong></p> <p>Fixed minor issues with:</p> <ol> <li>Conversion of bbp(443) to phytoplankton carbon.</li> <li>Missing values at ~180W.</li> </ol> <p> </p> <p><strong>Version 1.2</strong></p> <ol> <li>Updated to include 2023.</li> <li>Westberry-CbPM Nitracline depth updated with World Ocean Atlas 2023 Nitrate Data.</li> <li>Silsbe-CAFE bbwater calculations updated with World Ocean Atlas 2023 Salinity Data.</li> <li>File structure compressed using Zlib.</li> </ol> <p>Please note for Westberry-CbPM and Silsbe-CAFE all years were reprocessed.</p> <p> </p> <p><strong>Version 1.3</strong></p> <ol> <li>Updated to include 2024.</li> </ol> <p><strong>Version 1.3.1</strong></p> <ol> <li>Replaced Behrenfeld-VGPM file which had errors.</li> </ol>
Simulation results of routing algorithms for multilayer networks
<p>Most current routing protocols are based on path computation algorithms in graphs (e.g., Dijkstra, Bellman-Ford, etc.). These algorithms have been studied for a long time and are very well understood, both in a centralized and distributed context, as long as they are applied to a network having a single communication protocol. The problem becomes more complex in the multi-protocol case, where there is a possibility of encapsulation of some network protocols into others, therefore inducing nested tunnels. The classic algorithms cited above no longer work in this case because they cannot manage the protocol encapsulations and the corresponding protocol stacks. In this work, we propose a highly parallelizable algorithm that takes into account protocol encapsulations as well as protocol conversions in order to compute shortest paths in a multi-protocol network. To achieve this computation efficiently, we study the transitive closure between subpaths (i.e., the concatenation of two subpaths to obtain a longer one) in the case where each subpath induces a protocol stack, and thus tunnels. Leveraging on Software-Defined Networks with a controller having a highly parallel architecture enables us to compute the routing tables of all nodes in a very efficient way. Experimentation results on both random and realistic topologies show that our algorithm outperforms the previous solutions proposed in the literature.</p>
Dataset for paper 'McAN: a novel computational algorithm and platform for constructing and visualizing haplotype networks'
<p>The .zip file includes four datasets for testing the performance of McAN (doi: https://doi.org/10.1093/bib/bbad174).</p>
STAVER: A Standardized Benchmark Dataset-Based Algorithm for Effective Variation Reduction in Large-Scale DIA-MS Data
<p>This project focuses on developing and applying STAVER, an innovative DIA algorithm designed to eliminate non-biological noise and variability from the large-scale DIA-MS study dataset analyses. STAVER is a flexible framework that utilizes prior knowledge regarding peptide separation coordinates (RT) and fragment ion intensities from the standard benchmark datasets, which effectively mitigates non-biological noise potential during library searches, enhancing spectrum identifications and protein quantification accuracy. Furthermore, the robustness and broad applicability of STAVER were validated in multiple large-scale DIA datasets from different platforms and laboratories, demonstrating significantly improved precision and reproducibility of protein quantification. It facilitates the comparative and integrative analysis of DIA datasets across different platforms and laboratories, enhancing the consistency and reliability of findings in clinical research. The project aims to promote the adoption of hybrid library search and improve the sensitivity and quality of DIA proteomics data through the open-source STAVER software package.</p>
CBM algorithm TISMIR 2023: code and data for reproducing experiments
<p>This dataset contains the data necessary to reproduce experiments presented in the TISMIR article untitled "Barwise Music Structure Analysis with the Correlation Block-Matching Segmentation Algorithm" (under publication at the time of the upload, the link will be added after publication).</p><p>In details, this zenodo upload contains:</p><ul><li>Data, i.e. precomputed data and features (the self-similarity matrices in particular, along with beats and bars estimates) required to compute the CBM algorithm,</li><li>Code, i.e. the source files and the experiments (under the form of "Notebooks") used to compute results.</li></ul><p>This upload extends the version on git (https://gitlab.imt-atlantique.fr/a23marmo/autosimilarity_segmentation/-/tree/TISMIR).</p><p> </p>
Supplement of "Algorithm for continual monitoring of fog life cycles based on geostationary satellite imagery as a basis for solar energy forecasting"
<p>The file uploaded here is an animation that visually illustrates the outputs of the a newly developed machine learning based FLS (<strong>F</strong>og and <strong>L</strong>ow <strong>S</strong>tratus) detection algorithm for the SEVIRI (<strong>S</strong>pinning <strong>E</strong>nhanced <strong>V</strong>isible and <strong>I</strong>nfra<strong>R</strong>ed <strong>I</strong>mager) instrument onboard the MSG (<strong>M</strong>eteosat <strong>S</strong>econd <strong>G</strong>eneration) geo-stationary satellites over the 24hr cycle of the day for the day of <strong>02/March/2021</strong> and compares them with the corresponding raw channel values observed by SEVIRI. The proposed algorithm classifies each SEVIRI pixel as "clear-sky", "FLS", or "non-FLS-cloud" (identified with Khaki, Red, and Blue in the animation) based on the SEVIRI pixel values of BT12.0, BT8.7 - BT12.0, BT10.8 - BT12.0, and BT12.0 - BT13.4 plus the standard deviation of each of these variables in a spatial window sized 3x3 pixels with the central pixel being the target pixel. </p><p><br>In this animation, the left-hand panel shows a false-color RGB image constructed based on the SEVIRI raw channel data with the red, green, and blue channels being BT12.0- BT13.4, BT8.7 - BT12.0, and BT10.8 - BT12.0, respectively. In this panel, the green color represents the high clouds, and the light and dark red colors represent the clear-sky and FLS, respectively. The right-hand panel of this animation also shows the outputs of the ML FLS detection algorithm developed in the present study.</p>
Experiments with Frequency Fitness Assignment based Algorithms on the Traveling Salesperson Problem
<p><strong>1. Introduction</strong></p><p>In this archive, we provide the implementation and experimental results of eight different algorithms to solve Traveling Salesperson Problem (TSP) instances from <a href="http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/">TSPLIB</a>.</p><p>A TSP is defined by a fully-connected weighted graph of n cities. The goal is to find the overall shortest tour that visits each cities exactly once and returns to its starting point. The TSP is NP-hard. We consider 56 symmetric instances from the well-known <a href="http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/">TSPLIB</a>.</p><p>Solutions in our work are stored in the path representation, where such a tour is encoded as a permutation x of the numbers 1 to n, each identifying a city. If a city appears at index j in the permutation x, then it will be the jth city to be visited. This means that a tour x will pass the following edges: (x[1], x[2]), (x[2], x[3]), (x[3], x[4]), … (x[n-1], x[n]), (x[n], x[1]).</p><p><strong>2. Directory Structure</strong></p><p>This dataset is split into multiple separate <i>tar.xz</i> archives. These can be unpacked in the same folder and will produce the directory structure described below. Each archive contains this note and the license information, but apart from that, there is no redundancy.</p><p>This archive contains the following directories:</p><ul><li>source contains the Python source codes needed to run the experiment.<ul><li>moptipy-main is a local copy of the <a href="https://thomasweise.github.io/moptipy">moptipy</a> package used for our experiment.</li><li>tsplib contains the <a href="http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/">TSPLIB</a> data. This includes the instances used in our experiments as files in text format with suffix .tsp. If an optimal tour is given by TSPLib, it is stored in a text format file with suffix .opt.tour and name prefix identical to the instance file. In other words, the file eil51.tsp contains the TSP instance eil51 and the file eil51.opt.tour contains the corresponding optimal tour. Both the TSP instances and optimal tours can be downloaded from <a href="http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/tsp/">http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/tsp/</a>. We also include the documentation of TSPLIB in file <a href="http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/tsp95.pdf">tsp95.pdf</a> documenting them. We further include the <a href="http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/TSPFAQ.html">TSPLIB FAQ</a> both as HTML and PDF file (tsplib_faq.html and tsplib_faq.pdf) and the <a href="http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/STSP.html">list of known optimal tour lengths</a> as HTML and PDF file (optimal_tour_lengths_of_symmetric_tsps.html, optimal_tour_lengths_of_symmetric_tsps.pdf). Notice that, while the TSP instances we used are Euclidean, all distances are converted to integers as prescribed by the <a href="http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/tsp95.pdf">documentation</a>.</li></ul></li><li>results is the directory with the log files. Each log file contains information of one run, i.e., one execution of one algorithm on one problem instance. All improving moves of a run as well as the final solution are stored in the log file. The direct sub-folders results represent the algorithms and contain one folder per TSP instance, which, in turn, contain the log files.</li><li>evaluation is a folder with the extracted evaluation and figures</li><li>evaluation_edited is a folder with evaluation figures slightly edited for better visual appeal (but obviously without changing any result / scientific content)</li><li>evaluator is a folder with a Python script main.py that generates all the files in evaluation from the data it finds in results. It requires the <a href="https://thomasweise.github.io/moptipy">moptipy</a> package being installed for running in the version given in requirements.txt.</li></ul><p><strong>3. Algorithms</strong></p><p>The (1+1) EA is the most basic evolutionary algorithm and also be considered as a randomized local search. It starts with one random solution/permutation xc and computes its length yc=f(xc). In each iteration, it applies a unary search operator op to obtain a new tour xn=op(xc) and computes its length yn=f(xn). If yn<=yc, then it will accept the new tour and set xn=xn and yc=yn. The results of this algorithm are given in folder results/ea_revn.</p><p>FFA is a fitness assignment process that takes place before this last step in the EA. We integrate FFA into the (1+1) EA and obtain the (1+1) FEA. This algorithm uses an additional table H which counts, for any tour length y, how often it has been seen during the search so far. After the new tour xn is created and its objective value yn is computed, the (1+1) FEA sets H[yc] = H[yc] + 1 and H[yn] = H[yn] + 1. It will accept xn if and only if H[yn] <= H[yc] and, only in this case, set xn=xn and yc=yn. The results of this algorithm are given in folder results/fea_revn.</p><p>SA is the classical simulated annealing algorithm. In our experiment, it will accept the new solution xn with probability P. If the new solution is better, the acceptance probability P is 1. For worse solutions, the probability is between 0 and 1, i.e., sometimes, worse solution are also accepted. This algorithm has a temperature cooling schedule. It starts at an initial temperature and over time, the temperature decreases. The probability P of accepting the worse solution depends on the temperature and decreases as well. The results of this algorithm are given in folder results/sa_revn.</p><p>An FFA-based version of SA uses the frequency fitness instead of the objective values in all acceptance decisions. The results of this algorithm are given in folder results/fsa_revn.</p><p>EAFEA(A) is a hybrid which alternates between the EA and the FEA and copies a solution from the FEA to the EA if it has an entirely new objective value, i.e., if H[yn] = 1. The results of this algorithm are given in folder results/eafea2_revn.</p><p>SAFEA(A) is a hybrid which alternates between the SA and the FEA and copies a solution from the FEA to the SA if it has an entirely new objective value, i.e., if H[yn] = 1. The results of this algorithm are given in folder results/safea2_revn.</p><p>EAFEA(B) is a hybrid which alternates between the EA and the FEA and copies a solution from the FEA to the EA part if it has a better objective value. The results of this algorithm are given in folder results/eafea_revn.</p><p>SAFEA(B) is a hybrid which alternates between the SA and the FEA and copies a solution from the FEA to the SA part if it has a better objective value. The results of this algorithm are given in folder results/safea_revn.</p><p>We apply all algorithms with the same unary operator reverse, which reverses a randomly chosen subsequence of the tour. This operator is also often called a "2-opt move". It has the advantage that the new objective value of a new solution can be computed in O(1) if the objective value of the solution from which it is derived is known.</p><p><strong>4. How to Run the Experiment</strong></p><p>First, you need to make sure to have all the dependencies installed that this program requires. You can do this by executing the following command in the terminal:</p><p>pip install matplotlib numba numpy psutil scikit-learn moptipy moptipyapps</p><p>Now enter the source directory, i.e., the directory containing the run.py file, in your terminal. Depending on your system configuration and whether you run Windows or Linux, you can start the program with <i>one</i> of the commands below. (If running the first command returns with an error, just try the next one in the list.)</p><ul><li>python3 -m run</li><li>python -m run</li><li>python run.py</li><li>python3 run.py</li></ul><p>Then the experiment will run. It will automatically create a sub-folder results in source and place all log files that are generated into it. Be careful: The experiment will take a long time. However, if you have multiple CPUs, you can simply start several instances of this program in independent terminals. Each instance will then conduct different runs. This also works if this folder is shared over the network, in which case you can run multiple processes on multiple PCs.</p><p>Side note: This experiment uses the <a href="https://thomasweise.github.io/moptipy">moptipy</a> package for implementing its algorithms, running the experiments, and gathering their results. If you want to install moptipy on your system instead of using the version supplied here, you can install it via pip install moptipy. It also uses moptipyapps to load some data.</p><p><strong>5. Literature</strong></p><ul><li>Frequency Fitness Assignment (FFA):<ol><li>Thomas Weise, Zhize Wu, Xinlu Li, Yan Chen, and Jörg Lässig. Frequency Fitness Assignment: Optimization without Bias for Good Solutions can be Efficient. IEEE Transactions on Evolutionary Computation (TEVC). 2022. Early Access. doi:<a href="https:doi.org/10.1109/TEVC.2022.3191698">10.1109/TEVC.2022.3191698</a>.</li><li>Thomas Weise, Zhize Wu, Xinlu Li, and Yan Chen. Frequency Fitness Assignment: Making Optimization Algorithms Invariant under Bijective Transformations of the Objective Function Value. <i>IEEE Transactions on Evolutionary Computation</i> 25(2):307–319. April 2021. Preprint available at <a href="http://arxiv.org/abs/2001.01416">arXiv:2001.01416v5</a> [cs.NE] 15 Oct 2020. doi:<a href="http://dx.doi.org/10.1109/TEVC.2020.3032090">10.1109/TEVC.2020.3032090</a>. Experimental results and source code are available at doi:<a href="http://doi.org/10.5281/zenodo.3899474">10.5281/zenodo.3899474</a>.</li><li>Tianyu Liang, Zhize Wu, Jörg Lässig, Daan van den Berg, and Thomas Weise. Solving the Traveling Salesperson Problem using Frequency Fitness Assignment. In Hisao Ishibuchi, Chee-Keong Kwoh, Ah-Hwee Tan, Dipti Srinivasan, Chunyan Miao, Anupam Trivedi, and Keeley A. Crockett, editors, Proceedings of the IEEE Symposium on Foundations of Computational Intelligence (IEEE FOCI'22), part of the IEEE Symposium Series on Computational Intelligence (SSCI 2022). December 4–7, 2022, Singapore, pages 360–367. IEEE. doi:<a href="https://doi.org/10.1109/SSCI51031.2022.10022296">10.1109/SSCI51031.2022.10022296</a>.</li><li>Thomas Weise, Mingxu Wan, Ke Tang, Pu Wang, Alexandre Devert, and Xin Yao. Frequency Fitness Assignment. <i>IEEE Transactions on Evolutionary Computation (IEEE-EC)</i> 18(2):226-243, April 2014. doi:<a href="http://dx.doi.org/10.1109/TEVC.2013.2251885">10.1109/TEVC.2013.2251885</a>.</li><li>Thomas Weise, Xinlu Li, Yan Chen, and Zhize Wu. Solving Job Shop Scheduling Problems Without Using a Bias for Good Solutions. In <i>Genetic and Evolutionary Computation Conference Companion (GECCO'21 Companion),</i> July 10-14, 2021, Lille, France. ACM, New York, NY, USA. ISBN 978-1-4503-8351-6. doi:<a href="http://doi.org/10.1145/3449726.3463124">10.1145/3449726.3463124</a>.</li><li>Thomas Weise, Yan Chen, Xinlu Li, and Zhize Wu. Selecting a diverse set of benchmark instances from a tunable model problem for black-box discrete optimization algorithms. <i>Applied Soft Computing Journal (ASOC)</i>, 92:106269, June 2020. doi:<a href="http://dx.doi.org/10.1016/j.asoc.2020.106269">10.1016/j.asoc.2020.106269</a>.</li><li>Thomas Weise, Mingxu Wan, Ke Tang, and Xin Yao. Evolving Exact Integer Algorithms with Genetic Programming. In <i>Proceedings of the IEEE Congress on Evolutionary Computation (CEC'14), Proceedings of the 2014 World Congress on Computational Intelligence (WCCI'14)</i>, pages 1816-1823, Beijing, China, July 6-11, 2014. Los Alamitos, CA, USA: IEEE Computer Society Press. ISBN: 978-1-4799-1488-3. doi:<a href="http://dx.doi.org/10.1109/CEC.2014.6900292">10.1109/CEC.2014.6900292</a>.</li></ol></li><li>Traveling Salesperson Problem (TSP):<ol><li>Pedro Larrañaga, Cindy M. H. Kuijpers, Roberto H. Murga, I. Inza, and S. Dizdarevic. Genetic Algorithms for the Travelling Salesman Problem: A Review of Representations and Operators. <i>Artificial Intelligence Review,</i> 13(2):129–170, April 1999. Kluwer Academic Publishers, The Netherlands. doi:<a href="https://doi.org/10.1023/A:1006529012972">10.1023/A:1006529012972</a>.</li><li>Gerhard Reinelt. TSPLIB — A Traveling Salesman Problem Library. <i>ORSA Journal on Computing</i> 3(4):376-384. 1991. <a href="http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/">http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/</a>.</li><li>Gerhard Reinelt. TSPLIB95. 1995. Heidelberg, Germany: Universität Heidelberg, Institut für Angewandte Mathematik. <a href="http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/tsp95.pdf">http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/tsp95.pdf</a>.</li><li>Thomas Weise, Raymond Chiong, Ke Tang, Jörg Lässig, Shigeyoshi Tsutsui, Wenxiang Chen, Zbigniew Michalewicz, and Xin Yao. Benchmarking Optimization Algorithms: An Open Source Framework for the Traveling Salesman Problem. <i>IEEE Computational Intelligence Magazine (CIM)</i> 9(3):40-52, August 2014. doi:<a href="http://dx.doi.org/10.1109/MCI.2014.2326101">10.1109/MCI.2014.2326101</a>.</li><li>Eugene Leighton Lawler, Jan Karel Lenstra, Alexander Hendrik George Rinnooy Kan, and David B. Shmoys. <i>The Traveling Salesman Problem: A Guided Tour of Combinatorial Optimization.</i> Wiley Interscience. 1985.</li><li>David Lee Applegate, Robert E. Bixby, Vasek Chvatal, and William John Cook. <i>The Traveling Salesman Problem: A Computational Study.</i> Princeton University Press. 2007.</li><li>Gregory Z. Gutin and Abraham P. Punnen, editors. <i>The Traveling Salesman Problem and its Variations.</i> Volume 12 of Combinatorial Optimization. Kluwer Academic Publishers. 2002. doi:<a href="https://dx.doi.org/10.1007/b101971">10.1007/b101971</a>.</li></ol></li><li>Software:<ol><li>The Metaheuristic Optimization in Python Package <a href="https://thomasweise.github.io/moptipy">moptipy</a></li></ol></li></ul><p><strong>6. License</strong></p><p>The files in this repository are under the <a href="https://creativecommons.org/licenses/by/4.0/legalcode">Creative Commons Attribution 4.0 International</a>, with the exception of the files of <a href="http://comopt.ifi.uni-heidelberg.de/software/TSPLIB95/">TSPLIB</a> in directory source/tsplib, which are under copyright of their respective owner (we believe that they are in the public domain, as they are provided by many sources, included in many software packages under various open source licenses, and on many websites). The license is contained as file LICENSE.txt in this archive.</p><p><strong>7. Contact</strong></p><p>If you have any questions or suggestions, please contact</p><p>Mr. Tianyu LIANG (梁天宇) of the Institute of Applied Optimization (应用优化研究所, <a href="http://iao.hfuu.edu.cn">IAO</a>) of the School of Artificial Intelligence and Big Data (<a href="http://www.hfuu.edu.cn/aibd/">人工智能与大数据学院</a>) at <a href="http://www.hfuu.edu.cn/english/">Hefei University</a> (<a href="http://www.hfuu.edu.cn/">合肥学院</a>) in Hefei, Anhui, China (中国安徽省合肥市) via email to <a href="mailto:liangty@stu.hfuu.edu.cn">liangty@stu.hfuu.edu.cn</a>.</p>
INFLAMeR: a machine learning algorithm based on large-scale perturbation screening identified new lncRNAs regulating differentiation and survival of leukaemia cells
Open the record for dataset details and reuse information.
Designing Optimal Convolutional Neural Network Architecture Using Differential Evolution Algorithm
<p>Convolutional Neural Networks (CNNs) are widely used deep learning models for solving various tasks such as computer vision, speech recognition, among others. However, CNNs are developed manually based on problem-specific domain knowledge and tricky settings, which are laborious, time-consuming and challenging. To address these issues, this study proposes an Improved Differential Evolution of Convolutional Neural Network algorithm, namely IDECNN, to design CNN layer architectures for image classification task. </p>
Subseasonal to Seasonal (S2S) Prediction Algorithms using Hybrid Machine Learning Techniques
<p>< S2S dataset.zip ></p><p>1.ECMWF observations/hindcast realizations</p><ul><li>hindcast-like-observations_2000-2019_biweekly_deterministic.zarr</li><li>forecast-like-observations_2020_biweekly_deterministic.zarr</li><li>ecmwf_hindcast-input_2000-2019_biweekly_deterministic.zarr</li><li>ecmwf_forecast-input_2020_biweekly_deterministic.zarr</li><li>hindcast-like-observations_2000-2019_biweekly_tercile-edges.nc</li></ul><p>2. External variables</p><ul><li>"nino" folder -> nino12.long.anom.data, nino34.long.anom.data : El Niño data</li><li>"Oscillation" folder<ul><li>-> ersst.v5.pdo.dat.text : PDO (Pacific Decadal Oscillation)</li><li>-> norm.nao.monthly.b5001.current.ascii.table.txt : NAO (North Atlantic Oscillation)</li><li>-> qbo.dat : QBO (Quasi Biennial Oscillation)</li></ul></li><li>"great_lake" folder -> N_seaice_extent_daily_v3.0 : Great lakes ice cover</li><li>observed-solar-cycle-indices.json : Sunspot cycles (two variables: original value and smoothed value)</li></ul><p>3. Region.txt : Region and its bound</p><p>4. Biweekly historical statistics data</p><ul><li>biw_stat_w34 folder -> data (mean, standard deviation, median, skewness, kurtosis) for Week 3-4</li><li>biw_stat_w56 folder -> data (mean, standard deviation, median, skewness, kurtosis) for Week 5-6</li></ul><p> </p><p>< ML_code.zip ></p><ul><li>ML codes for training, testing, and calculating RPSS based on Python3</li><li>Check run_val.sh and run_2020.sh </li></ul><p> </p>
Absorbing Aerosol Optical Central Height (AOCH) retrieved from TROPOMI with UIowa's AOCH-O2AB algorithm
<p>Absorbing Aerosol Optical Centroid Height (AOCH) retrieved from TROPOMI with UIowa’s AOCH-O<sub>2</sub>AB algorithm. Dataset for analyzing dust and smoke cases over Asia during 2021-2023.</p> <p>More information about this dataset can be found in: </p> <p>Chen, X., Wang, J., Xu, X. G., Zhou, M., Zhang, H. X., Garcia, L. C., Colarco, P. R., Janz, S. J., Yorks, J., McGill, M., Reid, J. S., de Graaf, M., and Kondragunta, S.: First retrieval of absorbing aerosol height over dark target using TROPOMI oxygen B band: Algorithm development and application for surface particulate matter estimates, Remote Sensing of Environment, 265, 18, <a href="https://doi.org/10.1016/j.rse.2021.112674">https://doi.org/10.1016/j.rse.2021.112674</a>, 2021.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.