Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
318
datasets available to search
ShareScore release 0.7.1
Dataset results
318 results for “Data mining”
Figure 3. Data visualization-DATA MINING LEARNING MODELS AND ALGORITHMS ON A SCADA SYSTEM DATA REPOSITORY
<p>Data visualization is also a very useful technique because it helps to deter-<br> mine the di±culty of the learning problem. We visualized with Weka single<br> attributes (1-d) and pairs of attributes (2-d). The ¯gure 3 shows the variation<br> of the temperature in time.</p>
PubChem Data Mining of OXPHOS inhibitors: scripts, data, and models
<p>README doc, source, and data files from PubChem data mining project to identify OXPHOS inhibitory chemotypes.</p>
Appendix for Review Paper Entitled: "A literature review of "lawful" text and data mining."
<p><span>This appendix complements the review paper entitled ““A literature review of “lawful” text and data mining” with 8 Tables highlighting which scholarly works were used for each section of the literature review, but also how those works were used. </span></p>
Fig. 2 in Morphological and genetic data suggest a complex pattern of inter-island colonisation and differentiation for mining bees (Hymenoptera: Anthophila: Andrena) on the Macaronesian Islands
Fig. 2 Median Joining network of Andrena species (Micrandrena) of the Canary Islands and the Madeira Archipelago. Circle size is relative to number of haplotype copies present in dataset. A branch represents a single nucleotide change (mutation); bars on branches represent inferred missing haplotypes (single nucleotide changes). LG La Gomera, LP La Palma. The colours correspond to those used to represent the location of species on islands in Fig. 1
Fig. 3 in Morphological and genetic data suggest a complex pattern of inter-island colonisation and differentiation for mining bees (Hymenoptera: Anthophila: Andrena) on the Macaronesian Islands
Fig. 3 Dated species tree demonstrating the phylogenetic relationships of the different island populations calculated with *BEAST compared to the outgroup species Andrena enslinella, A. subopaca, A. minutuloides, and A. semilaevis (all Micrandrena); dating is based
Fig. 6 in Morphological and genetic data suggest a complex pattern of inter-island colonisation and differentiation for mining bees (Hymenoptera: Anthophila: Andrena) on the Macaronesian Islands
Fig. 6 Correlation of genetic distance (ΦST) and A squared Mahalanobis distance of morphometric data (r2= 0.18) and B Euclidian distance of qualitative morphological data (r2= 0.12)
Fig. 1 in Morphological and genetic data suggest a complex pattern of inter-island colonisation and differentiation for mining bees (Hymenoptera: Anthophila: Andrena) on the Macaronesian Islands
Fig. 1 Location of the Azores, the Archipelago of Madeira, the Selvagens Islands, the Canary Islands, and Cape Verde (a). Close view to the islands of the Madeira Archipelago (b) and the Western Canary Islands (below, right) (c). Tenerife is characterised by the regions of Anaga, Teno, Las Cañadas/Teide, and Dorsal Rift. The taxa of the A. wollastoni group (Kratochwil, 2020) and the centres of their distri-
BRAIN Journal-High Performance Data mining by Genetic Neural Network-Figure 8 . The Comparision of Run Times
<p>That is distinct that dynamic mutation rate or reduction idea for mutation operator is more<br> better of fixed rate. In fact obtain to high accuracy is result of our idea for mutation operator.<br> The number of hidden layer neurone is important problem for NN. The natural selection by<br> GA help finding the number of hidden layer neurone and it progress on duration generations.<br> The structured model of GANN finds better answer than NN but with much run time in<br> simulation. The learning of GA is much better than NN with back propagation because BP is a<br> method based on gradient descend and local optimum is a serious risk for that.<br> We hope that the number of training samples is more accurate without error, the new<br> algorithm is better. Tests show that the combination of genetic algorithms and neural networks to an<br> acceptable level solves the problem of overfitting.</p>
BRAIN Journal-High Performance Data mining by Genetic Neural Network-Figure 7 . Test Accuracy with prograess generation
<p>In the training phase, the neural network weights errors are minimized and network design<br> problem which the objective function to an acceptable level. In test step we have better results<br> because weights of neural network are adjusted by genetic algorithm and back propagation method.<br> Of course achievement to accuracy with 83.5% is reason using of good feature with minimum error.</p>
BRAIN Journal-High Performance Data mining by Genetic Neural Network-Figure 6. Training Accuracy with prograess generation
<p>There are many features will reduce the efficiency of the algorithm and its complexity.<br> Among the methods for selecting the appropriate features, the algorithm is a GA.<br> One of the important parameters for testing methods is accuracy rate on progress generation.<br> In fact accuracy is reverse error in algorithm results. As reader can compare the results of our paper<br> with another works. Figure 4 show that accuracy present for Training step. We achieve to best<br> answers of 800 generation to after generation.</p>
BRAIN Journal-High Performance Data mining by Genetic Neural Network-Figure 4. Structural Crossover
<p>Guided crossover operator is based on the two point separation from parents are selected<br> Left and right parts of them are related to each other by the condition to be meaningful With this<br> new child of his parents is that. But a new generation of the random choice to have reached this<br> stage. The crossover rate is fixed for our algorithm.</p>
BRAIN Journal-High Performance Data mining by Genetic Neural Network-Figure 5. Insertion and Deletion Hidden Layer in NN
<p>Change in NN structure is other method that we used to optimization of solution[18].<br> Insertion a hidden layer caused to mutation operator is much natural. As connection with father and<br> mother nodes is easily[20],[21]. Weights of node and errors automatically calculated.<br> For each stage of the implementation of the mutation operator in genetic algorithms, neural<br> networks, only one of the nodes in the hidden layer is selected and inserted. These layers are<br> inserted on condition that the definition does not harm the network structure and the action is<br> meaningful. As an added layer can adjust the weights and the connection to the parent node of a network<br> layer to be removed.</p>
BRAIN Journal-High Performance Data mining by Genetic Neural Network-Figure 3. The Structure of Neural Network
<p>A neural network (NN), in the case of artificial neurons called artificial neural<br> network (ANN) or simulated neural network (SNN), is an interconnected group of natural<br> or artificial neurons that uses a mathematical or computational model for information<br> processing based on a connectionist approach to computation. In most cases an ANN is an adaptive<br> system that changes its structure based on external or internal information that flows through the<br> network[9].<br> In more practical terms neural networks are nonlinear statistical data modelling or decision<br> making tools. They can be used to model complex relationships between inputs and outputs or<br> to find patterns in data.<br> Two neurons neural network active in memory (ON or 1) or disable (Off or 0), and each<br> edge (synapses or connections between nodes) is a weight. Edges with positive weight, stimulate or<br> activate next active node, and edges with negative weight, disable or inhibit the next connected<br> node (if it is active) ones.</p>
Data Pack fro VirSorter: mining viral signal from microbial genomic data
<p>This is the data pack for VirSorter, the publication of which by Roux et al. titled "<strong>VirSorter: mining viral signal from microbial genomic data</strong>" appeared in PeerJ on 2015-05-28 (<a href="https://doi.org/10.7717/peerj.985">doi:10.7717/peerj.985</a>).</p> <p>Most up-to-date tutorials and the code for VirSorter can be found at the GitHub repository <a href="https://github.com/simroux/VirSorter">https://github.com/simroux/VirSorter</a>.</p> <p>The original source of this data pack was here (last accessed 2018-02-03): <a href="http://datacommons.cyverse.org/browse/iplant/home/shared/imicrobe/VirSorter/virsorter-data.tar.gz">http://datacommons.cyverse.org/browse/iplant/home/shared/imicrobe/VirSorter/virsorter-data.tar.gz</a></p>
BRAIN Journal-Prediction of Thyroid Disease Using Data Mining Techniques-Figure 2. Attributes of the classification models used in the experiments
<p>The authors used for their experiments a data set (UCI, 2016) containing 756 records about persons with thyroid dysfunctions. The classification model has 22 attributes; the class attribute is the target and it has three possible values: hypothyroidism, hyperthyroidism and normal. The current data set was extracted and preprocessed from the original file. A description of the attributes used in the experiments is given in Figure 2 (an extract from thyroid.arff test file). </p>
BRAIN Journal-Prediction of Thyroid Disease Using Data Mining Techniques-Figure 3. KNIME Diagram
<p>The proposed KNIME diagram representing the data mining models is given in Figure 3. The nodes that constitute the model diagram are: ARFF Reader – the input node used to load the data set in arff format, Partitioning – the node with the role of data set partition (for training and for the validation of the classification model), Naive Bayes Learner and Decision Tree Learner – the nodes used to build the classification model, Naive Bayes Predictor and Decision Tree Predictor – the nodes used to validate the model, Scorer – the node reports a confusion matrix and the accompanying quality measures in its view, Normalizer – the data set are normalized to be able to apply the neural network models, Multilayer Perceptron and RBFNetwork – the nodes corresponding to the neural network classification models, Weka Predictor – a node implemented in Weka to validate the models. </p>
BRAIN Journal-Prediction of Thyroid Disease Using Data Mining Techniques-Figure 1. Factors that Affect Thyroid Function (The Institute for Functional Medicine, 2014)
<p> In Figure 1 are presented the main factors that affect the thyroid function. It is obvious that factors such as stress, infection, toxins, trauma and certain medication are directly responsible for the improper production of thyroid hormones. Symptoms identification and the early detection of abnormal values of thyroid hormones after clinical investigation will help in establishing the proper diagnostic and to prescribe the right medication. The patient must periodically evaluate his clinical state in order to receive the treatment as long as he needs it. </p>
BRAIN Journal-A New Challenge for Information Mining-Figure 5: New Approach Rich Data Exploration
<p>In the following figure (see Figure 5) it has represented a scheme of the new approach proposed to Rich Data Set's Exploration.</p> <p>Other experiments are running in order to validate our idea, both in order to optimize this clustering model by applying new algorithms and distance measures to the datasets presented here, and both applying these techniques to a different domain from the didactic one. Other experiments are also conducted to improve user exploration by skillfully combining multiple methods and exploration techniques through the application of a variety of models such as the Association Rule to extract hidden relationships and association rules between data and Artificial Neural Network mechanisms of learning applicable to classification and forecasting problems. </p>
Data set for the paper "What are the Effects of History Length and Age on Mining Software Change Impact?"
<p>Data set for the paper What are the Effects of History Length and Age on Mining Software Change Impact?<br> by Leon Moonen, Thomas Rolfsnes, David Binkley and Stefano di Alesio.<br> In Journal of Empirical Software Engineering (EMSE), 2018, Springer. https://doi.org/10.1007/s10664-017-9588-z<br> Available from https://evolveit.bitbucket.io/publications/emse2018/</p> <p>Please cite this work by referring to the corresponding journal publication (a preprint is included in this package).</p> <p>The goal of Software Change Impact Analysis is to identify artifacts (typically source-code files or individual methods therein) potentially affected by a change. Recently, there has been increased interest in <em>mining</em> software change impact based on evolutionary coupling. A particularly promising approach uses association rule mining to uncover potentially affected artifacts from patterns in the system’s change history. Two main considerations when using this approach are the <em>history length</em>, the number of transactions from the change history used to identify the impact of a change, and <em>history age</em>, the number of transactions that have occurred since patterns were last mined from the history. Although history length and age can significantly affect the quality of mining results, few guidelines exist on how to best select appropriate values for these two parameters.</p> <p>In this paper, we empirically investigate the effects of history length and age on the quality of change impact analysis using mined evolutionary coupling. Specifically, we report on a series of systematic experiments using three state-of-the-art mining algorithms that involve the change histories of two large industrial systems and 17 large open source systems. In these experiments, we vary the length and age of the history used to mine software change impact, and assess how this affects precision and applicability. Results from the study are used to derive practical guidelines for choosing history length and age when applying association rule mining to conduct software change impact analysis. </p>
Webis Simulation Data Mining Bridge Models Corpus 2012 (Webis-SDMbridge-12)
<p>This corpus provides the simulation data mining community with a collection of 14641 bridge models and simulated behavior.</p> <p><strong>1. Folder "1-designs"</strong></p> <p>The text files in this directory should contain all information for the<br> independent variables any machine learning experiment. For reference, all 14641 IFC models are supplied in subfolders 001 to 147.</p> <p><strong>2. Folder "2-simulation"</strong></p> <p>This folder contains samples of the simulation output that may be viewed in Paraview (http://www.paraview.org). The original model contains the "Org" filename fragment, and the maximum and minimum behaviors are indicated with "Max" and "Min" filename fragments. Displacement, strain, and stress behaviors are all given. Only three of the 14641 models are given as the file sizes are<br> around 1.4 to 2.2 megabytes each. The complete data (approximately 81 gigabytes) can be regenerated and provided if necessary on request (email webis@medien.uni-weimar.de).</p> <p><strong>3. Folder "3-aggregation"</strong></p> <p>Maximum displacement, strain, and stress measurements are given in the text files individually, and together in the files with the "vtk" filename fragment. This data should be sufficient for the dependent variables of any machine learning experiment.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.