Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
915
datasets available to search
ShareScore release 0.9.0
Dataset results
915 results for “graphs”
CompanyKG Dataset V2.0: A Large-Scale Heterogeneous Graph for Company Similarity Quantification
<p><strong>CompanyKG</strong> is a heterogeneous graph consisting of 1,169,931 nodes and 50,815,503 undirected edges, with each node representing a real-world company and each edge signifying a relationship between the connected pair of companies.</p> <p><strong>Edges</strong>: We model 15 different inter-company relations as undirected edges, each of which corresponds to a unique edge type. These edge types capture various forms of similarity between connected company pairs. Associated with each edge of a certain type, we calculate a real-numbered weight as an approximation of the similarity level of that type. It is important to note that the constructed edges do not represent an exhaustive list of all possible edges due to incomplete information. Consequently, this leads to a sparse and occasionally skewed distribution of edges for individual relation/edge types. Such characteristics pose additional challenges for downstream learning tasks. Please refer to our paper for a detailed definition of edge types and weight calculations.</p> <p><strong>Nodes</strong>: The graph includes all companies connected by edges defined previously. Each node represents a company and is associated with a descriptive text, such as "<em>Klarna is a fintech company that provides support for direct and post-purchase payments</em> ...". To comply with privacy and confidentiality requirements, we encoded the text into numerical embeddings using four different pre-trained text embedding models: <a href="https://huggingface.co/sentence-transformers/distiluse-base-multilingual-cased-v2">mSBERT</a> (multilingual Sentence BERT), <a href="https://platform.openai.com/docs/guides/embeddings/what-are-embeddings">ADA2</a>, <a href="https://github.com/princeton-nlp/SimCSE">SimCSE</a> (fine-tuned on the raw company descriptions) and <a href="https://github.com/EQTPartners/pause">PAUSE</a>.</p> <p><strong>Evaluation Tasks</strong>. The primary goal of CompanyKG is to develop algorithms and models for quantifying the similarity between pairs of companies. In order to evaluate the effectiveness of these methods, we have carefully curated three evaluation tasks:</p> <ul> <li><strong>Similarity Prediction (SP)</strong>. To assess the accuracy of pairwise company similarity, we constructed the SP evaluation set comprising 3,219 pairs of companies that are labeled either as positive (similar, denoted by "1") or negative (dissimilar, denoted by "0"). Of these pairs, 1,522 are positive and 1,697 are negative.</li> <li><strong>Competitor Retrieval (CR)</strong>. Each sample contains one <em>target company</em> and one of its direct competitors. It contains 76 distinct target companies, each of which has 5.3 competitors annotated in average. For a given target company A with <em>N</em> direct competitors in this CR evaluation set, we expect a competent method to retrieve all <em>N</em> competitors when searching for similar companies to A. </li> <li><strong>Similarity Ranking (SR)</strong> is designed to assess the ability of any method to rank <em>candidate companies</em> (numbered 0 and 1) based on their similarity to a <em>query company</em>. Paid human annotators, with backgrounds in engineering, science, and investment, were tasked with determining which candidate company is more similar to the query company. It resulted in an evaluation set comprising 1,856 rigorously labeled ranking questions. We retained 20% (368 samples) of this set as a validation set for model development. </li> <li><strong>Edge Prediction (EP)</strong> evaluates a model's ability to predict future or missing relationships between companies, providing forward-looking insights for investment professionals. The EP dataset, derived (and sampled) from new edges collected between April 6, 2023, and May 25, 2024, includes 40,000 samples, with edges not present in the pre-existing CompanyKG (a snapshot up until April 5, 2023).</li> </ul> <p><strong>Background and Motivation</strong></p> <p>In the investment industry, it is often essential to identify similar companies for a variety of purposes, such as market/competitor mapping and Mergers & Acquisitions (M&A). Identifying comparable companies is a critical task, as it can inform investment decisions, help identify potential synergies, and reveal areas for growth and improvement. The accurate quantification of inter-company similarity, also referred to as <strong>company similarity quantification</strong>, is the cornerstone to successfully executing such tasks. However, company similarity quantification is often a challenging and time-consuming process, given the vast amount of data available on each company, and the complex and diversified relationships among them.</p> <p>While there is no universally agreed definition of company similarity, researchers and practitioners in PE industry have adopted various criteria to measure similarity, typically reflecting the companies' operations and relationships. These criteria can embody one or more dimensions such as industry sectors, employee profiles, keywords/tags, customers' review, financial performance, co-appearance in news, and so on. Investment professionals usually begin with a limited number of companies of interest (a.k.a. seed companies) and require an algorithmic approach to expand their search to a larger list of companies for potential investment. </p> <p>In recent years, transformer-based Language Models (LMs) have become the preferred method for encoding textual company descriptions into vector-space embeddings. Then companies that are similar to the seed companies can be searched in the embedding space using distance metrics like cosine similarity. The rapid advancements in Large LMs (LLMs), such as GPT-3/4 and LLaMA, have significantly enhanced the performance of general-purpose conversational models. These models, such as ChatGPT, can be employed to answer questions related to similar company discovery and quantification in a Q&A format.</p> <p>However, graph is still the most natural choice for representing and learning diverse company relations due to its ability to model complex relationships between a large number of entities. By representing companies as nodes and their relationships as edges, we can form a <strong>Knowledge Graph (KG)</strong>. Utilizing this KG allows us to efficiently capture and analyze the network structure of the business landscape. Moreover, KG-based approaches allow us to leverage powerful tools from network science, graph theory, and graph-based machine learning, such as Graph Neural Networks (GNNs), to extract insights and patterns to facilitate similar company analysis. While there are various company datasets (mostly commercial/proprietary and non-relational) and graph datasets available (mostly for single link/node/graph-level predictions), there is a scarcity of datasets and benchmarks that combine both to create a large-scale KG dataset expressing rich pairwise company relations.</p> <p><strong>Source Code and Tutorial:<br></strong><a href="https://github.com/llcresearch/CompanyKG2"><strong>https://github.com/llcresearch/CompanyKG2</strong></a></p> <p><strong>Paper: to be published<br></strong></p>
NASA GES-DISC Knowledge Graph for Link Prediction
<p>This dataset includes a knowledge graph of NASA GES-DISC collections, featuring interconnected nodes for datasets, data center, projects, platforms, instruments, science keywords, and publications. Designed for link prediction tasks, it aids machine learning research in satellite observation, remote sensing, and climate change. The dataset is in CSV format, ready for graph databases and ML frameworks.</p> <p> </p>
Graph 1 in Gastroesophageal tube of the Iguana iguana (Iguanidae): histological description, histochemical and immunohistochemical analysis of 5-HT and SS cells
Graph 1. Number of 5-HT and SS-cells/µm2 per oesophageal and gastric region in I. iguana. Values are averages with their standard deviatin shown by vertical bars (p = 0,05). Cervical oesophagus (CER): 4.6x10-2 ± 2.0 [5-HT cells]/µm2. Celomatic oesophagus (CEL): 4.0x10-2 ± 1.0 [5-HT cells]/µm2. Cranial and middle regions of the stomach (C/MS): 6.18x10-2 ± 3.2 [5-HT cells]/µm2. Caudal Stomach (CS): 0.6x10-2 ± 0.2 [5-HT cells]/µm2 and 1.4x 10-2 ± 0.9 [SS-cells]/µm2.
Graph 1 in Toxicological assessments of basic blue 3 dye in fresh water bivalve Lamellidens marginalis
Graph 1: Changes in SODactivity in different tissues of fresh water bivalve, Lamellidens marginalis after acute exposure to basic blue 3(values are expressed in unit/mg protein/hour)
Graph 4 in Toxicological assessments of basic blue 3 dye in fresh water bivalve Lamellidens marginalis
Graph 4: DNA strand breaks in gill cells of Lamellidens marginalis after acute exposure to Basic blue 3.
Graph 2 in Toxicological assessments of basic blue 3 dye in fresh water bivalve Lamellidens marginalis
Graph 2: Changes in CAT activity in different tissues of fresh water bivalve, Lamellidens marginalis after acute exposure to basic blue 3 (values are inmmolH2O2/min/mg protein)
Graph 3 in Toxicological assessments of basic blue 3 dye in fresh water bivalve Lamellidens marginalis
Graph 3: Changes in GPx activityin different tissues of fresh water bivalve, Lamellidens marginalis after acute exposure to basic blue 3 (values are in mmol NADPH/min/mg protein)
ORBITAAL: cOmpRehensive BItcoin daTaset for temporAl grAph anaLysis
<h3>Dataset Construction</h3> <p>This dataset captures the temporal network of Bitcoin (BTC) flow exchanged between entities at the finest time resolution in UNIX timestamp. Its construction is based on the blockchain covering the period from January, 3rd of 2009 to January the 25th of 2021. The blockchain extraction has been made using bitcoin-etl (<a href="https://github.com/blockchain-etl/bitcoin-etl">https://github.com/blockchain-etl/bitcoin-etl</a>) Python package. The entity-entity network is built by aggregating Bitcoin addresses using the common-input heuristic [1] as well as popular Bitcoin users' addresses provided by <a href="https://www.walletexplorer.com/">https://www.walletexplorer.com/</a></p> <p>[1] M. Harrigan and C. Fretter, "The Unreasonable Effectiveness of Address Clustering," <em>2016 Intl IEEE Conferences on Ubiquitous Intelligence & Computing, Advanced and Trusted Computing, Scalable Computing and Communications, Cloud and Big Data Computing, Internet of People, and Smart World Congress (UIC/ATC/ScalCom/CBDCom/IoP/SmartWorld)</em>, Toulouse, France, 2016, pp. 368-373, doi: 10.1109/UIC-ATC-ScalCom-CBDCom-IoP-SmartWorld.2016.0071.<br>keywords: {Online banking;Merging;Protocols;Upper bound;Bipartite graph;Electronic mail;Size measurement;bitcoin;cryptocurrency;blockchain},</p> <p> </p> <h3>Dataset Description</h3> <p><strong>Bitcoin Activity Temporal Coverage</strong>: From 03 January 2009 to 25 January 2021</p> <h4>Overview:</h4> <p>This <strong>dataset </strong>provides a <strong>comprehensive</strong> representation of <strong>Bitcoin exchanges</strong> between entities over a s<strong>ignificant temporal span</strong>, spanning from the inception of Bitcoin to recent years. It encompasses <strong>various temporal resolutions</strong> and <strong>representations</strong> to <strong>facilitate Bitcoin transaction network analysis </strong>in the context of <strong>temporal graphs</strong>.</p> <p>Every dates have been retrieved from bloc UNIX timestamp and GMT timezone.</p> <h4>Contents:</h4> <p>The dataset is distributed across three compressed archives:</p> <p>All data are stored in the <strong>Apache Parquet file format</strong>, a columnar storage format optimized for analytical queries. It can be used with pyspark Python package.</p> <ol> <li> <p><strong>orbitaal-stream_graph.tar.gz</strong>:</p> <ul> <li>The root directory is <em>STREAM_GRAPH/</em></li> <li>Contains a <strong>stream graph</strong> representation of Bitcoin exchanges at the <strong>finest temporal scale</strong>, corresponding to the validation time of <strong>each block</strong> (averaging approximately 10 minutes).</li> <li>The stream graph is divided into 13 files, one for each year</li> <li>Files format is parquet</li> <li>Name format is <strong>orbitaal-stream_graph-date-[YYYY]-file-id-[ID].snappy.parquet,</strong> where <em>[YYYY]</em> stands for the corresponding <em>year</em> and <em>[ID]</em> is <em>an integer</em> from 1 to N (number of files here) such as sorting in increasing [ID] ordering is similar to sort by increasing year ordering</li> <li>These files are in the subdirectory <em>STREAM_GRAPH/EDGES/</em></li> </ul> </li> <li> <p><strong>orbitaal-snapshot-all.tar.gz</strong>:</p> <ul> <li>The root directory is <em>SNAPSHOT/</em></li> <li>Contains the <strong>snapshot</strong> network representing <strong>all transactions aggregated </strong>over the whole dataset period (from Jan. 2009 to Jan. 2021).</li> <li>Files format is parquet</li> <li>Name format is <strong>orbitaal-snapshot-all.snappy.parquet</strong>.</li> <li>These files are in the subdirectory <em>SNAPSHOT/EDGES/ALL/</em></li> </ul> </li> <li> <p><strong>orbitaal-snapshot-year.tar.gz</strong>:</p> <ul> <li>The root directory is <em>SNAPSHOT/</em></li> <li>Contains the <strong>yearly</strong><em> </em>resolution of <strong>snapshot</strong> networks</li> <li>Files format is parquet</li> <li>Name format is <strong>orbitaal-snapshot-date-[YYYY]-file-id-[ID].snappy.parquet</strong>, where <em>[YYYY]</em> stands for the corresponding <em>year </em>and <em>[ID]</em> is an <em>integer </em>from 1 to N (number of files here) such as sorting in increasing [ID] ordering is similar to sort by increasing year ordering</li> <li>These files are in the subdirectory <em>SNAPSHOT/EDGES/year/</em></li> </ul> </li> <li> <p><strong>orbitaal-snapshot-month.tar.gz</strong>:</p> <ul> <li>The root directory is <em>SNAPSHOT/</em></li> <li>Contains the <strong>monthly </strong>resoluted <strong>snapshot </strong>networks</li> <li>Files format is parquet</li> <li>Name format is <strong>orbitaal-snapshot-date-[YYYY]-[MM]-file-id-[ID].snappy.parquet</strong>, where</li> <li><em>[YYYY] </em>and <em>[MM] </em>stands for the corresponding <em>year </em>and <em>month, </em>and <em>[ID] </em>is an <em>integer </em>from 1 to N (number of files here) such as sorting in increasing [ID] ordering is similar to sort by increasing year and month ordering</li> <li>These files are in the subdirectory <em>SNAPSHOT/EDGES/month/</em></li> </ul> </li> <li> <p><strong>orbitaal-snapshot-day.tar.gz</strong>:</p> <ul> <li>The root directory is <em>SNAPSHOT/</em></li> <li>Contains the <strong>daily </strong>resoluted <strong>snapshot </strong>networks</li> <li>Files format is parquet</li> <li>Name format is <strong>orbitaal-snapshot-date-[YYYY]-[MM]-[DD]-file-id-[ID].snappy.parquet</strong>, where</li> <li><em>[YYYY]</em>, <em>[MM]</em>, and <em>[DD] </em>stand for the corresponding <em>year</em>, <em>month</em>, and <em>day</em>, and <em>[ID] </em>is an <em>integer </em>from 1 to N (number of files here) such as sorting in increasing [ID] ordering is similar to sort by increasing year, month, and day ordering</li> <li>These files are in the subdirectory <em>SNAPSHOT/EDGES/day/</em></li> </ul> </li> <li> <p><strong>orbitaal-snapshot-hour.tar.gz</strong>:</p> <ul> <li>The root directory is <em>SNAPSHOT/</em></li> <li>Contains the <strong>hourly </strong>resoluted <strong>snapshot </strong>networks</li> <li>Files format is parquet</li> <li>Name format is <strong>orbitaal-snapshot-date-[YYYY]-[MM]-[DD]-[hh]-file-id-[ID].snappy.parquet</strong>, where</li> <li><em>[YYYY]</em>, <em>[MM]</em>, <em>[DD]</em>, and <em>[hh]</em> stand for the corresponding <em>year, month, day, </em>and <em>hour</em>, and <em>[ID] </em>is an <em>integer </em>from 1 to N (number of files here) such as sorting in increasing [ID] ordering is similar to sort by increasing year, month, day and hour ordering</li> <li>These files are in the subdirectory <em>SNAPSHOT/EDGES/hour/</em></li> </ul> </li> <li> <p><strong>orbitaal-nodetable.tar.gz</strong>:</p> <ul> <li>The root directory is <em>NODE_TABLE/</em></li> <li>Contains two files in parquet format, the first one gives <strong>information </strong>related to <strong>nodes </strong>present in stream graphs and snapshots such as <strong>period of activity</strong> and associated global <strong>Bitcoin balance</strong>, and the other one contains the list of <strong>all associated Bitcoin addresses.</strong></li> </ul> </li> </ol> <p> </p> <p>Small samples in CSV format</p> <ol> <li> <p><strong>orbitaal-stream_graph-2016_07_08.csv</strong> and <strong>orbitaal-stream_graph-2016_07_09.csv</strong></p> <ul> <li>These two CSV files are related to stream graph representations of an halvening happening in 2016.</li> </ul> </li> <li> <p><strong>orbitaal-snapshot-2016_07_08.csv </strong>and<strong> orbitaal-snapshot-2016_07_09.csv</strong></p> <ul> <li>These two CSV files are related to daily snapshot representations of an halvening happening in 2016.</li> <li><strong> </strong></li> </ul> </li> </ol> <p> </p> <p> </p> <p> </p>
Conditional gradient for total variation regularization with PDE constraints: a graph cuts approach
This module solves a PDE constrained minimisation problem with TV-regularization, using the method described in the paper "Conditional gradient for total variation regularization with PDE constraints: a graph cuts approach"
TAXREF-LD: Knowledge Graph of the French taxonomic registry
<p>TAXREF-LD is a Linked Data knowledge graph representing <a href="https://inpn.mnhn.fr/programme/referentiel-taxonomique-taxref?lg=en">TAXREF</a>, the French national taxonomical register for fauna, flora and fungus, that covers mainland France and overseas territories.</p> <p>TAXREF-LD is a joint initiative of the <a href="http://www.patrinat.fr/">UMS Patrinat</a> of the <a href="http://www.mnhn.fr/">National Museum of Natural History</a>, and the <a href="http://www.i3s.unice.fr/">I3S laboratory</a>, <a href="https://univ-cotedazur.fr">University Côte d'Azur</a>, <a href="https://www.inria.fr">Inria</a>, <a href="https://www.cnrs.fr">CNRS</a>.</p> <p>Homepage: https://github.com/frmichel/taxref-ld/</p>
AQL queries and benchmark results from PhD thesis "ANNIS: A graph-based query system for deeply annotated text corpora"
<p>These are the queries, the benchmark results and the evaluation scripts of the thesis "ANNIS: A graph-based query system for deeply annotated text corpora" (Thomas Krause 2018, Humboldt-Universität zu Berlin)</p> <p><strong>diss_2018-01-12_v0.5.0.csv </strong><br> Results of all configurations of executed benchmarks for graphANNIS and also the baseline times of relANNIS.</p> <p><strong>queries.zip</strong><br> Contains folders for each corpus containing all queries used for the benchmark. Each file-name begins with the ID of the query. The extension denotes the type, which can be one of the following:</p> <ul> <li><em>".</em>aql" contains the original AQL (ANNIS query language) query which was collected</li> <li>".json" is JSON representation of the parsed AQL query</li> <li>".count" is the number of matches a query should have</li> <li>".time" is the average time in milliseconds that was needed to execute the query in relANNIS on the benchmark system</li> <li>".corpora" contains the name of the corpus the query belongs to (should be only one corpus and the same as the folder name in the selection of queries in this data set)</li> <li>".relplan" contains the PostgreSQL plan for the query</li> <li>".graphplan" contains the graphANNIS plan for the query</li> </ul> <p><strong>evaluation-scripts.py/evaluation-scripts.ipynb</strong><br> Python scripts to perform the evaluation and generate the images. This are both a Python-file and the original notebook file that can be used with the Jupyter Notebook application.</p> <p><strong>relannis_benchmark_scripts.zip </strong><br> The files in this zip-file can be used to execute the benchmarks in the relANNIS system by piping the into the "annis.sh" command line tool of relANNIS</p>
Figure 3. Graph created directly from the results.-Browsing Semantic Data in Slovakia
<p>Resulting graph is rather complex. There are 175 vertices and 201 edges, which were created directly, containing 158 persons and 17 companies. As we see on Figure 3, some filtering methods are required in order to create suitable overview of relations. Figure 4 thus shows the visualization of the same graph, but with tens of vertices merged. Now it contains 22 persons and 17 companies. A merging was performed for clarification and is built inside the visualization module. A simple condition says that a merging is performed if persons are unique, thus if a person is connected only to 1 firm. More formally, person vertices are merged, if their vertex degree equals 1 (each).</p>
BRAIN Journal-Automatic Anthropometric System Development Using Machine Learning-Figure 3. (3.a) – The flowchart of Graph cuts method; (3.b)- the result of Graph cuts image segmentation.
<p>Figure 3 describes the steps implemented Graph cuts algorithm for the segmentation of human body parts. The results obtained are 5 main sections that include the hands, the legs, the center of the body (chest, waist, hips), and the head. The result of the display image is taken from the human image database, which was collected by us (Нгуен, 2016). </p>
BRAIN Journal-Right-Linear Languages Generated in Systems of Knowledge Representation based on LSG-Right-Figure 3. Representation of the grammar G1 in the labelled graph G0 1
<p>If we take the labeled graph G0 1 given in Figure 3 and construct the stratified graph structure over (99) such that (100) we obtain (101), (102). </p> <p>In this paper, we proposed a new system for formal language generation by means of stratified graphs structures. This mechanism can generate languages of the first type and of the second type. More precisely, we propose a new system for formal language generation by means of a system of knowledge based on stratified graphs. We exemplified that, using an interpretation system specially defined for stratified graphs representations, a particular formal language can be obtained by means of the resulted accepted structured paths.</p>
BRAIN Journal-Auto-generative Learning Objects in Online Assessment of Data Structures Disciplines-Figure 3. Symbols definition for a graph-based test
<p>In this scenario, we intend to generate a random graph and compute a deep first-search node list. The first defined random symbol is n, namely the number of nodes in the graph as an integer from 5 to 9. The next symbol is named g and denotes the graph object created randomly using 3 parameters: the number of nodes, the minimum, and the maximum value for the weight. For the number of nodes, we used the previously computed value of n, whereas for the weights, we used two constants 0 and 1 since the graph is not weighted</p>
BRAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 10. Detailed accuracy separated by classes and a confusion matrix which belongs to the dataset of Coiflet 1 applied by our main method (ANNSVM)
<p>With regard to accuracy values of each class as presented in Figure 10, we observed that the accuracy of the two-dimensional chart class was the lowest (i.e., 0.875), while others were over 0.9. Results here suggested that both the bar and pie classes have their own unique characteristics, as opposed to the 2Dchart class. For example, the graph images that contained some rectangles were individually categorized in the bar graph class. A similar phenomenon occurred for circles in the pie chart class. In contrast, the 2Dchart class contained mixed types of graphs; hence, the graph characteristics belonging to the 2Dchart class varied. </p>
BRAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 8. Simulation of Coiflet 1 (PyWavelets discussion group, 2008), analyzing as one-dimensional images
<p>Using only the wavelet coefficients was inadequate for classification. For example, for the pie chart, we obtained large wavelet coefficients located in the low-frequency domain; however, if we changed a circle in the pie chart to other shapes, such as a radar chart, the wavelet transformation gave results that were similar to those of the original pie chart. The Hough transformation can solve this problem since it detects the shapes of objects</p>
BRAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 7b. Results from SVM, ANN, SVMANN, and ANNSVM that used WLHT
<p>From the results of ANNSVM_WL and ANNSVM_HT, we found that wavelet coefficients had a larger impact on classification than the Hough transformation data because the results from our proposed method applied to WL were more accurate than those of HT. The wavelet coefficients can capture the dominant characteristics from the graphs better than the Hough transformation. The one-dimensional image represented in the frequency domain had oscillations with different amplitudes depending on the graph types. For example, a dominant part of a pie chart should be in the low-frequency domain, because there is a large island of concatenated pixels in a onedimensional image, and it has only a few changes. Conversely, since the scatter plot contains many widely spread points, its dominant part should be located in the high-frequency domain. Performing the wavelet transformation, if a mother wavelet and a part of the wavelet function have a close match, the wavelet coefficient will be large. Assuming we use a suitable wavelet family with the example pie chart case, the wavelet coefficients in the low-frequency domain should be large as compared to other parts of the domain. </p>
BRAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 7c. Results from SVM, ANN, SVMANN, and ANNSVM that used WLHT
<p>From the results of ANNSVM_WL and ANNSVM_HT, we found that wavelet coefficients had a larger impact on classification than the Hough transformation data because the results from our proposed method applied to WL were more accurate than those of HT. The wavelet coefficients can capture the dominant characteristics from the graphs better than the Hough transformation. The one-dimensional image represented in the frequency domain had oscillations with different amplitudes depending on the graph types. For example, a dominant part of a pie chart should be in the low-frequency domain, because there is a large island of concatenated pixels in a onedimensional image, and it has only a few changes. Conversely, since the scatter plot contains many widely spread points, its dominant part should be located in the high-frequency domain. Performing the wavelet transformation, if a mother wavelet and a part of the wavelet function have a close match, the wavelet coefficient will be large. Assuming we use a suitable wavelet family with the example pie chart case, the wavelet coefficients in the low-frequency domain should be large as compared to other parts of the domain. </p>
BRAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 7. Results from SVM, ANN, SVMANN, and ANNSVM that used WLHT
<p>From the results of ANNSVM_WL and ANNSVM_HT, we found that wavelet coefficients had a larger impact on classification than the Hough transformation data because the results from our proposed method applied to WL were more accurate than those of HT. The wavelet coefficients can capture the dominant characteristics from the graphs better than the Hough transformation. The one-dimensional image represented in the frequency domain had oscillations with different amplitudes depending on the graph types. For example, a dominant part of a pie chart should be in the low-frequency domain, because there is a large island of concatenated pixels in a onedimensional image, and it has only a few changes. Conversely, since the scatter plot contains many widely spread points, its dominant part should be located in the high-frequency domain. Performing the wavelet transformation, if a mother wavelet and a part of the wavelet function have a close match, the wavelet coefficient will be large. Assuming we use a suitable wavelet family with the example pie chart case, the wavelet coefficients in the low-frequency domain should be large as compared to other parts of the domain. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.