Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “company graph”
CompanyKG Dataset V2.0: A Large-Scale Heterogeneous Graph for Company Similarity Quantification
<p><strong>CompanyKG</strong> is a heterogeneous graph consisting of 1,169,931 nodes and 50,815,503 undirected edges, with each node representing a real-world company and each edge signifying a relationship between the connected pair of companies.</p> <p><strong>Edges</strong>: We model 15 different inter-company relations as undirected edges, each of which corresponds to a unique edge type. These edge types capture various forms of similarity between connected company pairs. Associated with each edge of a certain type, we calculate a real-numbered weight as an approximation of the similarity level of that type. It is important to note that the constructed edges do not represent an exhaustive list of all possible edges due to incomplete information. Consequently, this leads to a sparse and occasionally skewed distribution of edges for individual relation/edge types. Such characteristics pose additional challenges for downstream learning tasks. Please refer to our paper for a detailed definition of edge types and weight calculations.</p> <p><strong>Nodes</strong>: The graph includes all companies connected by edges defined previously. Each node represents a company and is associated with a descriptive text, such as "<em>Klarna is a fintech company that provides support for direct and post-purchase payments</em> ...". To comply with privacy and confidentiality requirements, we encoded the text into numerical embeddings using four different pre-trained text embedding models: <a href="https://huggingface.co/sentence-transformers/distiluse-base-multilingual-cased-v2">mSBERT</a> (multilingual Sentence BERT), <a href="https://platform.openai.com/docs/guides/embeddings/what-are-embeddings">ADA2</a>, <a href="https://github.com/princeton-nlp/SimCSE">SimCSE</a> (fine-tuned on the raw company descriptions) and <a href="https://github.com/EQTPartners/pause">PAUSE</a>.</p> <p><strong>Evaluation Tasks</strong>. The primary goal of CompanyKG is to develop algorithms and models for quantifying the similarity between pairs of companies. In order to evaluate the effectiveness of these methods, we have carefully curated three evaluation tasks:</p> <ul> <li><strong>Similarity Prediction (SP)</strong>. To assess the accuracy of pairwise company similarity, we constructed the SP evaluation set comprising 3,219 pairs of companies that are labeled either as positive (similar, denoted by "1") or negative (dissimilar, denoted by "0"). Of these pairs, 1,522 are positive and 1,697 are negative.</li> <li><strong>Competitor Retrieval (CR)</strong>. Each sample contains one <em>target company</em> and one of its direct competitors. It contains 76 distinct target companies, each of which has 5.3 competitors annotated in average. For a given target company A with <em>N</em> direct competitors in this CR evaluation set, we expect a competent method to retrieve all <em>N</em> competitors when searching for similar companies to A. </li> <li><strong>Similarity Ranking (SR)</strong> is designed to assess the ability of any method to rank <em>candidate companies</em> (numbered 0 and 1) based on their similarity to a <em>query company</em>. Paid human annotators, with backgrounds in engineering, science, and investment, were tasked with determining which candidate company is more similar to the query company. It resulted in an evaluation set comprising 1,856 rigorously labeled ranking questions. We retained 20% (368 samples) of this set as a validation set for model development. </li> <li><strong>Edge Prediction (EP)</strong> evaluates a model's ability to predict future or missing relationships between companies, providing forward-looking insights for investment professionals. The EP dataset, derived (and sampled) from new edges collected between April 6, 2023, and May 25, 2024, includes 40,000 samples, with edges not present in the pre-existing CompanyKG (a snapshot up until April 5, 2023).</li> </ul> <p><strong>Background and Motivation</strong></p> <p>In the investment industry, it is often essential to identify similar companies for a variety of purposes, such as market/competitor mapping and Mergers & Acquisitions (M&A). Identifying comparable companies is a critical task, as it can inform investment decisions, help identify potential synergies, and reveal areas for growth and improvement. The accurate quantification of inter-company similarity, also referred to as <strong>company similarity quantification</strong>, is the cornerstone to successfully executing such tasks. However, company similarity quantification is often a challenging and time-consuming process, given the vast amount of data available on each company, and the complex and diversified relationships among them.</p> <p>While there is no universally agreed definition of company similarity, researchers and practitioners in PE industry have adopted various criteria to measure similarity, typically reflecting the companies' operations and relationships. These criteria can embody one or more dimensions such as industry sectors, employee profiles, keywords/tags, customers' review, financial performance, co-appearance in news, and so on. Investment professionals usually begin with a limited number of companies of interest (a.k.a. seed companies) and require an algorithmic approach to expand their search to a larger list of companies for potential investment. </p> <p>In recent years, transformer-based Language Models (LMs) have become the preferred method for encoding textual company descriptions into vector-space embeddings. Then companies that are similar to the seed companies can be searched in the embedding space using distance metrics like cosine similarity. The rapid advancements in Large LMs (LLMs), such as GPT-3/4 and LLaMA, have significantly enhanced the performance of general-purpose conversational models. These models, such as ChatGPT, can be employed to answer questions related to similar company discovery and quantification in a Q&A format.</p> <p>However, graph is still the most natural choice for representing and learning diverse company relations due to its ability to model complex relationships between a large number of entities. By representing companies as nodes and their relationships as edges, we can form a <strong>Knowledge Graph (KG)</strong>. Utilizing this KG allows us to efficiently capture and analyze the network structure of the business landscape. Moreover, KG-based approaches allow us to leverage powerful tools from network science, graph theory, and graph-based machine learning, such as Graph Neural Networks (GNNs), to extract insights and patterns to facilitate similar company analysis. While there are various company datasets (mostly commercial/proprietary and non-relational) and graph datasets available (mostly for single link/node/graph-level predictions), there is a scarcity of datasets and benchmarks that combine both to create a large-scale KG dataset expressing rich pairwise company relations.</p> <p><strong>Source Code and Tutorial:<br></strong><a href="https://github.com/llcresearch/CompanyKG2"><strong>https://github.com/llcresearch/CompanyKG2</strong></a></p> <p><strong>Paper: to be published<br></strong></p>
CorpWatch Companies Graph
<p>This dataset contains information about corporations listed in the U.S. stock exchange. The information concerns the ownership hierarchy of companies (i.e., holder-subsidiary company relations) along with the industry sector and location of each company. The dataset has the form of a graph. It has been produced by the SmartDataLake project (<a href="https://smartdatalake.eu">https://smartdatalake.eu</a>), using data collected from CorpWatch (<a href="https://corpwatch.org">https://corpwatch.org</a>).</p>
TaxGraph - A Knowledge Graph of Multi-National Companies and their Relationships
<p>The taxation of multi-national companies is a complex field, since it is influenced by the legislation of several states. Laws in different states may have unforeseen interaction effects, which can be exploited by allowing multi-national companies to minimize and avoid taxes.</p> <p>Thus, we created a knowledge graph of multi-national companies and their relationships. Many commonly known tax avoidance strategies can be formulated as subgraph queries to this graph, which allows for identifying companies using certain strategies. Moreover, we can identify anomalies in the graph which hint at potential tax avoidance strategies.</p>
Wikidata Companies Graph
<p>This dataset contains information about commercial organizations (companies) and their relations with other commercial organizations, persons, products, locations, groups and industries. The dataset has the form of a graph. It has been produced by the SmartDataLake project (<a href="https://smartdatalake.eu">https://smartdatalake.eu</a>), using data collected from Wikidata (<a href="https://www.wikidata.org">https://www.wikidata.org</a>).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.