Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,037
datasets available to search
ShareScore release 0.9.0
Dataset results
1,037 results for “large-scale”
DongTing: A Large-scale Dataset for Anomaly Detection of the Linux Kernel
<p>DongTing is the first large-scale dataset dedicated to Linux kernel anomaly detection. The dataset covers Linux kernels released in the last five years and includes a total of 18,966 well-labeled normal and attack sequences. The entire dataset is 85 GB in size (after decompression). The attack data covers 26 major kernel releases and contains a total of 12,116 system call sequences collected from running 17,855 bug-triggering programs. The normal data comes from 6,850 normal programs in four kernel regression test suites. We maintain the dataset and source code in Zenodo and Github, respectively, and back up the dataset and code in Baidu netdisk.</p> <h3><strong>Dataset</strong></h3> <p>The dataset is stored at <a href="http://doi.org/10.5281/zenodo.6627050">http://doi.org/10.5281/zenodo.6627050</a></p> <ul> <li>The data includes `abnormal_data`, `normal_data`, `models`, `npz` and baseline data, with a total volume of nearly 87 GB (including 85 GB for abnormal data and normal data, it's after decompression files size).</li> <li>The `Abnormal_data` directory contains 12,116 files containing system call sequence for 26 kernel releases, and the `Normal_data` directory contains 6,850 files containing system call sequences collected from four regression test suites. All of which are raw sequences.</li> <li>CNN/RNN, LSTM, and Wavenet (three sets of hyperparameters per model) machine learning models are selected, the ECOD model (without hyperparameters) was also chosen for the evaluation of DT. DT_abnormal, DT_normal, ADFA-LD, and PLAID are used for training respectively. The results of DT training models are stored in the directory `Models-DongTing`, and the results of ADFA-LD and PLAID training models are stored in the directory `Models-Comparison`.</li> <li>The directory `npz `stores the encoded dataset of DongTing, ADFA-LD, and PLAID (sequence length varies from 8 to 4495), according to syscall_64.tbl in Linux kernel 5.17, including the training set, validation set, and test set.</li> <li>The file `Baseline.xlsx` contains all the information about DongTing dataset, which can be used in training machine learning models. For example, the whole dataset is randomly divided into three sets with the ratio of 80%:10%:10% (training: validation: test). The implementation of dataset division can be found in the source code.</li> </ul> <h3><strong>Source Code</strong></h3> <p><br>The source code for dataset development is stored at <a href="https://github.com/HNUSystemsLab/DongTing">https://github.com/HNUSystemsLab/DongTing</a> and the following is a brief introduction.</p> <ul> <li>The source code contains three folders, i.e., `Source Code Files`, `Documents` and `DB`, where `Documents `stores the detailed documents related to development, `DB` stores samples data, and `Source Code Files` stores the source code related to the development of our dataset.</li> <li>The detailed description about the source code can be found in `Documents/Documentation.pdf`. The document consists of four parts: environment requirements, database, program structure and working steps, model training and evaluation (including training and evaluation). It details the preparation of the environment, data import method, functional description of each file in the source code directory, how model training and evaluation work and other related contents.</li> </ul> <p>We additionally maintain the dataset and source code on Baidu.com <a href="https://pan.baidu.com/s/1vu1WGZpf2DqMIoyGayNu3w?pwd=dtds">https://pan.baidu.com/s/1vu1WGZpf2DqMIoyGayNu3w?pwd=dtds</a> to facilitate the access from China.</p> <p> </p> <h3>Tips: </h3> <p>If you find DongTing useful for your research, please cite the article as "DongTing: A large-scale dataset for anomaly detection of the Linux kernel".</p> <blockquote> <p><br>@article{DUAN2023111745,<br>title = {DongTing: A large-scale dataset for anomaly detection of the Linux kernel},<br>journal = {Journal of Systems and Software},<br>volume = {203},<br>pages = {111745},<br>year = {2023},<br>issn = {0164-1212},<br>doi = {https://doi.org/10.1016/j.jss.2023.111745},<br>url = {https://www.sciencedirect.com/science/article/pii/S0164121223001401},<br>author = {Guoyun Duan and Yuanzhi Fu and Minjie Cai and Hao Chen and Jianhua Sun}<br>}<br><br></p> </blockquote>
A Large-Scale Empirical Study of Android Sports Apps in the Google Play Store
<p>This repository contains the dataset for our study "A Large-Scale Empirical Study of Android Sports Apps in the Google Play Store" and this will help to replicate our study, also the <a href="https://github.com/mooselab/Sports-Apps-Analysis">replication package</a> to direct you to help replicate it for your dataset too. </p> <p>Note: The dataset given are protected with password, and the password is available in our published paper</p>
WaterBench-Iowa: A Large-scale Benchmark Dataset for Data-Driven Streamflow Forecasting
<p>WaterBench-Iowa is a comprehensive benchmark dataset for streamflow forecasting. It follows FAIR data principles that are prepared with a focus on convenience for utilizing in data-driven and machine learning studies and provides benchmark performance for state-of-art deep learning architectures on the dataset for comparative analysis. By aggregating the datasets of streamflow, precipitation, watershed area, slope, soil types, and evapotranspiration from federal agencies and state organizations (i.e., NASA, NOAA, USGS, and Iowa Flood Center), we provided the WaterBench for hourly streamflow forecast studies. This dataset has a high temporal and spatial resolution with rich metadata and relational information, which can be used for varieties of deep learning and machine learning research. To some extent, WaterBench makes up for the lack of a unified benchmark in earth science research. We highly encourage researchers to use the WaterBench for deep learning research in hydrology.</p>
Flickr Africa: Examining Geo-Diversity in Large-Scale, Human-Centric Visual Data
<p>This dataset is provided for the paper "Flickr Africa: Examining Geo-Diversity in Large-Scale, Human-Centric Visual Data".</p> <p>Please refer to the readme in the zipped file for additional documentation.</p> <p>The zipped file contains two CSV files for every country in Africa obtained by queries "[country name]" and "[country name + people]".</p>
Data from: The State of OA: A large-scale analysis of the prevalence and impact of Open Access articles
<p>This is the raw data behind the publication: </p> <p><strong>The State of OA: A large-scale analysis of the prevalence and impact of Open Access articles.</strong></p> <p>Despite growing interest in Open Access (OA) to scholarly literature, there is an unmet need for large-scale, up-to-date, and reproducible studies assessing the prevalence and characteristics of OA. We address this need using oaDOI, an open online service that determines OA status for 67 million articles. We use three samples, each of 100,000 articles, to investigate OA in three populations: 1) all journal articles assigned a Crossref DOI, 2) recent journal articles indexed in Web of Science, and 3) articles viewed by users of Unpaywall, an open-source browser extension that lets users find OA articles using oaDOI. We estimate that at least 28% of the scholarly literature is OA (19M in total) and that this proportion is growing, driven particularly by growth in Gold and Hybrid. The most recent year analyzed (2015) also has the highest percentage of OA (45%). Because of this growth, and the fact that readers disproportionately access newer articles, we find that Unpaywall users encounter OA quite frequently: 47% of articles they view are OA. Notably, the most common mechanism for OA is not Gold, Green, or Hybrid OA, but rather an under-discussed category we dub Bronze: articles made free-to-read on the publisher website, without an explicit Open license. We also examine the citation impact of OA articles, corroborating the so-called open-access citation advantage: accounting for age and discipline, OA articles receive 18% more citations than average, an effect driven primarily by Green and Hybrid OA. We encourage further research using the free oaDOI service, as a way to inform OA policy and practice.</p>
Adaptive Behavior of Farmers Under Consecutive Droughts Results In More Vulnerable Farmers: A Large-Scale Agent-Based Modeling Analysis in the Bhima Basin, India
Open the record for dataset details and reuse information.
CompanyKG Dataset V2.0: A Large-Scale Heterogeneous Graph for Company Similarity Quantification
<p><strong>CompanyKG</strong> is a heterogeneous graph consisting of 1,169,931 nodes and 50,815,503 undirected edges, with each node representing a real-world company and each edge signifying a relationship between the connected pair of companies.</p> <p><strong>Edges</strong>: We model 15 different inter-company relations as undirected edges, each of which corresponds to a unique edge type. These edge types capture various forms of similarity between connected company pairs. Associated with each edge of a certain type, we calculate a real-numbered weight as an approximation of the similarity level of that type. It is important to note that the constructed edges do not represent an exhaustive list of all possible edges due to incomplete information. Consequently, this leads to a sparse and occasionally skewed distribution of edges for individual relation/edge types. Such characteristics pose additional challenges for downstream learning tasks. Please refer to our paper for a detailed definition of edge types and weight calculations.</p> <p><strong>Nodes</strong>: The graph includes all companies connected by edges defined previously. Each node represents a company and is associated with a descriptive text, such as "<em>Klarna is a fintech company that provides support for direct and post-purchase payments</em> ...". To comply with privacy and confidentiality requirements, we encoded the text into numerical embeddings using four different pre-trained text embedding models: <a href="https://huggingface.co/sentence-transformers/distiluse-base-multilingual-cased-v2">mSBERT</a> (multilingual Sentence BERT), <a href="https://platform.openai.com/docs/guides/embeddings/what-are-embeddings">ADA2</a>, <a href="https://github.com/princeton-nlp/SimCSE">SimCSE</a> (fine-tuned on the raw company descriptions) and <a href="https://github.com/EQTPartners/pause">PAUSE</a>.</p> <p><strong>Evaluation Tasks</strong>. The primary goal of CompanyKG is to develop algorithms and models for quantifying the similarity between pairs of companies. In order to evaluate the effectiveness of these methods, we have carefully curated three evaluation tasks:</p> <ul> <li><strong>Similarity Prediction (SP)</strong>. To assess the accuracy of pairwise company similarity, we constructed the SP evaluation set comprising 3,219 pairs of companies that are labeled either as positive (similar, denoted by "1") or negative (dissimilar, denoted by "0"). Of these pairs, 1,522 are positive and 1,697 are negative.</li> <li><strong>Competitor Retrieval (CR)</strong>. Each sample contains one <em>target company</em> and one of its direct competitors. It contains 76 distinct target companies, each of which has 5.3 competitors annotated in average. For a given target company A with <em>N</em> direct competitors in this CR evaluation set, we expect a competent method to retrieve all <em>N</em> competitors when searching for similar companies to A. </li> <li><strong>Similarity Ranking (SR)</strong> is designed to assess the ability of any method to rank <em>candidate companies</em> (numbered 0 and 1) based on their similarity to a <em>query company</em>. Paid human annotators, with backgrounds in engineering, science, and investment, were tasked with determining which candidate company is more similar to the query company. It resulted in an evaluation set comprising 1,856 rigorously labeled ranking questions. We retained 20% (368 samples) of this set as a validation set for model development. </li> <li><strong>Edge Prediction (EP)</strong> evaluates a model's ability to predict future or missing relationships between companies, providing forward-looking insights for investment professionals. The EP dataset, derived (and sampled) from new edges collected between April 6, 2023, and May 25, 2024, includes 40,000 samples, with edges not present in the pre-existing CompanyKG (a snapshot up until April 5, 2023).</li> </ul> <p><strong>Background and Motivation</strong></p> <p>In the investment industry, it is often essential to identify similar companies for a variety of purposes, such as market/competitor mapping and Mergers & Acquisitions (M&A). Identifying comparable companies is a critical task, as it can inform investment decisions, help identify potential synergies, and reveal areas for growth and improvement. The accurate quantification of inter-company similarity, also referred to as <strong>company similarity quantification</strong>, is the cornerstone to successfully executing such tasks. However, company similarity quantification is often a challenging and time-consuming process, given the vast amount of data available on each company, and the complex and diversified relationships among them.</p> <p>While there is no universally agreed definition of company similarity, researchers and practitioners in PE industry have adopted various criteria to measure similarity, typically reflecting the companies' operations and relationships. These criteria can embody one or more dimensions such as industry sectors, employee profiles, keywords/tags, customers' review, financial performance, co-appearance in news, and so on. Investment professionals usually begin with a limited number of companies of interest (a.k.a. seed companies) and require an algorithmic approach to expand their search to a larger list of companies for potential investment. </p> <p>In recent years, transformer-based Language Models (LMs) have become the preferred method for encoding textual company descriptions into vector-space embeddings. Then companies that are similar to the seed companies can be searched in the embedding space using distance metrics like cosine similarity. The rapid advancements in Large LMs (LLMs), such as GPT-3/4 and LLaMA, have significantly enhanced the performance of general-purpose conversational models. These models, such as ChatGPT, can be employed to answer questions related to similar company discovery and quantification in a Q&A format.</p> <p>However, graph is still the most natural choice for representing and learning diverse company relations due to its ability to model complex relationships between a large number of entities. By representing companies as nodes and their relationships as edges, we can form a <strong>Knowledge Graph (KG)</strong>. Utilizing this KG allows us to efficiently capture and analyze the network structure of the business landscape. Moreover, KG-based approaches allow us to leverage powerful tools from network science, graph theory, and graph-based machine learning, such as Graph Neural Networks (GNNs), to extract insights and patterns to facilitate similar company analysis. While there are various company datasets (mostly commercial/proprietary and non-relational) and graph datasets available (mostly for single link/node/graph-level predictions), there is a scarcity of datasets and benchmarks that combine both to create a large-scale KG dataset expressing rich pairwise company relations.</p> <p><strong>Source Code and Tutorial:<br></strong><a href="https://github.com/llcresearch/CompanyKG2"><strong>https://github.com/llcresearch/CompanyKG2</strong></a></p> <p><strong>Paper: to be published<br></strong></p>
Results from Interpreting Cis-Regulatory Interactions from Large-Scale Deep Neural Networks for Genomics
Open the record for dataset details and reuse information.
Figure 3 in Lessons learnt from large-scale eradication of Australian swamp stonecrop Crassula helmsii in a protected Natura 2000 site
Figure 3. Photo impressions of the large scale eradication of Crassula helmsii on the Island of Terschelling. A. The dominant infestation of C. helmsii in an artificial lake (location 5). B. Installing a mitigation fence in order to prevent entry by amphibians and reptiles (2.75 linear km). C. Drainage installation and steel road plates tracks. D. Excavation of 40 cm topsoil (Step 4 of the described eradication approach). E. Recolonization of characteristic native plant species in excavated area 2.5 years after the eradication of C. helmsii (March 2021). F. Guided tours in the study area to inform those concerned.
Figure 1 in Lessons learnt from large-scale eradication of Australian swamp stonecrop Crassula helmsii in a protected Natura 2000 site
Figure 1. Locations of the dune slacks (1–5) where measures for eradication of Crassula helmsii were taken.
Figure 2 in Lessons learnt from large-scale eradication of Australian swamp stonecrop Crassula helmsii in a protected Natura 2000 site
Figure 2. Flow chart of the process for large scale eradication of Crassula helmsii on the Island of Terschelling.
Linking the microarchitecture of neurotransmitter systems to large-scale MEG resting state networks
<p>Information processing and communication in neuronal circuits is enabled by dynamic networks of inter-areal coupling of neuronal oscillations in which hubs play a central role for regulation of communication. Oscillations are shaped by interactions between pyramidal cells and interneurons and are locally influenced by neuromodulatory systems. Here, we set out to investigate how sparial variability in neurotransmitter receptor and transporter density influences frequency-specific large-scale networks of phase-synchrony (PS) and amplitude-correlation (AC) in human magnetoencephalography data. We found that node centrality - indexing which individual brain regions function as hubs - covaried positively with GABA, NMDA, dopaminergic, and most serotonergic receptor and transporter densities in lower frequency bands (delta to low-alpha for PS, and delta for AC) and in the gamma band, but negatively in between. These results establish how local microarchitecture influences large-scale connectivity networks of neuronal oscillations in the human brain in frequency- and spatially-specific patterns.</p>
A Large-Scale Sensitivity Analysis on Latent Embeddings and Dimensionality Reductions for Text Spatializations
<p>Result Files for the Paper "A Large-Scale Sensitivity Analysis on Latent Embeddings and Dimensionality Reductions for Text Spatializations" to be published at IEEE Vis 2024</p>
Database: Influence of large-scale freestream turbulence on bypass transition in air and organic vapour flows
<p>This is an open-access database of numerical simulations of freestream-turbulence induced transitions of zero-pressure-gradient flat plates boundary layers. The data have been collected at Arts et Metiers Institute of Technology, at DynFluid laboratory, in the period 2021-2024 within the REGAL-ORC project- REal GAs effects on Loss mechanisms of ORC turbine flows, a project in collaboration with FH Münster in Germany.</p> <p>Work related to the database is published in Journal of Fluid Mechanics, under the title: "Influence of large-scale freestream turbulence on bypass transition in air and organic vapour flows". The aims of this work was to investigate freestream-turbulence induced transitions, in particular under relatively large integral length scale, of air and organic vapors boundary layers.</p> <p>This version contains evolutions of averaged quantities in the streamwise direction, along with profiles and spectra in the boundary layer for all the different simulations performed. This database is supplemented with python codes to read the data and descriptions of the data in the headers and README files.</p>
Large-scale and fine-grained phenological stage annotation of herbarium specimens datasets
<p>This upload is constituted of four datasets of specimens from American herbaria covering different levels of information precision and different floras - from temperate to equatorial.</p> <p>Three of these datasets consist of selected specimens from herbaria located in different geographic and environmental regions. Each specimen of these three datasets was annotated with the following fields: family, genus, species name, fertile / non-fertile, presence / absence of flower(s), presence / absence of fruit(s). The resulting dataset was composed of 163,233 herbarium specimens belonging to 7,782 species, 1,906 genera, and 236 families. Specimens were annotated as “fertile” if any reproductive structures were present, such as sporangia (ferns), cones (gymnosperms), flowers, or fruits (angiosperms). Non-fertile specimens were those that lacked any reproductive structures.</p> <p>The fourth dataset consists of 20,371 herbarium specimens from 11 genera in the sunflower family (<em>Asteraceae</em>). The main difference in this dataset is that it is annotated with fine-grained phenophase scores rather than presence/absence attributes (see description below).</p> <p>Each of these datasets is described below:</p> <ul> <li> <p>NEVP: this dataset of New England vascular plant (NEVP) specimens was produced by members of the Consortium of Northeastern Herbaria. The dataset comprises 42,658 digitized specimens that belong to 1,375 species and come from several North American institutions. Most of the specimens in this dataset are from the north-temperate region of the northeastern United States.</p> </li> <li> <p>FSU: this dataset was produced by the Florida State University's Robert K. Godfrey Herbarium (FSU), a collection that focuses on northern Florida and the U.S. Southeast Coastal Plain, one of North America's biodiversity hotspots. This dataset contains 54,263 digitized herbarium specimen records that belong to 3,870 species, making it the taxonomically richest dataset in this study. Most species in this dataset grow under subtropical or warm temperate conditions in the southeastern region of the United States.</p> </li> <li> <p>CAY: this dataset comes from the IRD’s Herbarium of French Guiana (CAY). CAY is dedicated to the Guayana Shield flora, with a strong focus on tropical tree species. This dataset is composed of 66,312 herbarium specimens that belong to 3,024 species. All digitized specimens of this herbarium are accessible online. Most specimens were collected in the tropical rainforests of French Guiana, with the remaining specimens coming mostly from Suriname and Guyana.</p> </li> <li> <p>PHENO: this dataset includes 20,371 herbarium specimens of 139 species in the <em>Asteraceae</em> produced in a study of phenological trends in the U.S. Southeast Coastal Plain. The dataset is composed of specimen records from 57 herbaria. Each recorded specimen was annotated for quartile percentages (0, 25, 50, 75, or 100%) of (i) closed buds, (ii) buds transformed into flowers, and (iii) fruits. According to the distribution of these three categories for each specimen, a phenophase code was computed.</p> </li> </ul> <p> </p> <p><strong>Datasets format</strong></p> <p>These datasets are grouped in 3 tasks:</p> <ol> <li>fertility detection</li> <li>flowers and/or fruit detection</li> <li>phenophase classification</li> </ol> <p>The first 2 tasks are carried on the first 3 previous datasets and thus are based on the same set of images, unlike the third task which has its own disjoint set of images. This is why the dataset is presented into two separated files, one for each set of images.</p> <p><em>Fertility detection & flower/fruit detection</em></p> <p>These tasks are contained into the <em>herbarium_fertility_annotations.zip</em> archive. It consists of 3 files:</p> <ul> <li><em>metadata.csv</em>: general information about all the herbarium specimens for these tasks <ul> <li><em>id</em>: specimen identifier</li> <li><em>collection</em>: which of NEVP, FSU or CAY does the specimen come from</li> <li><em>herbarium</em>: institution of origin of the specimen, especially for NEVP collection</li> <li><em>clade</em>,<em> family</em>,<em> genus</em>,<em> species</em>: classification of the specimen</li> <li><em>URL</em>: URL of the scan</li> </ul> </li> <li><em>fertility_task.csv</em>: specific information regarding the fertility detection task <ul> <li><em>id</em>: specimen identifier</li> <li><em>is_fertile</em>: <em>True</em> if the specimen has an expression of fertility, <em>False</em> otherwise</li> <li><em>train_test_set</em>: which subset does the specimen belong to; possible values are: <em>train</em>, <em>random_test</em>, <em>species_test</em> and <em>herbarium_test</em></li> </ul> </li> <li><em>flower_fruit_task.csv</em>: specific information regarding the flower/fruit detection task <ul> <li><em>id</em>: specimen identifier, note that in this case not all the specimen described in <em>metadata.csv</em> are included in this task</li> <li><em>has_flower</em>: <em>True</em> if the specimen has at least one flower, <em>False</em> otherwise</li> <li><em>has_fruit</em>: <em>True</em> if the specimen has at least one fruit, <em>False</em> otherwise</li> <li><em>train_test_set</em>: which subset does the specimen belong to; possible values are: <em>train</em>, <em>random_test</em>, <em>species_test</em> and <em>herbarium_test</em></li> </ul> </li> </ul> <p><em>Phenophase classification</em></p> <p>These tasks are contained into the <em>herbarium_asteraceae_phenophase_annotations.zip</em> archive. It consists of a single file:</p> <ul> <li><em>annotations.csv</em>: <ul> <li><em>id</em>: specimen identifier</li> <li><em>URL</em>: URL of the scan</li> <li><em>genus</em>: genus of the specimen</li> <li><em>phenophase</em>: integer from 1 to 9 describing the phenophase of the specimen</li> <li><em>train_test_set</em>: which subset does the specimen belong to; possible values are: <em>train</em> and <em>test</em></li> </ul> </li> </ul> <p> </p> <p><strong>Additional ressources</strong></p> <p>More information can be found in the related paper:<br> <em>Lorieul, T., K. D. Pearson, E. R. Ellwood, H. Goëau, J.-F. Molino, P. W. Sweeney, J. M. Yost, J. Sachs, E. Mata-Montero, G. Nelson, P. S. Soltis, P. Bonnet, and A. Joly. 2019. Toward a large-scale and deep phenological stage annotation of herbarium specimens: Case studies from temperate, tropical, and equatorial floras. Applications in Plant Sciences 7(3): e1233.</em></p> <p>For an example of usage of these datasets as well as a baseline, see: <a href="http://doi.org/10.5281/zenodo.2549996">http://doi.org/10.5281/zenodo.2549996</a></p> <p> </p>
FIG. 3. — Microphotina viridula n in Les mantes (Dictyoptera, Mantodea) du massif du Mitaraka (Guyane), in Touroult J. (ed.), "Our Planet Reviewed" 2015 large-scale biotic survey in Mitaraka, French Guiana.
FIG. 3. — Microphotina viridula n. sp., holotype mâle: A, vue dorsale; B, vue ventrale de l'avant-corps; C, étiquettes. Échelles: A, 1 cm; B, 5 mm. Photos M. Depraetere.
FIG. 4. — Microphotina viridula n in Les mantes (Dictyoptera, Mantodea) du massif du Mitaraka (Guyane), in Touroult J. (ed.), "Our Planet Reviewed" 2015 large-scale biotic survey in Mitaraka, French Guiana.
FIG. 4. — Microphotina viridula n. sp.: A, plaque sous-génitale d'un paratype; B, genitalia de l'holotype en vue ventrale avec en plus grand l'apex de l'épiphallus gauche (D) et sa variabilité chez un paratype (C). Échelles: A, B, 1 mm; C, D, 0,5 mm.
Supplementary data to Schmiester et al. *Efficient parameterization of large-scale dynamic models based on relative measurements*
<p>This archive contains Supplementary data to the manuscript <em>Efficient parameterization of large-scale dynamic models based on relative measurements</em> by Leonard Schmiester, Yannik Schälte, Fabian Fröhlich, Jan Hasenauer and Daniel Weindl.</p>
DATASET of Large-scale Neural Recordings for DENOISING Engine
<p><span>30 sec raw data (.brw) was recorded with BrainWave SW and detected LFP events and spikes were stored in (.bxr). These extracellular recordings were obtained from acute hippocampal-cortical slices and were collected at 14KHz/electrode sampling frequency.</span></p>
HiDy: A Large-scale Hierarchical Dynamic Financial Knowledge Base
<p>HiDy is a hierarchical, dynamic, robust, diverse, and large-scale financial KB that aims to provide various valuable financial knowledge as critical benchmarking data for fair model testing in different financial tasks. Specifically, HiDy currently contains 34 relation types, more than 506,444 relations, 17 entity types, and more than 51,095 entities. The scale of HiDy is steadily growing due to its continuous updates. To make HiDy easily accessible and retrieved, HiDy is organized in a well-formed financial hierarchy with four branches, <em>Macro</em>, <em>Meso</em>,<em> Micro</em>, and<em> Others</em>.</p> <p>We then give explanations on the various csv files as follows.</p> <ul> <li>"hidy.nodes.entity_type.csv" includes a mapping dictionary of entities with a specific entity type. The meta data is ID, name, (code), Label. For example, "hidy.nodes.company.csv" includes "0,东诚药业,002675.SZ,company", "1,大庆华科,000985.SZ,company".</li> <li>"hidy.relationships.relation_type.csv" includes quadruple knowledge with a specific relation type. The meta data is START_ID, END_ID, TYPE, time. For example, "hidy.relationships.cooperate.csv" includes "2197,245,cooperate,2019/9/17 12:00, "1165,756,cooperate,2020/8/2 17:10".</li> </ul> <p>For details of HiDy, please refer to our <a href="https://github.com/K-Quant/HiDy">GitHub</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.