Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
66
datasets available to search
ShareScore release 0.9.0
Dataset results
66 results for “graph model”
Graph Data: Hydrological impact of widespread afforestation in Great Britain using a large ensemble of modelled scenarios
<p>Data used for creating the figures in the paper: Hydrological impact of widespread afforestation in Great Britain using a large ensemble of modelled scenarios.</p> <p>It contains the flow exceedances (as mm day<sup>-1</sup>), flow duration slope, median elasticity and runoff ratio for the different afforestation scenarios. Also included is the information on the changes of broadleaf afforestation. </p> <p>If you have any questions, please email marcus.buechel@ouce.ox.ac.uk.</p>
Node2Vec model - Czech Wikidata (knowledge graph / concepts / l80 / rw40)
<p>Node2Vec embedding model trained on Czech wikidata (from October 2020) concepts using gensim implementation of Word2Vec with the following parameters for random walks:</p> <ul> <li>length of walk = 80</li> <li>number of random walks = 40</li> </ul>
Node2Vec model - Czech Wikidata (knowledge graph / concepts / l40 / rw10)
<p>Node2Vec embedding model trained on Czech wikidata (from October 2020) concepts using gensim implementation of Word2Vec with the following parameters for random walks:</p> <ul> <li>length of walk = 40</li> <li>number of random walks = 10</li> </ul>
Node2Vec model - Czech Wikidata (knowledge graph / labels / l80 / rw40)
<p>Node2Vec embedding model trained on Czech wikidata (from October 2020) labels using gensim implementation of Word2Vec with the following parameters for random walks:</p> <ul> <li>length of walk = 80</li> <li>number of random walks = 40</li> </ul>
Node2Vec model - Czech Wikidata (knowledge graph / labels / l160 / rw40)
<p>Node2Vec embedding model trained on Czech wikidata (from October 2020) labels using gensim implementation of Word2Vec with the following parameters for random walks:</p> <ul> <li>length of walk = 160</li> <li>number of random walks = 40</li> </ul>
scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data
<p>This repository contains the training data and source code to reproduce the results of our paper:<br>scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data</p> <p>More description can be also found in GitHub (https://github.com/LPH-BIG/scGraph2Vec).</p>
Deep learning models predicting gene functions and pathways using public DRKG knowledge graph and graph neural network
<p>The attached dataset contains pretrained link prediction models, as described in our paper 'Morphological Map of Under- and Over-Expression of Genes in Human Cells'.</p>
Is this bug severe? A text-cum-graph based model for bug severity prediction
<p>A snapshot of the dataset has been updated. For the time being, we are publishing a snapshot of the dataset where the bugs were reported after 2017.</p> <p>Paper link: <a href="https://arxiv.org/abs/2207.00623">https://arxiv.org/abs/2207.00623</a> (ECML-PKDD 2022)</p> <p>Cite our paper:</p> <p>@InProceedings{10.1007/978-3-031-26422-1_15,<br> author="Hazra, Rima<br> and Dwivedi, Arpit<br> and Mukherjee, Animesh",<br> editor="Amini, Massih-Reza<br> and Canu, St{\'e}phane<br> and Fischer, Asja<br> and Guns, Tias<br> and Kralj Novak, Petra<br> and Tsoumakas, Grigorios",<br> title="Is This Bug Severe? A Text-Cum-Graph Based Model for Bug Severity Prediction",<br> booktitle="Machine Learning and Knowledge Discovery in Databases",<br> year="2023",<br> publisher="Springer Nature Switzerland",<br> address="Cham",<br> pages="236--252",<br> isbn="978-3-031-26422-1"<br> }</p> <p><strong>*** Please see the new version. (10.5281/zenodo.5554978)</strong></p> <p>There is a total of six files.</p> <ul> <li><strong>bug_descriptions.csv:</strong> This file contains the bug id and its description.</li> <li><strong>bug_comments.csv:</strong> This file contains three columns. The columns are the bug ids, comments and timestamp of the comment.</li> <li><strong>bug_REPORTED_ON_details.csv:</strong> This file contains the bug id and the package name on which the bug is reported</li> <li><strong>affect_dataset.csv: </strong>This file contains the bug id and the affected packages along with the affect timestamp.</li> <li><strong>bug_heat_2019.csv:</strong> This file contains the bug ids and its bug heats crawled in November 2019.</li> <li><strong>bug_heat_2020.csv:</strong> This file contains the bug ids and its bug heats crawled in November 2020.</li> </ul>
OC-782K: Knowledge Graph of "Scientometrics" modelled according to the OpenCitations Data Model
<p>This dataset is a knowledge graph extracted from a <a href="https://static.aminer.cn/misc/na-data-kdd18.zip">t</a>riplestore covering information about the journal <em>Scientometrics</em> and modelled according to the OpenCitations Data Model. The original triplestore is available <a href="https://doi.org/10.5281/zenodo.5151264">here</a>. This KG was extracted for a research project on knowledge graph embeddings (KGEs) for author disambiguation. Structural triples of the knowledge graph are split into training, testing and validation for applying representation learning methods. Textual literals and numeric literals were stored separately in order to implement multimodal approaches for KGEs (see <a href="https://arxiv.org/abs/1802.00934">arXiv:1802.00934</a>). For the same reason, textual literals and numeric literals are already stored into sentence embeddings and a numeric matrix respectively in the files <em>textual_literals.npy </em>and <em>numeric_literals.npy</em>. The file <em>and_eval</em><em>.json </em>contains the evaluation dataset used for evaluating our AND architecture. For the script used to gather this dataset see the GitHub repository: <a href="https://github.com/sntcristian/and-kge/tree/main/aminer">https://github.com/sntcristian/and-kge/tree/main/open-citations</a>.</p>
Grid-graph modeling of emergent neuromorphic dynamics and heterosynaptic plasticity in memristive nanonetworks - Dataset
<p>This is the dataset of "Grid-graph modeling of emergent neuromorphic dynamics and heterosynaptic plasticity in memristive nanonetworks"</p>
A hierarchical graph-based model for mobility data representation and analysis
<p>Hierarchical representations of transportation networks should provide a better understanding of mobility patterns and the underlying structures at various abstraction levels. A hierarchical graph-based model allows representing moving objects and trajectories according to multiple spatial, temporal and semantic scales. The latter model is implemented here in a Neo4j graph database (version 4.4.0) and experimented with historical maritime data covering Brittany Bay in France.</p>
Figure 1: The ECG model-MAPPING BETWEEN SEMANTIC GRAPHS AND SENTENCES IN GRAMMAR INDUCTION SYSTEM
<p>The following Figure 1 shows a sample semantic graph that describes a<br> simple test world.<br> During the processing of the ECG, the base units of the graph are the ECG<br> atoms. An ECG atom corresponds to a primitive statements related to one<br> predicate. It has a structure of one-level deep tree, where the root of the tree<br> is the predicate and the concepts linked to it are the leaves. The child concept<br> of the root predicate may be not only a single concept but it can be another<br> ECG atom.</p>
Dataset and Model Weights for Plasma Sheet Model Graph Network Simulator
<p>This repository contains the simulation data and pre-trained Graph Neural Network (GNN) models produced in [1].</p> <p>Two *.zip files are provided:</p> <ul> <li>data.zip - contains the datasets of train/test simulations produced using the Sheet Model algorithm [1, 2]</li> <li>models.zip - contains the GNN model weights (<em>*.</em>pkl<em>) </em>+ relevant training information and model parameters <em>(*.</em>yml<em> and *</em>.txt)</li> </ul> <p>Dataset subfolders are named according to dataset/{'train' or 'test'}/{number of sheets}/{boundary condition}/. Each subfolder contains multiple simulations and a single info.yml file with relevant information regarding the overall setup. For each i-th simulation the following files are provided:</p> <ul> <li>x_{i}.npy - array with sheet trajectories (#time-steps, #sheets)</li> <li>v_{i}.npy - array with sheet velocities (#time-steps, #sheets)</li> <li>x_eq_{i}.npy - array with sheet equilibrium positions (#time-steps, #sheets)</li> </ul> <p> Model sub-folders are named according to :</p> <ul> <li>models/{time step}/{seed} - default architecture (preferred)</li> <li>models/{time step}/{'collisions', 'nosent' or 'equivariant'}/{seed} - alternative (less performing) architectures mentioned in the paper appendices.</li> </ul> <p>For each model we provide:</p> <ul> <li>params_best.pkl - model weights that performed the best during training on the validation set</li> <li>params_final.pkl - model weights at the end of training</li> <li>model_cfg.yml - GNN architecture metadata</li> <li>train_cfg.yml - training configuration metadata</li> <li>train_data.yml - training dataset metadata</li> <li>loss.txt - training and validation loss per epoch</li> <li>loss_i.txt - training loss per gradient update step</li> </ul> <h3>Source Code</h3> <p>The source code used to produce the data, train, and test the models can be found at: <a href="https://github.com/diogodcarvalho/gns-sheet-model">https://github.com/diogodcarvalho/gns-sheet-model</a></p> <h3>References</h3> <p>[1] D. D. Carvalho, D. R. Ferreira, L. O. Silva, "Learning the dynamics of a one-dimensional plasma model with graph neural networks<em>", Mach. Learn.: Sci. Technol. 5 025048 </em>(2024)</p> <p>[2] J. Dawson, "One‐Dimensional Plasma Model"<em>, The Physics of Fluids</em> 5.4 (1962): 445-459.</p> <p> </p>
Data and models for: Learning Ordering in Crystalline Materials with Symmetry-Aware Graph Neural Networks
<p>Data (ver 1.1) and trained models for our paper "<a href="https://arxiv.org/abs/2409.13851">Learning Ordering in Crystalline Materials with Symmetry-Aware Graph Neural Networks</a>". If you use such data or models, please cite our paper. These three directories need to be downloaded and copied into our source codes in order to reproduce our paper: <a href="https://github.com/learningmatter-mit/PerovskiteOrderingGCNNs">https://github.com/learningmatter-mit/PerovskiteOrderingGCNNs</a></p> <ul> <li>data: All data files for training and evaluating GCNNs, with a copy archived on the Materials Data Facility (<a href="https://doi.org/10.18126/ncqt-rh18">DOI: 10.18126/ncqt-rh18</a>)</li> <li>saved_models: All saved model files for evaluating GCNNs</li> <li>best_models: All best model files for evaluating GCNNs</li> </ul>
Data from: Predicting primate-parasite associations with exponential random graph models
<p>Ecological associations between hosts and parasites are influenced by host exposure and susceptibility to parasites, and by parasite traits, such as transmission mode. Advances in network analysis allow us to answer questions about the causes and consequences of traits in ecological networks in ways that could not be addressed in the past.</p> <p>We used a network-based framework (exponential random graph models, or ERGMs) to investigate the biogeographic, phylogenetic, and ecological characteristics of hosts and parasites characteristics that affect the probability of interactions among nonhuman primates and their parasites. Parasites included arthropods, bacteria, fungi, protozoa, viruses, and helminths.</p> <p>We investigated existing hypotheses, along with new predictors and an expanded host-parasite database that included 213 primate nodes, 763 parasite nodes, and 2,319 edges among them. Analyses also investigated phylogenetic relatedness, sampling effort, and spatial overlap among hosts.</p> <p>In addition to supporting some previous findings, our ERGM approach demonstrated that more threatened hosts had fewer parasites, and notably, that this effect was independent of threatened hosts also having a smaller geographic range. Despite having fewer parasites, threatened host species shared more parasites with other hosts, consistent with the loss of specialist parasites and threats arising from generalist parasites that can be maintained in other, non-threatened hosts. Viruses, protozoa, and helminths had broader host ranges than bacteria or fungi, and parasites that infect non-primates had a higher probability of infecting more primate species.</p> <p>The value of the ERGM approach for investigating the processes structuring host-parasite networks provided a more complete view of the biogeographic, phylogenetic, and ecological traits that influence parasite species richness and parasite sharing among hosts. The results supported some previous analyses and revealed new associations that warrant future research, thus revealing how hosts and parasites interact to form ecological networks.</p>
Dataset with the node discretisations employed for training advection models in "Multi-scale rotation-equivariant graph neural networks for unsteady Eulerian fluid dynamics"
<p>Dataset with the node discretisations employed for training advection models in "Multi-scale rotation-equivariant graph neural networks for unsteady Eulerian fluid dynamics" (https://doi.org/10.1063/5.0097679).</p> <p>The training code is available at https://github.com/mario-linov/graphs4cfd.</p>
Results and log of LLM-KG-Bench runs described in article "Developing a Scalable Benchmark for Assessing Large Language Models in Knowledge Graph Engineering", Meyer et al. 2023
<p>Results and logs of <a href="https://github.com/AKSW/LLM-KG-Bench">LLM-KG-Bench</a> runs described in article "Developing a Scalable Benchmark for Assessing Large Language Models in Knowledge Graph Engineering", Meyer et al., to appear in <a href="https://2023-eu.semantics.cc/page/accepted_posters">SEMANTICS 2023 poster track</a> proceedings.</p>
Surface Charge Boundary Condition Often Misused in CO2 Reduction Models (raw data for graphs)
<p>This contains the raw data, in text file format, for the figures in the journal article "Surface Charge Boundary Condition Often Misused in CO2 Reduction Models" published in Journal of Physical Chemistry C, DOI 10.1021/acs.jpcc.3c05364.</p> <p>The files are in directories corresponding to their figure number, and the files are named with the capacitance. </p> <p> </p> <p> </p>
Data from: Predicting primate-parasite associations with exponential random graph models
Open the record for dataset details and reuse information.
Feature attention graph neural network for estimating brain age and identifying important neural connections in mouse models of genetic risk for Alzheimer's disease
<p>Connectome, traits and behavior data for APOE234 mice.</p> <ul> <li>1. connectome.zip: mouse brain structural connectivity matrices from diffusion MRI.</li> <li>2. FAGNN_Phenotype.csv: a sheet of trait information of mice used in the study.</li> </ul> <p>columns: winding numbers, total distance, normalized NE time, normalized NE distance, normalized NW time, normalized NW distance, normalized SE time, normalized SE distance, normlaized SW time, normalized SW distance, island latency to first entry, island entries, normalized thigmataxis time, and normalized thigmotaxis distance</p> <div>rows: 4 trials for each day from day 1 to day 5 with 1 probing test each at day 5 and day 8</div> <ul> <li>3. mouse_anatomy.csv: brain region information regarding the connectivity matrix.</li> <li>4. behavior.zip: behavioral data for each mouse from Morris Water Maze experiments.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.