Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

43

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

43 results for “graph learning”

Learn how ShareScore rates datasets ↗
zenodo48/100

Challenges in Migrating Imperative Deep Learning Programs to Graph Execution: An Empirical Study

<p>Efficiency is essential to support responsiveness w.r.t. ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code that supports symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development tends to produce DL code that is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, less error-prone imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. While hybrid approaches aim for the &quot;best of both worlds,&quot; the challenges in applying them in the real world are largely unknown. We conduct a data-driven analysis of challenges&mdash;and resultant bugs&mdash;involved in writing reliable yet performant imperative DL code by studying 250 open-source projects, consisting of 19.7 MLOC, along with 470 and 446 manually examined code patches and bug reports, respectively. The results indicate that hybridization: (i) is prone to API misuse, (ii) can result in performance degradation&mdash;the opposite of its intention, and (iii) has limited application due to execution mode incompatibility. We put forth several recommendations, best practices, and anti-patterns for effectively hybridizing imperative DL code, potentially benefiting DL practitioners, API designers, tool developers, and educators.</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph Execution

<p>Efficiency is essential to support ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code---supporting symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. Though hybrid approaches aim for the "best of both worlds," using them effectively requires subtle considerations. Our key insight is that, while DL programs typically execute sequentially, hybridizing imperative DL code resembles parallelizing sequential code in traditional systems. Inspired by this, we present an automated refactoring approach that assists developers in determining which otherwise eagerly-executed imperative DL functions could be effectively and efficiently executed as graphs. The approach features novel static imperative tensor and side-effect analyses for Python. Due to its inherent dynamism, analyzing Python may be unsound; however, the conservative approach leverages a speculative (keyword-based) analysis for resolving difficult cases that informs developers of any assumptions made. The approach is: (i) implemented as a plug-in to the PyDev Eclipse IDE that integrates the WALA Ariadne analysis framework and (ii) evaluated on nineteen DL projects consisting of 132 KLOC. The results show that 326 of 766 candidate functions (42.56%) were refactorable, and an average relative speedup of 2.16x on performance tests was observed with negligible differences in model accuracy. The results indicate that the approach is useful in optimizing imperative DL code to its full potential.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Deep learning models predicting gene functions and pathways using public DRKG knowledge graph and graph neural network

<p>The attached dataset contains pretrained link prediction models, as described in our paper 'Morphological Map of Under- and Over-Expression of Genes in Human Cells'.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

CNN Wild Park - Graph Neural Networks for Learning Equivariant Representations of Neural Networks

<p>This repository contains the <strong>CNN Wild Park</strong> dataset from the paper:</p> <blockquote> <p><strong>Graph Neural Networks for Learning Equivariant Representations of Neural Networks</strong><br><a href="https://mkofinas.github.io/">Miltiadis Kofinas</a>*,&nbsp;<a href="https://bknyaz.github.io/">Boris Knyazev</a>, <a href="https://www.cyanogenoid.com/">Yan Zhang</a>,&nbsp;<a href="https://yunlu-chen.github.io/">Yunlu Chen</a>,&nbsp;<a href="https://gertjanburghouts.github.io/">Gertjan J. Burghouts</a>,&nbsp;<a href="https://egavves.com/">Efstratios Gavves</a>,&nbsp;<a href="https://www.ceessnoek.info/">Cees G. M. Snoek</a>,&nbsp;<a href="https://davzha.netlify.app/">David W. Zhang</a>*<br><em>ICLR 2024</em> (oral)<br><a href="https://arxiv.org/abs/2403.12143">https://arxiv.org/abs/2403.12143</a><br><a href="https://github.com/mkofinas/neural-graphs">https://github.com/mkofinas/neural-graphs</a><br>*Joint first and last authors</p> </blockquote> <p>We introduce a new dataset of CNNs, which we term <em>CNN Wild Park</em>.<br>The dataset consists of 117,241 checkpoints from 2,800 CNNs, trained for up to 1,000 epochs on CIFAR10.<br>The CNNs vary in the number of layers, kernel sizes, activation functions, and residual connections between arbitrary layers.</p> <p>More specifically, we construct the CNN Wild Park dataset by training 2,800 small CNNs with different architectures for 200 to 1,000 epochs on CIFAR10. We retain a checkpoint of its parameters every 10 steps and also record the test accuracy. The CNNs vary by:</p> <ul> <li>Number of layers L in [2, 3, 4, 5] (note that this does not count the input layer).</li> <li>Number of channels per layer c_l in [4, 8, 16, 32].</li> <li>Kernel size of each convolution k_l in [3, 5, 7].</li> <li>Activation functions at each layer are one of ReLU, GeLU, tanh, sigmoid, leaky ReLU, or the identity function.</li> <li>Skip connections between two layers with at least one layer in between. Each layer can have at most one incoming skip connection. We allow for skip connections even in the case when the number of channels differ, to increase the variety of architectures and ensure independence between different architectural choices. We enable this by adding the skip connection only to the min(c_n, c_m) nodes.</li> </ul> <p>We divide the dataset into train/val/test splits such that checkpoints from the same run are <strong>not</strong> contained in both the train and test splits.&nbsp;</p> <div>&nbsp;</div> <div>&nbsp;</div>

opencc-by-4.0May 2024View details →
zenodo40/100

PGB: A PubMed Graph Benchmark for Heterogeneous Network Representation Learning

<p>PubMed Graph Benchmark (PGB)&nbsp;aggregates&nbsp;the metadata associated with the biomedical articles from PubMed into a unified source.&nbsp;The benchmark contains metadata including&nbsp;title, abstract, authors, in/out citations, MeSH terms, MeSH hierarchy, venue, publication type, and chemicals.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

BRAIN Journal-Automatic Anthropometric System Development Using Machine Learning-Figure 3. (3.a) – The flowchart of Graph cuts method; (3.b)- the result of Graph cuts image segmentation.

<p>Figure 3 describes the steps implemented Graph cuts algorithm for the segmentation of human body parts. The results obtained are 5 main sections that include the hands, the legs, the center of the body (chest, waist, hips), and the head. The result of the display image is taken from the human image database, which was collected by us (Нгуен, 2016).&nbsp;</p>

opencc-by-4.0Aug 2016View details →
zenodo40/100

BRAIN Journal-Auto-generative Learning Objects in Online Assessment of Data Structures Disciplines-Figure 3. Symbols definition for a graph-based test

<p>In this scenario, we intend to generate a random graph and compute a deep first-search node list. The first defined random symbol is n, namely the number of nodes in the graph as an integer from 5 to 9. The next symbol is named g and denotes the graph object created randomly using 3 parameters: the number of nodes, the minimum, and the maximum value for the weight. For the number of nodes, we used the previously computed value of n, whereas for the weights, we used two constants 0 and 1 since the graph is not weighted</p>

opencc-by-4.0Sep 2017View details →
zenodo40/100

Data and models for: Learning Ordering in Crystalline Materials with Symmetry-Aware Graph Neural Networks

<p>Data (ver 1.1) and trained models for our paper "<a href="https://arxiv.org/abs/2409.13851">Learning Ordering in Crystalline Materials with Symmetry-Aware Graph Neural Networks</a>". If you use such data or models, please cite our paper. These three directories need to be downloaded and copied into our source codes in order to reproduce our paper:&nbsp;<a href="https://github.com/learningmatter-mit/PerovskiteOrderingGCNNs">https://github.com/learningmatter-mit/PerovskiteOrderingGCNNs</a></p> <ul> <li>data: All data files for training and evaluating GCNNs, with a copy archived on the Materials Data Facility (<a href="https://doi.org/10.18126/ncqt-rh18">DOI: 10.18126/ncqt-rh18</a>)</li> <li>saved_models: All saved model files for evaluating GCNNs</li> <li>best_models: All best model files for evaluating GCNNs</li> </ul>

opencc-by-4.0Sep 2024View details →
zenodo40/100

MTL-QA : A dataset and multi-task learning approach for knowledge graph and natural language question answering

<p>The dataset used for this project is created by enhancing the publicly available MetaQA (Movie Text Audio QA), which is primarily a KGQA dataset pertaining to movies, an extension of WikiMovies. This involves questions requiring 1, 2, and 3 hops which can be answered by using a MetaQA Knowledge Graph. The questions are available in text and audio format. The text has vanilla (original) and its paraphrased version, and is called ntm.&nbsp;</p> <p>In order to develop a dataset to support NLQA, a series of dataset augmentation steps has been performed.</p> <p>The dataset consists of natural language questions and a tagged topic entity as ground truth. This topic entity is used to retrieve textual information related to the question from Wikipedia. The introduction section of the entity&#39;s page is used as the context that is required for NLQA. Hence, this dataset has information related to both KGQA and NLQA. Certain preliminary checks and validations are done to only retain those data samples whose context can be used to answer a given question.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Evaluating the "Learning on Graphs" Conference Experience

<p>This is the reproducibility material for the following manuscript:</p> <p>Bastian Rieck and Corinna Coupette.&nbsp;<em>Evaluating the &quot;Learning on Graphs&quot; Conference Experience.</em> 2023.&nbsp;arXiv:&nbsp;<a href="https://doi.org/10.48550/arXiv.2306.00586">2306.00586 [cs.LG]</a>.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Lifelong Learning of Graph Neural Networks for Open-World Node Classification

<p>Three temporal graph datasets for node classification under distribution shift.</p> <p>DBLP-Easy and DBLP-Hard are citation graph datasets. PharmaBio is a collaboration graph dataset.</p> <p>Vertices are scientific publications, edges are either citations (DBLP) or at-least-one-common-author relationships (PharmaBio).</p> <p>The task is to classify the vertices of the graph into the respective conference/journal venues (DBLP) or journal categories (PharmaBio). In the DBLP datasets, new classes may appear over time.</p> <p>Each dataset follows the structure:</p> <p>- adjlist.txt -- the graph structure encoded as adjacency lists: in each row, the first entry is the source vertex, the remaining entries are adjacent vertices</p> <p>- X.npy -- numpy serialized format for node features indexed by node id corresponding to adjlist.txt</p> <p>- y.npy -- numpy serialized format for node labels indexed by node id corresponding to adjlist.txt</p> <p>- t.npy -- numpy serialized format for time steps indexed by node id corresponding to adjlist.txt</p> <p>A paper describing our incremental training and evaluation framework is published in IJCNN 2021 (Pre-print on arXiv:&nbsp;<a href="https://arxiv.org/abs/2006.14422">https://arxiv.org/abs/2006.14422</a>).</p> <p>If you use these datasets in your research, please cite the corresponding paper:</p> <pre><code>@inproceedings{galke2021lifelong, author={Galke, Lukas and Franke, Benedikt and Zielke, Tobias and Scherp, Ansgar}, booktitle={2021 International Joint Conference on Neural Networks (IJCNN)}, title={Lifelong Learning of Graph Neural Networks for Open-World Node Classification}, year={2021}, volume={}, number={}, pages={1-8}, doi={10.1109/IJCNN52387.2021.9533412} }</code></pre> <p><br> &nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

A supervised Graph-based deep learning algorithm to detect and quantify clustered particles

<p>In this data repository, we provide the necessary data for replicating results, including both simulated and biological datasets. Additionally, the repository includes trained models to infer from these datasets.</p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

TensorBoard runs for reinforcement learning on automated conjectured bounds on Laplacian spectral radius of graphs

<p>The zip archive contains separate folders with TensorBoard event files for the runs of our reinforcement learning implementation on the conjectured upper bounds for the Laplacian spectral radius, which are listed in Appendix B of the forthcoming paper: S. Al-Yakoob, M. Ghebleh, A. Kanso, D. Stevanović,&nbsp;Reinforcement learning for graph theory, I. Reimplementation of Wagner&rsquo;s approach, Art Discrete Appl. Math. (2024).</p> <p>To view the contained graphs and the evolution of rewards, unzip the archive in a folder of your choice, and run in the terminal the command&nbsp;"tensorboard --logdir runs" from the parent folder of the unzipped "runs" folder.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Source data for graphs and charts used in the paper "A machine learning-based estimator for real-time earthquake ground-shaking predictions in Southern California"

<h1>DESCRIPTION:</h1> <h3>This repository contains the source data for graphs and charts used in the paper "A machine learning-based estimator for real-time earthquake ground-shaking predictions in Southern California" submitted and accepted at "Communications Earth &amp; Environment journal"&nbsp;</h3> <h3>Marisol Monterrubio-Velasco, Scott Callaghan, David Modesto, Jose Carlos Carrasco, Rosa M. Badi , Pablo Pallares, Fernando V&aacute;zquez-Novoa, Enrique S. Quintana-Ortı́, Marta Pienkowska, and Josep de la Puente</h3> <h2><strong>DATA FOR FIGURES:&nbsp;</strong></h2> <h3><strong>Figure 1 :&nbsp;</strong></h3> <p>The data used in this figure comes from the CyberShake Study 15.4.&nbsp; Seismogram, intensity measure, and duration data from CyberShake Study 15.4 is available through the SCEC CyberShake Study 15.4 Globus Collection, served by the University of Southern California's Center for Advanced Research Computing.&nbsp; Direct link:<a href="https://g-46eaba.a78b8.36fe.data.globus.org/ACTN/3886/PeakVals_ACTN_10_0.bsa">https://g-46eaba.a78b8.36fe.data.globus.org</a>."</p> <h3><strong>Figure 2:</strong></h3> <p><strong>Model evaluation on the validation dataset for T = 2s</strong></p> <p>1. Random Forest predictions using the optimized hyperparameters depth=30, n_estimators=30 for the validation dataset at T=2s</p> <p><a href="../records/10640493/files/y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10.dat?download=1&amp;preview=1">y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10.dat</a></p> <p>2. Artificial Neural Network predictions using the optimized hyperparameters 9layer and 256 neurons for the validation dataset at T=2s</p> <p><a href="../api/records/10640493/draft/files/Prediction_NN_BS_256_9capas_DataSet_AllRuptureVariations_OneRuptureID_Period2.0_ALL_Validation_log10.csv/content" target="_blank" rel="noopener noreferrer">Prediction_NN_BS_256_9capas_DataSet_AllRuptureVariations_OneRuptureID_Period2.0_ALL_Validation_log10.csv</a></p> <p>3. True values for the validation dataset at T=2s</p> <p><a href="../records/10640493/files/y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10.dat?download=1&amp;preview=1">y_true_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10.dat</a></p> <h3><strong>Figure 3:</strong></h3> <p><strong>Error metrics obtained for each simulated scenario using:</strong></p> <p><strong>- Artificial Neural Networks</strong></p> <p><a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_2.0_NN.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_3.0_NN.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_5.0_NN.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_10.0_NN.csv</a></p> <p><strong>- Random Forest regressor</strong></p> <p><a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_2.0_RF.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_3.0_RF.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_5.0_RF.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_10.0_RF.csv</a></p> <p><strong>- ASK14 GMPE</strong></p> <p><a href="../api/records/10640493/draft/files/Results_metrics_paper_3.0_GMPE.csv/content" target="_blank" rel="noopener noreferrer">Results_metrics_paper_2.0_GMPE.csv</a>, <a href="../api/records/10640493/draft/files/Results_metrics_paper_3.0_GMPE.csv/content" target="_blank" rel="noopener noreferrer">Results_metrics_paper_3.0_GMPE.csv</a>,&nbsp;<a href="../api/records/10640493/draft/files/Results_metrics_paper_3.0_GMPE.csv/content" target="_blank" rel="noopener noreferrer">Results_metrics_paper_5.0_GMPE.csv, </a><a href="../api/records/10640493/draft/files/Results_metrics_paper_3.0_GMPE.csv/content" target="_blank" rel="noopener noreferrer">Results_metrics_paper_10.0_GMPE.csv</a></p> <h3>Figure 4:</h3> <p><strong>MLESmap RotD50 predictions on a validation event of magnitude 6.85</strong></p> <p><a href="../api/records/10640493/draft/files/RF_predictions_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">RF_predictions_T2s_map_2748.csv,&nbsp; </a><a href="../api/records/10640493/draft/files/ASK_14_prediction_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">ASK_14_prediction_T2s_map_2748.csv </a>,&nbsp;<a href="../api/records/10640493/draft/files/ANN_predictions_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">ANN_predictions_T2s_map_2748.csv</a></p> <p><a href="../api/records/10640493/draft/files/RF_predictions_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">RF_predictions_T3s_map_2748.csv,&nbsp; </a><a href="../api/records/10640493/draft/files/ASK_14_prediction_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">ASK_14_prediction_T3s_map_2748.csv </a>,&nbsp;<a href="../api/records/10640493/draft/files/ANN_predictions_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">ANN_predictions_T3s_map_2748.csv</a></p> <p><a href="../api/records/10640493/draft/files/RF_predictions_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">RF_predictions_T5s_map_2748.csv,&nbsp; </a><a href="../api/records/10640493/draft/files/ASK_14_prediction_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">ASK_14_prediction_T5s_map_2748.csv </a>,&nbsp;<a href="../api/records/10640493/draft/files/ANN_predictions_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">ANN_predictions_T5s_map_2748.csv</a></p> <p><a href="../api/records/10640493/draft/files/RF_predictions_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">RF_predictions_T10s_map_2748.csv,&nbsp; </a><a href="../api/records/10640493/draft/files/ASK_14_prediction_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">ASK_14_prediction_T10s_map_2748.csv&nbsp;</a>,&nbsp;<a href="../api/records/10640493/draft/files/ANN_predictions_T10s_map_2748.csv/content" target="_blank" rel="noopener noreferrer">ANN_predictions_T10s_map_2748.csv</a></p> <h3>Figure 5:</h3> <p><strong>Spatial configuration of five historical earthquakes and BBP stations also including the coordinates of synthetic stations&nbsp; from the CS_15_4 study</strong></p> <p><a href="../api/records/10640493/draft/files/SyntheticStationsCoordinates_CS_15.4.csv/content" target="_blank" rel="noopener noreferrer">SyntheticStationsCoordinates_CS_15.4.csv, </a><a href="../api/records/10640493/draft/files/Whittier_BBP_sites.csv/content" target="_blank" rel="noopener noreferrer">Whittier_BBP_sites.csv</a>, <a href="../api/records/10640493/draft/files/Northridge_BBP_sites.csv/content" target="_blank" rel="noopener noreferrer">Northridge_BBP_sites.csv</a>, <a href="../api/records/10640493/draft/files/North_Palm_Springs_BBP_sites.csv/content" target="_blank" rel="noopener noreferrer">North_Palm_Springs_BBP_sites.csv</a>,&nbsp;<a href="../api/records/10640493/draft/files/Landers_BBP_sites.csv/content" target="_blank" rel="noopener noreferrer">Landers_BBP_sites.csv</a>,&nbsp;<a href="../api/records/10640493/draft/files/Hector_Mine_BBP_sites.csv/content" target="_blank" rel="noopener noreferrer">Hector_Mine_BBP_sites.csv</a></p> <h3><strong>Figure 6:</strong></h3> <p><strong>RotD50 predictions for real events for the &lsquo;inside&rsquo; stations&nbsp;</strong></p> <p>NN_layers9North_Palm_Springs_IN_event_metrics-T_10.csv, NN_layers9Northridge_IN_event_metrics-T_10.csv, NN_layers9Landers_IN_event_metrics-T_10.csv, NN_layers9Hector_Mine_IN_event_metrics-T_10.csv, <a href="../api/records/10640493/draft/files/NN_layers9Whittier_IN_event_metrics-T_2.csv/content" target="_blank" rel="noopener noreferrer">NN_layers9Whittier_IN_event_metrics-T_10.csv</a></p> <p>NN_layers9North_Palm_Springs_IN_event_metrics-T_5.csv, NN_layers9Northridge_IN_event_metrics-T_5.csv, NN_layers9Landers_IN_event_metrics-T_5.csv, NN_layers9Hector_Mine_IN_event_metrics-T_5.csv,&nbsp;<a href="../api/records/10640493/draft/files/NN_layers9Whittier_IN_event_metrics-T_2.csv/content" target="_blank" rel="noopener noreferrer">NN_layers9Whittier_IN_event_metrics-T_5.csv</a></p> <p>NN_layers9North_Palm_Springs_IN_event_metrics-T_3.csv, NN_layers9Northridge_IN_event_metrics-T_3.csv, NN_layers9Landers_IN_event_metrics-T_3.csv, NN_layers9Hector_Mine_IN_event_metrics-T_3.csv, <a href="../api/records/10640493/draft/files/NN_layers9Whittier_IN_event_metrics-T_2.csv/content" target="_blank" rel="noopener noreferrer">NN_layers9Whittier_IN_event_metrics-T_3.csv</a></p> <p>NN_layers9North_Palm_Springs_IN_event_metrics-T_2.csv, NN_layers9Northridge_IN_event_metrics-T_2.csv, NN_layers9Landers_IN_event_metrics-T_2.csv, NN_layers9Hector_Mine_IN_event_metrics-T_2.csv, <a href="../api/records/10640493/draft/files/NN_layers9Whittier_IN_event_metrics-T_2.csv/content" target="_blank" rel="noopener noreferrer">NN_layers9Whittier_IN_event_metrics-T_2.csv</a></p> <h2>Supplementary material:</h2> <h3>Supplementary Fig 2</h3> <p><strong>Boxplots comparing ML models and "true" values</strong></p> <p><strong>-&nbsp; DNN</strong></p> <p><a href="10640493" target="_blank" rel="noopener noreferrer">Prediction_NN_BS_256_9capas_DataSet_AllRuptureVariations_OneRuptureID_Period2.0_ALL_Validation_log10.csv</a></p> <p><a href="10640493" target="_blank" rel="noopener noreferrer">Prediction_NN_BS_256_9capas_DataSet_AllRuptureVariations_OneRuptureID_Period3.0_ALL_Validation_log10.csv</a></p> <p><a href="10640493" target="_blank" rel="noopener noreferrer">Prediction_NN_BS_256_9capas_DataSet_AllRuptureVariations_OneRuptureID_Period5.0_ALL_Validation_log10.csv</a></p> <p><a href="10640493" target="_blank" rel="noopener noreferrer">Prediction_NN_BS_256_9capas_DataSet_AllRuptureVariations_OneRuptureID_Period10.0_ALL_Validation_log10.csv</a></p> <p>- RF</p> <p><a href="../api/records/10640493/draft/files/y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat/content" target="_blank" rel="noopener noreferrer">y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat</a></p> <p><a href="../api/records/10640493/draft/files/y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat/content" target="_blank" rel="noopener noreferrer">y_pred_dislib_T3s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat</a></p> <p><a href="../api/records/10640493/draft/files/y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat/content" target="_blank" rel="noopener noreferrer">y_pred_dislib_T5s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat</a></p> <p><a href="../api/records/10640493/draft/files/y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat/content" target="_blank" rel="noopener noreferrer">y</a><a href="../api/records/10640493/draft/files/y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat/content" target="_blank" rel="noopener noreferrer">_pred_dislib_T10s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat</a></p> <p>- TRUE VALUES FROM CYBERSHAKE SIMULATIONS</p> <p><a href="../api/records/10640493/draft/files/y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat/content" target="_blank" rel="noopener noreferrer">y_true_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat</a></p> <p><a href="../api/records/10640493/draft/files/y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat/content" target="_blank" rel="noopener noreferrer">y_true_dislib_T3s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat</a></p> <p><a href="../api/records/10640493/draft/files/y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat/content" target="_blank" rel="noopener noreferrer">y_true_dislib_T5s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat</a></p> <p><a href="../api/records/10640493/draft/files/y_pred_dislib_T2s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat/content" target="_blank" rel="noopener noreferrer">y_true_dislib_T10s_depth-30_n_estimators-30_val_MODEL_ORIGINAL_Dataset_ORIGINAL_ALL_log10_8Feat_Plus_RealData.dat</a></p> <p>&nbsp;</p> <h3>Supplementary Fig 3</h3> <p><strong>Error metrics obtained for each simulated scenario:</strong></p> <p><strong>Artificial Neural Networks:</strong></p> <p><a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_2.0_NN.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_3.0_NN.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_5.0_NN.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_10.0_NN.csv</a></p> <p><strong>&nbsp;Random Forest regressor:</strong></p> <p><a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_2.0_RF.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_3.0_RF.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_5.0_RF.csv</a>, <a href="../api/records/10640493/draft/files/Test_Results_metrics_paper_5.0_NN.csv/content" target="_blank" rel="noopener noreferrer">Test_Results_metrics_paper_10.0_RF.csv</a></p> <p><strong>ASK14 GMPE</strong></p> <p><a href="../api/records/10640493/draft/files/Results_metrics_paper_3.0_GMPE.csv/content" target="_blank" rel="noopener noreferrer">Results_metrics_paper_2.0_GMPE.csv</a>, <a href="../api/records/10640493/draft/files/Results_metrics_paper_3.0_GMPE.csv/content" target="_blank" rel="noopener noreferrer">Results_metrics_paper_3.0_GMPE.csv</a>,&nbsp;<a href="../api/records/10640493/draft/files/Results_metrics_paper_3.0_GMPE.csv/content" target="_blank" rel="noopener noreferrer">Results_metrics_paper_5.0_GMPE.csv, </a><a href="../api/records/10640493/draft/files/Results_metrics_paper_3.0_GMPE.csv/content" target="_blank" rel="noopener noreferrer">Results_metrics_paper_10.0_GMPE.csv</a></p> <h3>Supplementary Fig 4</h3> <p><strong>Predictions for the Synthetic event of magnitude 7.45</strong></p> <p>RF_predictions_T2s_map_3.csv, &nbsp;ASK_14_prediction_T2s_map_3.csv , ANN_predictions_T2s_map_3.csv</p> <p>RF_predictions_T3s_map_3.csv, &nbsp;ASK_14_prediction_T3s_map_3.csv , ANN_predictions_T3s_map_3.csv</p> <p>RF_predictions_T5s_map_3.csv, &nbsp;ASK_14_prediction_T5s_map_3.csv , ANN_predictions_T5s_map_3.csv</p> <p>RF_predictions_T10s_map_3.csv, &nbsp;ASK_14_prediction_T10s_map_3.csv , ANN_predictions_T10s_map_3.csv</p> <h3>Supplementary Fig 5</h3> <p><strong>Predictions for the Synthetic event of magnitude 8.05</strong></p> <p>RF_predictions_T2s_map_1240.csv, &nbsp;ASK_14_prediction_T2s_map_1240.csv , ANN_predictions_T2s_map_1240.csv</p> <p>RF_predictions_T1240s_map_1240.csv, &nbsp;ASK_14_prediction_T1240s_map_1240.csv , ANN_predictions_T1240s_map_1240.csv</p> <p>RF_predictions_T5s_map_1240.csv, &nbsp;ASK_14_prediction_T5s_map_1240.csv , ANN_predictions_T5s_map_1240.csv</p> <p>RF_predictions_T10s_map_1240.csv, &nbsp;ASK_14_prediction_T10s_map_1240.csv , ANN_predictions_T10s_map_1240.csv</p> <h3>Supplementary Fig 6</h3> <p><a href="../api/records/10640493/draft/files/NN_layers9North_Palm_Springs_OUT_event_metrics-T_10.csv/content" target="_blank" rel="noopener noreferrer">NN_layers9North_Palm_Springs_OUT_event_metrics-T_10.csv</a>, <a href="../api/records/10640493/draft/files/NN_layers9North_Palm_Springs_OUT_event_metrics-T_10.csv/content" target="_blank" rel="noopener noreferrer">NN_layers9Northridge_OUT_event_metrics-T_10.csv</a>, <a href="../api/records/10640493/draft/files/NN_layers9North_Palm_Springs_OUT_event_metrics-T_10.csv/content" target="_blank" rel="noopener noreferrer">NN_layers9Landers_OUT_event_metrics-T_10.csv</a>, <a href="../api/records/10640493/draft/files/NN_layers9North_Palm_Springs_OUT_event_metrics-T_10.csv/content" target="_blank" rel="noopener noreferrer">NN_layers9Hector_Mine_OUT_event_metrics-T_10.csv</a></p> <p>NN_layers9North_Palm_Springs_OUT_event_metrics-T_5.csv, NN_layers9Northridge_OUT_event_metrics-T_5.csv, NN_layers9Landers_OUT_event_metrics-T_5.csv, NN_layers9Hector_Mine_OUT_event_metrics-T_5.csv</p> <p>NN_layers9North_Palm_Springs_OUT_event_metrics-T_3.csv, NN_layers9Northridge_OUT_event_metrics-T_3.csv, NN_layers9Landers_OUT_event_metrics-T_3.csv, NN_layers9Hector_Mine_OUT_event_metrics-T_3.csv</p> <p>NN_layers9North_Palm_Springs_OUT_event_metrics-T_2.csv, NN_layers9Northridge_OUT_event_metrics-T_2.csv, NN_layers9Landers_OUT_event_metrics-T_2.csv, NN_layers9Hector_Mine_OUT_event_metrics-T_2.csv</p> <p>&nbsp;</p> <h3>Supplementary Fig 7</h3> <p>ASK-14 GMPE's</p> <p>df_InputDataPred_EQreal_Whittier_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_Landers_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_Northridge_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_North_Palm_Springs_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_Hector_Mine_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>&nbsp;</p> <p>RF and DNN</p> <p>Prediction_NN_BS_256_9capas_Northridge_3s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_Landers_3s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_Whittier_3s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_North_Palm_Springs_3s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_Hector_Mine_3s_log10_Review_ALL.csv</p> <h3>Suplementary Fig 8</h3> <p>ASK-14 GMPE's</p> <p>df_InputDataPred_EQreal_Whittier_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_Landers_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_Northridge_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_North_Palm_Springs_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_Hector_Mine_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>&nbsp;</p> <p>RF and DNN</p> <p>Prediction_NN_BS_256_9capas_Northridge_5s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_Landers_5s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_Whittier_5s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_North_Palm_Springs_5s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_Hector_Mine_5s_log10_Review_ALL.csv</p> <p>&nbsp;</p> <h3>Suplementary Fig 9</h3> <p>ASK-14 GMPE's</p> <p>df_InputDataPred_EQreal_Whittier_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_Landers_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_Northridge_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_North_Palm_Springs_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>df_InputDataPred_EQreal_Hector_Mine_plus_CS_sites_Scott_REVIEW_ASK_2014.csv</p> <p>&nbsp;</p> <p>RF and DNN</p> <p>Prediction_NN_BS_256_9capas_Northridge_10s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_Landers_10s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_Whittier_10s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_North_Palm_Springs_10s_log10_Review_ALL.csv</p> <p>Prediction_NN_BS_256_9capas_Hector_Mine_10s_log10_Review_ALL.csv</p> <p>&nbsp;</p> <h3>Suplementary Fig 10</h3> <p><a href="../api/records/10640493/draft/files/HyperParameters_T2s_log10.csv/content" target="_blank" rel="noopener noreferrer">HyperParameters_T2s_log10.csv</a></p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

Dataset for "Personalised Learning Environments Based on Knowledge Graphs and the Zone of Proximal Development"

<p>The dataset&nbsp;accompanying our paper &quot;Personalised Learning Environments Based on Knowledge Graphs and<br> the Zone of Proximal Development&quot; published in proceedings of CSEDU 2022.</p> <p>In the dataset you will find the raw csv results from both the explorative survey and the evaluation survey.<br> <br> The surveys were made in Google Forms and the results have also been exported as pdf files that are included as well.<br> <br> Finally, screenshots of the application, grouped by module can&nbsp;be found inside the screenshots.zip archive.</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Processed CODEX Datasets from - Graph deep learning for the characterization of tumour microenvironments from spatial protein profiles in tissue specimens

<p>This entry provides access to processed CODEX data files of three studies analyzed in the article "Graph deep learning for the characterization of tumour microenvironments from spatial protein profiles in tissue specimens". Details of datasets can be found in the Methods section of the article.</p> <p>For each dataset:</p> <ul> <li>A comma-separated values (CSV) file containing metadata of regions is included</li> <li>A zip file containing multiple CSV files is included: <ul> <li>`{region_id}.cell_data.csv`, a table containing three columns: "CELL_ID", "X", and "Y". This table provides centroid locations for all cells segmented in this region.</li> <li>`{region_id}.expression.csv`, a table containing multiple columns: "CELL_ID", "DAPI", "CD45", etc. This table provides detailed protein biomarker expression quantified and normalized for all cells in this region.</li> <li>`{region_id}.cell_types.csv`, a table containing two columns: "CELL_ID" and "CELL_TYPE". This table provides cell type annotations for all cells in this region.</li> <li>`{region_id}.cell_features.csv`, a table containing two columns: "CELL_ID" and "SIZE". This table provides morphology descriptors (only containing cell size for these studies) for all cells in this region.</li> </ul> </li> </ul> <p>These data files are also available through the Enable Medicine Public Study page: <a href="https://app.enablemedicine.com/portal/atlas-library/studies/92394a9f-6b48-4897-87de-999614952d94?sid=1168">https://app.enablemedicine.com/portal/atlas-library/studies/92394a9f-6b48-4897-87de-999614952d94?sid=1168</a>. Raw multiplexed immunofluorescence images will be accessible through the visualizer app of Enable Medicine Portal.</p> <p>Codes for this study are stored in <a href="https://gitlab.com/enable-medicine-public/space-gm">https://gitlab.com/enable-medicine-public/space-gm</a>. Please direct all further questions and/or issues to the gitlab repository or lead contact (A.E.T.).</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Dynamic Knowledge Graphs for Continual Learning of Embeddings

<p>These datasets are generated from real world usecases. They are treated as Knowledge graphs and include 20 snapshots, where between two snapshots there are 10% added links and 10% deleted links, making the first and last snapshot non-overlapping.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Semantic Web resources and Machine Learning systems - Knowledge Graph (SWeMLS-KG)

<p>This resource is part of our submission to ESWC 2023 resource track, which includes:</p> <p>Datasets:<br> - Folder &quot;pattern&quot; - a set of SWeMLS patterns represented based on OPMW and P-Plan ontology,<br> - Folder &quot;shapes&quot; - a set of SHACL constraints to check the conformance of SWeML Systems against SWeMLS patterns as well as a set of SHACL-AF rules to generate links between system components,<br> - File &quot;swemls-ontology.ttl&quot; - an ontology to represent Semantic Web resources and Machine Learning systems (SWeMLS),<br> - File &quot;swemls-instances.ttl&quot; - a set of triples representing the extracted metadata from 476 SWeML systems and papers,<br> - File &quot;swemls-kg.ttl&quot; - an integrated and validated KG containing all above files, including enrichment from SHACL-AF rules using &quot;swemls-toolkit&quot; [2].</p> <p>These resources are produced based on the result of the Systematic Mapping Study (SMS) reported in [1]. The latest SNAPSHOT-version of the resource can be accessed through our resource landing page: <a href="https://w3id.org/semsys/sites/swemls-kg/">https://w3id.org/semsys/sites/swemls-kg/</a></p> <p>[1] Breit, A., Waltersdorfer, L., Ekaputra, J.F., Sabou, M., Ekelhart, A., Iana, A., Paulheim, H., Portisch, J., Revenko, A., Ten Teije, A., van Harmelen, F.: Combining Machine Learning and Semantic Web -A Systematic Mapping Study (under review). ACM CSUR (2022)<br> [2] Source code of swemls-toolkit is available at: https://github.com/semanticsystems/swemls-toolkit</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Exploring the Global Reaction Coordinate for Retinal Photoisomerization: A Graph Theory-Based Machine Learning Approach

<p>This repository contains i. optimized geometry of the cis and trans retinal, ii. figure labelling the internal coordinates of retinal.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Graph Machine Learning Dataset LPWC

<p><strong>LPWC</strong> is a <strong>heterogenous graph machine learning dataset</strong> based on the RDF knowledge graph <a href="https://arxiv.org/pdf/2310.20475.pdf">Linked Papers With Code</a> (version from 2023-06-24).</p><p>LPWC contains four node types - papers (376,557 nodes), datasets (8,322 nodes), tasks (4,267 nodes) and methods (2,101), and six edge types.</p><p>Each node has rich semantic node features for node representation (content-based and topology-based node features are available).</p><p>&nbsp;</p><p>More information can be found in README.txt and on <a href="https://github.com/davidlamprecht/AutoRDF2GML">https://github.com/davidlamprecht/AutoRDF2GML</a></p>

opencc-by-sa-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record