Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
27
datasets available to search
ShareScore release 0.9.0
Dataset results
27 results for “Graph database”
Bar graphs of DEMIX database tile classifications
<p>The DEMIX database (Guth, 2023a) contains statistics from 6 test 1 arc second DEMs (ALOS, ASTER, CopDEM, FABDEM, NASADEM, and SRTM) compared to high resolution reference DEMs. The database contains 236 DEMIX tiles (Guth and others, 2023) and forms the basis for the ranking of global DEMs in Bielski and others (2023).</p> <p>A K-means clustering of the database using MICRODEM (Guth, 2023b, 2023c), and an additional set of 4 land cover and landform classifications (Table 1) for the 236 DEMIX tiles computed the percentage of each DEMIX tile in each classification category. Guth (2023d) has the raw data for the percentages of each category for each of the 236 tiles along with the K-means cluster assignments.</p> <p>This data set contains 3 figures for each of the 5 classification databases in Guth (2023d):</p> <ul> <li>Bar graph of the category percentages for each of the test areas. A composite version of these graphs is in Bielski and others (2023).</li> <li>Bar graph of the category percentages for each of the 236 test tiles. These graphs are too large to include on a single page with readable legends.</li> <li>Legend for the classification</li> </ul> <p> </p> <p>References:</p> <p>Bielski, C.; López-Vázquez, C.; Grohmann, C.H.; Guth. P.L.; and the TMSG DEMIX Working Group, 2023. DEMIX Method Ranks COPDEM, and FABDEM as Top 1” Global DEMs: <a href="https://arxiv.org/abs/2302.08425v3">https://arxiv.org/abs/2302.08425v3</a></p> <p>Guth, P. L., 2023a. DEMIX GIS Database Version 2 (2.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.8062008">https://doi.org/10.5281/zenodo.8062008</a></p> <p>Guth, P.L., 2023b. GIT-MICRODEM [Delphi source code, archived installation versions]. URL: https://github.com/prof-pguth/git_microdem </p> <p>Guth, P.L., 2023c, MICRODEM: Open-source GIS with a focus on Geomorphometry [download latest Win64 executable and CHM help file] URL: <a href="https://microdem.org/">https://microdem.org/</a></p> <p>Guth, P.L., 2023d, K-means clustering of the DEMIX data set (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.8283791</p> <p>Guth, Peter L., Peter Strobl, Kevin Gross, & Serge Riazanoff. (2023). DEMIX 10k Tile Data Set (1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7504791">https://doi.org/10.5281/zenodo.7504791</a></p>
Combat-TB-NeoDB: fostering Tuberculosis research through integrative analysis using graph database technologies.
<p>NeoDB is a free and open source integrated M.tuberculosis ‘omics’ knowledge-base. NeoDB is based on Neo4j and enables researchers to execute complex federated queries by linking well-known, curated and widely used biological data resources, and supplementary TB variants data from published literature.</p> <p>Documentation can be found at https://combat-tb-db.readthedocs.io</p>
[SZESLR] Semi-automated Systematic Literature Review (SLR) Design and P-graph (Process Graph Theory) SLR Database
<p>(c) Széchenyi István University and Artificial Intelligence National Laboratory (MILAB)</p> <p>A proposed semi-automated Systematic Literature Review (SLR) design has been deployed to conduct an exhaustive literature review on contributions concerning applications of the P-graph framework as a case study. The P-graph framework (<a href="https://p-graph.org/">https://p-graph.org</a>) was initially conceived for solving problems related to the synthesis of chemical processes; in particular, the framework has been useful for developing sustainable and resilient systems.</p> <p>The current contribution includes the P-graph SLR database itself, and additionally, the Python scripts used to generate such results. The case study has been selected in view of the authors’ familiarity with the P-graph framework, and because they maintain a manually-collected database of works on the subject. The contributions concerning the framework are easily distinguishable from others involving similar methodologies. It was observed that the implementation of the proposed SLR design extended prior review papers on the P-graph framework and provided additional information on authors and institutions devoted to research on this subject.</p> <p>Generally, it is expected that the dataset will assist researchers, scholars, and practitioners concerned with industrial sustainability in streamlining their literature reviews, thus accelerating the acquisition of pertinent publications that could advance and enrich their own research works on such a subject. Moreover, it is expected that, with proper adaptation, the proposed semi-automated SLR design can be applied to conduct literature reviews in other scientific fields whose results can be validated against existing reviews when available.</p> <p>The normalized P-graph database hereby presented thus can be used as a source of validated information for researchers. The P-graph database was uploaded in various file formats that support database management software (e.g. Zotero .RIS) as well as an "EasyHandle" MS Office (.XLSX) file.</p>
Information System for Cycling Navigation for Aveiro, Portugal - Database and ArcGIS Graph
<p>This data was collected within the scope of a master's thesis in mechanical engineering from University of Aveiro, Portugal. The goal was to develop a information system for cycling navigation for an urban area of the portuguese city of Aveiro. Therefore, data collection was achieved through an instrumented alluminium bicycle equipped with a 1) GNSS Data Logger, 2) wireless heart rate recorder device and 3) video camera. 120 km were covered to collect 8h of video and about 100000 second by second data points, which were analyzed and organized through an 449 link-map built in ArcGIS. Seven different bikeability indicators were built: travel time, energy expenditure, effort distribution, infrastructure performance, safety, comfort and emissions hotspots of two pollutants (CO2 and NOx).</p> <p>The spreadsheet dataset is divided into three sections:</p> <ol> <li>Final Attributes - all the collected and treated data of each link, both used in the indicators and in another set for system' network analysis;</li> <li>Attributes Statistics - reveals some statistics of the weighted final atrributes achieved, such as mean, standard error, standard deviation, range, confidence level, among others;</li> <li>Link Type Statistics - reveals the distribution of the different link-types along the map and the respective average speeds recorded.</li> </ol> <p>In order to associate the FID numbers with the respective links, the developed ArcGIS graph was also made available. </p>
Japanese Visual Media Graph - Anime Characters Database Ontology
<p>This group of files represents the RDF ontology used in the Japanese Visual Media Graph for the Anime Characters Database. Included in the upload are an explanatory PDF, the ontology in the Turtle serialization, and an HTML visualization of the ontology.</p>
Japanese Visual Media Graph - Visual Novel Database Ontology
<p>This group of files represents the RDF ontology used in the Japanese Visual Media Graph for the Visual Novel Database. Included in the upload are an explanatory PDF, the ontology in the Turtle serialization, and an HTML visualization of the ontology.</p>
Database Van der Meer (1988)-Stability graphs
<p>The Excel database with figures is the basis of the figures in: Van der Meer, J.W. (2021). Rock Armour Slope Stability under Wave Attack: the Van der Meer Formula revisited. Journal of Coastal and Hydraulic Structures, JCHS.</p> <p>Researchers may use the figures to compare with their own results on rock slope stability under wave attack.</p>
Graph database of the urban road network of Modena
<p>The file contains a dump of the neo4j instance of the road network of the city of Modena. The database contains both the primal graph and the dual graph and integrates traffic volume data and Points Of Interest (POI). The database is generated from Open Street Map data exploiting the code in <a href="https://anonymous.4open.science/r/roadRouting-2D96">this</a> git repository. The git repository shows how to employ the graph database for analysis, routing, and the simulation of road closure scenarios.</p>
Implementing Traceability Repositories as Graph Databases for Software Quality Improvement: Datasets used to test our methodology that is presented in the paper 10.1109/QRS.2018.00040
<p>The first dataset is the Event Based Traceability for Managing Evolutionary Change (EBT), it is a public dataset provided by CoEST, the original artifacts and trace links are represented in XML and text format. From the EBT dataset, we selected the 41<em> requirements </em>and 25<em> test case </em>artifacts, in addition to the answer set of 51 trace links which relates the <em>requirements </em>with the<em> test case. </em>Artifacts and trace links are prepared in XML format<strong>. </strong>The data set contains XML for each artifact such as RQ.xml, EBTrelations.xml is the answer set file, TradModel.xml which describes the defined model and TradTraceabilityRule.xml that includes the rules applied for trace link types.</p> <p> </p> <p>The second dataset AgileOERP is collected from commercial management tool to customize an open source ERP applying agile methodology. It contains 350<em> user stories (US), </em>1323<em> tasks (TS) </em>and 198<em> developer test (DT) </em>artifacts, in addition to answer set of trace links that manually generated by developers which relates the <em>user story </em>artifact with<em> task (</em>1304) artifact, as such relates the <em>task </em>artifact with<em> developer test </em>artifact (65). Artifacts and trace links are prepared in XML format<strong>. </strong>The data set contains XML for each artifact such as US.xml, ERPrelations.xml is the answer set file, AgileModel.xml which describes the defined model and AgileTraceabilityRule,xml that includes all rules applied for trace links type</p> <p> </p> <p>The original dataset of the last dataset is the Aqualush irrigation system which is used as a case study in “C. Fox, Introduction to Software Engineering Design: Processes, Principles and Patterns with UML2. Addison-Wesley, 2006”. The trace links are generated and provided in “E. Ben Charrada, D. Caspar, C. Jeanneret, and M. Glinz, towards a benchmark for traceability, in Joint EVOL and IWPSE 2011, pp. 21-30”, in HTML format. For our work, we selected the <em>software requirements specification (</em>396 SRS), <em>user level requirements (</em>48 ULR), <em>use case (</em>74 UC), <em>detailed design (85 DD) </em>and <em>software architecture(15 SArch) </em>artifacts in addition to the answer set of trace links that relate the SRS with other artifacts(4038) and thus relates the DD artifact with other artifacts (1719) . Artifacts and trace links are prepared in XML<strong>. </strong>The data set contains XML for each artifact such as SRS.xml, AqualushRelations.xml is the answer set file, TradModel.xml which describes the defined model and TradTraceabilityRule.xml that includes the rules applied for trace link types.</p>
Database file for Encyclopedia of Finite Graphs: Simple connected graphs, n<=10
<p>This initial release contains all simple connected graphs of order n<=10, and a collection of integer invariants. Up to order n<=6, there is a collection of "special" invariants that are stored in a custom table (see main project for details).</p>
DirectedSmallMoleculesNetwork (DSMN) graph database for Neo4j (3.x)
<p>Release of the Neo4j metabolic interaction database for species human (Homo sapiens) (first unzip before using!). The data is licensed under the <a href="https://creativecommons.org/share-your-work/public-domain/cc0/">CCZero waiver</a>. This file contains data from the following pathway databases: Reactome, LIPID MAPS, WikiPathways.</p> <p> </p>
The Ring: Worst-Case Optimal Joins in Graph Databases using (Almost) No Extra Space
<p><strong>wikidata-filtered-enumerated: </strong>Wikidata subgraph with 81,426,573 triples</p> <p><strong>wikidata-ring: </strong>Wikidata graph with 958,844,164 triples. The compressed file contains the mapping from subject/objects (.SO) and predicates (.P)</p>
Data for "Expanding the scope of a catalogue search to bioisosteric fragment merges using a graph database approach"
<p>Data for "Expanding the scope of a catalogue search to bioisosteric fragment merges using a graph database approach"</p> <p>Contains files with the SMILES retrieved from the database (fragnet_query_outputs.zip) and sdf files containing the full lists of scored compounds for Fragalysis target test cases (scored_output_sdfs.zip). The compound sets used in the bioisosteric and perfect comparisons are saved in a separate directory.</p>
Multivariate prediction on wake-affected wind turbines using graph neural networks (Eurodyn) database
<p>Database consisting of graphs generated using randomized layouts and PyWake simulations used in '<em>Multivariate prediction on wake-affected wind turbines using graph neural networks</em>', contribution to Eurodyn 2023. </p>
Data for "The use of a graph database is a complementary approach to a classical similarity search for identifying commercially available fragment merges"
<p>The input and output data to the filtering pipeline described in "The use of a graph database is a complementary approach to a classical similarity search for identifying commercially available fragment merges"; code available from https://github.com/oxpig/fragment_network_merges. </p>
Neo4j Database Dump of Corona-Warn-App Repository Provenance Graphs
<p>Provenance database (Neo4j 4.1) dumps of the following <a href="https://github.com/corona-warn-app/">Corona-Warn-App</a> repositories:</p> <ol> <li>cwa-app-android</li> <li>cwa-app-ios</li> <li>cwa-server</li> <li>cwa-documentation</li> </ol> <p>Username: covid</p> <p>Password: covid19</p>
An Orthology Graph Database
<p>The graph database used in the study....</p> <p>More info will be added following the publication...</p>
Database files 1 for network described in "Expanding the scope of a catalogue search to bioisosteric fragment merges using a graph database approach"
<p>Contains part 1/2 of the database files required for running the database described in "Expanding the scope of a catalogue search to bioisosteric fragment merges using a graph database approach" that can be run with https://github.com/stephwills/docker-neo4j.</p>
Database files 2 for network described in "Expanding the scope of a catalogue search to bioisosteric fragment merges using a graph database approach"
<p>Contains part 2/2 of the database files required for running the database described in "Expanding the scope of a catalogue search to bioisosteric fragment merges using a graph database approach" that can be run with https://github.com/stephwills/docker-neo4j.</p>
The Brill Knowledge Graph: A Database of Bibliographic References and Index Terms extracted from Books in Humanities and Social Sciences
<p>We present a complete dataset of linked bibliography and index data, partially disambiguated and augmented with references to external resources, extracted from the Brill’s archive in the field of Classics. Processed book identifiers are listed in a separate text file. Text fragments extracted from different books via this process are then parsed and compared using a string-based similarity metric to form clusters of bibliographic references to the same published work or (variants of) the same subjects discussed in these books. The entire set of references was then disambiguated using Google Books and Crossref APIs.</p> <p><a href="https://jdmdh.episciences.org/11062">Paper about extraction pipeline</a></p> <p><a href="https://www.nkokash.com/documents/KIEM-RDJ.pdf">Paper about extracted KG</a></p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.