Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2
datasets available to search
ShareScore release 0.9.0
Dataset results
2 results for “LUBM”
Lehigh University Benchmark (LUBM): Evolving Graph Simulation
<p>The Lehigh University Benchmark (LUBM) generates benchmark datasets containing people working at universities [1]. We use the Data Generator v1.7 to generate 10 versions of a graph containing 100 universities [2].<br> Thus, all versions are of similar size, but we emulate modifications by generating different vertex identifiers, i.e., each version is considered a timestamped graph. Each graph contains about 2.1 M vertices and 13 M edges.<br> Over all versions, the mean degree is 6.7 (+- 0.1), the mean in-degree is 6.8 (+- 0.1), and the mean out-degree is 5.1 (+- 0.1).</p> <p>1. <a href="https://dblp.uni-trier.de/pid/80/5390.html">Yuanbo Guo</a>, <a href="https://dblp.uni-trier.de/pid/48/6834.html">Zhengxiang Pan</a>, <a href="https://dblp.uni-trier.de/pid/94/1154.html">Jeff Heflin</a>: LUBM: A benchmark for OWL knowledge base systems. <a href="https://dblp.uni-trier.de/db/journals/ws/ws3.html#GuoPH05">J. Web Semant. 3(2-3)</a>: 158-182 (2005)</p> <p>2. <a href="https://dblp.uni-trier.de/pid/222/6353.html">Till Blume</a>, <a href="https://dblp.uni-trier.de/pid/r/DavidRicherby.html">David Richerby</a>, <a href="https://dblp.uni-trier.de/pid/06/2380.html">Ansgar Scherp</a>: Incremental and Parallel Computation of Structural Graph Summaries for Evolving Graphs. <a href="https://dblp.uni-trier.de/db/conf/cikm/cikm2020.html#BlumeRS20">CIKM 2020</a>: 75-84</p>
Automatically Extracted SHACL Shapes for WikiData, DBpedia, YAGO-4, and LUBM & Associated Coverage Statistics
<p>The uploaded datasets contain <strong>automatically extracted </strong>SHACL shapes for the following datasets:</p> <ul> <li>WikiData (the truthy dump from September 2021 filtered by removing non-English strings) [1]</li> <li>DBpedia [2]</li> <li>YAGO-4 [3] </li> <li>LUBM (scale factor 500) [4]</li> </ul> <p>The validating shapes for these datasets are generated by a program that parses the corresponding RDF files (in `.nt` format). The extracted shapes encode various SHACL constraints, e.g., sh:minCount, sh:path, sh:class, sh:datatype etc. For each shape we encode coverage in terms of number of entities satisfying such shape, this information is encoded using the <a href="http://vocab.deri.ie/void#entities">void:entities</a> predicate. </p> <p>We have provided as executable Jar file the program we developed to extract these SHACL shapes.<br> More details about the datasets used to extract these shapes and <em>how to run the Jar</em> are available on our GitHub repository <a href="https://github.com/dkw-aau/qse">https://github.com/dkw-aau/qse</a>.</p> <p>Read more about our Quality Shapes Extraction (QSE) tool on our website <a href="https://relweb.cs.aau.dk/qse/">https://relweb.cs.aau.dk/qse/</a></p> <p>[1] Vrandečić, Denny, and Markus Krötzsch. "Wikidata: a free collaborative knowledgebase." Communications of the ACM 57.10 (2014): 78-85.</p> <p>[2] Auer, Sören, et al. "Dbpedia: A nucleus for a web of open data." The semantic web. Springer, Berlin, Heidelberg, 2007. 722-735.</p> <p>[3] Pellissier Tanon, Thomas, Gerhard Weikum, and Fabian Suchanek. "Yago 4: A reason-able knowledge base." European Semantic Web Conference. Springer, Cham, 2020.</p> <p>[4] Guo, Yuanbo, Zhengxiang Pan, and Jeff Heflin. "LUBM: A benchmark for OWL knowledge base systems." Journal of Web Semantics 3.2-3 (2005): 158-182.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.