Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

915

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

915 results for “graphs”

Learn how ShareScore rates datasets ↗
zenodo40/100

BRAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 9. Illustration of three different wavelets with three waves that have high amplitude values, as indicated by the dashed red circles

<p>The mother wavelet of Coiflet 5 contained triple-high oscillation amplitude (i.e., Figure 9a). We considered that this mother wavelet was inappropriate for our data because overall our data possibly contained only a few matches with the mother wavelet of Coiflet 5. Moreover, the Symlet 10 (i.e., Figure 9b) and 20 (i.e., Figure 9c) also provided supportive results that were lower than others in ANNSVM_WLHT because their mother wavelets also had a similar shape as that of Coiflet 5. For similar reasons, the Haar wavelet was not proper because it is a step function.&nbsp;</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

BRAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 6. Results from ANNSVN that used WL and HT

<p>To identify which features of data influentially impacted data separability, we conducted experiments for ANNSVM with WL and HT (i.e., Figure 6). The WL contained only wavelet coefficients, whereas HT included only results of the Hough transformation. We found that, again, results obtained via the linear kernel were not significant; however, using the RBF kernel, accuracy&nbsp;for WL was higher than that of HT, indicating that wavelet coefficients provide influential features that make data separable.&nbsp;</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

BRAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 3. Demonstrating the process of classification by applying the ANN, then the SVM

<p>Essentially, if the number of nodes in the hidden layers increases, processing time increases, and the resultant ANN will suffer from over-fitting. Conversely, too small of a number of hidden layers will cause under-fitting for the ANN. In our setting, the number of hidden layers and the number of nodes in each hidden layer were fixed at five. Concerning the learning rate and momentum settings, these impact sensitive training performances are set to optimal values obtained via a grid search technique. The number of nodes in the output layer was three because there are three different class labels (i.e., 2Dchart, bar, and pie) in our datasets. We used the ANN here because our datasets have nonlinear separation, and the ANN is also highly applicable to nonlinear modeling. Thus the ANN with multiple hidden layers was an optimal candidate; however, since the ANN is a black box learning approach, it is difficult to interpret implicit relationships between inputs and outputs.</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

RAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 4. Processes of all experiments:

<p>In this study, accuracy values of each dataset showed the performance of each method. These values represent are the proportion of the total number of predictions that were correctly classified. Initially, we classified training instances into three classes, with approximately 300 images per class. The graphs had been selectively gathered from the Web. We manually normalized the collected images by eliminating unused areas, such as unnecessary text. Moreover, we evaluated the experiments with 10 folds cross-validation because such an approach can mitigate the problem of over-fitting.&nbsp;</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

BRAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 5. Results from CNN and ANNSVM that used 1Dimg and 2Dimg

<p>We compared the results of CNN_1Dimg, CNN_2Dimg, ANNSVM_1Dimg, and ANNSVM_2Dimg to confirm the validity of ANNSVM when applied to images. The 1Dimg represented the dataset of one-dimensional images, while 2Dimg represented the dataset of twodimensional images. Results are shown in Figure 5.&nbsp;</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

BRAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 2. Illustrating the core process of one-dimensional image construction by applying a DFT

<p>First, we collect graph images as raw data, which contain different scales and sizes, and therefore need to be normalized. We clean the images by omitting irrelevant areas. For example, we omit unnecessary text that has nothing to do with our classification procedure. Moreover, to standardize the sizes and shapes of the images, we resize and reshape them to be 64 x 64 squares. Second, we examine each image pixel, each of which contains one color value. After each pixel is projected along the x- and y-axes, we count the number of projected pixels with a color value greater than zero to reduce image dimensionality. We, therefore, obtain two one-dimensional images from the x- and y-axes.&nbsp;&nbsp;</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

BRAIN Journal-ANNSVM: A Novel Method for Graph-Type Classification by Utilization of Fourier Transformation, Wavelet Transformation, and Hough Transformation-Figure 1. Example displaying two scatter plots with different characteristics and patterns

<p>In addition to this introductory section, the remainder of this paper is organized as follows. In Section 2, we present previous work related to our present study. In Section 3, we describe details regarding the methodology used in this study. In Section 4, we describe our experiments and results, then discuss our findings. Finally, we summarize the key content of our study in Section 5.&nbsp;&nbsp;</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

Internet AS-Graph Edgelist for Network Analysis

<p>This edgelist serves network analyses.</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

BRAIN Journal-Swarm Robotics with Circular Formation Motion Including Obstacles Avoidance-Figure 5: Graph Illustrate the new points on the circumference of the circle

<p>Formulas (1) and (2) are used to determine the new points on the circumference of the circle to form the circular formation shown in Figure 5.&nbsp;</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

An example of citation graphs showing year by year dynamics

<p>This is an example of citation graphs showing year by year dynamics. On x-axis there are years of publication while on y-axis the total number of citing articles for a given paper. Size of node also shows the total number of citing articles. This visualization is made with help of the Gephi application.</p>

opencc-by-4.0Jun 2018View details →
zenodo40/100

Road network graphs for betweenness centrality algorithm

<p>Weighted graph representation of a road network&nbsp;in selected regions. Derived from Open Street Map&nbsp;<a href="https://www.openstreetmap.org/#map=8/49.817/15.478">https://www.openstreetmap.org</a>. The dataset can be used as input for&nbsp;the betweenness centrality algorithm implemented here:&nbsp;<a href="https://code.it4i.cz/ADAS/betweenness">https://code.it4i.cz/ADAS/betweenness</a>.</p> <p><strong>Archive contents</strong><br> The archive contains following folders.</p> <p><strong>CZE</strong><br> Static graphs of three major cities in the Czech Republic (Praha, Brno, Ostrava) and entire Czech road network.&nbsp;Weighted by length of the road segments in metres.</p> <p><strong>PT</strong><br> Static graphs of Lisbon, Porto and entire Portugese road network.&nbsp;Weighted by length of the road segments in metres.</p> <p><strong>Data format</strong><br> Standard UTF-8 encoded CSV files, separated by semicolon with the following columns:<br> <br> <em>id1: (Type: unsigned long) - </em>start node<br> <em>id2:&nbsp;(Type: unsigned long) - </em>end&nbsp;node<br> <em>dist: (Type: unsigned long) - </em>weight of the edge (length in metres, unless described otherwise)<br> <em>edge_id:&nbsp;(Type: unsigned long) - </em>unique edge identifier<br> <br> <br> <br> &nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-nc-nd-4.0Jun 2018View details →
zenodo40/100

Calculation verification and graph generation

<p>This spreadsheet is used to verify the results of the calculations for the selected best values from the programs 8075DegRise_CO2NMDPSIRR_V5.bas (DOI&nbsp;10.5281/zenodo.1418561) and 8075DegRise_NoNMDP_V1.bas (DOI&nbsp; 10.5281/zenodo.1419629).&nbsp;It is also used to determine the percentage of heat added by each component, and finally to generate graphical representations of the data results.</p>

opencc-by-4.0Sep 2018View details →
zenodo40/100

Supplementary files for the manuscript "Assembly of Long Error-Prone Reads Using Repeat Graphs"

<p>Supplementary files for the manuscript &quot;Assembly of Long Error-Prone Reads Using Repeat Graphs&quot;</p> <p>&nbsp;</p> <p>Contents<br> ---------</p> <p>* `human_assemblues` - Flye assemblies of the human ONT sequencing data + QUAST benchmarking<br> of Flye, Canu and MaSuRCA assemblies. Scripts for assembly graph analysis are also included.</p> <p>* `nctc_assemblis` - Flye assemblies of the NCTC 21 bacterial dataset.</p> <p>* `yeast_assemblies` - working directories Flye, Canu, Falcon, Hinge and Miniasm assemblies of&nbsp;<br> yeast PB and ONT datasets + final assemblies + quast report. Some large files&nbsp;<br> (such as read alignments) were deleted.</p> <p>* `worm_assemblies` - working directories Flye, Canu, Falcon, Hinge and Miniasm assemblies of&nbsp;<br> the c. elegans dataset + final assemblies + quast report. Some large files&nbsp;<br> (such as read alignments) were deleted. `tandem_misassemblies` directory contain<br> the detailed analysis of nine tandem misassemblies. We recommend &quot;gepard&quot; dot-plotter for visualization.</p> <p>* `metagenome_assemblies` - Flye and Canu assemblies of a PacBio mock metagenome dataset.<br> In addition to metagenome assemblies, each bacteria was reassembled separately to<br> estimate the rate of divergence between the target genomes and the available references.</p> <p>* `simulated_data` - two assemblies of the simulated data illustrating Figure 1 (from Appendix I),<br> as well as simulated unbridged repeats benchmark.</p> <p><br> Software versions and parameters<br> --------------------------------</p> <p>* Flye - 2.3.5 (commit 20afeda)<br> * Canu - 1.7.1 (commit dfa60b8)<br> * Falcon - 0.3.0 (FALCON-Integrate commit 7498ef9)<br> * HINGE - 0.5.0 (commit 79fdf66)<br> * Miniasm - &nbsp;0.2-r168-dirty (commit 40ec280) / Minimap2 2.8-r711-dirty (commit 8fc5f8d)<br> * Quast - 5.0.0 (commit de6973bb)</p> <p>Flye and Canu were run with the default parameters. The config files / scripts for<br> Falcon, HINGE and Miniasm could be found in the &#39;asm_config&#39; archive folder.</p> <p>The HUMAN (but not the HUMAN+) assembly was generated with the earlier&nbsp;<br> Flye version 2.3.2 (released on Feb 20 2018) to provide a fair comparison&nbsp;<br> with the Canu and MaSuRCA assemblies (which were not updated since the release of Flye 2.3.2).<br> We note that the HUMAN assembly using the latest Flye version 2.3.5 has&nbsp;<br> NGA50 = 7.3 Mb and improves over the Flye 2.3.2 assembly (NGA50 = 6.3Mb).&nbsp;<br> HUMAN+ was assembled using the latest Flye and Canu versions (as of September 2018).</p> <p>The code for unbridged repeat resolution is currently available&nbsp;<br> in a separate &#39;flye-trestle&#39; branch (commit 6100d32)</p>

opencc-by-4.0Jan 2018View details →
zenodo40/100

The development and evaluation of an online application to assist in the extraction of data from graphs for use in systematic reviews (data repository)

<p>These are the data we generated in our evaluation of the graphical user interface.</p> <p>Please see our publication on Wellcome Open Research for information about the evaluations.</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

Results from "Binary Reduction of Dependency Graphs"

<p>The raw data of the reduction of the&nbsp;238 bugs as extracted from the runs of the four algorithms ddmin, verify, closure, binary. An extra column also exists that describes the run of Binary Reduction on the list of classes directly.</p> <p>The columns in `deliverable.csv` are predicate (the name of the decompiler), name (the name of the project we ran on), size (number of classes), scc&nbsp;(number of strongly connected components). Then for each tool we have&nbsp;size, scc&nbsp;(final size and scc), iters (iterations to last success), total-iters (iterations before finishing),&nbsp;time (time to last success (s)), total-time (time before finishing), timeout (did the predicate timeout, or succeed), check (did the bug still exist after reduction)</p> <p>The `benchmarks.csv` covers basic statistics about the programs used in the results. Lib is the number of library classes, LOC is lines of source code in the program, classes are the number of classes, edges are the edges in the dependency graph, degree is the average in and out degree in the graph, scc&nbsp;is the number of strongly connected components, out_degree and in_degree is the median in and out degree in the graph.</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Task graphs for benchmarking schedulers

<p><strong>Workflow Task Graph Dataset</strong></p> <p>This dataset contains three sets of task graphs representing different types of task workflows:</p> <ul> <li><em>Elementary</em> - contains trivial graph shapes, such as tasks with no dependencies or simple fork-join graphs. This set should test how the scheduler heuristics react to basic graph scenarios that frequently form parts of larger workflows.</li> <li><em>IRW</em> - is inspired by real-world workflows, such as machine learning cross-validation or map-reduce.</li> <li>&nbsp;<em>Pegasus</em> - is derived from graphs created by Pegasus Synthetic Workflow Generators (https://github.com/pegasus-isi/WorkflowGenerator)</li> </ul> <p>All of the provided task graphs are generated and compatible with ESTEE (https://github.com/It4innovations/estee) that allows to simulate their execution on a distributed system using various scheduling heuristics and environment conditions.</p> <p><strong>Data Format</strong></p> <p>Task graphs are stored in {elementary, irw, pegasus}.zip files that contain JSON representation of respective task graphs with the following fields:</p> <ul> <li>`graph_name` - Task graph name</li> <li>`graph_id` - Unique task graph identifier</li> <li>&nbsp;`graph` - Task graph representation - list of tasks where each task is represented as a dictionary with the following keys:</li> <li>&nbsp;`d`: Actual task duration in seconds (float value)</li> <li>&nbsp;`e_d`: User estimated task duration in seconds (float value)</li> <li>&nbsp;`cpus`: Task CPU core requirements (integer value)</li> <li>&nbsp;`outputs`: List of task outputs (list of integers indicating sizes of task outputs in MiB)</li> <li>&nbsp;`inputs`: List of task inputs in format of list [task\_id, output\_index]}. Output index is zero-based.</li> </ul> <p>For example this task graph:</p> <p>[{&#39;d&#39;: 200, &#39;e_d&#39;: 180, &#39;cpus&#39;: 1, &#39;outputs&#39;: [100], &#39;inputs&#39;: []},</p> <p>{&#39;d&#39;: 50, &#39;e_d&#39;: 60, &#39;cpus&#39;: 2, &#39;outputs&#39;: [], &#39;inputs&#39;: [[0, 0]]}]</p> <p>contains two tasks. One requiring no input, single CPU core with estimated duration 180s, actual duration 200s and producing a single output of 100 MiB. And another one requiring as an input task0&#39;s 0-th output, requiring 2 CPU cores, producing no output with estimated duration 60s and actual duration 50s.</p> <p>&nbsp;</p> <p><strong>Parsing the data</strong></p> <p>In Python, to load the elementary task graph set run the following snippet:</p> <pre><code class="language-python">import pandas as pd graphs = pd.read_json("./elementary.zip")</code></pre> <p>&nbsp;</p> <p>If you have Estee installed, you can use its provided `json_deserialize`</p> <p>function to parse the JSON encoded graphs into Estee TaskGraph data structure.</p> <p>&nbsp;</p> <pre><code class="language-python">from estee.serialization.dask_json import json_deserialize graph_json = graphs.loc[0, "graph"] graph = json_deserialize(graph)</code></pre> <p>&nbsp;</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Ozymandias: A biodiversity knowledge graph

<p>Triples for the knowledge graph described in &quot;Ozymandias: A biodiversity knowledge graph&quot;&nbsp;<a href="https://doi.org/10.7717/peerj.6739">doi:10.7717/peerj.6739</a>&nbsp;Each set of triples corresponds to a named graph in the triple store:</p> <blockquote> <p>oz-ala.nt &nbsp; &nbsp; &nbsp; &nbsp; &lt;https://bie.ala.org.au&gt;<br> oz-publication.nt &lt;https://biodiversity.org.au/afd/publication&gt;<br> oz-zenodo.nt &nbsp; &nbsp; &nbsp;&lt;https://zenodo.org&gt;<br> oz-crossref.nt &nbsp; &nbsp;&lt;https://crossref.org&gt;<br> oz-orcid.nt &nbsp; &nbsp; &nbsp; &lt;https://orcid.org&gt;<br> oz-species.nt &nbsp; &nbsp; &lt;https://species.wikimedia.org&gt;<br> oz-gbif.nt &nbsp; &nbsp; &nbsp; &nbsp;&lt;https://gbif.org/species&gt;<br> oz-bold.nt &nbsp; &nbsp; &nbsp; &nbsp;&lt;http://boldsystems.org&gt;</p> </blockquote> <pre>&nbsp;</pre> <pre>&nbsp;</pre>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Linux Kernel 4.21 Call Graphs

<p>This is the<strong> Linux Kernel 4.21 Call Graphs</strong> created using <a href="http://github.com/dspinellis/cscout">CScout</a>&nbsp;containing the following graphs:</p> <ol> <li>File include graph (fgraph_I.txt)&nbsp;</li> <li>Compile Time Dependency Graph (fgraph_C.txt)</li> <li>Control Dependency Graph (through function calls) (fgraph_F_D.txt)</li> <li>Data Dependency Graph (through global variables) (fgraph_G.txt)</li> <li>Function and Macro Call Graph (cgraph.txt)</li> </ol> <p>Files are of the form</p> <p>foo.c boo.c</p> <p>which indicate a directed edge foo.c -&gt; boo.c.</p> <p>The call graphs refer to <strong>all </strong>(ending with _all.txt)&nbsp;files or only the <strong>writable files.</strong>&nbsp;</p> <p>These graphs were produced by processing the Linux Kernel Codebase consisting of 20.3 million lines of source code.&nbsp;</p> <p>The results were produced on an&nbsp;Intel(R) Xeon(R) CPU E5-1410 0 @ 2.80GHz server with 64GB of RAM.</p> <p><strong>References:&nbsp;</strong></p> <p>1. Papachristou, Marios. &quot;Software clusterings with vector semantics and the call graph.&quot;&nbsp;<em>Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering</em>. 2019.</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

ArCo Knowledge Graph

<p>The ArCo knowledge graph contains the ontology network and the data about the cultural properties catalogued&nbsp;by the Italian Institute of the General Catalogue and Documentation.</p> <p>Data are represented with RDF and&nbsp;by using&nbsp;N-Triples as syntax.</p> <p>The ontologies of the network are modelled with OWL 2 and serialised with the RDF/XML syntax.</p> <p>The ontology network is released along with alignments to other ontologies/vocabularies in the Semantic Web. Those alignments are provided within separate OWL files.</p> <p>The data are contained into a single RDF dump serialised as N-TRIPLES.</p> <p>Additionally, the release provides the links between ArCO entities and other entities published in other datasets in the Linked Open Data cloud. Such links are represented by using owl:sameAs axioms and serialised as N-TRIPLES into a separate file.</p>

opencc-by-sa-4.0Apr 2019View details →
zenodo40/100

A biodiversity dataset graph: Biodiversity Heritage Library (BHL)

<p>A biodiversity dataset graph: Biodiversity Heritage Library&nbsp;</p> <p>Biodiversity datasets, or descriptions of biodiversity datasets, are increasingly available through open digital data infrastructures such as the Biodiversity Heritage Library (BHL, https://biodiversitylibrary.org). &quot;The Biodiversity Heritage Library improves research methodology by collaboratively making biodiversity literature openly available to the world as part of a global biodiversity community.&quot; - https://biodiversitylibrary.org , June 2019. &nbsp;</p> <p>However, little is known about how these networks, and the data accessed through them, change over time. This dataset provide snapshots of all OCR item texts (e.g., individual items) available through BHL as tracked by Preston (https://github.com/bio-guoda/preston , https://doi.org/10.5281/zenodo.1410543 ) over period May - June 2019.</p> <p>This snapshot contains about 120GB of uncompressed OCR texts across 227k OCR BHL items. Also, a snapshot of the BHL item catalog at https://www.biodiversitylibrary.org/data/item.txt is included.</p> <p>The archive consists of 256 individual parts (e.g., preston-00.tar.gz, preston-01.tar.gz, ...) to allow for parallel file downloads. The archive contains three types of files: index files, provenance files and data files. Only two index and provenance files are included and have been individually included in this dataset publication. Index files provide a way to links provenance files in time to eestablish a versioning mechanism. Provenance files describe how, when and where the BHL OCR text items were retrieved. For more information, please visit https://preston.guoda.bio or https://doi.org/10.5281/zenodo.1410543). &nbsp;</p> <p>To retrieve and verify the downloaded BHL biodiversity dataset graph, first concatenate all the downloaded preston-*.tar.gz files (e.g., cat preston-*.tar.gz &gt; preston.tar.gz). Then, extract the archives into a &quot;data&quot; folder. After that, verify the index of the archive by reproducing the following result:</p> <p>$ java -jar preston.jar history<br> &lt;0659a54f-b713-4f86-a917-5be166a14110&gt; &lt;http://purl.org/pav/hasVersion&gt; &lt;hash://sha256/89926f33157c0ef057b6de73f6c8be0060353887b47db251bfd28222f2fd801a&gt; .<br> &lt;hash://sha256/41b19aa9456fc709de1d09d7a59c87253bc1f86b68289024b7320cef78b3e3a4&gt; &lt;http://purl.org/pav/previousVersion&gt; &lt;hash://sha256/89926f33157c0ef057b6de73f6c8be0060353887b47db251bfd28222f2fd801a&gt; .</p> <p>To check the integrity of the extracted archive, confirm that each line produce by the command &quot;preston verify&quot; produces lines as shown below, with each line including &quot;CONTENT_PRESENT_VALID_HASH&quot;. Depending on hardware capacity, this may take a while.</p> <p>$ java -jar preston.jar verify<br> hash://sha256/e0c131ebf6ad2dce71ab9a10aa116dcedb219ae4539f9e5bf0e57b84f51f22ca&nbsp;&nbsp; &nbsp;file:/home/preston/preston-bhl/data/e0/c1/e0c131ebf6ad2dce71ab9a10aa116dcedb219ae4539f9e5bf0e57b84f51f22ca&nbsp;&nbsp; &nbsp;OK&nbsp;&nbsp; &nbsp;CONTENT_PRESENT_VALID_HASH&nbsp;&nbsp; &nbsp;49458087<br> hash://sha256/1a57e55a780b86cff38697cf1b857751ab7b389973d35113564fe5a9a58d6a99&nbsp;&nbsp; &nbsp;file:/home/preston/preston-bhl/data/1a/57/1a57e55a780b86cff38697cf1b857751ab7b389973d35113564fe5a9a58d6a99&nbsp;&nbsp; &nbsp;OK&nbsp;&nbsp; &nbsp;CONTENT_PRESENT_VALID_HASH&nbsp;&nbsp; &nbsp;25745<br> hash://sha256/85efeb84c1b9f5f45c7a106dd1b5de43a31b3248a211675441ff584a7154b61c&nbsp;&nbsp; &nbsp;file:/home/preston/preston-bhl/data/85/ef/85efeb84c1b9f5f45c7a106dd1b5de43a31b3248a211675441ff584a7154b61c&nbsp;&nbsp; &nbsp;OK&nbsp;&nbsp; &nbsp;CONTENT_PRESENT_VALID_HASH&nbsp;&nbsp; &nbsp;519892</p> <p>Note that a copy of the java program &quot;preston&quot;, preston.jar, is included in this publication. The program runs on java 8+ virtual machine using &quot;java -jar preston.jar&quot;, or in short &quot;preston&quot;.&nbsp;</p> <p>Files in this data publication:</p> <p>README - this file</p> <p>preston-[00-ff].tar.gz - preston archives containing BHL OCR item texts, their provenance and a provenance index.</p> <p>9e8c86243df39dd4fe82a3f814710eccf73aa9291d050415408e346fa2b09e70 - preston index file<br> 2a5de79372318317a382ea9a2cef069780b852b01210ef59e06b640a3539cb5a - preston index file</p> <p>89926f33157c0ef057b6de73f6c8be0060353887b47db251bfd28222f2fd801a - preston provenance file<br> 41b19aa9456fc709de1d09d7a59c87253bc1f86b68289024b7320cef78b3e3a4 - preston provenance file<br> &nbsp;</p> <p>This work is funded in part by grant NSF OAC 1839201 from the National Science Foundation.</p>

opencc-zeroJun 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record