Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14
datasets available to search
ShareScore release 0.9.0
Dataset results
14 results for “database dump”
MongoDB database dump for the analysis of the current sustainability state of research software
<p>This data set is the MongoDB dump (bson files) of the data created and analyzed with the rsps framework. In the first step, a research subject is assigned to the research software repositories. Afterwards, the current sustainability state is evaluated. The data set comprises the following six bson files:</p> <p><strong>repositories: </strong>metadata, received from the GitHub REST API, for repositories containing the search terms "doi+10" or "doi+10+in:readme", additional information are the request date, the contained search term, and the repository hosting service, in this case for all repositories "github". For repositories the Readme files are available.</p> <p><strong>publications:</strong> metadata of publications, published on arXiv and ACM, that contain the search term "github.com".</p> <p><strong>rs_repositories:</strong> research software candidates containing a DOI or that are referenced by the publications contained in the publications data set.</p> <p><strong>rs_artifacts: </strong>research software artifacts that are referenced in the harvested GitHub repositories by a DOI and the harvested publications.</p> <p><strong>publication_subjects:</strong> All Science Journal Classification (ASJC) of Scopus combined with the Scopus source list and Scopus book title list (https://www.scopus.com/home.uri)</p> <p><strong>arxiv_subjects:</strong> arXiv taxonomy complemented with the ASJC research subject.</p> <p> </p>
Synergy database dump
<p>An SQL dump of the Synergy database. Synergy was first published in 2014, but the associated application has now reached the end of its life. This database contains the data that was presented in the publication, and that all tools in the web application were using.</p>
dbBact v2022.07.01 database dump
<p>A full database dump (postgres) for <a href="https://www.biorxiv.org/content/10.1101/2022.02.27.482174v2.abstract">dbBact</a>, the microbiome database (see <a href="https://dbbact.org">website</a>).</p> <p>This is version 2022.07.01, used in the paper examples.</p> <p>Latest version can be downloaded from:</p> <p><a href="https://dbbact.org/download">https://dbbact.org/download</a></p>
PIBASE.ligands database mysql dump
<p>This dataset is a mysql dump of the PIBASE.ligands database describing the overlap of small molecule and protein binding sites ( http://fredpdavis.com/pibase.ligands ) described in:</p> <p>The overlap of small molecule and protein binding sites within families of protein structures.<br /> Davis FP, Sali A. <em>PLoS Comput Biology</em> 2010 6(2): e1000668.<br /> doi:10.1371/journal.pcbi.1000668</p>
Database Dump of Weather Analytics Application
<p>This SQL file contains a PostgreSQL Database Dump of the Weather Analytics Application (https://doi.org/10.5281/zenodo.804027)</p> <p>The database contains the temperature differences reports of London from 1990 until 1999</p>
Open Context Database SQL Dump: Legacy Schema Tables and New Schema Tables
<p>Open Context (<a href="https://opencontext.org">https://opencontext.org</a>) publishes free and open access research data for archaeology and related disciplines. An open source (but bespoke) Django (Python) application supports these data publishing services. The software repository is here: <a href="https://github.com/ekansa/open-context-py">https://github.com/ekansa/open-context-py</a></p> <p>The Open Context team runs ETL (extract, transform, load) workflows to import data contributed by researchers from various source relational databases and spreadsheets. Open Context uses PostgreSQL (<a href="https://www.postgresql.org">https://www.postgresql.org</a>) relational database to manage these imported data in a graph style schema. The Open Context Python application interacts with the PostgreSQL database via the Django Object-Relational-Model (ORM).</p> <p>In 2023, the Open Context team finished migration of from a legacy database schema to a revised and refactored database schema with stricter referential integrity and better consistency across tables. During this process, the Open Context team de-duplicated records, cleaned some metadata, and redacted attribute data left over from records that had been incompletely deleted in the legacy schema.</p> <p>This database dump includes all Open Context data organized with the legacy schema (table names that start with the 'oc_' or 'link_' prefixes) along with all Open Context data after cleanup and migration to the new database schema (table names that start with 'oc_all_'). The binary media files referenced by these structured data records are stored elsewhere. Binary media files for some projects, still in preparation, are not yet archived with long term digital repositories.</p> <p>These data comprehensively reflect the structured data currently published and publicly available on Open Context. Other data (such as user and group information) used to run the Website are not included. </p> <p> </p> <p><strong>IMPORTANT</strong></p> <p>This database dump contains data from roughly 180 different projects. Each project dataset has its own metadata and citation expectations. If you use these data, you must cite each data contributor appropriately, not just this Zenodo archived database dump.</p> <p> </p> <p> </p>
Wikidata Dump Joconde database works
<p>RDF dump of wikidata produced with <a href="//wdumps.toolforge.org/">wdumper</a>.</p><p>Artworks with a Joconde ID (Joconde is a french database of artworks in the french cultural heritage)<br><a href="//wdumps.toolforge.org/dump/3269">View on wdumper</a></p><p><b>entity count<b>: 18099, <b>statement count</b>: 374647, <b>triple count</b>: 484207</b></b></p>
Wikidata Dump Joconde database works
<p>RDF dump of wikidata produced with <a href="//wdumps.toolforge.org/">wdumper</a>.</p><p>Artworks with a Joconde ID (Joconde is a french database of artworks in the french cultural heritage)<br><a href="//wdumps.toolforge.org/dump/3269">View on wdumper</a></p><p><b>entity count<b>: 18099, <b>statement count</b>: 374647, <b>triple count</b>: 484207</b></b></p>
Mongodb Database dump for TOSEM submission "Characterizing Deep Learning Package Supply Chains in PyPI: Domains, Clusters, and Disengagement"
<p>The Mongodb Database dump for TOSEM submission "Characterizing Deep Learning Package Supply Chains in PyPI: Domains, Clusters, and Disengagement"</p>
Neo4j Database Dump of Corona-Warn-App Repository Provenance Graphs
<p>Provenance database (Neo4j 4.1) dumps of the following <a href="https://github.com/corona-warn-app/">Corona-Warn-App</a> repositories:</p> <ol> <li>cwa-app-android</li> <li>cwa-app-ios</li> <li>cwa-server</li> <li>cwa-documentation</li> </ol> <p>Username: covid</p> <p>Password: covid19</p>
OCTOPUS database v.2.1 (SQL database dump)
<p>A full dump of <strong>OCTOPUS PostgreSQL database v.2.1</strong> as published upon</p> <ul> <li>a sematic database redesign (effective db v.2),</li> <li>the creation of a fully relational PostgreSQL database that uses the PostGIS spatial extension (effective db v.2),</li> <li>moving the database to GCP (effective db v.2),</li> <li>fostered FAIR, OPEN and CARE principles implementation (effective db v.2),</li> <li>the introduction of 'SahulSed' replacing 'OSL/TL Australia' (effective v.2(1)),</li> <li>the integration of the 'FosSahul' partner collection (effective v.2(1)),</li> <li>the integration of the 'ExpAge' partner collection (effective v.2(1)),</li> <li>major upgrades to the 'CRN INT' and 'CRN AUS' collections (effective v.2(2)),</li> <li><em>the integration of the 'SahulArch' collection (v.2.1(2))</em>.</li> </ul> <p>Accompanying publication: Codilean, A. T., Munack, H., Saktura, W. M., Cohen, T. J., Jacobs, Z., Ulm, S., Hesse, P. P., Heyman, J., Peters, K. J., Williams, A. N., Saktura, R. B. K., Rui, X., Chishiro-Dennelly, K., and Panta, A.: OCTOPUS database (v.2), Earth Syst. Sci. Data, 14, 3695–3713, <a href="https://doi.org/10.5194/essd-14-3695-2022" target="_blank" rel="noopener">https://doi.org/10.5194/essd-14-3695-2022</a>, 2022.</p>
OCTOPUS database v.2.2 (SQL database dump)
<p>A full dump of <strong>OCTOPUS PostgreSQL database v.2.2</strong> as published upon</p> <ul> <li>the integration of the 'SahulChar' collection (v.2.2(1)),</li> <li>the integration of the 'IPPD' collection (v.2.2(1)),</li> <li>major upgrades to the 'CRN INT' and 'CRN AUS' collections (effective v.2.2(3)),</li> <li>upgrades to the 'CRN XXL' and 'CRN UOW' collections (effective v.2.2(3)).</li> </ul> <p>Database frontend: https://octopusdata.org/</p>
Marktstammdatenregister Monthly Database Dump
<p>Since the dataset provided by <a href="https://zenodo.org/record/7387843">Open-MaStR</a> doesn't seem to get updated anymore and only uploads CSV files I will provide a monthly updated database instead.</p> <p>Like the original publication the data set contains all data from the <a href="https://www.marktstammdatenregister.de/MaStR">Marktstammdatenregister</a>.</p> <p>The data set is downloaded and processed with the software <a href="https://github.com/OpenEnergyPlatform/open-MaStR">open-MaStR</a>.</p> <p>If you don't have much experience dealing with databases directly a simple way to interact with the data set is using <a href="https://sqlitebrowser.org/">DB Browser for SQLite</a>. Simply open the file there and you are able to browse all tables included in the dataset. This also allows exporting individual tables as CSV files.</p> <p>License information:</p> <p><a href="https://www.marktstammdatenregister.de/MaStR">Marktstammdatenregister</a> - © Bundesnetzagentur für Elektrizität, Gas, Telekommunikation, Post und Eisenbahnen | <a href="https://www.govdata.de/dl-de/by-2-0">DL-DE-BY-2.0</a></p>
Dump of RDF dataset used by PO for a Graph Database benchmark, 2022
<p>This dataset represents a newer version of the NQUADS files in RDF from Publication Offices used for benchmarking graph databases. </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.