Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
468
datasets available to search
ShareScore release 0.9.0
Dataset results
468 results for “dump”
OER World Map data dump
<p>This is a data dump in CSV and JSON format from the OER World Map as of 2022-04-29, the date of its shutdown. It comprises 6615 metadata records which utilize the schema.org vocabulary to describe these kind of things:</p> <p> </p> <ul> <li>organizations (type "Organization")</li> <li>services (type "Service")</li> <li>persons (type "Person)</li> <li>projects (type "Action")</li> <li>events (type "Event")</li> <li>stories (type "Article")</li> <li>tools (type "Product")</li> <li>publications (type "WebPage")</li> <li>policies (type "Policy")</li> </ul> <p>The OER Worldmap was hosted and developed by the North Rhine-Westphalian Library Service Centre (hbz) from 2014 to 2022. See for more background information:</p> <ul> <li>Project blog: <a href="https://oerworldmap.wordpress.com">https://oerworldmap.wordpress.com</a></li> <li>Software repos & issue tracking: <a href="https://github.com/hbz/oerworldmap">https://github.com/hbz/oerworldmap</a> (backend), <a href="https://github.com/hbz/oerworldmap-ui">https://github.com/hbz/oerworldmap-ui</a> (frontend)</li> <li>Twitter: <a href="https://twitter.com/oerworldmap/">@oerworldmap</a></li> </ul>
Extracted external domains from Wikipedia dump - 20/03/2024
<p>This dataset contains 6,459,779 distinct domains derived from the external links section of Wikipedia pages.</p> <p>The external links section of a page such as <a href="https://en.wikipedia.org/wiki/OpenWeb" target="_blank" rel="noopener">OpenWeb</a> contains only one <a href="https://www.openweb.com/" target="_blank" rel="noopener">link</a>.</p> <p>The primary objective of assembling this dataset is to improve content prioritization and filtering in web crawling techniques.</p> <p>The dataset is structured as a text file, with each line representing a distinct domain.<br><br><br></p>
Helmholtz Knowledge Graph: RDF data dump
<p>Under this base DOI we regularly publish full RDF data dumps of the Helmholtz-Knowledge Graph.<br>Data dumps are typical associated with major releases, or major data updates.</p> <p>Dumps are serialized in .ttl and compressed with gzip.<br><br>For more information on deployment and documentation as well as data access see:<br>Search UI: https://search.unhide.helmholtz-metadaten.de/<br>SPARQL endpoint: https://sparql.unhide.helmholtz-metadaten.de/<br>Documentation: https://docs.unhide.helmholtz-metadaten.de/<br>Software: https://codebase.helmholtz.cloud/hmc/hmc-public/unhide</p>
Рис. 3. ÀенΑрограмма биоценотического схоΑства зоопΛанктона техногенных воΑоемов: 4–6 — ШерΛовогорское месторожΑение:4 — ШГ-10 — карьерное озеро, 5 — ШГ-8 — озеро поΑ отваΛами руΑного карьера, 6 — ШГ-9 — поΑпруΑное озеро у пгт. ШерΛовая Гора;7–8 — ОрΛовское месторожΑение: 7 — ОР-1, ОР-3 — хвостохраниΛище, 8 — ОР-7 — озеро ниже хвостохраниΛища; 9 — МаΛокуΛунΑинское месторожΑение: МК-2 — поΑпруΑное озеро р. МаΛая КуΛинΑа; 10 — Спокойнинское месторожΑение: ОР-8 — хвостохраниΛище; 11 — Жипкошинское месторожΑение: ЖП-2 — карьер Fig. 3. Dendrogram of zooplankton biocenotic similarity in technogenic reservoirs: 4–6 — Sherlovogorskoye deposit: 4 — ShG-10 pit lake, 5 — ShG-8, a lake under the dumps of an ore quarry, 6 — ShG-9 dammed lake near the village of Sherlovaya Gora;7 –8 — Orlovskoye deposit: O R-1, OR-3 — tailing dump, OR-7— lake below the tailing dump; 9 — Malokulundinskoye deposit: MK-2 — dammed lake on the Malaya Kulinda River; 10 — Spokoininskoye deposit: OR-8 — tailing dump; 11 — Zhipkoshinskoye deposit; ZhP-2 — pit lake in Zooplankton species diversity in technogenic reservoirs of the Southeastern Transbaikalia
Рис. 3. ÀенΑрограмма биоценотического схоΑства зоопΛанктона техногенных воΑоемов: 4–6 — ШерΛовогорское месторожΑение:4 — ШГ-10 — карьерное озеро, 5 — ШГ-8 — озеро поΑ отваΛами руΑного карьера, 6 — ШГ-9 — поΑпруΑное озеро у пгт. ШерΛовая Гора;7–8 — ОрΛовское месторожΑение: 7 — ОР-1, ОР-3 — хвостохраниΛище, 8 — ОР-7 — озеро ниже хвостохраниΛища; 9 — МаΛокуΛунΑинское месторожΑение: МК-2 — поΑпруΑное озеро р. МаΛая КуΛинΑа; 10 — Спокойнинское месторожΑение: ОР-8 — хвостохраниΛище; 11 — Жипкошинское месторожΑение: ЖП-2 — карьер Fig. 3. Dendrogram of zooplankton biocenotic similarity in technogenic reservoirs: 4–6 — Sherlovogorskoye deposit: 4 — ShG-10 pit lake, 5 — ShG-8, a lake under the dumps of an ore quarry, 6 — ShG-9 dammed lake near the village of Sherlovaya Gora;7 –8 — Orlovskoye deposit: O R-1, OR-3 — tailing dump, OR-7— lake below the tailing dump; 9 — Malokulundinskoye deposit: MK-2 — dammed lake on the Malaya Kulinda River; 10 — Spokoininskoye deposit: OR-8 — tailing dump; 11 — Zhipkoshinskoye deposit; ZhP-2 — pit lake
Wikidata dump 2017-12-27
<p>Wikidata dump retrieved from <a href="https://www.google.com/url?q=https://dumps.wikimedia.org/wikidatawiki/entities/latest-all.json.bz2&sa=D&ust=1522795489326000&usg=AFQjCNFKS4q0PRJ3VfsGUVQjno2fN9otlQ">https://dumps.wikimedia.org/wikidatawiki/entities/latest-all.json.bz2</a> on 27 Dec 2017</p>
20th Century Press Archives JSON-LD dump for CdV 2018 Rhein-Main: persons and companies
<p>Folder metadata for all person and company folders of PM20, which have publicly accessible documents. Published for the "Coding da Vinci" Hackathon 2018.</p> <p>For a preview and further information, please see https://github.com/zbw/cdv2018-pressemappe20 (mostly in German)</p> <p> </p>
Figure 6 in Rove beetle communities (Coleoptera: Staphylinidae) in the rock dumps after coal mining
Figure 6. Factors "Temperature" (A), "Year" (B) and environmental factors (C) contribution into rove beetles abundance on the study sites.
Figure 5 in Rove beetle communities (Coleoptera: Staphylinidae) in the rock dumps after coal mining
Figure 5. Dynamic density of rove beetles (A) and their dominant subfamilies (B) on the dumps of the Kedrovsky coal mine (mean ± SD).
Figure 3 in Rove beetle communities (Coleoptera: Staphylinidae) in the rock dumps after coal mining
Figure 3. Rank distribution of Staphylinidae species (the rank of species is along the abscissa axis; abundance, % is along the ordinate axis) in the rock dumps of the Kedrovsky coal mine for the entire period of research.
Figure 2 in Rove beetle communities (Coleoptera: Staphylinidae) in the rock dumps after coal mining
Figure 2. Numerical characteristics of some taxonomic categories of rove beetles in the studied area of the Kedrovsky coal mine (general number).
Figure 4 in Rove beetle communities (Coleoptera: Staphylinidae) in the rock dumps after coal mining
Figure 4. Similarity (according to Jaccard, IJ) of the rove beetle population in the sites of the Kedrovsky coal mine (the number at the nodes mean the bootstrap confidence intervals obtained based on 999 iterations are indicated).
Figure 6 in Diversity of ground-dwelling arthropods on overburden dumps after coal mining
Figure 6. The ratio of beetle families in the study site (with the exception of Carabidae and Staphylinidae).
Figure 5 in Species composition and ecological structure of ground beetle communities (Coleoptera, Carabidae) in reclaimed rock dumps in the south of Western Siberia
Figure 5. Ratio of species (first columns) and numerical (second columns) abundance of ground beetle trophic classes, %
MiMoTextBase RDF Dump
<p>This is the RDF-Dump of the knowledge graph "<a href="https://data.mimotext.uni-trier.de/wiki/Main_Page">MiMoTextBase</a>" created within the project "Mining and Modeling Text" (2019-2023, <a href="https://mimotext.uni-trier.de/">MiMoText</a>) at the University Trier. As the project focusses on French Enlightenment novels, the MiMoTextBase - a wikibase instance - contains items on novels and their respective authors of the years 1751-1800. The RDF-Dump is a copy of all entries within the MiMoTextBase and will be updated as the graph grows.</p>
DODO neo4j dump
<p>Neo4j dump of the public instance of DODO that can be loaded into neo4j/docker instance. </p> <p>See <a href="https://github.com/Elysheba/DODO">Elysheba/DODO</a> for more info.</p>
FDup deduplication software data benchmark: 10Mi OpenAIRE Publications Dump
<p>This dataset is a random subset of publications extracted from the OpenAIRE Research Graph (<a href="http://doi.org/10.5281/zenodo.4707307">http://doi.org/10.5281/zenodo.4707307</a>). The dataset contains ~10Mi JSON publications records. </p> <p>The file is a zip archive containing gz files, each with one JSON per line. Each JSON is compliant to the schema available at <a href="http://doi.org/10.5281/zenodo.4723403">http://doi.org/10.5281/zenodo.4723403</a>.<br> <br> Learn more about the OpenAIRE Research Graph at <a href="https://graph.openaire.eu/">https://graph.openaire.eu</a>.</p>
HackerNews Data Dump 2022-10-28
<p>A SQLITE database with every hackernews story, comment, and other item type compressed using the zstd library.</p> <p>Created like this:</p> <pre><code class="language-bash">tar cf hn.db.zst -I zstd hn.db</code></pre> <p>To uncompress run this:</p> <pre><code class="language-bash">unzstd hn.db.zst</code></pre> <p> </p>
Wikidata subset from 2018 dumps created with WDSub at Biohackathon 2022 - Turtle format
<p>Subset of Wikidata obtained using this Shape Expression: https://github.com/kg-subsetting/datasets-biohackathon2022/blob/main/GeneWiki/GeneWiki.shex</p> <p>And the wdsub tool version 0.0.28: https://github.com/weso/wdsub</p> <p>The input dump is: wikidata-20180115-all</p> <p>And the dumpformat is TURTLE</p>
Simple dataset obtained as a Wikidata subset from 2022 dump using entity schema about taxon
<p>The subset has been obtained using wdsub version 0.0.33 and the schema:</p> <p> </p> <pre>PREFIX p: <http://www.wikidata.org/prop/> PREFIX ps: <http://www.wikidata.org/prop/statement/> PREFIX prov: <http://www.w3.org/ns/prov#> PREFIX wd: <http://www.wikidata.org/entity/> PREFIX wdt: <http://www.wikidata.org/prop/direct/> start = @<taxon_by_wd_ontology> OR @<taxon_by_identifier> <taxon_by_wd_ontology> { wdt:P31 [wd:Q16521] ; } <taxon_by_identifier> {wdt:P685 . +;} OR # NCBI taxonomy ID {wdt:P846 . +;} OR # GBIF taxon ID {wdt:P3151 . +;} OR # iNaturalist taxon ID {wdt:P3444 . +;} # eBird taxon ID</pre>
DBLP-QuAD DBLP Dump
<p>This is the RDF dump of DBLP released on August 1, 2022. The DBLP RDF dump is published to allow fair and replicable evaluation of KGQA systems with the <a href="https://zenodo.org/record/7554379">DBLP-QuAD dataset</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.