Zenodo Open Metadata snapshot - Training dataset for records and communities classifier building
<p>This dataset contains Zenodo's published open access records and communities metadata, including entries marked by the Zenodo staff as spam and deleted.</p> <p>The datasets are gzipped compressed JSON-lines files, where each line is a JSON object representation of a Zenodo record or community.</p> <p><strong>Records dataset</strong></p> <p>Filename:<strong> </strong>zenodo_open_metadata_{ date of export }.jsonl.gz</p> <p>Each object contains the terms: <em>part_of, thesis, description, doi, meeting, imprint, references, recid, alternate_identifiers, resource_type, journal, related_identifiers, title, subjects, notes, creators, communities, access_right, keywords, contributors, publication_date</em></p> <p>which correspond to the fields with the same name available in Zenodo's record JSON Schema at <a href="https://zenodo.org/schemas/records/record-v1.0.0.json">https://zenodo.org/schemas/records/record-v1.0.0.json</a>.</p> <p>In addition, some terms have been altered:</p> <ul> <li>The term <strong>files</strong> contains a list of dictionaries containing <strong>filetype</strong>, <strong>size,</strong> and <strong>filename </strong>only.</li> <li>The term <strong>license</strong> contains a short Zenodo ID of the license (e.g. "cc-by").</li> </ul> <p><strong>Communities dataset</strong></p> <p>Filename:<strong> </strong>zenodo_community_metadata_{ date of export }.jsonl.gz</p> <p>Each object contains the terms: <em>id, title, description, curation_policy, page </em></p> <p>which correspond to the fields with the same name available in Zenodo's community creation form.</p> <p><strong>Notes for all datasets</strong></p> <p>For each object the term <strong>spam</strong> contains a boolean value, determining whether a given record/community was marked as spam content by Zenodo staff.</p> <p>Some values for the top-level terms, which were missing in the metadata may contain a <strong>null</strong> value.</p> <p>A smaller uncompressed random sample of 200 JSON lines is also included for each dataset to test and get familiar with the format without having to download the entire dataset.</p>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0