Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
71
datasets available to search
ShareScore release 0.9.0
Dataset results
71 results for “json”
A Quantum-Chemical Bonding Database for Solid-State Materials (JSONS: Part 1)
<p>This database consists of bonding data computed using Lobster for 1520 solid-state compounds consisting of insulators and semiconductors. It consists of two kinds of JSON files. Smaller lightweight JSONS consists of summarized bonding information for each of the compounds. The files are named as per ID numbers in the materials project database. </p><p>Here, we also provide the larger computational data JSON files for 700 compounds. This file consists of all important LOBSTER computation output file data stored as a dictionary.</p><p>Rest 820 computational data JSONs are provided as part of the following repository: A Quantum-Chemical Bonding Database for Solid-State Materials (Part 2): <a href="https://doi.org/10.5281/zenodo.8092187">https://doi.org/10.5281/zenodo.8092187.</a> </p><p>This dataset is published as part of our publication: <a href="https://www.nature.com/articles/s41597-023-02477-5">A Quantum-Chemical Bonding Database for Solid-State Materials. </a> Details about the data generation, validation, and metadata description can be found in our publication.</p><p>Additionally, all the scripts and tools used for curating (including metadata description), benchmarking, and reusing our data are documented in the openly accessible repository. This enables one to reproduce the data and results presented in our work fully. These scripts can be accessed from either of the following links:</p><p>Zenodo: <a href="https://doi.org/10.5281/zenodo.8172527">https://doi.org/10.5281/zenodo.8172527</a></p><p>Github: <a href="https://github.com/naik-aakash/lobster-database-paper-analysis-scripts/tree/v1.0.6">https://github.com/naik-aakash/lobster-database-paper-analysis-scripts/ (v1.0.6)</a></p><p>The dataset will also be available through <a href="https://next-gen.materialsproject.org/">The Materials Project</a> soon.</p>
A Quantum-Chemical Bonding Database for Solid-State Materials (JSONS: Part 2)
<p>This database consists of bonding data computed using Lobster for 1520 solid-state compounds consisting of insulators and semiconductors. The files are named as per ID numbers in the materials project database. </p><p>Here, we provide the larger computational data JSON files for the rest of the 820 compounds. This file consists of all important LOBSTER computation output file data stored as a dictionary.</p><p>This dataset is published as part of our publication: <a href="https://www.nature.com/articles/s41597-023-02477-5">A Quantum-Chemical Bonding Database for Solid-State Materials. </a>Details about the data generation, validation, and metadata description can be found in our publication.</p><p>Additionally, all the scripts and tools used for curating (including metadata description), benchmarking, and reusing our data are documented in the openly accessible repository. This enables one to reproduce the data and results presented in our work fully. These scripts can be accessed from either of the following links:</p><p>Zenodo: <a href="https://doi.org/10.5281/zenodo.8172527">https://doi.org/10.5281/zenodo.8172527</a></p><p>Github: <a href="https://github.com/naik-aakash/lobster-database-paper-analysis-scripts/tree/v1.0.6">https://github.com/naik-aakash/lobster-database-paper-analysis-scripts/ (v1.0.6)</a></p><p>The dataset will also be available through <a href="https://next-gen.materialsproject.org/">The Materials Project</a> soon.</p>
EBSCO articles dataset (domain knowledge: rehabilitation medicine) + JSON of every article
<p><strong>EBSCO articles dataset (domain knowledge: rehabilitation medicine) + JSON of every article.</strong></p> <p><em>EBSCO articles dataset (parsed into JSON).zip</em> - an archive with structuring of each article in JSON format.</p> <p>This study would not have been possible without the financial support of the National Research Foundation of Ukraine (FundRef ID: 10.13039/100018227).</p> <p>Our work was funded by Grant contracts:</p> <p>– “Transdisciplinary intelligent information and analytical system for the re-habilitation processes support in a pandemic (TISP)”, application ID: 2020.01/0245 (2020–2021, project was success-fully completed);</p> <p>– “Development of the cloud-based platform for patient-centered telerehabilitation of oncology patients with mathematical-related modeling”, application ID: 2021.01/0136 (2022–2023, project is still in progress).</p>
Small version of other JSON benchmarking datasets
<p>10% size version of bestbuy, google, twitter, and walmart datasets used for benchmarking JSON engines.<br> <br> Full versions:<br> <br> https://zenodo.org/record/7607865</p> <p>https://zenodo.org/record/7607889</p> <p>https://zenodo.org/record/7607891</p> <p>https://zenodo.org/record/7607882</p>
Example FAIRtracks JSON document - augmented
<p><strong>Background</strong></p> <p>Many types of data from genomic analyses can be represented as genomic tracks, i.e. features linked to the genomic coordinates of a reference genome. Examples of such data are epigenetic DNA methylation data, ChIP-seq peaks, germline or somatic DNA variants, or RNA-seq expression levels. Researchers often face difficulties in locating, accessing and combining relevant tracks from external sources, as well as locating the raw data, reducing the value of the generated information. </p> <p><strong>FAIRtracks software ecosystem</strong></p> <p>We have, as an output of the ELIXIR Implementation Study "FAIRification of Genomic Tracks", developed a basic set of recommendations for genomic track metadata together with an implementation called FAIRtracks in the form of a JSON Schema. We propose FAIRtracks as a draft standard for genomic track metadata in order to advance the application of FAIR data principles (Findable, Accessible, Interoperable, and Reusable). We have demonstrated practical usage of this approach by designing a software ecosystem around the FAIRtracks draft standard, integrating globally identifiable metadata from various track hubs in the Track Hub Registry and other relevant repositories into a novel track search service, called TrackFind. The software ecosystem also includes the FAIRtracks augmentation service, which assists metadata producers by automatically augmenting minimal machine-readable metadata with their human-readable counterparts, as well as the FAIRtracks validation service, which extends basic JSON Schema validation to include FAIR-related features (global identifiers, ontology terms, and object references). Finally, we have implemented track metadata search and import functionality into relevant analytical tools: EPICO and the GSuite HyperBrowser. For an overview of the FAIRtracks software ecosystem, please visit: <a href="http://fairtracks.github.io/">http://fairtracks.github.io/</a></p> <p><strong>Example FAIRtracks JSON document - augmented</strong></p> <p>The "Example FAIRtracks JSON document - augmented" is generated as part of the build process of the FAIRtracks draft standard JSON Schema (source code: <a href="https://github.com/fairtracks/fairtracks_standard/">https://github.com/fairtracks/fairtracks_standard/</a>). The example FAIRtracks document contains a small selection of tracks and objects from the ENCODE project metadata (<a href="https://www.encodeproject.org/">https://www.encodeproject.org/</a>), adapted to align with the FAIRtracks draft standard. In addition to being available in the above-mentioned GitHub repository, the "Example FAIRtracks JSON document - augmented" is also published here on Zenodo in order for the document to be globally uniquely identifiable by a Digital Object Identifier (DOI).</p>
Wikidata dump from 2018-12-17 in JSON
<p>This is a dump from Wikidata from 2018-12-17 in JSON. This one is not avavailable anymore from Wikidata. It was downloaded originally from <a href="https://dumps.wikimedia.org/other/wikidata/20181217.json.gz">https://dumps.wikimedia.org/other/wikidata/20181217.json.gz</a> and recompressed to fit on Zenodo.</p>
Europe Pubmed Central Lite metadata (JSON object stream)
<p>The raw JSON object stream associated with https://zenodo.org/deposit/107781/</p>
Json file from Twitter API used for benchmarking Jsonpath
<p>A JSON file used as an example to illustrate queries and to benchmark some tool.</p>
Json file from Twitter API used for benchmarking Jsonpath
<p>A JSON file used as an example in several software (such as SIMDJson) to illustrate queries and to benchmark some tools.</p>
Dataset for interface calculations as BSON mongodump and JSON formats for the publication: "High-throughput generation of potential energy surfaces for solid interfaces"
<p>This dataset that contains the results presented in the journal article entitled "High-throughput generation of potential energy surfaces for solid interfaces" that was published in Volume 207 of the Elsevier journal Computational Materials Science.</p> <p>The dataset consists of a dump of a MongoDB database with a single collection that contains data on 6 solid interfaces including the generalized stacking fault energies, corrugation, interface distances and adhesion sites for the film and the substrate at minimum and maximum adhesion energy configurations, and images of the full potential energy surface (PES).</p> <p>The data is served in two different formats; a BSON mongodump folder that can be restored to a MongoDB instance using the mongorestore tool, and additionally as simple .json files. The contents are identical and the users are encouraged to choose the format that is convenient for them.</p>
GeneSeqToFamily: Gene JSON
<p>Gene duplication is a major factor contributing to evolutionary novelty, and the contraction or expansion of gene families has often been associated with morphological, physiological and environmental adaptations. The study of homologous genes helps us to understand the evolution of gene families. It plays a vital role in finding ancestral gene duplication events as well as identifying genes that have diverged from a common ancestor under positive selection. There are various tools available, such as MSOAR, OrthoMCL and HomoloGene, to identify gene families and visualise syntenic information between species, providing an overview of syntenic regions evolution at the family level. Unfortunately, none of them provide information about structural changes within genes, such as the conservation of ancestral exon boundaries amongst multiple genomes. The Ensembl GeneTrees computational pipeline generates gene trees based on coding sequences and provides details about exon conservation, and is used in the Ensembl Compara project to discover gene families. </p>
User Generated Content (EOL v2): user activity collections (json format)
<p></p>https://eol-jira.bibalex.org/browse/DATA-1780 Data as of Oct 28, 2018 For questions or use cases calling for large, multi-use aggregate data files, please visit the EOL Services forum at <p></p>http://discuss.eol.org/c/eol-services
National Statistics Post-code Lookup: Json file for benchmarking purpose
<p>Dataset used in several JsonPath engines</p>
CSV Dataset Files and JSON OpenRefine Recipes for Alignment of the Schoenberg Dataset of Manuscripts (SDBM) Name Authority with Wikidata
<p>Dataset CSVs and JSON recipe files for OpenRefine for a project to align Name Authority records in the Schoenberg Dataset of Manuscripts (SDBM) with Wikidata Items</p>
CSV and JSON data describing the quantity and content of uploads to Thingiverse for 2015-2020
<p>The COVID-19 pandemic profoundly affected various aspects of daily life, particularly the supply and demand of essential goods, resulting in critical shortages. This included personal protective equipment (PPE) for medical professionals and the general public. To address these shortages, online "maker communities" emerged, aiming to develop and locally manufacture critical products. While some organized efforts existed, the majority of initiatives originated from individuals and groups on platforms like Thingiverse. This paper presents a longitudinal analysis of Thingiverse, one of the largest maker community websites, to examine the pandemic's effects. Our findings reveal a surge in community output during the initial lockdown periods in major contributing nations (primarily those in the western-hemisphere), followed by a subsequent decline. Additionally, throughout 2020, pandemic-related products dominated uploads and interactions during this period. Based on these observations, we propose recommendations to expedite the community's ability to support local, national, and international responses to future disasters.</p>
Five star movie ratings form the MovieLens 25M dataset, grouped by user id, in JSON format
<p>From the MovieLens 25M dataset, I have extracted the five star ratings and grouped them by user ID. The original source files can be found here: </p> <p>https://grouplens.org/datasets/movielens/</p> <p> </p>
CSV and JSON data describing the quantity and content of uploads to Thingiverse for 2015-2020
Open the record for dataset details and reuse information.
Wikidata subset with revision history information [JSON]
<p>This dataset consists the complete revision history of every instance of the 100 most important classes in Wikidata. It contains 9.3 million classes and around 450 million revisions made to those classes. This dataset was exported from a MongoDB database. After decompressing the files, the resulting JSON files can be imported into MongoDB using the following commands:</p> <pre><code>mongoimport --db=db_name --collection=wd_entities --file=wd_entities.json mongoimport --db=db_name --collection=wd_revisions --file=wd_revisions.json </code></pre> <p>Make sure that <em>db_name</em> is replaced by the database where this data will be imported.</p> <p>Documents within the <em>wd_entities</em> collection have the following schema:</p> <ul> <li><strong>id</strong>: Internal id of the entity used by Wikidata (e.g. 8195238).</li> <li><strong>entity_id</strong>: Public id of the entity in Wikidata (e.g. 'Q42')</li> <li><strong>class_ids</strong>: List of classes that the entity belongs to (e.g. ['Q5', 'Q100'])</li> <li><strong>entity_json</strong>: JSON contents of the entity, following Wikidata's JSON data model (https://doc.wikimedia.org/Wikibase/master/php/md_docs_topics_json.html).</li> </ul> <p>Documents within the <em>wd_revisions</em> collection have the following schema:</p> <ul> <li><strong>id</strong>: Identifier of the revision (e.g. 15921539)</li> <li><strong>entity_id</strong>: Public id of the entity in Wikidata affected by this revision (e.g. 'Q42')</li> <li><strong>class_ids</strong>: List of classes that the entity affected by this revision belongs to (e.g. ['Q5', 'Q100'])</li> <li><strong>parent_id</strong>: Identifier of the previous revision to this one, if it exists (e.g. 15921214)</li> <li><strong>timestamp</strong>: Date where the revision was made, following the ISO 8601 format (e.g. +2019-05-27T09:31:10Z)</li> <li><strong>username</strong>: Username of the user that made the revision.</li> <li><strong>comment</strong>: Comments made by the user in the revision, if any.</li> <li><strong>entity_diff</strong>: List of operations made in this revision, following the JSON Patch format.</li> </ul>
MongoDB JSON Logs
<p>A sample of JSON log events from a MongoDB instance. The logs were generated by repeatedly running YCSB workloads A-E.</p>
huARdb Database V2 JSON Files TCR Partition 7 [P-S]
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.