Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3
datasets available to search
ShareScore release 0.9.0
Dataset results
3 results for “Schema.org”
Listing of data repositories that embed schema.org metadata in dataset landing pages
<p>Machine-readable metadata available from landing pages for datasets facilitate data citation by enabling easy integration with reference managers and other tools used in a data citation workflow. Embedding these metadata using the schema.org standard with the JSON-LD is emerging as the community standard. This dataset is a listing of data repositories that have implemented this approach or are in the progress of doing so.</p> <p>This is the first version of this dataset and was generated via community consultation. We expect to update this dataset, as an increasing number of data repositories adopt this approach, and we hope to see this information added to registries of data repositories such as re3data and FAIRsharing.</p> <p>In addition to the listing of data repositories we provide information of the schema.org properties supported by these data repositories, focussing on the required and recommended properties from the "Data Citation Roadmap for Scholarly Data Repositories".</p>
Schema.org mark-up data for named entities
<p>This dataset contains two files: original_data.zip, and website_5folds.zip</p> <p><strong>original_data.zip </strong>will unpack into three .csv files, Place.csv, CreativeWork.csv, and LocalBusiness.csv. Each file contains one entity on each row, and this entity belongs to a subclass of the class indicated by the file name. There are 8 columns:</p> <ul> <li>the first 2 columns are simply the index of the row</li> <li>description_t: the long textual description of the entity</li> <li>schemaorg_class: the schema.org class assigned to the entity</li> <li>name_tpage_domain: always empty</li> <li>name_t: the name of the entity</li> <li>page_domain: the website where the entity mark-up data is found</li> <li>label: an index for the schemaorg_class</li> <li>description: this is the name of the entity (name_t) plus the first sentence of its description (from description_t)</li> </ul> <p><strong>website_5folds.zip</strong> is a transformation of the original_data.zip. It unzips into three folders, Place, LocalBusiness, and CreativeWork. Inside each folder, there are five folders: 0, 1, 2, 3 and 4 indicating five folds. Inside each of the numbered sub-folder there is a train.csv and test.csv file. Then each csv file contains one entity on each row, with the following columns:</p> <ul> <li>the first column is simply the index of the row</li> <li>schemaorg_class: the schema.org class assigned to the entity</li> <li>name_t: the name of the entity</li> <li>description: this is the name of the entity (name_t) plus the first sentence of its description (from description_t)</li> <li>page_domain: the name of the entity plus the processed domain name. The process includes parsing the domain URL, extract the host name, applying word segmentation (tescobank -> tesco bank), and removing stopwords and TLDs (co, uk, com, fr)</li> </ul> <p>As mentioned, website_5folds.zip is a transformation of the original_data.zip and in fact contains multiple replications of original_data.zip. It is created for 5 fold validation experiment while ensuring that there are no overlap in the page_domain of entities in training and test sets. </p>
Schema.org Characteristic Sets computed from the JSON-LD subset of Web Data Commons dataset (October 2021 release)
<p>This dataset reports the computation of Characteristic Sets from the JSON-LD subset of Web Data Commons dataset (October 2021 release). Each row consists in a combination of Schema.org properties and its cardinality. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.