Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

110

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

110 results for “schema”

Learn how ShareScore rates datasets ↗
zenodo48/100

The Red Queen in the Repository: metadata quality in an ever-changing environment (preprint of paper, presentation slides and dataset collection with validation schemas to IDCC2019 conference paper)

<p>This fileset contains a preprint version of the conference paper (.pdf), presentation slides (as .pptx) and the dataset(s) and validation schema(s) for the IDCC 2019 (Melbourne) conference paper: <em>The Red Queen in the Repository: metadata quality in an ever-changing environment. </em>Datasets and schemas are&nbsp; in .xml, .xsd , Excel (.xlsx) and .csv&nbsp; (two files representing two different sheets in the .xslx -file). The <em>validationSchemas.zip</em> holds the additional validation schemas (.xsd), that were not found in the schemaLocations of the metadata xml-files to be validated. The schemas must all be placed in the same folder, and are to be used for validating the Dataverse <em>dcterms</em> records (with <em>metadataDCT.xsd</em>) and the Zenodo <em>oai_datacite</em> feeds respectively (<em>schema.datacite.org_oai_oai-1.0_oai.xsd</em>). In the latter case, a simpler way of doing it might be to replace the incorrect URL &quot;<em>http://schema.datacite.org/oai/oai-1.0/ oai_datacite.xsd</em>&quot; in the <em>schemaLocation </em>of these xml-files by the CORRECT:&nbsp; <em>schemaLocation=&quot;http://schema.datacite.org/oai/oai-1.0/ http://schema.datacite.org/oai/oai-1.0/oai.xsd&quot;</em>&nbsp; as has been done already in the sample files here. The sample file folders <em>testDVNcoll.zip </em>(Dataverse), <em>testFigColl.zip </em>(Figshare)<em> </em>and <em>testZenColl.zip </em>(Zenodo)<em> </em>contain all the metadata files tested and validated that are registered in the spreadsheet with objectIDs.<br> In the case of Zenodo, one original file feed,<br> <em>zen2018oai_datacite3orig-https%20_zenodo.org_oai2d%20verb=ListRecords%26metadata<br> Prefix=oai_datacite%26from=2018-11-29%26until=2018-11-30.xml</em> ,<br> is also supplied to show what was necessary to change in order to perform validation as indicated in the paper.</p> <p>For Dataverse, a corrected version of a file,<br> <em>dvn2014ddi-27595<strong>Corr</strong>_https%20_dataverse.harvard.edu_api_datasets_export%20<br> exporter=ddi%26persistentId=doi%253A10.7910_DVN_27595<strong>Corr</strong>.xml</em> ,<br> is also supplied in order to show the changes it would take to make the file validate without error.</p>

opencc-by-4.0Feb 2019View details →
zenodo44/100

An Empirical Characterization of Event Sourced Systems and Their Schema Evolution - Lessons from Industry - Accompanying Anonymized Transcripts

<p>Anonymized interviews with 25 engineers on their experience applying Event Sourcing, with accompanying classifications.&nbsp;These transcripts are used in our publication &quot;An Empirical Characterization of Event Sourced Systems and Their Schema Evolution - Lessons from Industry&quot;.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Supplemental material of the Streptococcus pyogenes whole genome MLST schema deposited in Chewie-NS

<p>This supplemental material includes the lists of accession numbers for the Blackwell et al. and NCBI RefSeq assemblies used to populate the whole genome MLST schema for <em>Streptococcus pyogenes</em>, the UniProt identifiers of the reference proteomes used for schema annotation and the set of complete genomes, and associated metadata, used for schema creation.</p> <p>The wgMLST schema was created with <a href="https://github.com/B-UMMI/chewBBACA">chewBBACA</a> and is publicly available at <a href="https://chewbbaca.online/species/1/schemas/1">chewie-NS</a>, where a more detailed description of schema creation, annotation and curation can be found.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Crosswalk between CESSDA Data Catalogue (CDC) Metadata Profile and ECRIN Metadata Schema.

<p>This dataset contains two files: (1) a crosswalk between CESSDA Data Catalogue (CDC) DDI2.5 Metadata Profile (<a href="https://cmv.cessda.eu/profiles/cdc/ddi-2.5/1.0.4/profile.html">https://cmv.cessda.eu/profiles/cdc/ddi-2.5/1.0.4/profile.html</a>)&nbsp; and ECRIN Metadata Schema for Clinical Research Data Objects Version 6.0 (August 2021) (<a href="https://zenodo.org/record/5554961">https://zenodo.org/record/5554961</a>) with an extension to &ldquo;geographical data&rdquo; and (2) vice versa.</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

EvoBench: Benchmarking Schema Evolution in NoSQL

<p>Docker containers for reproducing the proof of concept measurements with our NoSQL Schema Evolution Benchmark.</p>

opengpl-2.0Jun 2021View details →
zenodo40/100

D2.1: Artefact, Contributor, and Organisation Relationship Data Schema - Appendix A

<p>Comparison of metadata schema for ORCID, DataCite, Dublin Core, CASRAI, MODS&nbsp;and DDI&nbsp;regarding contributors, organizations and artefacts.</p>

opencc-zeroSep 2015View details →
zenodo40/100

XML-Schema (Gemein-Nachrichten)

<p>Mit den&nbsp;<em><strong>Gemein-Nachrichten</strong></em>&nbsp;stellt das Unit&auml;tsarchiv Herrnhut der weltweiten Evangelischen Br&uuml;der-Unit&auml;t - Herrnhuter Br&uuml;dergemeine (Unitas Fratrum / Moravian Church) das &auml;lteste und umfangreichste Mitteilungsblatt der Br&uuml;dergemeine digital zur Verf&uuml;gung. Es enth&auml;lt Berichte aus Gemeinden sowie dem Missions- und Diasporawerk der Br&uuml;dergemeine sowie Reden und Lebensl&auml;ufe. Die&nbsp;<em>Gemein-Nachrichten</em>&nbsp;wurden ab 1765 in Fortsetzung des&nbsp;<em>J&uuml;ngerhaus-Diariums</em>&nbsp;(1747-1764) ausschlie&szlig;lich handschriftlich vervielf&auml;ltigt. In Druck gingen 1817 und 1818 die&nbsp;<em>Beytr&auml;ge aus der Br&uuml;der-Gemeine</em>&nbsp;und zwischen 1819 und 1894 die&nbsp;<em>Nachrichten aus der Br&uuml;der-Gemeine</em>. Das Nachrichtenblatt fand in den&nbsp;<em>Mitteilungen aus der Br&uuml;der-Gemeine zur F&ouml;rderung christlicher Gemeinschaft</em>&nbsp;ab 1895 bis 1941 seine Fortsetzung.</p> <p>Mit dem hier vorliegenden <strong>XML-Schema</strong> werden erschlossene Transkripte mit standardisierten Metadaten angereichert.</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

TEI-Schema für das Patristische Textarchiv (PTA)

This dataset contains the TEI-Schema and its documentation

opengpl-3.0Jun 2023View details →
zenodo40/100

An annotated corpus of clinical trial publications supporting schema-based relational information extraction

<p>Repository of an annotated corpus of clinical trial abstracts supporting schema-based relational information extraction and the code for the inter-annotation agreement calculation and the baseline information extraction method.</p>

openother-openMar 2022View details →
zenodo40/100

Mean, standard deviation, and percentiles of the schema and domain scores of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3) in a German opportunity sample (n=1,150)

<p>Mean, standard deviation, and percentiles of the schema and domain scores of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3) in a German opportunity sample (n=1,150). Details are reported in: Kriston L, Sch&auml;fer J, Jacob GA, H&auml;rter M, H&ouml;lzel LP. Reliability and validity of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3). <em>Eur J Psychol Assess</em> 2013; 29: 205-212.</p> <p>IMPORTANT: This is an opportunity sample that is not representative of any well-defined population. Accordingly, the values should not be used as reference or norm values for the German version of the YSQ-S3.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Mittelwert, Standardabweichung und Perzentile der Schema- und Domänen-Scores der deutschen Version des Young Schema Questionnaire - Short Form 3 (YSQ-S3) in einer deutschen Gelegenheitsstichprobe (n=1150)

<p>Mittelwert, Standardabweichung und Perzentile der Schema- und Dom&auml;nen-Scores der deutschen Version des Young Schema Questionnaire - Short Form 3 (YSQ-S3) in einer deutschen Gelegenheitsstichprobe (n=1150). Details sind beschrieben in: Kriston L, Sch&auml;fer J, Jacob GA, H&auml;rter M, H&ouml;lzel LP. Reliability and validity of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3). <em>Eur J Psychol Assess</em> 2013; 29: 205-212.</p> <p>WICHTIG: Es handelt sich um eine Gelegenheitsstichprobe, die f&uuml;r keine gut definierbare Population repr&auml;sentativ ist. Dementsprechend sollten die Werte nicht als Referenz- oder Normwerte f&uuml;r die deutsche Version des YSQ-S3 verwendet werden.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

EFSA cgMLST gene lists for Escherichia coli and Salmonella enterica chewieNS schema

<p>Annex A contains the list of cgMLST loci of <em>Escherichia coli</em> and <em>Salmonella enterica</em> used in the&nbsp;cgMLST analysis in the EFSA One Health WGS System.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Text-fig. 2. E-W cross section of the Urema Graben from Gorongosa to Inhaminga adapted from Flores (1973: fig. 5). Note that in this schema the Mazamba Sandstone directly overlies the Cheringoma Limestone. I.P.CO No. 5 is a bore hole. Vertical exaggeration ×10. in Stratigraphy, Chronology And Palaeontology Of The Tertiary Rocks Of The Cheringoma Plateau, Mozambique

Text-fig. 2. E-W cross section of the Urema Graben from Gorongosa to Inhaminga adapted from Flores (1973: fig. 5). Note that in this schema the Mazamba Sandstone directly overlies the Cheringoma Limestone. I.P.CO No. 5 is a bore hole. Vertical exaggeration ×10.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 1. Schema for a data-integration solution-A Proposed Data Driven Architecture for Cardiology Network Application

<p>Data integration has favored loosening the coupling between data. This may involve<br> providing a uniform query interface over a mediated schema (see figure 1), thus transforming<br> a query into specialized queries over the original databases. One can also term this process<br> &quot;view-based query-answering&quot; because each of the data sources functions as a view over the<br> (nonexistent) mediated schema.</p>

opencc-by-4.0Apr 2010View details →
zenodo40/100

SchemaPile: A Large Collection of Relational Database Schemas

<p>Access to fine-grained schema information is crucial for understanding how relational databases are designed and used in practice, and for building systems that help users interact with them. Furthermore, such information is required as training data to leverage the potential of large language models (LLMs) for improving data preparation, data integration and natural language querying.<br><br>Existing single-table corpora such as GitTables provide insights into how tables are structured in-the-wild, but lack detailed schema information about how tables relate to each other, as well as metadata like data types or integrity constraints. On the other hand, existing multi-table (or database schema) datasets are rather small and attribute-poor, leaving it unclear to what extent they actually represent typical real-world database schemas.&nbsp;</p> <p>In order to address these challenges, we present SchemaPile, a corpus of 221,171 database schemas, extracted from SQL files on GitHub. It contains 1.7 million tables with 10 million column definitions, 700 thousand foreign key relationships, seven million integrity constraints, and data content for more than 340 thousand tables. We conduct an in-depth analysis on the millions of schema metadata properties in our corpus, as well as its highly diverse language and topic distribution. In addition, we showcase the potential of SchemaPile to improve a variety of data management applications, e.g., fine-tuning LLMs for schema-only foreign key detection, improving CSV header detection and evaluating multi-dialect SQL parsers. We publish the code and data for recreating SchemaPile and a permissively licensed subset SchemaPile-Perm.&nbsp;</p> <p><a href="http://schemapile.github.io">http://schemapile.github.io</a></p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

BRAIN Journal-Intelligent Continuous Double Auction method For Service Allocation in Cloud Computing-Figure 1. Resource allocation schema in proposed method

<p>We assume that the resources allocation satisfies the following conditions:<br> &bull; The quantity of a resource can be measured in arbitrary units (e.g. 60 units of resource<br> A).<br> &bull; A resource can be divided into an arbitrary fraction (e.g. a resource of 60 units is divided<br> into 20 units for consumer 1 and 40 units for consumer 2).<br> &bull; A resource request of a service can be divided into sub-requests and acquired from<br> multiple providers (e.g. a resource request of 40 units utilized as 10 units from provider<br> 1 and 30 units from provider 2).<br> Figure 1 shows a cloud computing environment with the proposed mechanism.</p>

opencc-by-4.0Oct 2013View details →
zenodo40/100

Figure 1. Schema of clientside application for semantic browsing.-Browsing Semantic Data in Slovakia

<p>With the aim primary on unstructured information extraction and refining, relationship discovery and visualization, we propose our solution for SBR in the first place. The reason for this is, primary, that HTML formatted results of SBR are very jerky and uncertainty regarding the structure of information is very high. Readers can also be pointed by J. Suchal and P. Vojtek (2009), that care should be taken towards type errors. We discuss that later. In this work, we try to fill&amp;up the gap of visualization and, somehow limited data access offered by SBR, adapting to the problems disclaimed above. We suggest a new client&amp; side paradigm, which does not depend on a particular website like foaf.sk. Figure 1 describes the schema briefly and the key elements are parsers with other tools on the top and structured formats, for datastore, on the bottom.</p>

opencc-by-4.0Nov 2015View details →
zenodo40/100

Figure 3. A brief schema of the CoDOA-SVM approach-Cognitive Development Optimization Algorithm Based Support Vector Machines for Determining Diabetes

<p>In the Equation 23, TP stands for true classified diabetes positive individuals; TN stands for true classified diabetes negative individuals; FP stands for false classified diabetes positive individuals and finally, FN stands for false classified diabetes negative individuals. &bull; After determining good (optimum) particles, default CoDOA steps are run. &bull; After achieving the total iteration number, it is allowed to train the SVM via optimum Gauss (RBF) kernel function parameters, by using the optimum particle value [sigma (&sigma;) value]. &bull; The trained SVM is now ready for the classification and so is diabetes determination process.<br> A brief schema of the CoDOA-SVM approach is also provided in Figure 3.</p>

opencc-by-4.0Jan 2016View details →
zenodo40/100

Text-fig. 5. Original material of Peziza sulphurea (syntype L 910,261-594), the substrate and schema with location of six apothecia and two fragments of apothecia. Apothecia marked with numbers 1–3 were examined microscopically. in A Revision Of Trichopeziza Lizonii, T. Sulphurea And T. Violascens (Ascomycota, Helotiales) From The Herbarium Prm With Notes On Type Material Of Peziza Sulphurea

Text-fig. 5. Original material of Peziza sulphurea (syntype L 910,261-594), the substrate and schema with location of six apothecia and two fragments of apothecia. Apothecia marked with numbers 1–3 were examined microscopically.

opencc-by-4.0Sep 2013View details →
zenodo40/100

Text-fig. 3. Schema of radial section of G. rudolphii (sample 99/04). t – tracheid, r – ray, bp – bordered pit, tp – taxodioid cross-field pit, gp – glyptostroboid cross-field pit. in New Fossil Woods From The Paleogene Of Doupovské Hory And České Středohoří Mts. (Bohemian Massif, Czech Republic)

Text-fig. 3. Schema of radial section of G. rudolphii (sample 99/04). t – tracheid, r – ray, bp – bordered pit, tp – taxodioid cross-field pit, gp – glyptostroboid cross-field pit.

opencc-by-4.0Dec 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record