Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
110
datasets available to search
ShareScore release 0.9.0
Dataset results
110 results for “schema”
The Red Queen in the Repository: metadata quality in an ever-changing environment (preprint of paper, presentation slides and dataset collection with validation schemas to IDCC2019 conference paper)
<p>This fileset contains a preprint version of the conference paper (.pdf), presentation slides (as .pptx) and the dataset(s) and validation schema(s) for the IDCC 2019 (Melbourne) conference paper: <em>The Red Queen in the Repository: metadata quality in an ever-changing environment. </em>Datasets and schemas are in .xml, .xsd , Excel (.xlsx) and .csv (two files representing two different sheets in the .xslx -file). The <em>validationSchemas.zip</em> holds the additional validation schemas (.xsd), that were not found in the schemaLocations of the metadata xml-files to be validated. The schemas must all be placed in the same folder, and are to be used for validating the Dataverse <em>dcterms</em> records (with <em>metadataDCT.xsd</em>) and the Zenodo <em>oai_datacite</em> feeds respectively (<em>schema.datacite.org_oai_oai-1.0_oai.xsd</em>). In the latter case, a simpler way of doing it might be to replace the incorrect URL "<em>http://schema.datacite.org/oai/oai-1.0/ oai_datacite.xsd</em>" in the <em>schemaLocation </em>of these xml-files by the CORRECT: <em>schemaLocation="http://schema.datacite.org/oai/oai-1.0/ http://schema.datacite.org/oai/oai-1.0/oai.xsd"</em> as has been done already in the sample files here. The sample file folders <em>testDVNcoll.zip </em>(Dataverse), <em>testFigColl.zip </em>(Figshare)<em> </em>and <em>testZenColl.zip </em>(Zenodo)<em> </em>contain all the metadata files tested and validated that are registered in the spreadsheet with objectIDs.<br> In the case of Zenodo, one original file feed,<br> <em>zen2018oai_datacite3orig-https%20_zenodo.org_oai2d%20verb=ListRecords%26metadata<br> Prefix=oai_datacite%26from=2018-11-29%26until=2018-11-30.xml</em> ,<br> is also supplied to show what was necessary to change in order to perform validation as indicated in the paper.</p> <p>For Dataverse, a corrected version of a file,<br> <em>dvn2014ddi-27595<strong>Corr</strong>_https%20_dataverse.harvard.edu_api_datasets_export%20<br> exporter=ddi%26persistentId=doi%253A10.7910_DVN_27595<strong>Corr</strong>.xml</em> ,<br> is also supplied in order to show the changes it would take to make the file validate without error.</p>
An Empirical Characterization of Event Sourced Systems and Their Schema Evolution - Lessons from Industry - Accompanying Anonymized Transcripts
<p>Anonymized interviews with 25 engineers on their experience applying Event Sourcing, with accompanying classifications. These transcripts are used in our publication "An Empirical Characterization of Event Sourced Systems and Their Schema Evolution - Lessons from Industry".</p>
Supplemental material of the Streptococcus pyogenes whole genome MLST schema deposited in Chewie-NS
<p>This supplemental material includes the lists of accession numbers for the Blackwell et al. and NCBI RefSeq assemblies used to populate the whole genome MLST schema for <em>Streptococcus pyogenes</em>, the UniProt identifiers of the reference proteomes used for schema annotation and the set of complete genomes, and associated metadata, used for schema creation.</p> <p>The wgMLST schema was created with <a href="https://github.com/B-UMMI/chewBBACA">chewBBACA</a> and is publicly available at <a href="https://chewbbaca.online/species/1/schemas/1">chewie-NS</a>, where a more detailed description of schema creation, annotation and curation can be found.</p>
Crosswalk between CESSDA Data Catalogue (CDC) Metadata Profile and ECRIN Metadata Schema.
<p>This dataset contains two files: (1) a crosswalk between CESSDA Data Catalogue (CDC) DDI2.5 Metadata Profile (<a href="https://cmv.cessda.eu/profiles/cdc/ddi-2.5/1.0.4/profile.html">https://cmv.cessda.eu/profiles/cdc/ddi-2.5/1.0.4/profile.html</a>) and ECRIN Metadata Schema for Clinical Research Data Objects Version 6.0 (August 2021) (<a href="https://zenodo.org/record/5554961">https://zenodo.org/record/5554961</a>) with an extension to “geographical data” and (2) vice versa.</p>
EvoBench: Benchmarking Schema Evolution in NoSQL
<p>Docker containers for reproducing the proof of concept measurements with our NoSQL Schema Evolution Benchmark.</p>
D2.1: Artefact, Contributor, and Organisation Relationship Data Schema - Appendix A
<p>Comparison of metadata schema for ORCID, DataCite, Dublin Core, CASRAI, MODS and DDI regarding contributors, organizations and artefacts.</p>
XML-Schema (Gemein-Nachrichten)
<p>Mit den <em><strong>Gemein-Nachrichten</strong></em> stellt das Unitätsarchiv Herrnhut der weltweiten Evangelischen Brüder-Unität - Herrnhuter Brüdergemeine (Unitas Fratrum / Moravian Church) das älteste und umfangreichste Mitteilungsblatt der Brüdergemeine digital zur Verfügung. Es enthält Berichte aus Gemeinden sowie dem Missions- und Diasporawerk der Brüdergemeine sowie Reden und Lebensläufe. Die <em>Gemein-Nachrichten</em> wurden ab 1765 in Fortsetzung des <em>Jüngerhaus-Diariums</em> (1747-1764) ausschließlich handschriftlich vervielfältigt. In Druck gingen 1817 und 1818 die <em>Beyträge aus der Brüder-Gemeine</em> und zwischen 1819 und 1894 die <em>Nachrichten aus der Brüder-Gemeine</em>. Das Nachrichtenblatt fand in den <em>Mitteilungen aus der Brüder-Gemeine zur Förderung christlicher Gemeinschaft</em> ab 1895 bis 1941 seine Fortsetzung.</p> <p>Mit dem hier vorliegenden <strong>XML-Schema</strong> werden erschlossene Transkripte mit standardisierten Metadaten angereichert.</p>
TEI-Schema für das Patristische Textarchiv (PTA)
This dataset contains the TEI-Schema and its documentation
An annotated corpus of clinical trial publications supporting schema-based relational information extraction
<p>Repository of an annotated corpus of clinical trial abstracts supporting schema-based relational information extraction and the code for the inter-annotation agreement calculation and the baseline information extraction method.</p>
Mean, standard deviation, and percentiles of the schema and domain scores of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3) in a German opportunity sample (n=1,150)
<p>Mean, standard deviation, and percentiles of the schema and domain scores of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3) in a German opportunity sample (n=1,150). Details are reported in: Kriston L, Schäfer J, Jacob GA, Härter M, Hölzel LP. Reliability and validity of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3). <em>Eur J Psychol Assess</em> 2013; 29: 205-212.</p> <p>IMPORTANT: This is an opportunity sample that is not representative of any well-defined population. Accordingly, the values should not be used as reference or norm values for the German version of the YSQ-S3.</p>
Mittelwert, Standardabweichung und Perzentile der Schema- und Domänen-Scores der deutschen Version des Young Schema Questionnaire - Short Form 3 (YSQ-S3) in einer deutschen Gelegenheitsstichprobe (n=1150)
<p>Mittelwert, Standardabweichung und Perzentile der Schema- und Domänen-Scores der deutschen Version des Young Schema Questionnaire - Short Form 3 (YSQ-S3) in einer deutschen Gelegenheitsstichprobe (n=1150). Details sind beschrieben in: Kriston L, Schäfer J, Jacob GA, Härter M, Hölzel LP. Reliability and validity of the German version of the Young Schema Questionnaire - Short Form 3 (YSQ-S3). <em>Eur J Psychol Assess</em> 2013; 29: 205-212.</p> <p>WICHTIG: Es handelt sich um eine Gelegenheitsstichprobe, die für keine gut definierbare Population repräsentativ ist. Dementsprechend sollten die Werte nicht als Referenz- oder Normwerte für die deutsche Version des YSQ-S3 verwendet werden.</p>
EFSA cgMLST gene lists for Escherichia coli and Salmonella enterica chewieNS schema
<p>Annex A contains the list of cgMLST loci of <em>Escherichia coli</em> and <em>Salmonella enterica</em> used in the cgMLST analysis in the EFSA One Health WGS System.</p>
Text-fig. 2. E-W cross section of the Urema Graben from Gorongosa to Inhaminga adapted from Flores (1973: fig. 5). Note that in this schema the Mazamba Sandstone directly overlies the Cheringoma Limestone. I.P.CO No. 5 is a bore hole. Vertical exaggeration ×10. in Stratigraphy, Chronology And Palaeontology Of The Tertiary Rocks Of The Cheringoma Plateau, Mozambique
Text-fig. 2. E-W cross section of the Urema Graben from Gorongosa to Inhaminga adapted from Flores (1973: fig. 5). Note that in this schema the Mazamba Sandstone directly overlies the Cheringoma Limestone. I.P.CO No. 5 is a bore hole. Vertical exaggeration ×10.
Figure 1. Schema for a data-integration solution-A Proposed Data Driven Architecture for Cardiology Network Application
<p>Data integration has favored loosening the coupling between data. This may involve<br> providing a uniform query interface over a mediated schema (see figure 1), thus transforming<br> a query into specialized queries over the original databases. One can also term this process<br> "view-based query-answering" because each of the data sources functions as a view over the<br> (nonexistent) mediated schema.</p>
SchemaPile: A Large Collection of Relational Database Schemas
<p>Access to fine-grained schema information is crucial for understanding how relational databases are designed and used in practice, and for building systems that help users interact with them. Furthermore, such information is required as training data to leverage the potential of large language models (LLMs) for improving data preparation, data integration and natural language querying.<br><br>Existing single-table corpora such as GitTables provide insights into how tables are structured in-the-wild, but lack detailed schema information about how tables relate to each other, as well as metadata like data types or integrity constraints. On the other hand, existing multi-table (or database schema) datasets are rather small and attribute-poor, leaving it unclear to what extent they actually represent typical real-world database schemas. </p> <p>In order to address these challenges, we present SchemaPile, a corpus of 221,171 database schemas, extracted from SQL files on GitHub. It contains 1.7 million tables with 10 million column definitions, 700 thousand foreign key relationships, seven million integrity constraints, and data content for more than 340 thousand tables. We conduct an in-depth analysis on the millions of schema metadata properties in our corpus, as well as its highly diverse language and topic distribution. In addition, we showcase the potential of SchemaPile to improve a variety of data management applications, e.g., fine-tuning LLMs for schema-only foreign key detection, improving CSV header detection and evaluating multi-dialect SQL parsers. We publish the code and data for recreating SchemaPile and a permissively licensed subset SchemaPile-Perm. </p> <p><a href="http://schemapile.github.io">http://schemapile.github.io</a></p>
BRAIN Journal-Intelligent Continuous Double Auction method For Service Allocation in Cloud Computing-Figure 1. Resource allocation schema in proposed method
<p>We assume that the resources allocation satisfies the following conditions:<br> • The quantity of a resource can be measured in arbitrary units (e.g. 60 units of resource<br> A).<br> • A resource can be divided into an arbitrary fraction (e.g. a resource of 60 units is divided<br> into 20 units for consumer 1 and 40 units for consumer 2).<br> • A resource request of a service can be divided into sub-requests and acquired from<br> multiple providers (e.g. a resource request of 40 units utilized as 10 units from provider<br> 1 and 30 units from provider 2).<br> Figure 1 shows a cloud computing environment with the proposed mechanism.</p>
Figure 1. Schema of clientside application for semantic browsing.-Browsing Semantic Data in Slovakia
<p>With the aim primary on unstructured information extraction and refining, relationship discovery and visualization, we propose our solution for SBR in the first place. The reason for this is, primary, that HTML formatted results of SBR are very jerky and uncertainty regarding the structure of information is very high. Readers can also be pointed by J. Suchal and P. Vojtek (2009), that care should be taken towards type errors. We discuss that later. In this work, we try to fill&up the gap of visualization and, somehow limited data access offered by SBR, adapting to the problems disclaimed above. We suggest a new client& side paradigm, which does not depend on a particular website like foaf.sk. Figure 1 describes the schema briefly and the key elements are parsers with other tools on the top and structured formats, for datastore, on the bottom.</p>
Figure 3. A brief schema of the CoDOA-SVM approach-Cognitive Development Optimization Algorithm Based Support Vector Machines for Determining Diabetes
<p>In the Equation 23, TP stands for true classified diabetes positive individuals; TN stands for true classified diabetes negative individuals; FP stands for false classified diabetes positive individuals and finally, FN stands for false classified diabetes negative individuals. • After determining good (optimum) particles, default CoDOA steps are run. • After achieving the total iteration number, it is allowed to train the SVM via optimum Gauss (RBF) kernel function parameters, by using the optimum particle value [sigma (σ) value]. • The trained SVM is now ready for the classification and so is diabetes determination process.<br> A brief schema of the CoDOA-SVM approach is also provided in Figure 3.</p>
Text-fig. 5. Original material of Peziza sulphurea (syntype L 910,261-594), the substrate and schema with location of six apothecia and two fragments of apothecia. Apothecia marked with numbers 1–3 were examined microscopically. in A Revision Of Trichopeziza Lizonii, T. Sulphurea And T. Violascens (Ascomycota, Helotiales) From The Herbarium Prm With Notes On Type Material Of Peziza Sulphurea
Text-fig. 5. Original material of Peziza sulphurea (syntype L 910,261-594), the substrate and schema with location of six apothecia and two fragments of apothecia. Apothecia marked with numbers 1–3 were examined microscopically.
Text-fig. 3. Schema of radial section of G. rudolphii (sample 99/04). t – tracheid, r – ray, bp – bordered pit, tp – taxodioid cross-field pit, gp – glyptostroboid cross-field pit. in New Fossil Woods From The Paleogene Of Doupovské Hory And České Středohoří Mts. (Bohemian Massif, Czech Republic)
Text-fig. 3. Schema of radial section of G. rudolphii (sample 99/04). t – tracheid, r – ray, bp – bordered pit, tp – taxodioid cross-field pit, gp – glyptostroboid cross-field pit.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.